A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo)
Details
- External ID
- 49242475
- Source
- HN
- Company
- —
- Product
- A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo)
- Website domain
- mikeayles.com
- Launched
- Aug. 10, 2026
- Cohort
- —
- Upvotes
- 79
- Upvotes percentile
- 0.907258064516129
- Tags
- —
- Fetched at
- Sept. 10, 2026, 5:32 a.m.
- Updated at
- Sept. 10, 2026, 5:32 a.m.
Enrichment
- Theme
- ML inference and model optimization
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- efficient llm inference on fpga
- Manually corrected
- False
Could you build this?
No This project implements custom hardware HDL to execute an LLM directly on an FPGA chip with on-chip weights reaching tens of thousands of tokens per second. Building FPGA-based deep learning accelerators requires specialized digital hardware design (Verilog/VHDL/Chisel) and computer architecture expertise.
What it would actually take: The system requires designing a custom hardware matrix-vector multiplication pipeline in RTL (Verilog/VHDL or Bluespec), memory controller interfaces to FPGA BRAM/URAM, and custom quantization/weight-packing schemes. The core difficulty is timing closure, pipeline stalling mitigation, and resource utilization optimization under hardware constraints on a budget FPGA. This demands expert hardware and embedded systems engineers.
Discussion
20 comments analyzed.
Competitors mentioned: Cactus models, Alveo V80, GPUs for inference, ASICs, CIM/analog-compute startups
Concerns raised: FPGAs are niche with small total market volume, Can't cost-effectively compete with other options for LLM inference, Model too large for time series/sensor data processing, HLS is a dead-end approach, FPGA vendors optimize for larger chips, not cost-effectiveness
Feature requests: Make four ROMs boot-loadable for hot-swapping models without reconfiguration, Dynamic partial reconfiguration to avoid full bitstream reloads, Support for novel memory technologies for weight storage, Better tooling than HLS and SystemVerilog
Competitors
Other products that read as similar to this one — 331 launches clear the similarity bar, closest 8 shown.
Attention rank: #35 of 332 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 280 days after the earliest competitor.
- Minimal LLM Post-Training Experiments on an 8GB GPU (SFT, DPO, GRPO) · hn · 2026-08-01 · 21 upvotes · similarity 0.57
- PocketOrca-LLM · github · 2026-09-14 · 8 upvotes · similarity 0.56
- A tiny C program where an LLM rewires its DAG while running · hn · 2026-05-05 · 15 upvotes · similarity 0.54
- BootLife · github · 2026-09-20 · 55 upvotes · similarity 0.49
- I built a portable Yahtzee device with custom PCB and WASM simulator · hn · 2025-12-31 · 8 upvotes · similarity 0.48
- Local AI · hn · 2026-02-05 · 5 upvotes · similarity 0.48
- LLM Inference Calculator · hn · 2026-08-28 · 6 upvotes · similarity 0.47
- I wrote a 1-bit WebGPU runtime to run a 1.7B LLM in the browser · hn · 2026-07-07 · 5 upvotes · similarity 0.46
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.