Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo)

Details

External ID
49242475
Source
HN
Company
—
Product
A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo)
Website domain
mikeayles.com
Launched
Aug. 10, 2026
Cohort
—
Upvotes
79
Upvotes percentile
0.907258064516129
Tags
—
Fetched at
Sept. 10, 2026, 5:32 a.m.
Updated at
Sept. 10, 2026, 5:32 a.m.

Enrichment

Theme
ML inference and model optimization
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
efficient llm inference on fpga
Manually corrected
False

Could you build this?

No This project implements custom hardware HDL to execute an LLM directly on an FPGA chip with on-chip weights reaching tens of thousands of tokens per second. Building FPGA-based deep learning accelerators requires specialized digital hardware design (Verilog/VHDL/Chisel) and computer architecture expertise.

What it would actually take: The system requires designing a custom hardware matrix-vector multiplication pipeline in RTL (Verilog/VHDL or Bluespec), memory controller interfaces to FPGA BRAM/URAM, and custom quantization/weight-packing schemes. The core difficulty is timing closure, pipeline stalling mitigation, and resource utilization optimization under hardware constraints on a budget FPGA. This demands expert hardware and embedded systems engineers.

Discussion

20 comments analyzed.

Competitors mentioned: Cactus models, Alveo V80, GPUs for inference, ASICs, CIM/analog-compute startups

Concerns raised: FPGAs are niche with small total market volume, Can't cost-effectively compete with other options for LLM inference, Model too large for time series/sensor data processing, HLS is a dead-end approach, FPGA vendors optimize for larger chips, not cost-effectiveness

Feature requests: Make four ROMs boot-loadable for hot-swapping models without reconfiguration, Dynamic partial reconfiguration to avoid full bitstream reloads, Support for novel memory technologies for weight storage, Better tooling than HLS and SystemVerilog

Competitors

Other products that read as similar to this one — 331 launches clear the similarity bar, closest 8 shown.

Attention rank: #35 of 332 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 280 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.