Running Gemma-4 26B at 124 tokens/SEC on a CPU, no GPU
Details
- External ID
- 48731643
- Source
- HN
- Company
- —
- Product
- Running Gemma-4 26B at 124 tokens/SEC on a CPU, no GPU
- Website domain
- apeg.dev
- Launched
- June 30, 2026
- Cohort
- —
- Upvotes
- 10
- Upvotes percentile
- 0.5935792349726776
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
I wanted to know how fast a 26B mixture-of-experts model could run on a desktop CPU with no GPU. Got ~40 tok/s single-stream (lossless) and ~124 batched. The surprising part was the byte budget: for this model you compress the output head (32% of per-token bytes), not the experts (16%). The writeup has the bandwidth roofline and the dead-ends; the repo has the reproducible recipe. Happy to answer questions.Repo: https://github.com/arun-prasath2005/gemma4-cpu-moe
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- cpu-based gemma-4 26b inference
- Manually corrected
- False
Could you build this?
No Achieving 124 tokens/sec for a 26B MoE model on a CPU without a GPU requires high-performance systems engineering, low-level CPU vector assembly (AVX-512/AMX), and novel quantization/compression schemes.
What it would actually take: Requires custom C++/assembly or high-performance C inference kernels tailored to specific CPU cache architectures and SIMD instruction sets (AVX-512, AMX, or ARM NEON). The core difficulty lies in writing custom quantization kernels specifically optimizing output-head matrix multiplication alongside memory-bandwidth-bound MoE routing without hardware accelerators, far beyond general LLM API wrappers.
Discussion
1 comment analyzed.
Concerns raised: Output head byte budget is surprisingly high
Feature requests: More aggressive compression of head while keeping experts mostly untouched
Competitors
Other products that read as similar to this one — 93 launches clear the similarity bar, closest 8 shown.
Attention rank: #53 of 94 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 244 days after the earliest competitor.
- Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac · hn · 2026-07-29 · 919 upvotes · similarity 0.60
- I made a Gemma 4 Mac app that names screenshots with local AI · hn · 2026-05-31 · 7 upvotes · similarity 0.51
- RunNburn · hn · 2026-07-30 · 11 upvotes · similarity 0.51
- A new engine to run Kimi K3 on a laptop · hn · 2026-07-29 · 7 upvotes · similarity 0.47
- Cactus Hybrid: We taught Gemma 4 to know when it's wrong · hn · 2026-07-22 · 191 upvotes · similarity 0.47
- ExANS · hn · 2026-08-05 · 15 upvotes · similarity 0.46
- Google Gemma 4 12B · ph · 2026-06-04 · 309 upvotes · similarity 0.45
- I embedded 685M public texts in 32 minutes (on 8x A100, Rust, TensorRT) · hn · 2026-06-04 · 7 upvotes · similarity 0.43
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.