Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Running Gemma-4 26B at 124 tokens/SEC on a CPU, no GPU

Details

External ID
48731643
Source
HN
Company
—
Product
Running Gemma-4 26B at 124 tokens/SEC on a CPU, no GPU
Website domain
apeg.dev
Launched
June 30, 2026
Cohort
—
Upvotes
10
Upvotes percentile
0.5935792349726776
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

I wanted to know how fast a 26B mixture-of-experts model could run on a desktop CPU with no GPU. Got ~40 tok/s single-stream (lossless) and ~124 batched. The surprising part was the byte budget: for this model you compress the output head (32% of per-token bytes), not the experts (16%). The writeup has the bandwidth roofline and the dead-ends; the repo has the reproducible recipe. Happy to answer questions.Repo: https://github.com/arun-prasath2005/gemma4-cpu-moe

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
Not AI
Project type
Hobby / open-source project
Normalized one-liner
cpu-based gemma-4 26b inference
Manually corrected
False

Could you build this?

No Achieving 124 tokens/sec for a 26B MoE model on a CPU without a GPU requires high-performance systems engineering, low-level CPU vector assembly (AVX-512/AMX), and novel quantization/compression schemes.

What it would actually take: Requires custom C++/assembly or high-performance C inference kernels tailored to specific CPU cache architectures and SIMD instruction sets (AVX-512, AMX, or ARM NEON). The core difficulty lies in writing custom quantization kernels specifically optimizing output-head matrix multiplication alongside memory-bandwidth-bound MoE routing without hardware accelerators, far beyond general LLM API wrappers.

Discussion

1 comment analyzed.

Concerns raised: Output head byte budget is surprisingly high

Feature requests: More aggressive compression of head while keeping experts mostly untouched

Competitors

Other products that read as similar to this one — 93 launches clear the similarity bar, closest 8 shown.

Attention rank: #53 of 94 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 244 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.