Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Details
- External ID
- 49098510
- Source
- HN
- Company
- —
- Product
- Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
- Website domain
- github.com
- Launched
- July 29, 2026
- Cohort
- —
- Upvotes
- 919
- Upvotes percentile
- 0.994026284348865
- Tags
- —
- Fetched at
- Sept. 8, 2026, 8:28 p.m.
- Updated at
- Sept. 8, 2026, 8:28 p.m.
Description
Hi HN,I built a specialized inference engine for running 4-bit Gemma 4 26B-A4B-IT on any M-series Mac using about 2 GB of RAM. It is called TurboFieldfare and is written in Swift and Metal.I have always adored on-device AI. It feels like magic that you can run a powerful NN on your Mac or iPhone. So I wanted to push the limits a bit and run a model whose weights don’t fit in memory.The model’s 4-bit quantized weights occupy roughly 14 GB, which makes running it with conventional inference tools almost impossible on an 8 GB or even 16 GB Mac once the OS, applications, and KV cache are included.The trick is to keep the shared part of the model and the KV cache in RAM, then stream only the routed experts needed for each token from SSD. An SSD is way slower than RAM, so the runtime uses a small expert cache and bounded parallel `pread`. While those reads are in flight, the GPU runs the shared part of the layer.I ran more than 100 experiments. Most didn’t work. A few got me here. The experiments are described in the GitHub repo.It currently generates 5–6 tok/s on an 8 GB M2 MacBook Air and 31–35 tok/s on an M5 MacBook Pro.I also added an experimental OpenAI-compatible local server. It supports streaming and tool calls, and reuses one prompt prefix from the KV cache.Try it! The Mac app is easy to install. On the first run, it will download 15 GB of weights from Hugging Face. The model is surprisingly capable.I would love any kind of feedback!
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- open-source engine running gemma 4 26b on m-series mac with 2gb ram
- Manually corrected
- False
Could you build this?
No Designing an ultra-low-memory inference engine in Swift and Metal for a 26B model requires deep GPU kernel programming and proprietary memory-efficient execution strategies.
What it would actually take: The architecture requires hand-written Metal Shading Language (MSL) compute shaders, 4-bit dequantization kernels, custom unified memory allocation, and KV-cache offloading techniques. It also demands intimate knowledge of Apple Silicon's GPU architecture, memory bandwidth constraints, and matrix coprocessors (AMX). Specialized skills in GPU systems engineering and low-level ML compiler optimization are essential.
Discussion
20 comments analyzed.
Competitors mentioned: Asahi Linux, Qwen models, Ollama, Gemma 4, DeepSeek
Concerns raised: Performance slow on M1 Max and 8GB RAM devices, Storage bandwidth limitations on Mac devices, Thinking loops in quantized variants, Model response quality differences from unconstrained versions, Tool-calling issues in some model variants
Feature requests: Streaming chat via command line for multi-user SSH access, Non-streaming terminal client for interactive chat, Support for new DeepSeek models, Branch predictor for expert weight selection, Increased flash storage bandwidth via parallel chips
Competitors
Other products that read as similar to this one — 149 launches clear the similarity bar, closest 8 shown.
Attention rank: #7 of 150 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 263 days after the earliest competitor.
- I made a Gemma 4 Mac app that names screenshots with local AI · hn · 2026-05-31 · 7 upvotes · similarity 0.61
- Running Gemma-4 26B at 124 tokens/SEC on a CPU, no GPU · hn · 2026-06-30 · 10 upvotes · similarity 0.60
- OpenGraviton · hn · 2026-03-07 · 13 upvotes · similarity 0.54
- Google Gemma 4 12B · ph · 2026-06-04 · 309 upvotes · similarity 0.53
- A new engine to run Kimi K3 on a laptop · hn · 2026-07-29 · 7 upvotes · similarity 0.52
- Run 500B+ Parameter LLMs Locally on a Mac Mini · hn · 2026-03-09 · 17 upvotes · similarity 0.52
- RunNburn · hn · 2026-07-30 · 11 upvotes · similarity 0.51
- Cactus Hybrid: We taught Gemma 4 to know when it's wrong · hn · 2026-07-22 · 191 upvotes · similarity 0.50
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.