Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac

Details

External ID
49098510
Source
HN
Company
—
Product
Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Website domain
github.com
Launched
July 29, 2026
Cohort
—
Upvotes
919
Upvotes percentile
0.994026284348865
Tags
—
Fetched at
Sept. 8, 2026, 8:28 p.m.
Updated at
Sept. 8, 2026, 8:28 p.m.

Description

Hi HN,I built a specialized inference engine for running 4-bit Gemma 4 26B-A4B-IT on any M-series Mac using about 2 GB of RAM. It is called TurboFieldfare and is written in Swift and Metal.I have always adored on-device AI. It feels like magic that you can run a powerful NN on your Mac or iPhone. So I wanted to push the limits a bit and run a model whose weights don’t fit in memory.The model’s 4-bit quantized weights occupy roughly 14 GB, which makes running it with conventional inference tools almost impossible on an 8 GB or even 16 GB Mac once the OS, applications, and KV cache are included.The trick is to keep the shared part of the model and the KV cache in RAM, then stream only the routed experts needed for each token from SSD. An SSD is way slower than RAM, so the runtime uses a small expert cache and bounded parallel `pread`. While those reads are in flight, the GPU runs the shared part of the layer.I ran more than 100 experiments. Most didn’t work. A few got me here. The experiments are described in the GitHub repo.It currently generates 5–6 tok/s on an 8 GB M2 MacBook Air and 31–35 tok/s on an M5 MacBook Pro.I also added an experimental OpenAI-compatible local server. It supports streaming and tool calls, and reuses one prompt prefix from the KV cache.Try it! The Mac app is easy to install. On the first run, it will download 15 GB of weights from Hugging Face. The model is surprisingly capable.I would love any kind of feedback!

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
open-source engine running gemma 4 26b on m-series mac with 2gb ram
Manually corrected
False

Could you build this?

No Designing an ultra-low-memory inference engine in Swift and Metal for a 26B model requires deep GPU kernel programming and proprietary memory-efficient execution strategies.

What it would actually take: The architecture requires hand-written Metal Shading Language (MSL) compute shaders, 4-bit dequantization kernels, custom unified memory allocation, and KV-cache offloading techniques. It also demands intimate knowledge of Apple Silicon's GPU architecture, memory bandwidth constraints, and matrix coprocessors (AMX). Specialized skills in GPU systems engineering and low-level ML compiler optimization are essential.

Discussion

20 comments analyzed.

Competitors mentioned: Asahi Linux, Qwen models, Ollama, Gemma 4, DeepSeek

Concerns raised: Performance slow on M1 Max and 8GB RAM devices, Storage bandwidth limitations on Mac devices, Thinking loops in quantized variants, Model response quality differences from unconstrained versions, Tool-calling issues in some model variants

Feature requests: Streaming chat via command line for multi-user SSH access, Non-streaming terminal client for interactive chat, Support for new DeepSeek models, Branch predictor for expert weight selection, Increased flash storage bandwidth via parallel chips

Competitors

Other products that read as similar to this one — 149 launches clear the similarity bar, closest 8 shown.

Attention rank: #7 of 150 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 263 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.