Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
Details
- External ID
- 49158333
- Source
- HN
- Company
- —
- Product
- Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
- Website domain
- github.com
- Launched
- Aug. 3, 2026
- Cohort
- —
- Upvotes
- 312
- Upvotes percentile
- 0.9879032258064516
- Tags
- —
- Fetched at
- Sept. 10, 2026, 5:32 a.m.
- Updated at
- Sept. 10, 2026, 5:32 a.m.
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- run large language models on resource-constrained devices
- Manually corrected
- False
Could you build this?
No Fitting an 80B model into 4.3 GB of RAM requires extreme sub-1-bit quantization research (such as BitNet, AQLM, or custom ternary/binary matrix multiplication kernels) and deep hardware-level optimization for Apple Silicon/Metal.
What it would actually take: This project requires custom low-level C++/Metal kernels implementing extreme quantization schemes (e.g., 0.5 to 1.58 bits per weight) alongside tailored activation memory management. The developer must have deep expertise in low-precision numerical methods, SIMD/NEON/Metal assembly optimization, and transformer inference runtime internals. Off-the-shelf vibe coding cannot produce correct or performant custom GPU/NPU kernels at this level of algorithmic complexity.
Discussion
20 comments analyzed.
Competitors mentioned: iPhone local LLMs, Ollama, TurboFieldfare, Chinese LLMs
Concerns raised: Prefill latency takes too long, Frontier labs will prevent open-weight model adoption, SSD storage solution won't be cost-effective within 10 years, Greedy pricing stalls progress and adoption, Noise and performance concerns with local inference
Feature requests: Support for Android/Linux/Windows platforms, Gemma model support, Improved prefill performance
Competitors
Other products that read as similar to this one — 365 launches clear the similarity bar, closest 8 shown.
Attention rank: #9 of 366 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 269 days after the earliest competitor.
- Samosa Chat · hn · 2026-07-15 · 6 upvotes · similarity 0.64
- Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s · hn · 2026-09-01 · 240 upvotes · similarity 0.58
- Run Full Kimi K3 with 29 GB of RAM · hn · 2026-07-30 · 9 upvotes · similarity 0.57
- Qwen3.6-35B-A3B on a 16 GB M1 Pro with SSD-streamed MoE · hn · 2026-07-17 · 24 upvotes · similarity 0.56
- qwen3.6-35b-a3b-144T-S · github · 2026-09-20 · 15 upvotes · similarity 0.54
- Synapse · hn · 2026-06-02 · 18 upvotes · similarity 0.51
- Maple-Preview · hn · 2026-08-04 · 173 upvotes · similarity 0.51
- emu · github · 2026-09-25 · 34 upvotes · similarity 0.50
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.