Maple-Preview
Ternary 20B MoE running at 120 tok/s on a iPhone
Details
- External ID
- 49173984
- Source
- HN
- Company
- —
- Product
- Maple-Preview
- Website domain
- deepgrove.ai
- Launched
- Aug. 4, 2026
- Cohort
- —
- Upvotes
- 173
- Upvotes percentile
- 0.9657258064516129
- Tags
- —
- Fetched at
- Sept. 10, 2026, 5:32 a.m.
- Updated at
- Sept. 10, 2026, 5:32 a.m.
Enrichment
- Theme
- security exploits and system hacking tools
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- efficient moe language model on mobile
- Manually corrected
- False
Could you build this?
No Running a 20B parameter ternary Mixture-of-Experts (MoE) LLM at 120 tokens/second on an iPhone requires novel research in ternary quantization and custom low-level Apple Metal/ANE kernel engineering.
What it would actually take: Requires an inference engine written in C++/Metal utilizing customized 1.58-bit (ternary) matrix multiplication kernels tailored specifically for Apple Silicon GPU and Neural Engine registers. The hard parts are novel model weight quantization without catastrophic perplexity loss, efficient dynamic routing of MoE sparse experts within mobile memory limits, and hand-tuned SIMD vector instructions for ternary arithmetic.
Discussion
20 comments analyzed.
Competitors mentioned: Qwen 35B-A3B, Llama 3.5 35B, Llama 3.6 35B, Binary Bonsai 27B, Ternary Bonsai 27B
Concerns raised: Models don't know when to search internet, think they already know answers, Can't compress human knowledge into few bits without significant quality loss, Knowledge accuracy degrades fastest with quantization, Small models generate surface-level answers, not reliably accurate facts, App size (5.9 GB) too large for iOS, would trigger jetsam on most iPhones
Feature requests: Add tool/API integration for web search capability, Enable local Wikipedia or knowledge base querying via embedding search, Support for vision capabilities (VLM not text-only), Better reasoning and tool-use for conversational requirements
Competitors
Other products that read as similar to this one — 141 launches clear the similarity bar, closest 8 shown.
Attention rank: #9 of 142 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 269 days after the earliest competitor.
- Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone · hn · 2026-08-03 · 312 upvotes · similarity 0.51
- Qwen3.6-35B-A3B on a 16 GB M1 Pro with SSD-streamed MoE · hn · 2026-07-17 · 24 upvotes · similarity 0.44
- bonsai-apple-silicon · github · 2026-09-18 · 24 upvotes · similarity 0.42
- aiistream-q3.6 · github · 2026-09-29 · 10 upvotes · similarity 0.42
- emu · github · 2026-09-25 · 34 upvotes · similarity 0.41
- wloc · github · 2026-09-10 · 18 upvotes · similarity 0.39
- S3 compatible store with 1M IOPS(4K-R,p99~5ms), BYOC in 5min with rust · hn · 2025-12-07 · 23 upvotes · similarity 0.39
- SwiftIPA · github · 2026-09-22 · 9 upvotes · similarity 0.39
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.