Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Maple-Preview

Ternary 20B MoE running at 120 tok/s on a iPhone

Details

External ID
49173984
Source
HN
Company
—
Product
Maple-Preview
Website domain
deepgrove.ai
Launched
Aug. 4, 2026
Cohort
—
Upvotes
173
Upvotes percentile
0.9657258064516129
Tags
—
Fetched at
Sept. 10, 2026, 5:32 a.m.
Updated at
Sept. 10, 2026, 5:32 a.m.

Enrichment

Theme
security exploits and system hacking tools
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
efficient moe language model on mobile
Manually corrected
False

Could you build this?

No Running a 20B parameter ternary Mixture-of-Experts (MoE) LLM at 120 tokens/second on an iPhone requires novel research in ternary quantization and custom low-level Apple Metal/ANE kernel engineering.

What it would actually take: Requires an inference engine written in C++/Metal utilizing customized 1.58-bit (ternary) matrix multiplication kernels tailored specifically for Apple Silicon GPU and Neural Engine registers. The hard parts are novel model weight quantization without catastrophic perplexity loss, efficient dynamic routing of MoE sparse experts within mobile memory limits, and hand-tuned SIMD vector instructions for ternary arithmetic.

Discussion

20 comments analyzed.

Competitors mentioned: Qwen 35B-A3B, Llama 3.5 35B, Llama 3.6 35B, Binary Bonsai 27B, Ternary Bonsai 27B

Concerns raised: Models don't know when to search internet, think they already know answers, Can't compress human knowledge into few bits without significant quality loss, Knowledge accuracy degrades fastest with quantization, Small models generate surface-level answers, not reliably accurate facts, App size (5.9 GB) too large for iOS, would trigger jetsam on most iPhones

Feature requests: Add tool/API integration for web search capability, Enable local Wikipedia or knowledge base querying via embedding search, Support for vision capabilities (VLM not text-only), Better reasoning and tool-use for conversational requirements

Competitors

Other products that read as similar to this one — 141 launches clear the similarity bar, closest 8 shown.

Attention rank: #9 of 142 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 269 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.