Bonsai 1.7B ternary model at 442T/s on M4 Max
Details
- External ID
- 48010204
- Source
- HN
- Company
- —
- Product
- Bonsai 1.7B ternary model at 442T/s on M4 Max
- Website domain
- agents2agents.ai
- Launched
- May 4, 2026
- Cohort
- —
- Upvotes
- 13
- Upvotes percentile
- 0.654281098546042
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
We took a recently released Bonsai 1.7B ternary model from PrismML (https://github.com/PrismML-Eng/Bonsai-demo) and ran our agentic evolution search on it for 6 hours to optimize the Metal kernels. The search was fully autonomous.Measured against unmodified upstream llama.cpp at the same Bonsai/Q2_0 commit, same M4 Max:- tg128: 309.82 → 442.42 t/s (+42.0%)- pp512: 4250.32 → 4622.63 t/s (+8.8%)
Enrichment
- Theme
- niche developer utilities and guides
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- ternary language model
- Manually corrected
- False
Could you build this?
No Engineering ultra-fast ternary LLM inference on Apple Silicon involves writing custom low-level Metal Shading Language (MSL) kernels, threadgroup memory management, and SIMD-level matrix multiplication optimizations.
What it would actually take: This requires modifying the llama.cpp or MLX runtime with hand-tuned Metal Compute Shaders (MSL) specifically designed for 1.58-bit / ternary weight packing. The developer needs deep GPU microarchitecture expertise to optimize SIMD-group matrix multiply-accumulate (mma) instructions, utilize Apple Silicon unified memory caching, and minimize bandwidth bottlenecks through register pressure tuning and automated kernel search loops.
Discussion
3 comments analyzed.
Competitors
Other products that read as similar to this one — 30 launches clear the similarity bar, closest 8 shown.
Attention rank: #11 of 31 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 167 days after the earliest competitor.
- bonsai2-small-gpu · github · 2026-09-19 · 54 upvotes · similarity 0.56
- lmstudio-prism-bonsai · github · 2026-09-21 · 9 upvotes · similarity 0.54
- ninfer-ternary-bonsai-ada · github · 2026-09-20 · 23 upvotes · similarity 0.54
- Running PrismML's Bonsai inside DRAM by breaking DDR4 timing rules · hn · 2026-07-23 · 23 upvotes · similarity 0.51
- Bonsai-27B-NInfer · github · 2026-09-24 · 10 upvotes · similarity 0.50
- mlxfast-bonsai2-27b-engine · github · 2026-09-24 · 30 upvotes · similarity 0.47
- OrcaBonsai-27B-Uncensored · github · 2026-09-18 · 526 upvotes · similarity 0.46
- Qwen3.8-Max · ph · 2026-08-03 · 276 upvotes · similarity 0.36
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.