OpenGraviton
Run 500B+ parameter models on a consumer Mac Mini
Details
- External ID
- 47289127
- Source
- HN
- Company
- —
- Product
- OpenGraviton
- Website domain
- github.io
- Launched
- March 7, 2026
- Cohort
- —
- Upvotes
- 13
- Upvotes percentile
- 0.6660516605166051
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Hi HN,I built OpenGraviton, an open-source AI inference engine designed to push the limits of running extremely large models on consumer hardware.The system combines several techniques to drastically reduce memory and compute requirements:• 1.58-bit ternary quantization ({-1, 0, +1}) for ~10x compression • dynamic sparsity with Top-K pruning and MoE routing • mmap-based layer streaming to load weights directly from NVMe SSDs • speculative decoding to improve generation throughputThese allow models far larger than system RAM to run locally.In early benchmarks, OpenGraviton reduced TinyLlama-1.1B from ~2.05GB (FP16) to ~0.24GB using ternary quantization. Synthetic stress tests at the 140B scale show that models which would normally require ~280GB FP16 can fit within ~35GB when packed with the ternary format.The project is optimized for Apple Silicon and currently uses custom Metal + C++ tensor unpacking.Benchmarks, architecture, and details: https://opengraviton.github.ioGitHub: https://github.com/opengraviton
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- run large language models on consumer hardware
- Manually corrected
- False
Could you build this?
No Implementing a custom local inference engine supporting 1.58-bit ternary quantization, layer-by-layer SSD streaming, and custom GPU/Metal compute kernels requires deep low-level systems and ML compilation expertise.
What it would actually take: Requires high-performance C++/Rust or Metal/CUDA engineering to write custom matrix multiplication kernels optimized for ternary or low-bit weights. The system needs asynchronous disk I/O pipelines (io_uring on Linux, direct I/O on macOS) to stream multi-gigabyte layer weights continuously into GPU unified memory without stalling compute. Building speculative decoding and dynamic sparsity into an engine requires expert-level understanding of transformer internals and hardware cache architectures.
Discussion
5 comments analyzed.
Concerns raised: Hardware detection issues, engine.generate() not implemented, Missing implementation yields empty string
Feature requests: Implement engine.generate() method, Improve hardware detection
Competitors
Other products that read as similar to this one — 232 launches clear the similarity bar, closest 8 shown.
Attention rank: #83 of 233 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 123 days after the earliest competitor.
- Run 500B+ Parameter LLMs Locally on a Mac Mini · hn · 2026-03-09 · 17 upvotes · similarity 0.88
- Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac · hn · 2026-07-29 · 919 upvotes · similarity 0.54
- Serve 100 Large AI models on a single GPU with low impact to TTFT · hn · 2025-11-08 · 7 upvotes · similarity 0.47
- IronMule · ph · 2026-09-17 · 1 upvotes · similarity 0.46
- Moonshine Open-Weights STT models · hn · 2026-02-24 · 316 upvotes · similarity 0.46
- Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload · hn · 2026-07-26 · 22 upvotes · similarity 0.45
- Mini-AGI · hn · 2026-09-21 · 277 upvotes · similarity 0.45
- Tiny-vLLM · hn · 2026-05-29 · 205 upvotes · similarity 0.44
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.