Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

OpenGraviton

Run 500B+ parameter models on a consumer Mac Mini

Details

External ID
47289127
Source
HN
Company
—
Product
OpenGraviton
Website domain
github.io
Launched
March 7, 2026
Cohort
—
Upvotes
13
Upvotes percentile
0.6660516605166051
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Hi HN,I built OpenGraviton, an open-source AI inference engine designed to push the limits of running extremely large models on consumer hardware.The system combines several techniques to drastically reduce memory and compute requirements:• 1.58-bit ternary quantization ({-1, 0, +1}) for ~10x compression • dynamic sparsity with Top-K pruning and MoE routing • mmap-based layer streaming to load weights directly from NVMe SSDs • speculative decoding to improve generation throughputThese allow models far larger than system RAM to run locally.In early benchmarks, OpenGraviton reduced TinyLlama-1.1B from ~2.05GB (FP16) to ~0.24GB using ternary quantization. Synthetic stress tests at the 140B scale show that models which would normally require ~280GB FP16 can fit within ~35GB when packed with the ternary format.The project is optimized for Apple Silicon and currently uses custom Metal + C++ tensor unpacking.Benchmarks, architecture, and details: https://opengraviton.github.ioGitHub: https://github.com/opengraviton

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
run large language models on consumer hardware
Manually corrected
False

Could you build this?

No Implementing a custom local inference engine supporting 1.58-bit ternary quantization, layer-by-layer SSD streaming, and custom GPU/Metal compute kernels requires deep low-level systems and ML compilation expertise.

What it would actually take: Requires high-performance C++/Rust or Metal/CUDA engineering to write custom matrix multiplication kernels optimized for ternary or low-bit weights. The system needs asynchronous disk I/O pipelines (io_uring on Linux, direct I/O on macOS) to stream multi-gigabyte layer weights continuously into GPU unified memory without stalling compute. Building speculative decoding and dynamic sparsity into an engine requires expert-level understanding of transformer internals and hardware cache architectures.

Discussion

5 comments analyzed.

Concerns raised: Hardware detection issues, engine.generate() not implemented, Missing implementation yields empty string

Feature requests: Implement engine.generate() method, Improve hardware detection

Competitors

Other products that read as similar to this one — 232 launches clear the similarity bar, closest 8 shown.

Attention rank: #83 of 233 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 123 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.