Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Run 500B+ Parameter LLMs Locally on a Mac Mini

Details

External ID
47305561
Source
HN
Company
—
Product
Run 500B+ Parameter LLMs Locally on a Mac Mini
Website domain
github.com
Launched
March 9, 2026
Cohort
—
Upvotes
17
Upvotes percentile
0.7103321033210332
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Hi HN, I built OpenGraviton, an open-source AI inference engine that pushes the limits of running extremely large LLMs on consumer hardware. By combining 1.58-bit ternary quantization, dynamic sparsity with Top-K pruning and MoE routing, and mmap-based layer streaming, OpenGraviton can run models far larger than your system RAM—even on a Mac Mini. Early benchmarks: TinyLlama-1.1B drops from ~2GB (FP16) to ~0.24GB with ternary quantization. At 140B scale, models that normally require ~280GB fit within ~35GB packed. Optimized for Apple Silicon with Metal + C++ tensor unpacking, plus speculative decoding for faster generation. Check benchmarks, architecture, and details here: https://opengraviton.github.io GitHub: https://github.com/opengraviton This project isn’t just about squeezing massive models onto tiny hardware—it’s about democratizing access to giant LLMs without cloud costs. Feedback, forks, and ideas are very welcome!

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
run large language models locally on mac mini
Manually corrected
False

Could you build this?

No Creating an inference engine capable of running 500B+ parameter models on consumer hardware demands expert low-level systems programming, custom Metal/GPU kernels, ternary quantization math, and layer streaming.

What it would actually take: This requires writing custom C++/Metal kernels for 1.58-bit ternary matrix-vector multiplications, building dynamic Top-K routing routines for Mixture-of-Experts, and managing OS virtual memory via customized mmap layer-streaming architectures to overcome physical RAM constraints. Developers need deep expertise in Apple Silicon hardware architecture, memory bandwidth saturation, and low-bit quantized model compilation.

Discussion

10 comments analyzed.

Competitors mentioned: Gemini API, Cloud-based LLM APIs

Concerns raised: Quantization to 1.58-bit creates only 'mini' versions of models, not true full models, Benchmarks lack real-world results for larger-than-memory models beyond TinyLlama, Memory pressure and thermal throttling on sustained inference with Mac Mini 16GB, Latency degradation when layers stream from disk during inference at 140B scale, Sloppy or incomplete benchmark documentation

Feature requests: Benchmarks for 140B models on base Mac Mini 16GB configuration, Performance data for M1 Max and M3 Ultra systems, Latency metrics for layer streaming from disk mid-inference, Real-world token/sec results for larger models beyond TinyLlama

Competitors

Other products that read as similar to this one — 213 launches clear the similarity bar, closest 8 shown.

Attention rank: #74 of 214 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 126 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.