Shoehorn, a library to quantize an LLM to fit your Mac's VRAM
Details
- External ID
- 49299386
- Source
- HN
- Company
- —
- Product
- Shoehorn, a library to quantize an LLM to fit your Mac's VRAM
- Website domain
- github.com
- Launched
- Aug. 14, 2026
- Cohort
- —
- Upvotes
- 6
- Upvotes percentile
- 0.3125
- Tags
- —
- Fetched at
- Sept. 10, 2026, 5:32 a.m.
- Updated at
- Sept. 10, 2026, 5:32 a.m.
Description
I made this after seeing someone posit the idea online yesterday over lunch then spent some time refining it. So far it's pretty impressive IMO! Right now I am running Qwen3-30B-A3B on my 24gb unified memory m4 MacBook Pro at 50 tok/sec and this should definitely not be working for such a large model on my middling hardware.Things are detailed in the README to get up and running and DESIGN.md has details on all the choices and such made along the way.
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- library to quantize llm for mac vram
- Manually corrected
- False
Could you build this?
Partial While wrapping existing quantization libraries (like llama.cpp / gguf / MLX) can be scripted, writing a tailored quantization and memory-profiling pipeline that dynamically fits large models into Apple Silicon unified memory at high inference speeds requires solid low-level ML engineering.
What it would actually take: A proper solution requires deep integration with Apple's Metal Performance Shaders (MPS) or MLX, writing custom quantization kernels (such as 2-bit to 4-bit AWQ or GGUF formats), and profiling unified memory bandwidth to dynamically adjust layer offloading and activation caches.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 246 launches clear the similarity bar, closest 8 shown.
Attention rank: #177 of 247 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 287 days after the earliest competitor.
- Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s · hn · 2026-09-01 · 240 upvotes · similarity 0.56
- KV-psi, using Linux PSI to to trim an LLM KV cache · hn · 2026-06-27 · 8 upvotes · similarity 0.54
- IronMule · ph · 2026-09-17 · 1 upvotes · similarity 0.50
- I run 30B 22tok/s, 109tok/s not novel,6GB/16GB RAM overcoming llama.cpp · hn · 2026-07-29 · 5 upvotes · similarity 0.50
- L88 · hn · 2026-02-24 · 12 upvotes · similarity 0.49
- I built a version of Omarchy that runs on Apple Silicon · hn · 2026-09-02 · 27 upvotes · similarity 0.48
- Shoehorn · hn · 2026-08-18 · 97 upvotes · similarity 0.48
- A beautiful and local-first PDF reader for studying dense things · hn · 2026-06-06 · 6 upvotes · similarity 0.48
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.