Rapid-MLX
Run local LLMs on Mac, 2-3x faster than alternatives
Details
- External ID
- 47816238
- Source
- HN
- Company
- —
- Product
- Rapid-MLX
- Website domain
- github.com
- Launched
- April 18, 2026
- Cohort
- —
- Upvotes
- 9
- Upvotes percentile
- 0.5501285347043702
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Enrichment
- Theme
- ML inference and model optimization
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- fast local llm runtime for mac
- Manually corrected
- False
Could you build this?
No Beating existing optimized inference engines (like Apple MLX or llama.cpp) by 2-3x requires expert-level Metal Shading Language (MSL) kernel programming, GPU memory layout optimization, and hardware-level quantization schemes.
What it would actually take: The stack requires Apple Silicon MLX, C++, and custom Metal compute kernels. The hard parts involve low-latency attention mechanisms (e.g., FlashAttention customized for Unified Memory Architecture), kernel fusion for matrix multiplication and activations, and fine-tuned KV-cache paging. This requires deep GPU systems programming expertise and low-level Apple silicon architecture tuning.
Discussion
4 comments analyzed.
Competitors mentioned: Ollama, omlx, fast-mlx, existing inference servers
Concerns raised: Most models fail at structured tool calling, Existing servers are slow on MLX, Non-Qwen models have inconsistent tool calling (40-100% depending on framework)
Feature requests: Benchmarks against omlx and fast-mlx, Support for other open source projects beyond coding
Competitors
Other products that read as similar to this one — 627 launches clear the similarity bar, closest 8 shown.
Attention rank: #287 of 628 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 171 days after the earliest competitor.
- local-llms-with-mlx · github · 2026-09-13 · 18 upvotes · similarity 0.81
- Local AI · hn · 2026-02-05 · 5 upvotes · similarity 0.59
- PocketOrca-LLM · github · 2026-09-14 · 8 upvotes · similarity 0.58
- INXM // local` OSS for using LLM as compiler and not as runtime · hn · 2026-08-19 · 5 upvotes · similarity 0.56
- Self-host open-source LLMs on AWS with scale-to-zero · hn · 2026-09-09 · 7 upvotes · similarity 0.55
- Find the best local LLM for your hardware, ranked by benchmarks · hn · 2026-05-15 · 283 upvotes · similarity 0.54
- Timber (iOS) and Timber Tabs (Mac) · hn · 2026-08-26 · 12 upvotes · similarity 0.54
- Wafer Pass: flat-rate access to the fastest open-source LLMs · yc · 2026-04-30 · 7 upvotes · similarity 0.54
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.