lightweight and on-device AI runtimes
Recent window: last 5.2 months (2026-04-24 → 2026-09-28), compared with the prior 5.2 months.
These products provide compact models, distillation techniques, and optimized runtime engines designed to run AI locally on consumer hardware and edge devices. They are built for developers, hardware hackers, and embedded systems engineers seeking private, low-latency machine learning execution without expensive cloud GPUs. Unlike general cloud-hosted LLM APIs, this cluster focuses strictly on extreme memory efficiency and resource-constrained local inference.
Metrics
- Stage
- crowded
- Recent count
- 134
- Prior count
- 111
- Total count
- 264
- Momentum
- 20.72
- Attention
- 0.54
- Crowding
- 0.70
- Concentration
- 0.83
- Opportunity
- 0.33
Opportunity components
- Attention
- 0.54
- Low crowding
- 0.30
- Momentum (normalized)
- 0.30
- Low concentration
- 0.17
Monthly trajectory
Source split
- github
- 10 (0.07)
- hn
- 91 (0.68)
- ph
- 33 (0.25)
- yc
- 0 (0.00)
Dominant source: hn · Divergence: 0.68
Similar themes
- ML inference and model optimization (0.40)
- efficient local AI inference tools (0.38)
- gpu compute and acceleration tools (0.37)
- AI infrastructure and inference optimization (0.34)
- systems utilities and sandbox infrastructure (0.30)
- OpenAI-compatible AI API gateways (0.30)
- systems tools and desktop utilities (0.30)
- DeepSeek model deployment and inference (0.29)
Members
| Name | Source | Upvotes ▲ | Launched |
|---|---|---|---|
| General Compute | PH | 314 | 2026-05-22 |
| Moonshine Open-Weights STT models | HN | 316 | 2026-02-24 |
| LocalGPT | HN | 331 | 2026-02-08 |
| KiDoom | HN | 362 | 2025-11-25 |
| Ollama v0.19 | PH | 412 | 2026-04-01 |
| Google Gemma 4 | PH | 439 | 2026-04-03 |
| Echo | HN | 484 | 2026-07-23 |
| How I topped the HuggingFace open LLM leaderboard on two gaming GPUs | HN | 495 | 2026-03-10 |
| Z80-μLM, a 'Conversational AI' That Fits in 40KB | HN | 514 | 2025-12-29 |
| Needle2: 14MB agentic LLM for phones, wearables, smart home and robots | HN | 537 | 2026-08-10 |
| Three new Kitten TTS models | HN | 561 | 2026-03-19 |
| Needle: We Distilled Gemini Tool Calling into a 26M Model | HN | 776 | 2026-05-12 |
| Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac | HN | 919 | 2026-07-29 |
| Getting GLM 5.2 running on my slow computer | HN | 937 | 2026-07-09 |