ML inference and model optimization
Recent window: last 5.1 months (2026-04-24 → 2026-09-27), compared with the prior 5.1 months.
These products provide specialized runtimes, compression techniques, and acceleration engines to deploy and benchmark machine learning models efficiently. They are designed for machine learning engineers, systems developers, and AI researchers working with constrained hardware or demanding latency requirements. The cluster focuses directly on low-level compute, quantization, and runtime performance rather than high-level application wrappers or end-user workflow tools.
Metrics
- Stage
- crowded
- Recent count
- 151
- Prior count
- 49
- Total count
- 223
- Momentum
- 208.16
- Attention
- 0.46
- Crowding
- 0.79
- Concentration
- 0.84
- Opportunity
- 0.40
Opportunity components
- Attention
- 0.46
- Low crowding
- 0.21
- Momentum (normalized)
- 0.77
- Low concentration
- 0.16
Monthly trajectory
Source split
- github
- 71 (0.47)
- hn
- 61 (0.40)
- ph
- 10 (0.07)
- yc
- 9 (0.06)
Dominant source: github · Divergence: 0.41
Similar themes
- AI infrastructure and inference optimization (0.72)
- voice AI and speech tools (0.61)
- gpu compute and acceleration tools (0.61)
- decision model runtimes and tools (0.56)
- persistent memory for AI agents (0.55)
- systems tools and desktop utilities (0.54)
- low-level systems and developer tools (0.54)
- autonomous agent research and evaluation (0.54)
Members
| Name | Source | Upvotes ▼ | Launched |
|---|---|---|---|
| CLM | GITHUB | 1815 | 2026-09-23 |
| laya-coreml | GITHUB | 1381 | 2026-09-19 |
| deepopen | GITHUB | 1016 | 2026-09-21 |
| I built a tiny LLM to demystify how language models work | HN | 915 | 2026-04-06 |
| mini-AGI | GITHUB | 692 | 2026-09-19 |
| splash | GITHUB | 610 | 2026-09-18 |
| OrcaBonsai-27B-Uncensored | GITHUB | 526 | 2026-09-18 |
| 1-Bit Bonsai, the First Commercially Viable 1-Bit LLMs | HN | 430 | 2026-03-31 |
| TurboQuant | PH | 295 | 2026-03-25 |
| Find the best local LLM for your hardware, ranked by benchmarks | HN | 283 | 2026-05-15 |
| Duplicate 3 layers in a 24B LLM, logical deduction .22→.76. No training | HN | 265 | 2026-03-18 |
| cek-probe-model | GITHUB | 260 | 2026-09-10 |
| Data Engineering Book | HN | 251 | 2026-02-13 |
| routeVSCODE | GITHUB | 239 | 2026-09-10 |
| Timber | HN | 207 | 2026-03-02 |
| Tiny-vLLM | HN | 205 | 2026-05-29 |
| Mellum by JetBrains | PH | 196 | 2026-06-20 |
| DeepEP-Ascend | GITHUB | 174 | 2026-09-30 |
| LLM Attention Visualization | HN | 173 | 2026-09-08 |
| JevRouter | GITHUB | 163 | 2026-09-18 |
| Pollen | HN | 137 | 2026-04-30 |
| ThoughtDAG | HN | 136 | 2026-08-15 |
| llmwiki-operational | GITHUB | 136 | 2026-09-16 |
| tpu-megakernels | GITHUB | 119 | 2026-09-23 |
| Tacopy | HN | 95 | 2025-11-30 |
| Inference-Engineering | GITHUB | 95 | 2026-09-14 |
| DBOSify | HN | 91 | 2026-06-24 |
| open-jev-fast | GITHUB | 87 | 2026-09-27 |
| A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo) | HN | 79 | 2026-08-10 |
| ResolveHQ | HN | 72 | 2026-09-11 |
| YuE2-Turbo | GITHUB | 66 | 2026-09-21 |
| Mu | HN | 65 | 2025-11-24 |
| Lamb Labs: Custom Chips for AI Inference | YC | 64 | 2026-08-03 |
| causilo | GITHUB | 64 | 2026-09-13 |
| B-IR | HN | 62 | 2026-01-12 |
| Mnemo | HN | 60 | 2026-06-03 |
| Reame | HN | 59 | 2026-07-11 |
| QQ-Network-Protocol | GITHUB | 59 | 2026-09-13 |
| Reducing LLM input tokens by 70% | HN | 56 | 2026-05-12 |
| An LLM-Powered Tool to Catch PCB Schematic Mistakes | HN | 55 | 2025-11-28 |
| Modeloop | HN | 46 | 2026-06-15 |
| livenerf | GITHUB | 44 | 2026-09-22 |
| runntime | GITHUB | 43 | 2026-09-14 |
| orukeet | GITHUB | 40 | 2026-09-09 |
| dev-0.4b | GITHUB | 39 | 2026-09-21 |
| VLCoT | GITHUB | 37 | 2026-09-28 |
| TypeLLM | GITHUB | 35 | 2026-09-17 |
| LMMOCK | GITHUB | 35 | 2026-09-20 |
| agentic-cuda-optimizer | GITHUB | 35 | 2026-09-24 |
| muse2api | GITHUB | 35 | 2026-09-27 |