ML inference and model optimization
Recent window: last 5.2 months (2026-04-24 → 2026-09-28), compared with the prior 5.2 months.
These products provide specialized runtimes, compression techniques, and acceleration engines to deploy and benchmark machine learning models efficiently. They are designed for machine learning engineers, systems developers, and AI researchers working with constrained hardware or demanding latency requirements. The cluster focuses directly on low-level compute, quantization, and runtime performance rather than high-level application wrappers or end-user workflow tools.
Metrics
- Stage
- crowded
- Recent count
- 156
- Prior count
- 50
- Total count
- 223
- Momentum
- 212.00
- Attention
- 0.46
- Crowding
- 0.79
- Concentration
- 0.84
- Opportunity
- 0.40
Opportunity components
- Attention
- 0.46
- Low crowding
- 0.21
- Momentum (normalized)
- 0.78
- Low concentration
- 0.16
Monthly trajectory
Source split
- github
- 76 (0.49)
- hn
- 61 (0.39)
- ph
- 10 (0.06)
- yc
- 9 (0.06)
Dominant source: github · Divergence: 0.43
Similar themes
- AI infrastructure and inference optimization (0.72)
- voice AI and speech tools (0.61)
- gpu compute and acceleration tools (0.61)
- decision model runtimes and tools (0.56)
- persistent memory for AI agents (0.55)
- systems tools and desktop utilities (0.54)
- low-level systems and developer tools (0.54)
- autonomous agent research and evaluation (0.54)
Members
| Name | Source | Upvotes ▲ | Launched |
|---|---|---|---|
| Fusion-MoA-Pioneer | GITHUB | 19 | 2026-09-10 |
| aide | GITHUB | 20 | 2026-09-23 |
| Otari: your open-source LLM control plane | HN | 20 | 2026-07-06 |
| Quant Picker | HN | 20 | 2026-06-13 |
| FactorForge | GITHUB | 20 | 2026-09-21 |
| Piris Labs: Inference at Light Speed | YC | 21 | 2026-02-12 |
| Conduct, open-source guardrails for LLM and MCP tool calls | HN | 22 | 2026-08-28 |
| ZigFormer | HN | 22 | 2025-11-27 |
| Mdarena | HN | 22 | 2026-04-05 |
| VeriTile | GITHUB | 24 | 2026-09-24 |
| LLM Thought Visualization | HN | 24 | 2026-07-06 |
| The Token Company: Intelligent compression for LLM context bloat | YC | 25 | 2026-03-03 |
| CrossDomainAdjust | GITHUB | 26 | 2026-09-25 |
| dexgpt | GITHUB | 26 | 2026-09-09 |
| UMM-Reflection | GITHUB | 26 | 2026-09-27 |
| We cut RAG latency ~2× by switching embedding model | HN | 27 | 2025-11-25 |
| ttd-capa-cpp | GITHUB | 28 | 2026-09-20 |
| captains-deck | GITHUB | 29 | 2026-09-24 |
| mlxfast-bonsai2-27b-engine | GITHUB | 30 | 2026-09-24 |
| PrunedCTC | GITHUB | 30 | 2026-09-27 |
| STEPQuant | GITHUB | 32 | 2026-09-29 |
| reflex | GITHUB | 34 | 2026-09-23 |
| Continual Learning with .md | HN | 34 | 2026-04-13 |
| muse2api | GITHUB | 35 | 2026-09-27 |
| LMMOCK | GITHUB | 35 | 2026-09-20 |
| agentic-cuda-optimizer | GITHUB | 35 | 2026-09-24 |
| TypeLLM | GITHUB | 35 | 2026-09-17 |
| dev-0.4b | GITHUB | 39 | 2026-09-21 |
| orukeet | GITHUB | 40 | 2026-09-09 |
| runntime | GITHUB | 43 | 2026-09-14 |
| VLCoT | GITHUB | 43 | 2026-09-28 |
| livenerf | GITHUB | 44 | 2026-09-22 |
| Modeloop | HN | 46 | 2026-06-15 |
| An LLM-Powered Tool to Catch PCB Schematic Mistakes | HN | 55 | 2025-11-28 |
| Reducing LLM input tokens by 70% | HN | 56 | 2026-05-12 |
| QQ-Network-Protocol | GITHUB | 59 | 2026-09-13 |
| Reame | HN | 59 | 2026-07-11 |
| Mnemo | HN | 60 | 2026-06-03 |
| B-IR | HN | 62 | 2026-01-12 |
| causilo | GITHUB | 64 | 2026-09-13 |
| Lamb Labs: Custom Chips for AI Inference | YC | 64 | 2026-08-03 |
| Mu | HN | 65 | 2025-11-24 |
| YuE2-Turbo | GITHUB | 66 | 2026-09-21 |
| ResolveHQ | HN | 72 | 2026-09-11 |
| A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo) | HN | 79 | 2026-08-10 |
| open-jev-fast | GITHUB | 88 | 2026-09-27 |
| DBOSify | HN | 91 | 2026-06-24 |
| Inference-Engineering | GITHUB | 95 | 2026-09-14 |
| Tacopy | HN | 95 | 2025-11-30 |
| tpu-megakernels | GITHUB | 119 | 2026-09-23 |