lightweight and on-device AI runtimes
Recent window: last 5.2 months
(2026-04-24 → 2026-09-28), compared with the prior
5.2 months.
horizontal · 264 members ·
Data as of 2026-09-30
These products provide compact models, distillation techniques, and optimized runtime engines designed to run AI locally on consumer hardware and edge devices. They are built for developers, hardware hackers, and embedded systems engineers seeking private, low-latency machine learning execution without expensive cloud GPUs. Unlike general cloud-hosted LLM APIs, this cluster focuses strictly on extreme memory efficiency and resource-constrained local inference.
Metrics
- Stage
- crowded
- Recent count
- 134
- Prior count
- 111
- Total count
- 264
- Momentum
- 20.72
- Attention
- 0.54
- Crowding
- 0.70
- Concentration
- 0.83
- Opportunity
- 0.33
Opportunity components
- Attention
- 0.54
- Low crowding
- 0.30
- Momentum (normalized)
- 0.30
- Low concentration
- 0.17
Source split
- github
- 10 (0.07)
- hn
- 91 (0.68)
- ph
- 33 (0.25)
- yc
- 0 (0.00)
Dominant source: hn · Divergence: 0.68
Members
|
Name
|
Source
|
Upvotes ▲
|
Launched
|
| CJIT, a single-binary C compiler that can self host |
HN |
7 |
2026-04-13 |
| Veil |
HN |
7 |
2026-03-30 |
| RISCY-V02: A 16-bit 2-cycle RISC-V-ish CPU in the 6502 footprint |
HN |
7 |
2026-03-06 |
| Composable middleware for LLM inference Optimization Passes |
HN |
7 |
2026-03-04 |
| We built a <60ms, open-source alternative to E2B using RustVMM and KVM |
HN |
7 |
2026-04-22 |
| Tsplat |
HN |
7 |
2026-04-14 |
| Linear RNN/Reservoir hybrid generative model, one C file (no deps.) |
HN |
7 |
2026-04-09 |
| RiceVM |
HN |
7 |
2026-04-02 |
| I made a Gemma 4 Mac app that names screenshots with local AI |
HN |
7 |
2026-05-31 |
| LyneCode Beats AntiGravity and Codex (and It's Open Source) |
HN |
7 |
2025-12-16 |
| Serve 100 Large AI models on a single GPU with low impact to TTFT |
HN |
7 |
2025-11-08 |
| Selora |
HN |
7 |
2026-06-17 |
| I embedded 685M public texts in 32 minutes (on 8x A100, Rust, TensorRT) |
HN |
7 |
2026-06-04 |
| agent-gpu-calculator |
GITHUB |
7 |
2026-09-24 |
| Python running on the Super Nintendo (in-browser demo) |
HN |
7 |
2026-07-06 |
| Hekate |
HN |
8 |
2026-01-18 |
| eBook to audiobook narration with realistic AI voices |
HN |
8 |
2026-06-24 |
| I build a strace clone for macOS |
HN |
8 |
2025-11-17 |
| Llm.sql |
HN |
8 |
2026-04-24 |
| local-enough |
GITHUB |
8 |
2026-09-29 |
| Lockstep |
HN |
8 |
2026-03-16 |
| I benchmarked Gemma 4 E2B |
HN |
8 |
2026-04-13 |
| OpenCode Senses, An insanely fast and highly accurate vision plugin |
HN |
8 |
2026-08-13 |
| Lunar, a "fast", memory-efficient Lua 5.1 VM written in Go |
HN |
8 |
2026-08-03 |
| KV-psi, using Linux PSI to to trim an LLM KV cache |
HN |
8 |
2026-06-27 |
| DeltaGlider |
HN |
8 |
2025-11-12 |
| Ant |
HN |
8 |
2026-05-09 |
| Self-growing neural networks via a custom Rust-to-LLVM compiler |
HN |
8 |
2025-12-28 |
| Plasmite |
HN |
9 |
2026-03-25 |
| Lumabri |
HN |
9 |
2026-08-09 |
| FEVER Multimodal DB |
PH |
9 |
2026-09-28 |
| N0x |
HN |
9 |
2026-03-18 |
| Litelink |
HN |
9 |
2026-09-03 |
| Hodor |
HN |
9 |
2026-05-27 |
| Determinstic LLM inference for lowest price Gemma 4, with Windows XP |
HN |
9 |
2026-09-12 |
| Unsiloed AI |
HN |
9 |
2026-05-25 |
| Running Gemma-4 26B at 124 tokens/SEC on a CPU, no GPU |
HN |
10 |
2026-06-30 |
| siliconflow-dev.github.io |
GITHUB |
10 |
2026-09-21 |
| aiistream-q3.6 |
GITHUB |
10 |
2026-09-29 |
| Netra Runtime |
PH |
10 |
2026-09-14 |
| AI-Augmented Memory for Groups |
HN |
10 |
2025-12-16 |
| Compile English specs into 22 MB neural functions that run locally |
HN |
11 |
2026-04-15 |
| RunNburn |
HN |
11 |
2026-07-30 |
| NanoRL |
HN |
11 |
2026-08-13 |
| Oodle |
HN |
11 |
2025-11-04 |
| Clone, a small Rust VMM, forks VMs in under 20ms via CoW |
HN |
11 |
2026-04-19 |
| Off Grid: On-device AI-web browsing, tools vision,image,voice–3x faster |
HN |
12 |
2026-02-24 |
| local-ai-recipe-kit |
GITHUB |
12 |
2026-09-29 |
| L88 |
HN |
12 |
2026-02-24 |
| OpenGraviton |
HN |
13 |
2026-03-07 |