lightweight and on-device AI runtimes
Recent window: last 5.2 months
(2026-04-24 → 2026-09-28), compared with the prior
5.2 months.
horizontal · 264 members ·
Data as of 2026-09-30
These products provide compact models, distillation techniques, and optimized runtime engines designed to run AI locally on consumer hardware and edge devices. They are built for developers, hardware hackers, and embedded systems engineers seeking private, low-latency machine learning execution without expensive cloud GPUs. Unlike general cloud-hosted LLM APIs, this cluster focuses strictly on extreme memory efficiency and resource-constrained local inference.
Metrics
- Stage
- crowded
- Recent count
- 134
- Prior count
- 111
- Total count
- 264
- Momentum
- 20.72
- Attention
- 0.54
- Crowding
- 0.70
- Concentration
- 0.83
- Opportunity
- 0.33
Opportunity components
- Attention
- 0.54
- Low crowding
- 0.30
- Momentum (normalized)
- 0.30
- Low concentration
- 0.17
Source split
- github
- 10 (0.07)
- hn
- 91 (0.68)
- ph
- 33 (0.25)
- yc
- 0 (0.00)
Dominant source: hn · Divergence: 0.68
Members
|
Name
|
Source
|
Upvotes ▼
|
Launched
|
| ChunkBack |
HN |
6 |
2025-11-19 |
| Model-agnostic cognitive architecture for LLMs |
HN |
6 |
2025-11-18 |
| FASHN VTON v1.5 |
HN |
6 |
2026-01-28 |
| Go-CoreML |
HN |
6 |
2026-01-14 |
| Project AELLA |
HN |
6 |
2025-11-11 |
| Microterm runs Linux VM in any browser tab via WASM, RISCV64 emulation |
HN |
6 |
2026-02-22 |
| Breathe-Memory |
HN |
6 |
2026-03-26 |
| Running BitNet b1.58 inside DRAM by breaking DDR4 timing rules |
HN |
6 |
2026-05-23 |
| Makes local LLMs faster and more reliable by optimizing for your device |
HN |
6 |
2026-06-30 |
| Numax |
HN |
6 |
2026-06-16 |
| Omni |
HN |
6 |
2026-06-05 |
| Samosa Chat |
HN |
6 |
2026-07-15 |
| Fixing LLM memory degradation in long coding sessions |
HN |
5 |
2025-11-27 |
| m6502, a 6502 CPU for FPGAs and Tiny Tapeout |
HN |
5 |
2026-02-18 |
| I built a local Elixir/Python pipeline to curate 14,000 RAW photos |
HN |
5 |
2026-04-20 |
| go-binsync |
HN |
5 |
2026-08-27 |
| Revibing nanochat's inference model in C++ with ggml |
HN |
5 |
2026-01-10 |
| ClawMem |
HN |
5 |
2026-03-22 |
| Compute:Arena |
HN |
5 |
2026-09-17 |
| RamScout |
HN |
5 |
2025-12-08 |
| Sipp |
HN |
5 |
2026-06-24 |
| Sub-microsecond (890 ns) trading execution research system |
HN |
5 |
2025-12-15 |
| CoreTrace, a visual 16-bit CPU simulator |
HN |
5 |
2026-08-13 |
| Run open-weight OCR, VLM and vision models behind one API |
HN |
5 |
2026-09-04 |
| On the edge of Apple Silicon memory speeds |
HN |
5 |
2026-01-17 |
| Symbolic regression as an MCP tool (SINDy and PySR, free, no install) |
HN |
5 |
2026-04-02 |
| Anchor Engine |
HN |
5 |
2026-03-06 |
| A new language for COBOL workloads, built on Go |
HN |
5 |
2025-11-05 |
| Axiom |
HN |
5 |
2026-02-02 |
| TurboBench, the Compression Lie Detector, 100 Codecs, Daily Update |
HN |
5 |
2026-09-16 |
| Rust-split |
HN |
5 |
2026-09-12 |
| I run 30B 22tok/s, 109tok/s not novel,6GB/16GB RAM overcoming llama.cpp |
HN |
5 |
2026-07-29 |
| TabPFN Scaling Mode |
HN |
5 |
2025-12-03 |
| Fast NF4 dequantization Triton kernel (1.41x faster than bitsandbytes) |
HN |
5 |
2026-07-15 |
| MiniVim a Minimal Neovim Configuration |
HN |
5 |
2026-02-24 |
| Building a full agentic harness around a 4B model is hard |
HN |
5 |
2026-08-19 |
| Scope-structured arena memory for C, O(1) cleanup, no GC/borrow checker |
HN |
5 |
2026-04-15 |
| Sofka |
HN |
5 |
2026-08-19 |
| Clawbernetes |
HN |
5 |
2026-02-20 |
| Aurion OS, A 1.8MB OS with a browser, try it live (C/x86 ASM) |
HN |
5 |
2026-04-03 |
| Mercury 2.5 |
PH |
3 |
2026-09-09 |
| TRLoom |
PH |
3 |
2026-09-15 |
| builtwithlaya |
PH |
3 |
2026-09-24 |
| ElideDB. Database for Physical AI |
PH |
2 |
2026-09-18 |
| Efficio - AI Harness for speed, memory |
PH |
2 |
2026-09-10 |
| UsingOpen |
PH |
2 |
2026-09-26 |
| ROCmFix & InferBench |
PH |
2 |
2026-09-20 |
| Lifeboat |
PH |
2 |
2026-09-23 |
| BestLLMfor |
PH |
2 |
2026-09-09 |
| Autotune Doctor |
PH |
2 |
2026-09-20 |