RunNburn
Run a 295B Moe from a 98GB GGUF on a 64GB RAM Desktop
Details
- External ID
- 49105154
- Source
- HN
- Company
- —
- Product
- RunNburn
- Website domain
- github.com
- Launched
- July 30, 2026
- Cohort
- —
- Upvotes
- 11
- Upvotes percentile
- 0.5764635603345281
- Tags
- —
- Fetched at
- Sept. 10, 2026, 5:32 a.m.
- Updated at
- Sept. 10, 2026, 5:32 a.m.
Description
runNburn is an Apache-2.0 Rust inference engine for quantized GGUF models that are too big for your fast memory.The core idea: weights stay file-backed (mmap), host residency stays under an explicit byte budget (--ram-budget), and GPU caches are sized from detected free/total VRAM — never from device-name presets. There is no conversion step, no sidecar cache files, no silent requantization. The GGUF on disk is the single source of truth.The result that made me want to post this: Tencent's Hy3 (295B total / 21B active sparse MoE, a single 97.8 GiB Q2_K GGUF) runs on my desktop with 64 GB of RAM and one consumer NVIDIA GPU. The file is larger than RAM and VRAM combined; the selected experts for each token are pulled on demand (the newest path batches O_DIRECT reads through io_uring), while the pretrained routing is left untouched. On the same machine, same prompt, same decode length, a warm-run median gave ~5.5 tok/s decode vs ~2.0 tok/s for llama.cpp.To be upfront about scope: for models that fit comfortably in VRAM, llama.cpp is still faster than runNburn today — its CUDA kernels have years of tuning and we measure against it honestly (interleaved A/B runs, medians, and any "speedup" that changes output quality is rejected). runNburn's lane is the model that doesn't fit.What's in the box:- CLI, interactive chat, and an OpenAI-compatible server (chat/completions + responses + conversations, SSE streaming, stateful continuation with KV/SSM snapshot reuse). It's built as a single-owner personal server — one active generation is the optimization unit; continuous batching and multi-tenant throughput are explicit non-goals. - Architecture-aware paths: Llama family, Phi, Gemma, Qwen dense/hybrid/MoE (including GatedDeltaNet layers), Nemotron-H MoE, Hy3, GLM — plus in-model multi-token prediction (self-speculative decoding) with device-side verification where the GGUF ships a drafter. - Backends: CPU is the default (x86 AVX2, ARM NEON), CUDA and Metal are active, Vulkan/OpenCL are experimental. Android works through a small C ABI (rnb.h). - Native quantized kernels for Q2_K–Q6_K, Q4_0, Q8_0 — including the low-bit K-quants that big-MoE builds actually ship in.It's pre-1.0 and rough in places; recognition of an architecture doesn't mean every community variant works. But if you've got a model file bigger than your machine and you'd rather it run slowly than not at all, that's exactly the case it was built for.Happy to answer questions about the offloading design, the expert-streaming path, or the measurement protocol.
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- run large language model on limited ram
- Manually corrected
- False
Could you build this?
No Building a custom Rust inference engine that runs a 295B MoE model from file-backed mmap with precise RAM budgeting and dynamic VRAM caching requires world-class systems, memory management, and GPU programming expertise.
What it would actually take: The architecture requires custom Rust systems programming with manual virtual memory manipulation (POSIX mmap, madvise, and fine-grained page-cache eviction) paired with custom compute kernels (CUDA, ROCm, Vulkan, or Metal). The core challenge is orchestrating prefetching of sparse MoE expert matrices from disk/RAM directly into VRAM asynchronously ahead of layer execution to hide I/O latency while strictly adhering to a hard RAM cap. This demands deep specialized knowledge in high-performance computing, kernel I/O subsystem internals, and deep learning runtime design.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 199 launches clear the similarity bar, closest 8 shown.
Attention rank: #101 of 200 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 274 days after the earliest competitor.
- Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac · hn · 2026-07-29 · 919 upvotes · similarity 0.51
- Running Gemma-4 26B at 124 tokens/SEC on a CPU, no GPU · hn · 2026-06-30 · 10 upvotes · similarity 0.51
- Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload · hn · 2026-07-26 · 22 upvotes · similarity 0.49
- ChonkLM · hn · 2026-05-09 · 6 upvotes · similarity 0.48
- Mini-AGI · hn · 2026-09-21 · 277 upvotes · similarity 0.47
- collabosm · github · 2026-09-25 · 90 upvotes · similarity 0.46
- ExANS · hn · 2026-08-05 · 15 upvotes · similarity 0.46
- A new engine to run Kimi K3 on a laptop · hn · 2026-07-29 · 7 upvotes · similarity 0.46
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.