I run 30B 22tok/s, 109tok/s not novel,6GB/16GB RAM overcoming llama.cpp
Details
- External ID
- 49093202
- Source
- HN
- Company
- —
- Product
- I run 30B 22tok/s, 109tok/s not novel,6GB/16GB RAM overcoming llama.cpp
- Website domain
- github.com
- Launched
- July 29, 2026
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.1081242532855436
- Tags
- —
- Fetched at
- Sept. 8, 2026, 8:28 p.m.
- Updated at
- Sept. 8, 2026, 8:28 p.m.
Description
Democratisation of local AI is key. I've been working on pushing the limits of commercial hardware, squeezing any extra bit possible. My Scientific Agentic AI hareness helped me to reallocate every single bit of it. I rewrote the Kernel, I went down the CUDA rabbit hole until I have been able to explain any bit and any ms of computational power involved in the process pushing the Qwen 30B-A3B from 8 tok7s to 19 tok/s with llama.cpp up to 22.2 tok/s with my project and 109 tok/s on not novel content and speeding up the prefill by 5-9X
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- 30b language model inference optimization
- Manually corrected
- False
Could you build this?
No Rewriting GPU inference kernels, custom memory paging, and optimizing low-level CUDA execution to outperform libraries like llama.cpp requires elite systems programming and GPU architecture expertise.
What it would actually take: This demands writing custom CUDA/C++ kernels, highly optimized quantized matrix multiplication (e.g., Marlin, AWQ, FP4/FP8 routines), tensor parallelism, and manual unified memory/VRAM-host streaming management. It requires specialized engineers deeply versed in NVIDIA hardware architecture (SMs, tensor cores, memory hierarchy) and modern LLM runtime internals.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 281 launches clear the similarity bar, closest 8 shown.
Attention rank: #256 of 282 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 272 days after the earliest competitor.
- OS Megakernel that match M5 Max Tok/w at 2x the Throughput on RTX 3090 · hn · 2026-04-08 · 6 upvotes · similarity 0.51
- Shoehorn, a library to quantize an LLM to fit your Mac's VRAM · hn · 2026-08-14 · 6 upvotes · similarity 0.50
- GhostBox · hn · 2026-05-01 · 126 upvotes · similarity 0.50
- Llama 3.1 70B on a single RTX 3090 via NVMe-to-GPU bypassing the CPU · hn · 2026-02-21 · 395 upvotes · similarity 0.48
- I'm tired of my LLM bullshitting. So I fixed it · hn · 2026-01-22 · 5 upvotes · similarity 0.47
- Anos · hn · 2026-04-04 · 115 upvotes · similarity 0.46
- Serve 100 Large AI models on a single GPU with low impact to TTFT · hn · 2025-11-08 · 7 upvotes · similarity 0.45
- I used AI to recreate a $4000 piece of audio hardware as a plugin · hn · 2026-01-03 · 160 upvotes · similarity 0.45
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.