Llama 3.1 70B on a single RTX 3090 via NVMe-to-GPU bypassing the CPU
Details
- External ID
- 47104667
- Source
- HN
- Company
- —
- Product
- Llama 3.1 70B on a single RTX 3090 via NVMe-to-GPU bypassing the CPU
- Website domain
- github.com
- Launched
- Feb. 21, 2026
- Cohort
- —
- Upvotes
- 395
- Upvotes percentile
- 0.9865229110512129
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Hi everyone, I'm kinda involved in some retrogaming and with some experiments I ran into the following question: "It would be possible to run transformer models bypassing the cpu/ram, connecting the gpu to the nvme?"This is the result of that question itself and some weekend vibecoding (it has the linked library repository in the readme as well), it seems to work, even on consumer gpus, it should work better on professional ones tho
Enrichment
- Theme
- graphics rendering and visual tools
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- run llama 3.1 70b on single gpu
- Manually corrected
- False
Could you build this?
No Bypassing the CPU/system RAM to stream LLM weights directly from NVMe to GPU requires low-level kernel driver manipulation and GPU Direct Storage (GDS) / SPDK engineering.
What it would actually take: This architecture requires low-level C++/CUDA systems engineering utilizing NVIDIA GPUDirect Storage (cuFile API) or custom user-space NVMe drivers (SPDK) over PCIe to transfer model layers directly into VRAM on demand. It involves designing asynchronous prefetching pipelines that overlap layer compute with NVMe DMA reads, precisely synchronized to memory bandwidth and GPU compute streams without CPU/OS page cache intervention.
Discussion
20 comments analyzed.
Competitors mentioned: MemeRadar, llama.cpp, vLLM, SGLang
Concerns raised: Cloud latency breaks real-time signal processing, Memory bandwidth bottleneck limits performance more than raw TFLOPS, Quantization loses precision needed for subtle pattern detection, Local hardware TCO beats cloud rentals for 24/7 production, MoE expert routing requires keeping all parameters in memory or waiting for slower storage loads
Feature requests: Predictive MoE expert loading to VRAM via PCIe before inference, CPU optimization for 50% speed improvement potential, Research on MoE scheduling and optimization, CXL memory support for improved bandwidth across system, Selective layer fine-tuning instead of full LoRA on all layers
Competitors
Other products that read as similar to this one — 197 launches clear the similarity bar, closest 8 shown.
Attention rank: #4 of 198 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 115 days after the earliest competitor.
- LLMKube · hn · 2025-11-18 · 5 upvotes · similarity 0.49
- VRAMGlass · ph · 2026-09-19 · 1 upvotes · similarity 0.49
- fast-rtxvsr · github · 2026-09-09 · 15 upvotes · similarity 0.48
- I run 30B 22tok/s, 109tok/s not novel,6GB/16GB RAM overcoming llama.cpp · hn · 2026-07-29 · 5 upvotes · similarity 0.48
- CUDA/graphics in QEMU-KVM VMs without passing the Nvidia card to them · hn · 2026-09-08 · 5 upvotes · similarity 0.48
- OS Megakernel that match M5 Max Tok/w at 2x the Throughput on RTX 3090 · hn · 2026-04-08 · 6 upvotes · similarity 0.46
- Fine-tune an 8B model on a 4 GB laptop GPU · hn · 2026-08-04 · 139 upvotes · similarity 0.45
- Llama.cpp Tutorial 2026: Run GGUF Models Locally on CPU and GPU · hn · 2026-04-18 · 13 upvotes · similarity 0.44
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.