Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Llama 3.1 70B on a single RTX 3090 via NVMe-to-GPU bypassing the CPU

Details

External ID
47104667
Source
HN
Company
—
Product
Llama 3.1 70B on a single RTX 3090 via NVMe-to-GPU bypassing the CPU
Website domain
github.com
Launched
Feb. 21, 2026
Cohort
—
Upvotes
395
Upvotes percentile
0.9865229110512129
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Hi everyone, I'm kinda involved in some retrogaming and with some experiments I ran into the following question: "It would be possible to run transformer models bypassing the cpu/ram, connecting the gpu to the nvme?"This is the result of that question itself and some weekend vibecoding (it has the linked library repository in the readme as well), it seems to work, even on consumer gpus, it should work better on professional ones tho

Enrichment

Theme
graphics rendering and visual tools
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
Not AI
Project type
Hobby / open-source project
Normalized one-liner
run llama 3.1 70b on single gpu
Manually corrected
False

Could you build this?

No Bypassing the CPU/system RAM to stream LLM weights directly from NVMe to GPU requires low-level kernel driver manipulation and GPU Direct Storage (GDS) / SPDK engineering.

What it would actually take: This architecture requires low-level C++/CUDA systems engineering utilizing NVIDIA GPUDirect Storage (cuFile API) or custom user-space NVMe drivers (SPDK) over PCIe to transfer model layers directly into VRAM on demand. It involves designing asynchronous prefetching pipelines that overlap layer compute with NVMe DMA reads, precisely synchronized to memory bandwidth and GPU compute streams without CPU/OS page cache intervention.

Discussion

20 comments analyzed.

Competitors mentioned: MemeRadar, llama.cpp, vLLM, SGLang

Concerns raised: Cloud latency breaks real-time signal processing, Memory bandwidth bottleneck limits performance more than raw TFLOPS, Quantization loses precision needed for subtle pattern detection, Local hardware TCO beats cloud rentals for 24/7 production, MoE expert routing requires keeping all parameters in memory or waiting for slower storage loads

Feature requests: Predictive MoE expert loading to VRAM via PCIe before inference, CPU optimization for 50% speed improvement potential, Research on MoE scheduling and optimization, CXL memory support for improved bandwidth across system, Selective layer fine-tuning instead of full LoRA on all layers

Competitors

Other products that read as similar to this one — 197 launches clear the similarity bar, closest 8 shown.

Attention rank: #4 of 198 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 115 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.