Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload

Details

External ID
49057767
Source
HN
Company
—
Product
Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload
Website domain
github.com
Launched
July 26, 2026
Cohort
—
Upvotes
22
Upvotes percentile
0.7353643966547192
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Hey HN, we’re the developers of OpenLake, an open source storage engine for offloading LLM KV caches from GPU memory into a shared tier of RAM and NVMe. We built OpenLake because KV caches are outgrowing GPU memory.A single 256K token conversation on Gemma 4 31B produces approximately 43GB of KV state, more than half the memory of an 80GB H100. The problem becomes even harder across a cluster: a prefix cached on one GPU host is unavailable when the next request lands on a different GPU, forcing the new GPU to repeat work the fleet has already completed.Once the KV cache is offloaded, network bandwidth becomes a major constraint on read latency. To move less data across the wire, we built deferred materialization: a custom CUDA kernel that losslessly compresses KV blocks before they leave GPU memory and decompresses them on the GPU after retrieval. In our tests, this achieved:- 1.72× lossless KV compression. - Approximately 600GB/s decompression throughput on an H100 - 80GB/s of effective KV throughput over a physical 50GB/s linkAt 128K context, retrieving cached KV reduces TTFT from 44 seconds to 0.6 seconds, a 66× improvement. Across the complete workload, GPU time reduces from 1,169 seconds to 606 seconds, saving 48.2% of GPU cost.OpenLake is written in Rust and uses io_uring with one pinned runtime per physical core. We provide connectors for vLLM and SGLang so the cache can be enabled without modifying the inference engine itself.I would love to hear how others are handling KV reuse across GPU hosts, especially for long contexts, and get to know your thoughts.Thanks!GitHub: https://github.com/openlake-project/openlakeHere is our blog: https://cloud.theopenlake.com/blog/taming-the-beast-managing...

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
kv cache optimization for llm inference
Manually corrected
False

Could you build this?

No Creating a low-latency shared KV cache offload engine across GPU VRAM, host RAM, and NVMe requires deep low-level systems programming, CUDA kernel optimization, and ML systems research.

What it would actually take: Requires deep integration with inference engines like vLLM, SGLang, or TensorRT-LLM using C++, CUDA, and custom PyTorch extensions. It involves writing asynchronous zero-copy DMA memory transfer pipelines, SPDK/io_uring for direct NVMe access, and cache eviction/prefetch algorithms tailored to LLM attention mechanisms without stalling inference generation loops.

Discussion

No comments on this launch.

Competitors

Other products that read as similar to this one — 306 launches clear the similarity bar, closest 8 shown.

Attention rank: #85 of 307 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 270 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.