ExANS
Lossless KV cache compression at 622 GB/s on H100
Details
- External ID
- 49185576
- Source
- HN
- Company
- —
- Product
- ExANS
- Website domain
- theopenlake.com
- Launched
- Aug. 5, 2026
- Cohort
- —
- Upvotes
- 15
- Upvotes percentile
- 0.6975806451612904
- Tags
- —
- Fetched at
- Sept. 10, 2026, 5:32 a.m.
- Updated at
- Sept. 10, 2026, 5:32 a.m.
Description
Hi HN,We are the developers of OpenLake, an open source storage engine for KV cache offloading to remote disk and memory.Once we offloaded to local disk, we realized the bottleneck is the PCIe or NIC bandwidth. We wondered whether on GPU lossless compression is viable for fast reads and lower TTFT.BF16 is usually very hard to compress, (high entropy of sign/mantissa). What surprised us is that real world KV blocks are very different. The exponent byte has a very low entropy and barely populated. Instead of compressing the whole tensor, we compress only the exponent stream on the GPU.We see the following results: (H100, production KV snapshot):- 1.51× lossless compression - 622 GB/s median GPU decodeDecompression is ~10× faster than a 400 Gb/s NIC bandwidth delivering data losslessly without quality change.We've are open sourcing this as: ExANS which will be available through our vLLM and SGLang connectors on OpenLake v0.8 version. No changes are required in the inference engine.I'm curious how others are handling KV transfer today. Are you using KV compression or is bandwidth not a bottleneck yet?Thanks!GitHub: https://github.com/openlake-project/openlake Technical Blog: https://theopenlake.com/blog/exans-lossless-gpu-compression-...
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI feature
- Project type
- Commercial product
- Normalized one-liner
- kv cache compression for llms
- Manually corrected
- False
Could you build this?
No Lossless KV cache compression at 622 GB/s on NVIDIA H100 requires deep low-level CUDA/GPU systems engineering, custom ANS entropy coding implementations, and hardware-specific memory optimization.
What it would actually take: Building ExANS requires deep expertise in GPU architecture, high-performance CUDA/C++, and asymmetric numeral systems (ANS) compression algorithms. The developer must write custom GPU kernels that balance compression ratio against high warp-level parallelism and memory coalescing to saturate PCIe/memory bus bandwidths. Profiling with Nsight Compute and rigorous validation against LLM KV cache memory layouts are required.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 176 launches clear the similarity bar, closest 8 shown.
Attention rank: #59 of 177 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 280 days after the earliest competitor.
- Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload · hn · 2026-07-26 · 22 upvotes · similarity 0.70
- Smaller than WinRaR but 4x faster · ph · 2026-09-18 · 1 upvotes · similarity 0.54
- Overfitted a 900KB Transformer to Compress a 100MB CSV into 7MB · hn · 2026-06-23 · 112 upvotes · similarity 0.52
- UL-SMF · hn · 2026-08-17 · 13 upvotes · similarity 0.50
- KV-psi, using Linux PSI to to trim an LLM KV cache · hn · 2026-06-27 · 8 upvotes · similarity 0.47
- Running Gemma-4 26B at 124 tokens/SEC on a CPU, no GPU · hn · 2026-06-30 · 10 upvotes · similarity 0.46
- Taliesin · hn · 2026-06-04 · 10 upvotes · similarity 0.46
- RunNburn · hn · 2026-07-30 · 11 upvotes · similarity 0.46
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.