KV-psi, using Linux PSI to to trim an LLM KV cache
Details
- External ID
- 48702538
- Source
- HN
- Company
- —
- Product
- KV-psi, using Linux PSI to to trim an LLM KV cache
- Website domain
- github.com
- Launched
- June 27, 2026
- Cohort
- —
- Upvotes
- 8
- Upvotes percentile
- 0.4952185792349727
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
I thought it'd be interesting to use Linux PSI (Pressure Stall Information) for an LLM runtime to trim the KV cache. This is mainly useful imo for edge devices like the Jetson Orin super nano kit which have unified memory. I haven't benched much, but plan to do so more over time and see if I can make a real use of it as I run local LLMs. Let me know if it makes sense :P (I of course vibed this idea)
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI feature
- Project type
- Hobby / open-source project
- Normalized one-liner
- optimize llm kv cache with linux psi
- Manually corrected
- False
Could you build this?
No Integrating Linux kernel Pressure Stall Information (PSI) triggers with low-level LLM inference engines (such as llama.cpp or vLLM) to dynamically evict KV cache tokens requires deep kernel and GPU memory architecture expertise.
What it would actually take: This requires modifying low-level C++/CUDA inference runtimes (like vLLM or llama.cpp) to interface with Linux `/proc/pressure/` epoll triggers. The difficult challenges involve designing non-disruptive, latency-critical KV cache quantization/pruning algorithms that preserve context coherence without causing GPU pipeline stalls on unified memory architectures like Jetson. This demands deep systems programming and ML systems hardware optimization skills.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 151 launches clear the similarity bar, closest 8 shown.
Attention rank: #87 of 152 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 239 days after the earliest competitor.
- Shoehorn, a library to quantize an LLM to fit your Mac's VRAM · hn · 2026-08-14 · 6 upvotes · similarity 0.54
- Makes local LLMs faster and more reliable by optimizing for your device · hn · 2026-06-30 · 6 upvotes · similarity 0.51
- Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload · hn · 2026-07-26 · 22 upvotes · similarity 0.50
- Korean-llm-v4 · github · 2026-09-09 · 21 upvotes · similarity 0.49
- ExANS · hn · 2026-08-05 · 15 upvotes · similarity 0.47
- GhostBox · hn · 2026-05-01 · 126 upvotes · similarity 0.44
- Lowfat · hn · 2026-06-05 · 156 upvotes · similarity 0.44
- A "Cram tests" script for windows shells · hn · 2025-11-26 · 6 upvotes · similarity 0.44
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.