Llmtop
Htop for LLM Inference Clusters (vLLM, SGLang, Ollama, llama)
Details
- External ID
- 47421736
- Source
- HN
- Company
- —
- Product
- Llmtop
- Website domain
- github.com
- Launched
- March 18, 2026
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.1070110701107011
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
I work on inference scheduling — KV cache-aware routing, load balancing across GPU workers, that kind of thing. I wanted something like k9s but for my inference stack. Nothing existed, so I built it.llmtop is a real-time terminal dashboard for LLM inference workers. It scrapes the Prometheus /metrics endpoints that vLLM, SGLang, and LMCache already expose and shows everything in one view: KV cache usage, queue depth, TTFT/ITL latencies (P50/P99 from histogram buckets), token throughput, prefix cache hit rates. Color-coded — red means go fix it.``` brew install InfraWhisperer/tap/llmtop Or go install github.com/InfraWhisperer/llmtop/cmd/llmtop@latest. ```Single binary, no Prometheus server needed, no Grafana, no config. Just run llmtop and it auto-discovers local workers.Written in Go with Bubbletea. Working on Kubernetes pod auto-discovery and a GPU metrics view next.
Enrichment
- Theme
- ML inference and model optimization
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Commercial product
- Normalized one-liner
- monitoring tool for llm inference clusters
- Manually corrected
- False
Could you build this?
Yes This is a terminal user interface (TUI) dashboard that simply scrapes Prometheus `/metrics` endpoints from existing inference engines and displays the figures in formatted terminal widgets.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 102 launches clear the similarity bar, closest 8 shown.
Attention rank: #97 of 103 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 130 days after the earliest competitor.
- LLMKube · hn · 2025-11-18 · 5 upvotes · similarity 0.54
- Docker Model Runner Integrates vLLM for High-Throughput Inference · hn · 2025-11-20 · 7 upvotes · similarity 0.51
- Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload · hn · 2026-07-26 · 22 upvotes · similarity 0.49
- Composable middleware for LLM inference Optimization Passes · hn · 2026-03-04 · 7 upvotes · similarity 0.48
- Tiny-vLLM · hn · 2026-05-29 · 205 upvotes · similarity 0.46
- Kairo · hn · 2026-09-14 · 5 upvotes · similarity 0.45
- KV-psi, using Linux PSI to to trim an LLM KV cache · hn · 2026-06-27 · 8 upvotes · similarity 0.43
- Llm.sql · hn · 2026-04-24 · 8 upvotes · similarity 0.43
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.