LLMKube
Kubernetes for Local LLMs with GPU Acceleration
Details
- External ID
- 45968719
- Source
- HN
- Company
- —
- Product
- LLMKube
- Website domain
- github.com
- Launched
- Nov. 18, 2025
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.0982532751091703
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Hi HN! I built LLMKube, a Kubernetes operator for deploying GPU-accelerated LLMs in production. One command gets you from zero to inference with full observability.Why this exists: Regulated industries (healthcare, defense, finance) need air-gapped LLM deployments, but existing tools are either single-node only (Ollama) or lack GPU optimization and SLO enforcement. LLMKube bridges the gap.What's working:- 17x speedup with NVIDIA GPUs (64 tok/s on Llama 3.2 3B vs 4.6 tok/s CPU)- One command: llmkube deploy llama-3b --gpu (auto CUDA setup, scheduling, layer offloading)- Production observability: Prometheus + Grafana + DCGM GPU metrics out of the box- OpenAI-compatible API endpoints- Terraform configs for GKE GPU clusters with auto-scale to zeroTech: Kubernetes CRDs, llama.cpp with CUDA, NVIDIA GPU Operator, cost-optimized spot instances (~$50-150/mo dev workloads).Status: v0.2.0 production-ready for single-GPU deployments on standard K8s clusters. Multi-GPU and multi-node model sharding on the roadmap.Apache 2.0 licensed. Would love feedback from anyone running LLMs in production!Website: https://llmkube.comGitHub: https://github.com/Defilan/LLMKube
Enrichment
- Theme
- ML inference and model optimization
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- kubernetes for local llms with gpu acceleration
- Manually corrected
- False
Could you build this?
Partial While basic Kubernetes operator scaffolding with Kubebuilder is doable, managing multi-node GPU slicing, vLLM/Ollama orchestration, and air-gapped security configurations is difficult.
What it would actually take: A production version requires writing a Go-based Kubernetes Custom Resource Definition (CRD) and controller using the operator-sdk. The controller must interface directly with NVIDIA container runtime / GPU operator resources, handle dynamic VRAM allocation, manage local storage drivers for model weights, and enforce strict air-gapped network policies. This necessitates deep Kubernetes infrastructure, container runtime, and GPU hardware engineering experience.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 146 launches clear the similarity bar, closest 8 shown.
Attention rank: #138 of 147 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 20 days after the earliest competitor.
- Llmtop · hn · 2026-03-18 · 5 upvotes · similarity 0.54
- Llama.cpp Tutorial 2026: Run GGUF Models Locally on CPU and GPU · hn · 2026-04-18 · 13 upvotes · similarity 0.53
- Clawbernetes · hn · 2026-02-20 · 5 upvotes · similarity 0.53
- Docker Model Runner Integrates vLLM for High-Throughput Inference · hn · 2025-11-20 · 7 upvotes · similarity 0.50
- Llama 3.1 70B on a single RTX 3090 via NVMe-to-GPU bypassing the CPU · hn · 2026-02-21 · 395 upvotes · similarity 0.49
- GPULlama3.java Llama Compilied to PTX/OpenCL Now Integrated in Quarkus · hn · 2025-12-11 · 24 upvotes · similarity 0.48
- VMetal · hn · 2026-03-19 · 12 upvotes · similarity 0.47
- VRAMGlass · ph · 2026-09-19 · 1 upvotes · similarity 0.45
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.