ZSE
Open-source LLM inference engine with 3.9s cold starts
Details
- External ID
- 47160526
- Source
- HN
- Company
- —
- Product
- ZSE
- Website domain
- github.com
- Launched
- Feb. 26, 2026
- Cohort
- —
- Upvotes
- 58
- Upvotes percentile
- 0.8315363881401617
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
I've been building ZSE (Z Server Engine) for the past few weeks — an open-source LLM inference engine focused on two things nobody has fully solved together: memory efficiency and fast cold starts.The problem I was trying to solve: Running a 32B model normally requires ~64 GB VRAM. Most developers don't have that. And even when quantization helps with memory, cold starts with bitsandbytes NF4 take 2+ minutes on first load and 45–120 seconds on warm restarts — which kills serverless and autoscaling use cases.What ZSE does differently:Fits 32B in 19.3 GB VRAM (70% reduction vs FP16) — runs on a single A100-40GBFits 7B in 5.2 GB VRAM (63% reduction) — runs on consumer GPUsNative .zse pre-quantized format with memory-mapped weights: 3.9s cold start for 7B, 21.4s for 32B — vs 45s and 120s with bitsandbytes, ~30s for vLLMAll benchmarks verified on Modal A100-80GB (Feb 2026)It ships with:OpenAI-compatible API server (drop-in replacement)Interactive CLI (zse serve, zse chat, zse convert, zse hardware)Web dashboard with real-time GPU monitoringContinuous batching (3.45× throughput)GGUF support via llama.cppCPU fallback — works without a GPURate limiting, audit logging, API key authInstall:----- pip install zllm-zse zse serve Qwen/Qwen2.5-7B-Instruct For fast cold starts (one-time conversion):----- zse convert Qwen/Qwen2.5-Coder-7B-Instruct -o qwen-7b.zse zse serve qwen-7b.zse # 3.9s every timeThe cold start improvement comes from the .zse format storing pre-quantized weights as memory-mapped safetensors — no quantization step at load time, no weight conversion, just mmap + GPU transfer. On NVMe SSDs this gets under 4 seconds for 7B. On spinning HDDs it'll be slower.All code is real — no mock implementations. Built at Zyora Labs. Apache 2.0.Happy to answer questions about the quantization approach, the .zse format design, or the memory efficiency techniques.
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- open-source llm inference engine with fast cold starts
- Manually corrected
- False
Could you build this?
No Developing an LLM inference engine with novel memory paging and fast cold starts requires specialized systems and GPU engineering.
What it would actually take: The system requires custom CUDA or Triton kernels, bespoke KV-cache management (similar to PagedAttention), and low-latency weight streaming via direct NVMe-to-GPU transfer (GPUDirect Storage). This requires deep expertise in systems programming, Linux memory paging, and GPU hardware architectures.
Discussion
9 comments analyzed.
Competitors mentioned: AutoModelForCausalLM.from_pretrained(), Regular quantized models
Concerns raised: Model loading fails with 'vocab_size' attribute error on certain models, GPU not detected on Apple M1/M1 Max despite available memory, Unclear how this differs from dynamic quantization, Cold start timing measurements lack clarity on baseline conditions
Feature requests: Apple M series GPU support, Clarify cold start timing conditions with other models loaded
Competitors
Other products that read as similar to this one — 117 launches clear the similarity bar, closest 8 shown.
Attention rank: #24 of 118 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 110 days after the earliest competitor.
- Llm.sql · hn · 2026-04-24 · 8 upvotes · similarity 0.49
- Hekate · hn · 2026-01-18 · 8 upvotes · similarity 0.46
- Run 500B+ Parameter LLMs Locally on a Mac Mini · hn · 2026-03-09 · 17 upvotes · similarity 0.44
- Serve 100 Large AI models on a single GPU with low impact to TTFT · hn · 2025-11-08 · 7 upvotes · similarity 0.44
- RunNburn · hn · 2026-07-30 · 11 upvotes · similarity 0.43
- Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload · hn · 2026-07-26 · 22 upvotes · similarity 0.43
- OpenGraviton · hn · 2026-03-07 · 13 upvotes · similarity 0.42
- IronMule · ph · 2026-09-17 · 1 upvotes · similarity 0.42
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.