Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

ZSE

Open-source LLM inference engine with 3.9s cold starts

Details

External ID
47160526
Source
HN
Company
—
Product
ZSE
Website domain
github.com
Launched
Feb. 26, 2026
Cohort
—
Upvotes
58
Upvotes percentile
0.8315363881401617
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

I've been building ZSE (Z Server Engine) for the past few weeks — an open-source LLM inference engine focused on two things nobody has fully solved together: memory efficiency and fast cold starts.The problem I was trying to solve: Running a 32B model normally requires ~64 GB VRAM. Most developers don't have that. And even when quantization helps with memory, cold starts with bitsandbytes NF4 take 2+ minutes on first load and 45–120 seconds on warm restarts — which kills serverless and autoscaling use cases.What ZSE does differently:Fits 32B in 19.3 GB VRAM (70% reduction vs FP16) — runs on a single A100-40GBFits 7B in 5.2 GB VRAM (63% reduction) — runs on consumer GPUsNative .zse pre-quantized format with memory-mapped weights: 3.9s cold start for 7B, 21.4s for 32B — vs 45s and 120s with bitsandbytes, ~30s for vLLMAll benchmarks verified on Modal A100-80GB (Feb 2026)It ships with:OpenAI-compatible API server (drop-in replacement)Interactive CLI (zse serve, zse chat, zse convert, zse hardware)Web dashboard with real-time GPU monitoringContinuous batching (3.45× throughput)GGUF support via llama.cppCPU fallback — works without a GPURate limiting, audit logging, API key authInstall:----- pip install zllm-zse zse serve Qwen/Qwen2.5-7B-Instruct For fast cold starts (one-time conversion):----- zse convert Qwen/Qwen2.5-Coder-7B-Instruct -o qwen-7b.zse zse serve qwen-7b.zse # 3.9s every timeThe cold start improvement comes from the .zse format storing pre-quantized weights as memory-mapped safetensors — no quantization step at load time, no weight conversion, just mmap + GPU transfer. On NVMe SSDs this gets under 4 seconds for 7B. On spinning HDDs it'll be slower.All code is real — no mock implementations. Built at Zyora Labs. Apache 2.0.Happy to answer questions about the quantization approach, the .zse format design, or the memory efficiency techniques.

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
open-source llm inference engine with fast cold starts
Manually corrected
False

Could you build this?

No Developing an LLM inference engine with novel memory paging and fast cold starts requires specialized systems and GPU engineering.

What it would actually take: The system requires custom CUDA or Triton kernels, bespoke KV-cache management (similar to PagedAttention), and low-latency weight streaming via direct NVMe-to-GPU transfer (GPUDirect Storage). This requires deep expertise in systems programming, Linux memory paging, and GPU hardware architectures.

Discussion

9 comments analyzed.

Competitors mentioned: AutoModelForCausalLM.from_pretrained(), Regular quantized models

Concerns raised: Model loading fails with 'vocab_size' attribute error on certain models, GPU not detected on Apple M1/M1 Max despite available memory, Unclear how this differs from dynamic quantization, Cold start timing measurements lack clarity on baseline conditions

Feature requests: Apple M series GPU support, Clarify cold start timing conditions with other models loaded

Competitors

Other products that read as similar to this one — 117 launches clear the similarity bar, closest 8 shown.

Attention rank: #24 of 118 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 110 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.