Self-host open-source LLMs on AWS with scale-to-zero
Details
- External ID
- 49620813
- Source
- HN
- Company
- —
- Product
- Self-host open-source LLMs on AWS with scale-to-zero
- Website domain
- github.com
- Launched
- Sept. 9, 2026
- Cohort
- —
- Upvotes
- 7
- Upvotes percentile
- 0.4393939393939394
- Tags
- —
- Fetched at
- Sept. 13, 2026, 5:56 p.m.
- Updated at
- Sept. 13, 2026, 5:56 p.m.
Enrichment
- Theme
- systems tools and desktop utilities
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Commercial product
- Normalized one-liner
- self-hosted open-source llm deployment on aws
- Manually corrected
- False
Could you build this?
Partial While basic AWS CDK/Terraform scripts can spin up vLLM on EC2 or ECS, reliable scale-to-zero with cold-start mitigation and request buffering requires non-trivial cloud architecture.
What it would actually take: Architecture requires an API gateway or proxy layer (e.g., Envoy or Lambda) that holds incoming requests while triggering AWS Auto Scaling groups or ECS tasks, coordinating with GPU spot instances or EC2 GPU capacity, and pulling multi-gigabyte model weights rapidly from S3 or EBS. The hard part is managing cold-start latency, queue persistence during boot, and graceful scaling down without dropping active streaming inferences. Requires strong AWS infrastructure and systems engineering experience.
Discussion
5 comments analyzed.
Competitors mentioned: Modal, Baseten, AWS Auto Scaling Groups (ASG), Kubernetes
Concerns raised: CPU-only example doesn't represent real production GPU workloads, Lack of pricing/cost benchmarks for larger models (7B, 70B), Bugs and issues to be resolved, AWS Spot instance interruption handling clarity
Feature requests: Pricing table for different model sizes and instance types, Benchmarks for larger LLM models in README
Competitors
Other products that read as similar to this one — 856 launches clear the similarity bar, closest 8 shown.
Attention rank: #451 of 857 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 315 days after the earliest competitor.
- Self hosting a modern LLM stack · hn · 2026-06-29 · 5 upvotes · similarity 0.64
- Kyushu · hn · 2026-06-07 · 89 upvotes · similarity 0.59
- Nat-zero · hn · 2026-04-28 · 5 upvotes · similarity 0.57
- jevify · github · 2026-09-19 · 35 upvotes · similarity 0.57
- Gopher · hn · 2026-09-02 · 6 upvotes · similarity 0.56
- Fakecloud · hn · 2026-04-15 · 37 upvotes · similarity 0.56
- Rapid-MLX · hn · 2026-04-18 · 9 upvotes · similarity 0.55
- opensend · github · 2026-09-09 · 119 upvotes · similarity 0.55
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.