Lumina
Open-source observability for LLM applications
Details
- External ID
- 46751546
- Source
- HN
- Company
- —
- Product
- Lumina
- Website domain
- github.com
- Launched
- Jan. 25, 2026
- Cohort
- —
- Upvotes
- 6
- Upvotes percentile
- 0.2549407114624506
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Hey HN! I built Lumina – an open-source observability platform for AI/LLM applications. Self-host it in 5 minutes with Docker Compose, all features included.The Problem:I've been building LLM apps for the past year, and I kept running into the same issues: - LLM responses would randomly change after prompt tweaks, breaking things - Costs would spike unexpectedly (turns out a bug was hitting GPT-4 instead of 3.5) - No easy way to compare "before vs after" when testing prompt changes - Existing tools were either too expensive or missing features in free tiersWhat I Built:Lumina is OpenTelemetry-native, meaning: - Works with your existing OTEL stack (Datadog, Grafana, etc.) - No vendor lock-in – standard trace format - Integrates in 3 lines of codeKey features: - Cost & quality monitoring – Automatic alerts when costs spike or responses degrade - Replay testing – Capture production traces, replay them after changes, see diffs - Semantic comparison – Not just string matching – uses Claude to judge if responses are "better" or "worse" - Self-hosted tier – 50k traces/day, 7-day retention, ALL features included (alerts, replay, semantic scoring)How it works:Start Luminagit clone https://github.com/use-lumina/Lumina cd Lumina/infra/docker docker-compose up -d// Add to your app (no API key needed for self-hosted!)import { Lumina } from '@uselumina/sdk';const lumina = new Lumina({ endpoint: 'http://localhost:8080/v1/traces', });// Wrap your LLM call const response = await lumina.traceLLM( async () => await openai.chat.completions.create({...}), { provider: 'openai', model: 'gpt-4', prompt: '...' } );That's it. Every LLM call is now tracked with cost, latency, tokens, and quality scores.What makes it different:1. Free self-hosted with limits that work – 50k traces/day and 7-day retention (resets daily at midnight UTC). All features included: alerts, replay testing, semantic scoring. Perfect for most development and small production workloads. Need more? Upgrade to managed cloud.2. OpenTelemetry-native – Not another proprietary format. Use standard OTEL exporters, works with existing infra. Can send traces to both Lumina AND Datadog simultaneously.3. Replay testing – The killer feature. Capture 100 production traces, change your prompt, replay them all, get a semantic diff report. Like snapshot testing for LLMs.4. Fast – Built with Bun, Postgres, Redis, NATS. Sub-500ms from trace to alert. Handles 10k+ traces/min on a single machine.What I'm looking for:- Feedback on the approach (is OTEL the right foundation?) - Bug reports (tested on Mac/Linux/WSL2, but I'm sure there are issues) - Ideas for what features matter most (alerts? replay? cost tracking?) - Help with the semantic scorer (currently uses Claude, want to make it pluggable)Why open source:I want this to be the standard for LLM observability. That only works if it's: - Free to use and modify (Apache 2.0) - Easy to self-host (Docker Compose, no cloud dependencies) - Open to contributions (good first issues tagged)The business model is managed hosting for teams who don't want to run infrastructure. But the core product is and always will be free.Try it: - GitHub: https://github.com/use-lumina/Lumina - Demo video: [YouTube link] - Docs: https://docs.uselumina.io - Quick start: 5 minutes from `git clone` to dashboardI'd love to hear what you think! Especially interested in: - What observability problems you're hitting with LLMs - Missing features that would make this useful for you - Any similar tools you're using (and what they do better)Thanks for reading!
Enrichment
- Theme
- AI agent frameworks and developer tools
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- observability for llm applications
- Manually corrected
- False
Could you build this?
Yes Building an LLM observability tool involves wrapping API calls with an SDK proxy/middleware, collecting token/latency metrics, and storing them in a database with a Next.js/Docker dashboard.
Discussion
2 comments analyzed.
Competitors mentioned: toran.sh (HTTP-level observability for agents)
Concerns raised: Low-level HTTP details like retries/timeouts/failovers can be hidden at SDK/trace level, Gap between agent decision layer and what actually went over the wire, Opaque magic in agent frameworks makes debugging difficult
Feature requests: Capture raw HTTP layer alongside traces for production debugging, See underlying HTTP calls and requests/responses in addition to trace level
Competitors
Other products that read as similar to this one — 75 launches clear the similarity bar, closest 8 shown.
Attention rank: #61 of 76 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 83 days after the earliest competitor.
- Torrix, self hosted, LLM Observability,(no Postgres, no Redis) · hn · 2026-05-13 · 74 upvotes · similarity 0.51
- AI Observability by OpenObserve · ph · 2026-09-10 · 304 upvotes · similarity 0.46
- Logify360 · ph · 2026-09-10 · 2 upvotes · similarity 0.46
- Lumina · ph · 2026-09-25 · 2 upvotes · similarity 0.44
- Linnix · hn · 2025-11-11 · 21 upvotes · similarity 0.43
- Clx · hn · 2026-07-11 · 143 upvotes · similarity 0.42
- Oodle.ai · hn · 2026-07-14 · 31 upvotes · similarity 0.42
- Lumina · ph · 2026-09-16 · 1 upvotes · similarity 0.41
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.