Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Kelet

Root Cause Analysis agent for your LLM apps

Details

External ID
47767606
Source
HN
Company
—
Product
Kelet
Website domain
kelet.ai
Launched
April 14, 2026
Cohort
—
Upvotes
47
Upvotes percentile
0.833547557840617
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

I've spent the past few years building 50+ AI agents in prod (some reached 1M+ sessions/day), and the hardest part was never building them — it was figuring out why they fail.AI agents don't crash. They just quietly give wrong answers. You end up scrolling through traces one by one, trying to find a pattern across hundreds of sessions.Kelet automates that investigation. Here's how it works:1. You connect your traces and signals (user feedback, edits, clicks, sentiment, LLM-as-a-judge, etc.) 2. Kelet processes those signals and extracts facts about each session 3. It forms hypotheses about what went wrong in each case 4. It clusters similar hypotheses across sessions and investigates them together 5. It surfaces a root cause with a suggested fix you can review and applyThe key insight: individual session failures look random. But when you cluster the hypotheses, failure patterns emerge.The fastest way to integrate is through the Kelet Skill for coding agents — it scans your codebase, discovers where signals should be collected, and sets everything up for you. There are also Python and TypeScript SDKs if you prefer manual setup.It’s currently free during beta. No credit card required. Docs: https://kelet.ai/docs/I'd love feedback on the approach, especially from anyone running agents in prod. Does automating the manual error analysis sound right?

Enrichment

Theme
AI agent frameworks and developer tools
Vertical
Horizontal
Function
Agent / copilot
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
root cause analysis for llm apps
Manually corrected
False

Could you build this?

Partial Basic OpenTelemetry trace viewers are easy to build, but clustering failure patterns and automatically identifying true root causes across millions of agentic traces requires specialized telemetry and ML pipelines.

What it would actually take: The architecture requires high-volume telemetry ingestion (OTel/ClickHouse), distributed embedding and trace graph analyzers, and automated pattern clustering to group nondeterministic failure states. The core challenge is synthesizing semantic diffs across long multi-step agent graphs to isolate actual decision failures from downstream cascades. This necessitates expertise in distributed observability pipelines and agent evaluation science.

Discussion

20 comments analyzed.

Competitors mentioned: LangSmith, Langfuse, Braintrust

Concerns raised: Generated hypotheses may not surface multi-factor issues (retrieval, tool selection, state interactions), LitLLM has been recently heavily compromised, RCA agents struggle with proprietary code understanding and lack world models, Evaluators often unqualified to review output quality in production environments

Feature requests: Bootstrap from user feedback signals without requiring SDK instrumentation first, Dimensional analysis before hypothesis generation to reduce feature space

Competitors

Other products that read as similar to this one — 197 launches clear the similarity bar, closest 8 shown.

Attention rank: #42 of 198 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 164 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a agent / copilot tool for Agriculture yet.