Verse AI
Catch the AI failures your evals miss
Details
- External ID
- 45914386
- Source
- HN
- Company
- —
- Product
- Verse AI
- Website domain
- tryverse.ai
- Launched
- Nov. 13, 2025
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.0982532751091703
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Our AI recruitment pipeline was auto-rejecting anyone who'd worked at companies founded after 2023. It didn't recognize names like Harvey or Snorkel AI, or didn't realize how important they'd become because of training data cutoffs.We had traces, evals, Langfuse dashboards - everything looked fine - but we kept finding failures we should have caught earlier.The pattern kept repeating:- ship an improvement - it works for a while - hit an edge case that breaks it - don't notice until we've lost good candidatesThat's when we realized - the problem wasn't just our recruitment pipeline - almost every AI product has blind spots that evals miss.So we built Verse, a tool that surfaces issues directly from real AI interactions - whether that's candidates talking to your recruitment pipeline, users interacting with your agent, or any AI making decisions.Instead of relying solely on evals, we cluster conversations, identify the key ones to review, and flag the ones that show failure patterns. We use OpenTelemetry for trace ingestion, so it's compatible with Langfuse, Langsmith, Braintrust, and other AI observability tools - you can add it right alongside your existing setup.I'm posting this because I'm curious whether other teams are hitting the same wall. If you want, I'm happy to audit your AI implementation for free and show you where things commonly break - even if you never use Verse.Happy to answer any technical questions.
Enrichment
- Theme
- AI agent frameworks and developer tools
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- B2B
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- ai failure detection and evaluation
- Manually corrected
- False
Could you build this?
Partial The front-end and telemetry tracking are straightforward, but building an adversarial/automated LLM failure detection engine that uncovers edge-case hallucinations and distribution shifts requires sophisticated synthetic testing methodology.
What it would actually take: The architecture needs high-throughput trace collection, semantic clustering, automated adversarial test case generation (red-teaming), and custom meta-evaluator models. The difficult aspect is detecting unknown unknowns (like out-of-distribution entity drops or subtle reasoning failures) without relying solely on static ground truth labels. This requires machine learning eval specialization and continuous pipeline testing infrastructure.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 158 launches clear the similarity bar, closest 8 shown.
Attention rank: #149 of 159 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 9 days after the earliest competitor.
- Agnost AI · ph · 2026-08-25 · 289 upvotes · similarity 0.50
- Memograph CLI- A tool to diagnose 'memory failures' in AI agents · hn · 2026-02-25 · 6 upvotes · similarity 0.44
- Agent-evals · hn · 2026-05-04 · 9 upvotes · similarity 0.43
- AI-Evals.io · hn · 2026-02-15 · 5 upvotes · similarity 0.42
- I Built a Debugging Challenge for the AI Coding Age · hn · 2026-05-25 · 5 upvotes · similarity 0.42
- Progress AI Observability · ph · 2026-08-07 · 168 upvotes · similarity 0.42
- Spec27 · hn · 2026-04-30 · 13 upvotes · similarity 0.41
- We post-trained a model that pen tests instead of refusing · hn · 2026-06-20 · 93 upvotes · similarity 0.41
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.