Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Selfship.ai

Surface and fix isues with your agentic applications 24x7

Details

External ID
49522967
Source
HN
Company
—
Product
Selfship.ai
Website domain
selfship.ai
Launched
Sept. 1, 2026
Cohort
—
Upvotes
10
Upvotes percentile
0.5853269537480064
Tags
—
Fetched at
Sept. 10, 2026, 5:31 a.m.
Updated at
Sept. 10, 2026, 5:31 a.m.

Description

We've been building an AI chat based trading system for last 3 years. The biggest issue was that when our Agent would mess up, we wouldn't know until a user reported. Agent traces helped us uncover what's wrong. But surfacing issues was almost always a manual trigger. So from those learnings, we built Selfship.ai. It's an autonomous system that observes every trace/turn/multi turn convo to find out issues. If a user got what they wanted, if a tool call is failing repeatedly, if users have to always reframe their questions, if the agent is taking optimal paths and many more. It's a loop - group failures by user intent, evaluate them, and ship fixes as PRs. After a fix is deployed, it evaluates if it worked or not. We recently opened it up as a SaaS. If you have an agentic product in production, we would love for you to try it out.

Enrichment

Theme
task-specific ai agents and assistants
Vertical
Horizontal
Function
Observability & eval
Audience
Developer
AI stance
AI feature
Project type
Commercial product
Normalized one-liner
monitoring and debugging for agentic applications
Manually corrected
False

Could you build this?

Partial A web dashboard and GitHub PR generator are straightforward, but building automated root-cause failure clustering on agent execution traces that produces verified code fixes is difficult.

What it would actually take: Requires an ingestion pipeline that parses heterogeneous LLM/agent traces (OpenTelemetry, Langfuse format), performs semantic clustering on trace error topologies, and isolates failure-inducing prompt/code steps. An evaluation sandbox spins up headless Docker containers to reproduce the failed session, runs a coding model against the repo to patch it, and verifies the fix via regression tests before opening a PR. This demands solid DevOps/sandboxing engineering and robust agent evaluation methodologies.

Discussion

10 comments analyzed.

Competitors mentioned: Domino (systematic error/slice discovery), AgentBoard (trajectory/progress evaluation), τ-bench (goal/outcome correctness), MAST (failure taxonomies)

Concerns raised: How to distinguish actual agent failures from normal user drop-offs, Risk of overfitting fixes to specific conversations rather than solving root problems, Uncertainty about optimal conversation turn count for effectiveness, Unclear if approach has domain-specific limitations or breaks down in certain areas

Feature requests: On-premise deployment option, Identify domain-specific sweet spots for best performance

Competitors

Other products that read as similar to this one — 561 launches clear the similarity bar, closest 8 shown.

Attention rank: #244 of 562 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 301 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a observability & eval tool for Media & entertainment yet.