Selfship.ai
Surface and fix isues with your agentic applications 24x7
Details
- External ID
- 49522967
- Source
- HN
- Company
- —
- Product
- Selfship.ai
- Website domain
- selfship.ai
- Launched
- Sept. 1, 2026
- Cohort
- —
- Upvotes
- 10
- Upvotes percentile
- 0.5853269537480064
- Tags
- —
- Fetched at
- Sept. 10, 2026, 5:31 a.m.
- Updated at
- Sept. 10, 2026, 5:31 a.m.
Description
We've been building an AI chat based trading system for last 3 years. The biggest issue was that when our Agent would mess up, we wouldn't know until a user reported. Agent traces helped us uncover what's wrong. But surfacing issues was almost always a manual trigger. So from those learnings, we built Selfship.ai. It's an autonomous system that observes every trace/turn/multi turn convo to find out issues. If a user got what they wanted, if a tool call is failing repeatedly, if users have to always reframe their questions, if the agent is taking optimal paths and many more. It's a loop - group failures by user intent, evaluate them, and ship fixes as PRs. After a fix is deployed, it evaluates if it worked or not. We recently opened it up as a SaaS. If you have an agentic product in production, we would love for you to try it out.
Enrichment
- Theme
- task-specific ai agents and assistants
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI feature
- Project type
- Commercial product
- Normalized one-liner
- monitoring and debugging for agentic applications
- Manually corrected
- False
Could you build this?
Partial A web dashboard and GitHub PR generator are straightforward, but building automated root-cause failure clustering on agent execution traces that produces verified code fixes is difficult.
What it would actually take: Requires an ingestion pipeline that parses heterogeneous LLM/agent traces (OpenTelemetry, Langfuse format), performs semantic clustering on trace error topologies, and isolates failure-inducing prompt/code steps. An evaluation sandbox spins up headless Docker containers to reproduce the failed session, runs a coding model against the repo to patch it, and verifies the fix via regression tests before opening a PR. This demands solid DevOps/sandboxing engineering and robust agent evaluation methodologies.
Discussion
10 comments analyzed.
Competitors mentioned: Domino (systematic error/slice discovery), AgentBoard (trajectory/progress evaluation), τ-bench (goal/outcome correctness), MAST (failure taxonomies)
Concerns raised: How to distinguish actual agent failures from normal user drop-offs, Risk of overfitting fixes to specific conversations rather than solving root problems, Uncertainty about optimal conversation turn count for effectiveness, Unclear if approach has domain-specific limitations or breaks down in certain areas
Feature requests: On-premise deployment option, Identify domain-specific sweet spots for best performance
Competitors
Other products that read as similar to this one — 561 launches clear the similarity bar, closest 8 shown.
Attention rank: #244 of 562 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 301 days after the earliest competitor.
- ReplyPilot · ph · 2026-09-17 · 1 upvotes · similarity 0.50
- AgentInflux · ph · 2026-09-07 · 1 upvotes · similarity 0.47
- Robinhood Agentic Trading · ph · 2026-05-28 · 165 upvotes · similarity 0.47
- Logic · ph · 2026-04-27 · 274 upvotes · similarity 0.46
- Agent Checker · ph · 2026-09-06 · 2 upvotes · similarity 0.46
- Artis · ph · 2026-09-30 · 2 upvotes · similarity 0.46
- AMA2, messenger built for AI agent · hn · 2026-06-30 · 5 upvotes · similarity 0.45
- AgentRow · ph · 2026-09-25 · 2 upvotes · similarity 0.44
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.