Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Meta-agent: self-improving agent harnesses from live traces

Details

External ID
47665630
Source
HN
Company
—
Product
Meta-agent: self-improving agent harnesses from live traces
Website domain
github.com
Launched
April 6, 2026
Cohort
—
Upvotes
14
Upvotes percentile
0.6876606683804627
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

We built meta-agent: an open-source library that automatically and continuously improves agent harnesses from production traces.Point it at an existing agent, a stream of unlabeled production traces, and a small labeled holdout set.An LLM judge scores unlabeled production traces as they stream.A proposer reads failed traces and writes one targeted harness update at a time, such as changes to prompts, hooks, tools, or subagents. The update is kept only if it improves holdout accuracy.On tau-bench v3 airline, meta-agent improved holdout accuracy from 67% to 87%.We open-sourced meta-agent. It currently supports Claude Agent SDK, with more frameworks coming soon.Try it here: https://github.com/canvas-org/meta-agent

Enrichment

Theme
AI agent frameworks and developer tools
Vertical
Horizontal
Function
Agent / copilot
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
self-improving ai agent framework
Manually corrected
False

Could you build this?

Partial The basic framework of tracing prompts and feeding outputs to an LLM judge is approachable, but building an automated optimizer that continuously refines agent harnesses without regressing requires complex evaluation pipelines and prompt optimization theory.

What it would actually take: The system requires a distributed tracing and observability pipeline (OpenTelemetry, ClickHouse/PostgreSQL), an automated LLM-as-a-judge scoring harness, and an iterative optimization engine using algorithms like DSPy or GEPO (Genetic Prompt Optimization). The hard problem is preventing reward hacking, managing prompt drift, and maintaining statistical rigor across agent harness mutations without human-in-the-loop validation, requiring specialized AI evaluation engineering.

Discussion

No comments on this launch.

Competitors

Other products that read as similar to this one — 698 launches clear the similarity bar, closest 8 shown.

Attention rank: #215 of 699 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 159 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a agent / copilot tool for Agriculture yet.