RepoGym
Validate the task before you benchmark the coding agent
Details
- External ID
- 1258166
- Source
- PH
- Company
- —
- Product
- RepoGym
- Website domain
- producthunt.com
- Launched
- Sept. 23, 2026
- Cohort
- —
- Upvotes
- 1
- Upvotes percentile
- 0.30815693820825313
- Tags
- Developer Tools, GitHub
- Fetched at
- Sept. 25, 2026, 1:02 a.m.
- Updated at
- Sept. 25, 2026, 1:02 a.m.
Description
Turn repository tasks into repeatable coding-agent evaluations. Define a task, check that doing nothing fails and the reference fix passes, then inspect scores and trajectories. Includes Python and JavaScript examples, composable graders, task mining and HTML reports. Free, MIT-licensed Python library. Start from the source checkout; local execution is for trusted code.
Enrichment
- Theme
- coding agent interfaces and environments
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- benchmark task validator for coding agents
- Manually corrected
- False
Could you build this?
Yes It is an open-source Python evaluation harness that runs test suites inside Docker containers and reports diff/pass/fail results for AI coding agents.
Competitors
Other products that read as similar to this one — 243 launches clear the similarity bar, closest 8 shown.
Attention rank: #170 of 244 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 327 days after the earliest competitor.
- RepoGauntlet · ph · 2026-09-28 · 1 upvotes · similarity 0.51
- Cheddar-bench · hn · 2026-02-22 · 9 upvotes · similarity 0.46
- Score your GitHub repo for AI coding agents · hn · 2026-03-17 · 7 upvotes · similarity 0.45
- repro-lens · github · 2026-09-13 · 9 upvotes · similarity 0.43
- Recursive-Mode for Coding Agents · hn · 2026-04-11 · 6 upvotes · similarity 0.43
- Agentic coding workflows built on Git worktrees and task evidence · hn · 2026-06-20 · 10 upvotes · similarity 0.43
- formwork · github · 2026-09-09 · 12 upvotes · similarity 0.43
- Real-SWE: A coding benchmark built from private company codebases · yc · 2026-09-10 · 11 upvotes · similarity 0.43
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.