Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

RepoGym

Validate the task before you benchmark the coding agent

Details

External ID
1258166
Source
PH
Company
—
Product
RepoGym
Website domain
producthunt.com
Launched
Sept. 23, 2026
Cohort
—
Upvotes
1
Upvotes percentile
0.30815693820825313
Tags
Developer Tools, GitHub
Fetched at
Sept. 25, 2026, 1:02 a.m.
Updated at
Sept. 25, 2026, 1:02 a.m.

Description

Turn repository tasks into repeatable coding-agent evaluations. Define a task, check that doing nothing fails and the reference fix passes, then inspect scores and trajectories. Includes Python and JavaScript examples, composable graders, task mining and HTML reports. Free, MIT-licensed Python library. Start from the source checkout; local execution is for trusted code.

Enrichment

Theme
coding agent interfaces and environments
Vertical
Horizontal
Function
Observability & eval
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
benchmark task validator for coding agents
Manually corrected
False

Could you build this?

Yes It is an open-source Python evaluation harness that runs test suites inside Docker containers and reports diff/pass/fail results for AI coding agents.

Competitors

Other products that read as similar to this one — 243 launches clear the similarity bar, closest 8 shown.

Attention rank: #170 of 244 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 327 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a observability & eval tool for Media & entertainment yet.