Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Cheddar-bench

unsupervised benchmark for coding agents

Details

External ID
47110494
Source
HN
Company
—
Product
Cheddar-bench
Website domain
github.com
Launched
Feb. 22, 2026
Cohort
—
Upvotes
9
Upvotes percentile
0.4865229110512129
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

I built a small benchmark to test CLI coding agents on blind bug detection.A challenger agent injects bugs and writes ground truth (`bugs.json`). A different reviewer agent audits the repo without seeing ground truth, and an LLM matcher scores bug-to-finding assignments.Current run: 50 repos, 150 challenges, 450 reviews, 2,603 injected bugs.Weighted detection: Claude 58.05%, Codex 37.84%, Gemini 27.81%.LLM-judge benchmarks are easy to get wrong, so I’d really appreciate critical feedback on benchmark fairness, scoring/matching methodology, and obvious failure modes I’m missing.Full dataset is linked in the docs.

Enrichment

Theme
coding agent interfaces and environments
Vertical
Horizontal
Function
Analytics & BI
Audience
Developer
AI stance
Not AI
Project type
Hobby / open-source project
Normalized one-liner
benchmark for coding agents
Manually corrected
False

Could you build this?

Yes This is a benchmarking harness that scripts LLM API calls to inject code changes, prompts another agent to find them, and uses an LLM-as-a-judge matcher to calculate metrics.

Discussion

No comments on this launch.

Competitors

Other products that read as similar to this one — 324 launches clear the similarity bar, closest 8 shown.

Attention rank: #158 of 325 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 114 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a analytics & bi tool for Legal yet.