Cheddar-bench
unsupervised benchmark for coding agents
Details
- External ID
- 47110494
- Source
- HN
- Company
- —
- Product
- Cheddar-bench
- Website domain
- github.com
- Launched
- Feb. 22, 2026
- Cohort
- —
- Upvotes
- 9
- Upvotes percentile
- 0.4865229110512129
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
I built a small benchmark to test CLI coding agents on blind bug detection.A challenger agent injects bugs and writes ground truth (`bugs.json`). A different reviewer agent audits the repo without seeing ground truth, and an LLM matcher scores bug-to-finding assignments.Current run: 50 repos, 150 challenges, 450 reviews, 2,603 injected bugs.Weighted detection: Claude 58.05%, Codex 37.84%, Gemini 27.81%.LLM-judge benchmarks are easy to get wrong, so I’d really appreciate critical feedback on benchmark fairness, scoring/matching methodology, and obvious failure modes I’m missing.Full dataset is linked in the docs.
Enrichment
- Theme
- coding agent interfaces and environments
- Vertical
- Horizontal
- Function
- Analytics & BI
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- benchmark for coding agents
- Manually corrected
- False
Could you build this?
Yes This is a benchmarking harness that scripts LLM API calls to inject code changes, prompts another agent to find them, and uses an LLM-as-a-judge matcher to calculate metrics.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 324 launches clear the similarity bar, closest 8 shown.
Attention rank: #158 of 325 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 114 days after the earliest competitor.
- Pencil Puzzle Bench · hn · 2026-03-03 · 5 upvotes · similarity 0.49
- Cua-Bench · hn · 2026-01-26 · 40 upvotes · similarity 0.47
- OpenBenchmarks · hn · 2026-07-11 · 6 upvotes · similarity 0.47
- RepoGym · ph · 2026-09-23 · 1 upvotes · similarity 0.46
- LLM agents that write Python to analyze execution traces at scale · hn · 2026-03-07 · 5 upvotes · similarity 0.45
- Open Benchmarks Grants– a $3M commitment to close the AI eval gap · hn · 2026-02-11 · 6 upvotes · similarity 0.44
- Autofix Bot · hn · 2025-12-11 · 37 upvotes · similarity 0.43
- SOCBench · hn · 2026-07-07 · 6 upvotes · similarity 0.43
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a analytics & bi tool for Legal yet.