Benchmark: AI doesn't find bugs unless you tell it what's wrong
This is 1 of 584 launches in ai agent infrastructure and tooling — see how it stacks up on momentum and crowding →
146 other launches read as similar to this one →
Details
- External ID
- 49923102
- Source
- HN
- Company
- —
- Product
- Benchmark: AI doesn't find bugs unless you tell it what's wrong
- Website domain
- swesweep.com
- Launched
- Oct. 1, 2026
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.11764705882352941
- Tags
- —
- Fetched at
- Oct. 2, 2026, 1:01 a.m.
- Updated at
- Oct. 2, 2026, 1:01 a.m.
Description
New benchmark from researchers at Meta, Stanford, Harvard, UW, including the researchers who've worked on SWE-bench, ProgramBench etc.Most benchmarks just test if AI can fix a problem you've already pointed out.But obviously it would be much better to fix problems before you or any user runs into it. Like, isn't it crazy that we still have to wait for people to open tickets before a lot of obvious bugs get found?We wanted to test that capability at scale. Turns out that models still are terrible at it (best setup we tested still fixed < 5% of bugs)We have 100 repos of 22 languages and 4k bugs between. All the bugs are real-world bugs from github. We do a lot of filtering to ensure everything can be solved in this setting.===========================================Sol 5.6 (xhigh) 4.7% $7,230Luna 5.6 (xhigh) 2.5% $224Terra 5.6 (xhigh) 1.5% $357Luna 5.6 (high) 1.4% $28Opus 5 (xhigh) 1.3% $5,363Kimi K3 0.6% $2,451Luna 5.6 0.5% $4GPT-5.4 Mini (high) 0.5% $122GPT-5.4 Mini 0.2% $5Gemini 3.5 Flash Lite 0.1% $6===========================================Also the best model is very expensive.We have a lot more FAQ on the website https://swesweep.com/ Oh and we're all open-source (MIT license) at https://github.com/facebookresearch/swe-sweepCurious what you all think!
Enrichment
- Theme
- ai agent infrastructure and tooling
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- benchmark for evaluating ai bug-finding capabilities
- Manually corrected
- False
Could you build this?
No SWE-sweep is a massive research benchmark spanning 100 repositories and over 4,000 verified bugs, requiring deep program analysis, harness execution sandboxes, and academic rigour.
What it would actually take: Constructing this benchmark requires automated repository mining, regression test generation, and complex isolated Dockerized execution environments to reliably verify multi-bug codebases. It demands significant systems and software engineering expertise, alongside thousands of dollars in cloud/LLM evaluation compute to benchmark models.
Discussion
2 comments analyzed.
Concerns raised: Models hallucinating fixes on unbroken code
Feature requests: Show precision alongside recall on the leaderboard
Competitors
Other products that read as similar to this one — 146 launches clear the similarity bar, closest 8 shown.
Attention rank: #124 of 147 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 334 days after the earliest competitor.
- Detail, a Bug Finder · hn · 2025-12-09 · 67 upvotes · similarity 0.56
- Statewright · hn · 2026-05-12 · 126 upvotes · similarity 0.49
- Compute:Arena · hn · 2026-09-17 · 5 upvotes · similarity 0.48
- New Benchmark from SWE-bench team is 0% solved · hn · 2026-05-05 · 24 upvotes · similarity 0.47
- Pencil Puzzle Bench · hn · 2026-03-03 · 5 upvotes · similarity 0.46
- Cheddar-bench · hn · 2026-02-22 · 9 upvotes · similarity 0.46
- Don't ask if devs cheat with AI, test if they're good with it · hn · 2026-06-30 · 5 upvotes · similarity 0.45
- Benchmark your eng team's AI agent maturity in 5 minutes · hn · 2026-07-14 · 14 upvotes · similarity 0.45
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.