Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Benchmark: AI doesn't find bugs unless you tell it what's wrong

This is 1 of 584 launches in ai agent infrastructure and tooling — see how it stacks up on momentum and crowding →

146 other launches read as similar to this one →

Details

External ID
49923102
Source
HN
Company
—
Product
Benchmark: AI doesn't find bugs unless you tell it what's wrong
Website domain
swesweep.com
Launched
Oct. 1, 2026
Cohort
—
Upvotes
5
Upvotes percentile
0.11764705882352941
Tags
—
Fetched at
Oct. 2, 2026, 1:01 a.m.
Updated at
Oct. 2, 2026, 1:01 a.m.

Description

New benchmark from researchers at Meta, Stanford, Harvard, UW, including the researchers who've worked on SWE-bench, ProgramBench etc.Most benchmarks just test if AI can fix a problem you've already pointed out.But obviously it would be much better to fix problems before you or any user runs into it. Like, isn't it crazy that we still have to wait for people to open tickets before a lot of obvious bugs get found?We wanted to test that capability at scale. Turns out that models still are terrible at it (best setup we tested still fixed < 5% of bugs)We have 100 repos of 22 languages and 4k bugs between. All the bugs are real-world bugs from github. We do a lot of filtering to ensure everything can be solved in this setting.===========================================Sol 5.6 (xhigh) 4.7% $7,230Luna 5.6 (xhigh) 2.5% $224Terra 5.6 (xhigh) 1.5% $357Luna 5.6 (high) 1.4% $28Opus 5 (xhigh) 1.3% $5,363Kimi K3 0.6% $2,451Luna 5.6 0.5% $4GPT-5.4 Mini (high) 0.5% $122GPT-5.4 Mini 0.2% $5Gemini 3.5 Flash Lite 0.1% $6===========================================Also the best model is very expensive.We have a lot more FAQ on the website https://swesweep.com/ Oh and we're all open-source (MIT license) at https://github.com/facebookresearch/swe-sweepCurious what you all think!

Enrichment

Theme
ai agent infrastructure and tooling
Vertical
Horizontal
Function
Observability & eval
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
benchmark for evaluating ai bug-finding capabilities
Manually corrected
False

Could you build this?

No SWE-sweep is a massive research benchmark spanning 100 repositories and over 4,000 verified bugs, requiring deep program analysis, harness execution sandboxes, and academic rigour.

What it would actually take: Constructing this benchmark requires automated repository mining, regression test generation, and complex isolated Dockerized execution environments to reliably verify multi-bug codebases. It demands significant systems and software engineering expertise, alongside thousands of dollars in cloud/LLM evaluation compute to benchmark models.

Discussion

2 comments analyzed.

Concerns raised: Models hallucinating fixes on unbroken code

Feature requests: Show precision alongside recall on the leaderboard

Competitors

Other products that read as similar to this one — 146 launches clear the similarity bar, closest 8 shown.

Attention rank: #124 of 147 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 334 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a observability & eval tool for Media & entertainment yet.