Trying to fix the web scraping industry's benchmark problem
Details
- External ID
- 48935383
- Source
- HN
- Company
- —
- Product
- Trying to fix the web scraping industry's benchmark problem
- Website domain
- github.com
- Launched
- July 16, 2026
- Cohort
- —
- Upvotes
- 18
- Upvotes percentile
- 0.6893667861409797
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
We've been trying to evaluate web scraping companies, but when you look at their benchmarks, you can't verify anything, and they mostly exist to prove the company is successful. They put somewhere between 98% and 100% because they pick their own urls, define success their own way, and don't publish the harness. We also saw companies like scrapfly astroturf websites like scrapeway and call them independent.So, we built an open source benchmark that we want to represent the frontier of web data. We're trying to look across all major anti-bot providers and industries, to build a comprehensive hard-target test set. It's all in the repo, you can run it, and we pledge to maintain it monthly.Curious what people think we could do to improve it? Would you change any of the conditions, add sites, etc. Our target list and pass criteria are something we're trying to improve, so would love to get feedback from the community on how we can actually provide something useful for people to test against!
Enrichment
- Theme
- proxy, dns, and networking tools
- Vertical
- —
- Function
- Analytics & BI
- Audience
- B2B
- AI stance
- Not AI
- Project type
- Commercial product
- Normalized one-liner
- web scraping industry benchmark
- Manually corrected
- False
Could you build this?
Partial The benchmarking dashboard and test orchestration UI can be vibe-coded, but running automated scraping evaluations against hostile anti-bot defenses requires massive proxy infrastructure and reverse-engineering skills.
What it would actually take: The architecture requires a distributed worker cluster executing headless browsers across rotating residential and mobile proxy pools, paired with automated fingerprint analysis. The hard part is continuously bypassing and accurately measuring responses against advanced bot mitigations (Cloudflare, Akamai, Datadome, TLS/JA4 fingerprinting, CAPTCHA challenges) without burning proxy fleets. This requires deep web scraping and reverse-engineering domain expertise.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 172 launches clear the similarity bar, closest 8 shown.
Attention rank: #49 of 173 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 256 days after the earliest competitor.
- 12K Reddit posts scraped and AI-scored for startup ideas · hn · 2025-11-22 · 5 upvotes · similarity 0.50
- WhiskeySour · hn · 2026-04-25 · 8 upvotes · similarity 0.48
- StackScope · hn · 2026-06-12 · 67 upvotes · similarity 0.48
- Proxy Tester by ScrapeOps · ph · 2026-08-09 · 208 upvotes · similarity 0.47
- scrape.land · ph · 2026-09-21 · 11 upvotes · similarity 0.47
- I built an SDK that scrambles HTML so scrapers get garbage · hn · 2026-03-12 · 16 upvotes · similarity 0.46
- Stop AI scrapers from hammering your self-hosted blog (using porn) · hn · 2025-12-16 · 373 upvotes · similarity 0.45
- OpenBenchmarks · hn · 2026-07-11 · 6 upvotes · similarity 0.44
Other launches for this product
- No other launches for this product.