Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Trying to fix the web scraping industry's benchmark problem

Details

External ID
48935383
Source
HN
Company
—
Product
Trying to fix the web scraping industry's benchmark problem
Website domain
github.com
Launched
July 16, 2026
Cohort
—
Upvotes
18
Upvotes percentile
0.6893667861409797
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

We've been trying to evaluate web scraping companies, but when you look at their benchmarks, you can't verify anything, and they mostly exist to prove the company is successful. They put somewhere between 98% and 100% because they pick their own urls, define success their own way, and don't publish the harness. We also saw companies like scrapfly astroturf websites like scrapeway and call them independent.So, we built an open source benchmark that we want to represent the frontier of web data. We're trying to look across all major anti-bot providers and industries, to build a comprehensive hard-target test set. It's all in the repo, you can run it, and we pledge to maintain it monthly.Curious what people think we could do to improve it? Would you change any of the conditions, add sites, etc. Our target list and pass criteria are something we're trying to improve, so would love to get feedback from the community on how we can actually provide something useful for people to test against!

Enrichment

Theme
proxy, dns, and networking tools
Vertical
—
Function
Analytics & BI
Audience
B2B
AI stance
Not AI
Project type
Commercial product
Normalized one-liner
web scraping industry benchmark
Manually corrected
False

Could you build this?

Partial The benchmarking dashboard and test orchestration UI can be vibe-coded, but running automated scraping evaluations against hostile anti-bot defenses requires massive proxy infrastructure and reverse-engineering skills.

What it would actually take: The architecture requires a distributed worker cluster executing headless browsers across rotating residential and mobile proxy pools, paired with automated fingerprint analysis. The hard part is continuously bypassing and accurately measuring responses against advanced bot mitigations (Cloudflare, Akamai, Datadome, TLS/JA4 fingerprinting, CAPTCHA challenges) without burning proxy fleets. This requires deep web scraping and reverse-engineering domain expertise.

Discussion

No comments on this launch.

Competitors

Other products that read as similar to this one — 172 launches clear the similarity bar, closest 8 shown.

Attention rank: #49 of 173 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 256 days after the earliest competitor.

Other launches for this product