Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

CatchAll

slowest web search API that outperforms everything on recall

Details

External ID
47863906
Source
HN
Company
—
Product
CatchAll
Website domain
newscatcherapi.com
Launched
April 22, 2026
Cohort
—
Upvotes
7
Upvotes percentile
0.38817480719794345
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Hey HN,Artem and Maksym from NewsCatcher here.Some of you know us as we started six years ago as two freshly graduated economics students who decided to build the best news API product.We started NewsCatcher thinking the market for news APIs was so big that we could build a self-serve platform and get millions of $29 users.Obviously, it was a wrong assumption. We pivoted to serve enterprises and had success with it.But we are hackers at heart, and we want to serve hackers.We haven't used our Launch HN yet, so consider this our smoke test.We're looking for feedback and power users rather than revenue. So, happy to provide enough credits for any HN user who finds CatchAll useful.CatchAll is built for one thing: retrieving every matching event from the web. The use cases that fit it are ones where missing events have real consequences — funding and M&A monitoring, regulatory and compliance feeds (FDA approvals, SEC filings, policy changes), cybersecurity incident tracking, supply chain signals.If your pipeline consumes structured records and the answer to your query is "find all of them," that's where it works. It's not the right tool for small, bounded queries that return 5 high-precision results.The 15-minute job time is a direct consequence of the pipeline depth: analyze, fetch, cluster, validate, extract, deduplicate. You're not getting a ranked list of links; you're getting a verified record set.Our latest benchmark run: https://newscatcherapi.com/blog-posts/web-search-api-benchma...

Enrichment

Theme
Hacker News clients, datasets, and tools
Vertical
Horizontal
Function
Search & retrieval
Audience
Developer
AI stance
Not AI
Project type
Commercial product
Normalized one-liner
web search api with high recall
Manually corrected
False

Could you build this?

No Operating a massive web-scale news crawler and search engine with high recall across billions of articles requires massive infrastructure, continuous scraping, anti-bot circumvention, and distributed index management.

What it would actually take: The system requires a distributed web crawling cluster ingesting tens of millions of RSS feeds, sitemaps, and web pages daily, bypassing anti-bot shields and paywalls at scale. Data processing pipelines use distributed Kafka/Flink clusters for deduplication, content extraction, and entity recognition, indexing into petabyte-scale Elasticsearch/OpenSearch or hybrid vector indices. The primary obstacle is the massive ongoing capital expenditure and systems engineering needed to maintain continuous web ingestion and high-recall indexing.

Discussion

1 comment analyzed.

Competitors

Other products that read as similar to this one — 57 launches clear the similarity bar, closest 8 shown.

Attention rank: #38 of 58 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 166 days after the earliest competitor.

Other launches for this product