CatchAll
slowest web search API that outperforms everything on recall
Details
- External ID
- 47863906
- Source
- HN
- Company
- —
- Product
- CatchAll
- Website domain
- newscatcherapi.com
- Launched
- April 22, 2026
- Cohort
- —
- Upvotes
- 7
- Upvotes percentile
- 0.38817480719794345
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Hey HN,Artem and Maksym from NewsCatcher here.Some of you know us as we started six years ago as two freshly graduated economics students who decided to build the best news API product.We started NewsCatcher thinking the market for news APIs was so big that we could build a self-serve platform and get millions of $29 users.Obviously, it was a wrong assumption. We pivoted to serve enterprises and had success with it.But we are hackers at heart, and we want to serve hackers.We haven't used our Launch HN yet, so consider this our smoke test.We're looking for feedback and power users rather than revenue. So, happy to provide enough credits for any HN user who finds CatchAll useful.CatchAll is built for one thing: retrieving every matching event from the web. The use cases that fit it are ones where missing events have real consequences — funding and M&A monitoring, regulatory and compliance feeds (FDA approvals, SEC filings, policy changes), cybersecurity incident tracking, supply chain signals.If your pipeline consumes structured records and the answer to your query is "find all of them," that's where it works. It's not the right tool for small, bounded queries that return 5 high-precision results.The 15-minute job time is a direct consequence of the pipeline depth: analyze, fetch, cluster, validate, extract, deduplicate. You're not getting a ranked list of links; you're getting a verified record set.Our latest benchmark run: https://newscatcherapi.com/blog-posts/web-search-api-benchma...
Enrichment
- Theme
- Hacker News clients, datasets, and tools
- Vertical
- Horizontal
- Function
- Search & retrieval
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Commercial product
- Normalized one-liner
- web search api with high recall
- Manually corrected
- False
Could you build this?
No Operating a massive web-scale news crawler and search engine with high recall across billions of articles requires massive infrastructure, continuous scraping, anti-bot circumvention, and distributed index management.
What it would actually take: The system requires a distributed web crawling cluster ingesting tens of millions of RSS feeds, sitemaps, and web pages daily, bypassing anti-bot shields and paywalls at scale. Data processing pipelines use distributed Kafka/Flink clusters for deduplication, content extraction, and entity recognition, indexing into petabyte-scale Elasticsearch/OpenSearch or hybrid vector indices. The primary obstacle is the massive ongoing capital expenditure and systems engineering needed to maintain continuous web ingestion and high-recall indexing.
Discussion
1 comment analyzed.
Competitors
Other products that read as similar to this one — 57 launches clear the similarity bar, closest 8 shown.
Attention rank: #38 of 58 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 166 days after the earliest competitor.
- CatchAll: Recall-first web search API · yc · 2026-04-16 · 5 upvotes · similarity 0.43
- I built a fast RSS reader in Zig · hn · 2025-12-16 · 90 upvotes · similarity 0.42
- HN Buffer · hn · 2025-11-22 · 5 upvotes · similarity 0.39
- Trying to fix the web scraping industry's benchmark problem · hn · 2026-07-16 · 18 upvotes · similarity 0.39
- I measured the half-life of 41,301 Show HN launches. It's 7 hours · hn · 2026-07-02 · 37 upvotes · similarity 0.38
- Seltz · hn · 2026-04-20 · 5 upvotes · similarity 0.37
- Feed.news · hn · 2026-03-26 · 5 upvotes · similarity 0.37
- Hackobar · hn · 2026-05-25 · 5 upvotes · similarity 0.36
Other launches for this product
- No other launches for this product.