Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Infini-News

1.36B news articles from Common Crawl, queryable in ms

Details

External ID
48746018
Source
HN
Company
—
Product
Infini-News
Website domain
uni-graz.at
Launched
July 1, 2026
Cohort
—
Upvotes
7
Upvotes percentile
0.3972520908004779
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Infini-News is ten years of CC-NEWS (the news subset of Common Crawl), cleaned, enriched and turned into a full-text index so you can count any keyword or phrase across 1.36B articles in sub-second time (ok, now maybe a few seconds, but circumstantial), without downloading anything. It's free and open on Hugging Face. I did it because I was sick of having to manually scrape news websites and the like for research purposes and because it felt interesting personally to tackle a project of this scale. On top of data cleaning, we have run language, country (via TLDs and some other heuristics) and topic tagging over all the articles and I have indexed all of them using a recent new n-gram indexing technology that I consider akin to magic. I would encourage you to read the blogpost and play with the interactive viz I made for it. Also, of course, happy to answer questions. Blog: https://cs2.uni-graz.at/blog/infini-news/ Dataset: https://huggingface.co/datasets/ruggsea/infini-news-corpus Index: https://huggingface.co/datasets/ruggsea/infini-news-index Preprint: https://arxiv.org/abs/2605.18337

Enrichment

Theme
Hacker News clients, datasets, and tools
Vertical
Horizontal
Function
Search & retrieval
Audience
Developer
AI stance
Not AI
Project type
Commercial product
Normalized one-liner
queryable news article database
Manually corrected
False

Could you build this?

No Indexing 1.36 billion articles from Common Crawl for millisecond-latency full-text querying involves petabyte-scale big data engineering, custom distributed inverted indexing, and substantial computing infrastructure that vibe coding cannot deliver.

What it would actually take: The architecture requires a distributed data processing pipeline (Apache Spark/Flink) over Common Crawl WARC archives, coupled with a distributed search engine index (such as heavily sharded Tantivy, Lucene, or ClickHouse with custom tokenization). The hard parts are massive scale ETL (clearing duplicates, boilerplate extraction, compression) and architecting memory-mapped indexes capable of sub-second aggregations across billions of records without astronomical cloud hosting bills. This requires veteran data infrastructure and search engineering expertise.

Discussion

1 comment analyzed.

Competitors

Other products that read as similar to this one — 41 launches clear the similarity bar, closest 8 shown.

Attention rank: #26 of 42 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 242 days after the earliest competitor.

Other launches for this product