Infini-News
1.36B news articles from Common Crawl, queryable in ms
Details
- External ID
- 48746018
- Source
- HN
- Company
- —
- Product
- Infini-News
- Website domain
- uni-graz.at
- Launched
- July 1, 2026
- Cohort
- —
- Upvotes
- 7
- Upvotes percentile
- 0.3972520908004779
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Infini-News is ten years of CC-NEWS (the news subset of Common Crawl), cleaned, enriched and turned into a full-text index so you can count any keyword or phrase across 1.36B articles in sub-second time (ok, now maybe a few seconds, but circumstantial), without downloading anything. It's free and open on Hugging Face. I did it because I was sick of having to manually scrape news websites and the like for research purposes and because it felt interesting personally to tackle a project of this scale. On top of data cleaning, we have run language, country (via TLDs and some other heuristics) and topic tagging over all the articles and I have indexed all of them using a recent new n-gram indexing technology that I consider akin to magic. I would encourage you to read the blogpost and play with the interactive viz I made for it. Also, of course, happy to answer questions. Blog: https://cs2.uni-graz.at/blog/infini-news/ Dataset: https://huggingface.co/datasets/ruggsea/infini-news-corpus Index: https://huggingface.co/datasets/ruggsea/infini-news-index Preprint: https://arxiv.org/abs/2605.18337
Enrichment
- Theme
- Hacker News clients, datasets, and tools
- Vertical
- Horizontal
- Function
- Search & retrieval
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Commercial product
- Normalized one-liner
- queryable news article database
- Manually corrected
- False
Could you build this?
No Indexing 1.36 billion articles from Common Crawl for millisecond-latency full-text querying involves petabyte-scale big data engineering, custom distributed inverted indexing, and substantial computing infrastructure that vibe coding cannot deliver.
What it would actually take: The architecture requires a distributed data processing pipeline (Apache Spark/Flink) over Common Crawl WARC archives, coupled with a distributed search engine index (such as heavily sharded Tantivy, Lucene, or ClickHouse with custom tokenization). The hard parts are massive scale ETL (clearing duplicates, boilerplate extraction, compression) and architecting memory-mapped indexes capable of sub-second aggregations across billions of records without astronomical cloud hosting bills. This requires veteran data infrastructure and search engineering expertise.
Discussion
1 comment analyzed.
Competitors
Other products that read as similar to this one — 41 launches clear the similarity bar, closest 8 shown.
Attention rank: #26 of 42 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 242 days after the earliest competitor.
- European tech news in 6 languages · hn · 2025-11-14 · 46 upvotes · similarity 0.46
- NewsBite · ph · 2026-09-25 · 1 upvotes · similarity 0.45
- Real-time system that tracks how news spreads across 200k websites · hn · 2025-11-26 · 256 upvotes · similarity 0.42
- Alexandria, free open source news aggregation and classification suite · hn · 2026-03-26 · 6 upvotes · similarity 0.41
- The Crawl Times · hn · 2026-03-17 · 5 upvotes · similarity 0.40
- Trader News · hn · 2026-09-24 · 25 upvotes · similarity 0.40
- Large Scale Article Extract of Newspapers 1730s-1960s · hn · 2026-05-02 · 57 upvotes · similarity 0.39
- Feed.news · hn · 2026-03-26 · 5 upvotes · similarity 0.37
Other launches for this product
- No other launches for this product.