Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Large Scale Article Extract of Newspapers 1730s-1960s

Details

External ID
47984614
Source
HN
Company
—
Product
Large Scale Article Extract of Newspapers 1730s-1960s
Website domain
snewpapers.com
Launched
May 2, 2026
Cohort
—
Upvotes
57
Upvotes percentile
0.8481421647819063
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Hello HN, over the past 7 months I've spent nearly 3,000 hours on building SNEWPAPERS, the first historical newpaper archive with full-text extractions, nearly perfect OCR, a vast categorization taxonomy and of course with semantic and agentic search capabilities.Problem: I wanted to search through newspaper archives, but when I tried every service only lets you search for keywords and dates, and gives you back raw images of the papers, and too many of them with no context. A sea of noise.Solution: I taught machines how to read the newspapers and so far I've extracted the content from > 600k pages (about 5TB) from the Chronicling America collection. Problems I had to deal with were an infinite variety of layouts, font sizes, image scan qualities, resolutions, aspect ratios, navigating around the images on the page. I also had to figure out how to get OCR to be nearly perfect so people wouldn't hate reading the extracts. I stitched together a multi-model pipeline (layout tech, ocr tech, llm, vllm) with heuristics to go from layout -> segmentation -> classification. I put it all in OpenSearch / Postgres and made it semantically searchable and also put an agentic search tool on top that knows how to use the API really well and helps you write queries to find what you're looking for. Happy to discuss AWS architecture and scaling as well, that was tough!If you have five minutes and you just want to jump in and have your own personalized experience, what I would suggest is:Before searching for anything, go to the Sleuth page Ask it about anything from 1736 to 1963, maybe 1 or 2 follow up questions Then go to the search page so you can see the queries it wrote for you (bottom left "saved queries") and uncover more info on whatever it is you're interested inIf you think it's cool and you want to learn more, then there's about 10 minutes of video guides on the various capabilities in "Guide" on the nav barSome other people have also taken a crack at this, notably:https://dell-research-harvard.github.io/resources/americanst... (very good attempt) https://labs.loc.gov/work/experiments/newspaper-navigator/ (focused on images)

Enrichment

Theme
searchable public records and archives
Vertical
Media & entertainment
Function
Data infrastructure
Audience
B2B
AI stance
Not AI
Project type
Hobby / open-source project
Normalized one-liner
historical newspaper article database
Manually corrected
False

Could you build this?

No Building an archive spanning centuries requires ingesting petabytes of historical scans, training custom OCR models capable of reading antique typography and deteriorated paper, and operating massive search infrastructure.

What it would actually take: Requires high-throughput distributed processing pipelines (e.g., Apache Spark/Flink, Ray) processing millions of digitized newspaper TIFF/JPEG images. The hard parts are custom vision/OCR models trained on historical 18th-20th century fonts, article segmentation algorithms to disentangle multi-column layouts, and petabyte-scale semantic vector indexes combined with full-text search (Elasticsearch/OpenSearch).

Discussion

20 comments analyzed.

Competitors mentioned: Arcanum (newspaper segmentation/OCR), Paddle Paddle (document layout analysis models), ProQuest (historical newspaper archive), vLLMs (text extraction from images)

Concerns raised: OCR struggles with magazine layouts, multi-column text, headers, and text over images, Complex newspaper layouts make reading order prediction difficult even with SOTA models, Pricing information not visible before signup/free trial, LLM cleanup passes miss obvious errors/typos in OCR output, Concerns about backend database load from free search access

Feature requests: Free searchable subset of data (e.g., one year of specific topic) without registration, Trending analysis across time (monthly/yearly top headlines, left/right publisher swings), Category-specific limited free searches for authenticated users, Better handling of old-time spellings and historical text variants, Improved reading order for complex newspaper layouts

Competitors

Other products that read as similar to this one — 74 launches clear the similarity bar, closest 8 shown.

Attention rank: #16 of 75 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 182 days after the earliest competitor.

Other launches for this product