Large Scale Article Extract of Newspapers 1730s-1960s
Details
- External ID
- 47984614
- Source
- HN
- Company
- —
- Product
- Large Scale Article Extract of Newspapers 1730s-1960s
- Website domain
- snewpapers.com
- Launched
- May 2, 2026
- Cohort
- —
- Upvotes
- 57
- Upvotes percentile
- 0.8481421647819063
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Hello HN, over the past 7 months I've spent nearly 3,000 hours on building SNEWPAPERS, the first historical newpaper archive with full-text extractions, nearly perfect OCR, a vast categorization taxonomy and of course with semantic and agentic search capabilities.Problem: I wanted to search through newspaper archives, but when I tried every service only lets you search for keywords and dates, and gives you back raw images of the papers, and too many of them with no context. A sea of noise.Solution: I taught machines how to read the newspapers and so far I've extracted the content from > 600k pages (about 5TB) from the Chronicling America collection. Problems I had to deal with were an infinite variety of layouts, font sizes, image scan qualities, resolutions, aspect ratios, navigating around the images on the page. I also had to figure out how to get OCR to be nearly perfect so people wouldn't hate reading the extracts. I stitched together a multi-model pipeline (layout tech, ocr tech, llm, vllm) with heuristics to go from layout -> segmentation -> classification. I put it all in OpenSearch / Postgres and made it semantically searchable and also put an agentic search tool on top that knows how to use the API really well and helps you write queries to find what you're looking for. Happy to discuss AWS architecture and scaling as well, that was tough!If you have five minutes and you just want to jump in and have your own personalized experience, what I would suggest is:Before searching for anything, go to the Sleuth page Ask it about anything from 1736 to 1963, maybe 1 or 2 follow up questions Then go to the search page so you can see the queries it wrote for you (bottom left "saved queries") and uncover more info on whatever it is you're interested inIf you think it's cool and you want to learn more, then there's about 10 minutes of video guides on the various capabilities in "Guide" on the nav barSome other people have also taken a crack at this, notably:https://dell-research-harvard.github.io/resources/americanst... (very good attempt) https://labs.loc.gov/work/experiments/newspaper-navigator/ (focused on images)
Enrichment
- Theme
- searchable public records and archives
- Vertical
- Media & entertainment
- Function
- Data infrastructure
- Audience
- B2B
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- historical newspaper article database
- Manually corrected
- False
Could you build this?
No Building an archive spanning centuries requires ingesting petabytes of historical scans, training custom OCR models capable of reading antique typography and deteriorated paper, and operating massive search infrastructure.
What it would actually take: Requires high-throughput distributed processing pipelines (e.g., Apache Spark/Flink, Ray) processing millions of digitized newspaper TIFF/JPEG images. The hard parts are custom vision/OCR models trained on historical 18th-20th century fonts, article segmentation algorithms to disentangle multi-column layouts, and petabyte-scale semantic vector indexes combined with full-text search (Elasticsearch/OpenSearch).
Discussion
20 comments analyzed.
Competitors mentioned: Arcanum (newspaper segmentation/OCR), Paddle Paddle (document layout analysis models), ProQuest (historical newspaper archive), vLLMs (text extraction from images)
Concerns raised: OCR struggles with magazine layouts, multi-column text, headers, and text over images, Complex newspaper layouts make reading order prediction difficult even with SOTA models, Pricing information not visible before signup/free trial, LLM cleanup passes miss obvious errors/typos in OCR output, Concerns about backend database load from free search access
Feature requests: Free searchable subset of data (e.g., one year of specific topic) without registration, Trending analysis across time (monthly/yearly top headlines, left/right publisher swings), Category-specific limited free searches for authenticated users, Better handling of old-time spellings and historical text variants, Improved reading order for complex newspaper layouts
Competitors
Other products that read as similar to this one — 74 launches clear the similarity bar, closest 8 shown.
Attention rank: #16 of 75 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 182 days after the earliest competitor.
- Doomscrolling Research Papers · hn · 2025-12-02 · 14 upvotes · similarity 0.47
- The Federalist Papers, typeset as the 1787 newspapers they ran in · hn · 2026-07-29 · 58 upvotes · similarity 0.46
- I indexed the academic papers buried in the DOJ Epstein Files · hn · 2026-02-20 · 7 upvotes · similarity 0.46
- Forty.News · hn · 2025-11-22 · 443 upvotes · similarity 0.44
- Scanned 1927-1945 Daily USFS Work Diary · hn · 2026-02-16 · 121 upvotes · similarity 0.43
- Garden of Flowers · hn · 2026-06-16 · 162 upvotes · similarity 0.42
- Research Papers as Memes · hn · 2025-11-28 · 10 upvotes · similarity 0.41
- Browsing forgotten digicam photos from early Flickr (2000–2012) · hn · 2026-04-15 · 5 upvotes · similarity 0.41
Other launches for this product
- No other launches for this product.