Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

I indexed the academic papers buried in the DOJ Epstein Files

Details

External ID
47083160
Source
HN
Company
—
Product
I indexed the academic papers buried in the DOJ Epstein Files
Website domain
jeescholar.com
Launched
Feb. 20, 2026
Cohort
—
Upvotes
7
Upvotes percentile
0.38207547169811323
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

The DOJ released ~3.5M pages of Epstein documents across 12 datasets. Buried in them are 207 academic papers and 14 books that nobody was really talking about. From what I understand these papers aren't usually freely accesible, but since they are public documents, now they are.I don't know, thought it was interesting to see what this dude was reading. You can check it out at jeescholar.com Pipeline: 1. Downloaded all 12 DOJ datasets + House Oversight Committee release 2. Heuristic pre-filter (abstract detection, DOI regex, citation block patterns, affiliation strings) to cut noise 3. LLM classifier to confirm and extract metadata 4. CrossRef and Semantic Scholar APIs for DOI matching, citation counts, abstracts 5. 87 of 207 papers got DOI matches; the rest are identified but not in major indexes Stack: FastAPI + SQLite (FTS5 for full-text search) + Cloudflare R2 for PDFs + nginx/Docker on Hetzner. The fields represented are genuinely iteresting: there's a cluster of child abuse/grooming research, but also quantum gravity, AGI safety, econophysics, and regenerative medicine. Each paper links back to its original government PDF and Bates number. For sure not an exhaustive list. Would be happy to add more if anyone finds them.

Enrichment

Theme
searchable public records and archives
Vertical
Legal
Function
Search & retrieval
Audience
B2B
AI stance
Not AI
Project type
Hobby / open-source project
Normalized one-liner
academic paper indexing from legal documents
Manually corrected
False

Could you build this?

Yes The project is a standard pipeline: extracting text/PDFs, running heuristic/LLM filters to classify papers, querying public scholarly APIs (Crossref/Semantic Scholar), and serving them in a full-text search web UI.

Discussion

No comments on this launch.

Competitors

Other products that read as similar to this one — 94 launches clear the similarity bar, closest 8 shown.

Attention rank: #64 of 95 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 101 days after the earliest competitor.

Other launches for this product