Visualizing How Books Reference Each Other Across 3k Years
Details
- External ID
- 46978805
- Source
- HN
- Company
- —
- Product
- Visualizing How Books Reference Each Other Across 3k Years
- Website domain
- github.io
- Launched
- Feb. 11, 2026
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.10512129380053908
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
There are two parts for this project:1) The LLM-powered pipeline to extract citations (books + authors) from books and resolve them using both Wikipedia and Goodreads with offline copies I have. The result is data associating Books/Authors to other Books/Authors with accurate bibliographical information spanning centuries.2) A WebGPU + D3.js powered visualization tool written by Claude Code so I'm able to deal with all this data on the browser on a more or less comfortable experience for the viewer.I spent some months on a off with this project, and definitely the most challenging part was dealing with accurate bibliographical information across centuries, with original publication dates and etc. For that I wrote what is now a very complex pipeline with LLMs (I used DeepSeek V3.2) wired on offline Goodreads and Wikipedia databases + a fallback that actually uses the internet.Hope you enjoy it! Open to suggestions on how to improve the system :)Code is here: https://github.com/ThiagoLira/bookgraph-revisited
Enrichment
- Theme
- browser utilities and bookmark managers
- Vertical
- Media & entertainment
- Function
- Analytics & BI
- Audience
- B2C
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- visualization of historical book references
- Manually corrected
- False
Could you build this?
Partial Building the interactive 3D/timeline frontend is straightforward, but processing centuries of book texts and resolving entity citations against Wikipedia/Goodreads offline dumps requires serious data pipeline work.
What it would actually take: The backend requires an ingestion and OCR pipeline across thousands of historical texts, NER extraction using local or API-based LLMs, and an entity disambiguation engine matching against Wikidata and Goodreads dumps. The frontend requires a WebGL/Three.js or D3 force-directed 3D network visualization optimized for high-density historical graph rendering. A developer needs strong graph data modeling and large-scale data engineering experience.
Discussion
3 comments analyzed.
Competitors mentioned: DeepSeek V3.2
Concerns raised: High token costs for processing large datasets, Data accuracy issues (incorrect historical dating)
Competitors
Other products that read as similar to this one — 146 launches clear the similarity bar, closest 8 shown.
Attention rank: #138 of 147 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 100 days after the earliest competitor.
- I used Claude Code to discover connections between 100 books · hn · 2026-01-10 · 524 upvotes · similarity 0.58
- I scraped 3B Goodreads reviews to train a better recommendation model · hn · 2025-11-05 · 606 upvotes · similarity 0.54
- litsurvey · github · 2026-09-11 · 7 upvotes · similarity 0.47
- Wikigraph · hn · 2026-06-02 · 12 upvotes · similarity 0.45
- Librario, a book metadata API that aggregates G Books, ISBNDB, and more · hn · 2026-01-10 · 140 upvotes · similarity 0.45
- A better LLM-wiki with multi-path research [550 stars] · hn · 2026-06-11 · 5 upvotes · similarity 0.45
- What 180k words look like as a temporal knowledge graph (Oz series) · hn · 2026-07-26 · 22 upvotes · similarity 0.44
- genpark-deep-research-cross-citation-graph-skill · github · 2026-09-12 · 7 upvotes · similarity 0.44
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a analytics & bi tool for Legal yet.