I scraped 3B Goodreads reviews to train a better recommendation model
Details
- External ID
- 45825733
- Source
- HN
- Company
- —
- Product
- I scraped 3B Goodreads reviews to train a better recommendation model
- Website domain
- book.sv
- Launched
- Nov. 5, 2025
- Cohort
- —
- Upvotes
- 606
- Upvotes percentile
- 0.9934497816593887
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Hi everyone,For the past couple months I've been working on a website with two main features:- https://book.sv - put in a list of books and get recommendations on what to read next from a model trained on over a billion reviews- https://book.sv/intersect - put in a list of books and find the users on Goodreads who have read them all (if you don't want to be included in these results, you can opt-out here: https://book.sv/remove-my-data)Technical info available here: https://book.sv/how-it-worksNote 1: If you only provide one or two books, the model doesn't have a lot to work with and may include a handful of somewhat unrelated popular books in the results. If you want recommendations based on just one book, click the "Similar" button next to the book after adding it to the input book list on the recommendations page.Note 2: This is uncommon, but if you get an unexpected non-English titled book in the results, it is probably not a mistake and it very likely has an English edition. The "canonical" edition of a book I use for display is whatever one is the most popular, which is usually the English version, but this is not the case for all books, especially those by famous French or Russian authors.
Enrichment
- Theme
- media discovery and streaming tools
- Vertical
- Media & entertainment
- Function
- Analytics & BI
- Audience
- B2C
- AI stance
- AI feature
- Project type
- Hobby / open-source project
- Normalized one-liner
- book recommendation model trained on goodreads reviews
- Manually corrected
- False
Could you build this?
Partial The web interface and vector search endpoints are easily vibe-coded, but scraping 3 billion Goodreads reviews and training a collaborative filtering or embedding model at that scale requires massive data pipelines.
What it would actually take: Building this requires a distributed web scraping cluster running across thousands of residential proxies to bypass Goodreads/Amazon rate limits and Cloudflare protections for billions of pages. Ingesting and cleaning terabytes of raw text demands distributed data processing frameworks like Apache Spark or Ray. Finally, training modern matrix factorization or graph neural network recommendation models (e.g., LightGCN, two-tower models) on billions of interactions requires specialized ML systems engineering and high-memory GPU/TPU instances.
Discussion
20 comments analyzed.
Competitors mentioned: Netflix Cinematch (SVD-based recommendation algorithm), Netflix Prize winning algorithm
Concerns raised: Missing negative feedback/low ratings in model training, Firefox AbortError with concurrency timeout issues, Too many repetitive recommendations from same series/author, Profile access issues with large shelves (5000+ books), User confusion about input format (URL vs numeric ID)
Feature requests: Filter recommendations to exclude already-read authors, Click-through recommendations (chain recommendations from output books), Include negative ratings (1-2 stars) as model features, Reduce author/series repetition in results
Competitors
Other products that read as similar to this one — 100 launches clear the similarity bar, closest 8 shown.
Attention rank: #3 of 101 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Looks like the first mover among its competitors.
- See what readers who loved your favorite book/author also loved to read · hn · 2025-12-29 · 134 upvotes · similarity 0.57
- Visualizing How Books Reference Each Other Across 3k Years · hn · 2026-02-11 · 5 upvotes · similarity 0.54
- Extension to See Rating from Google Book, Amazon,StoryGraph on Goodread · hn · 2026-03-21 · 5 upvotes · similarity 0.52
- Zenòdot · hn · 2026-03-09 · 15 upvotes · similarity 0.50
- Bookrank · ph · 2026-09-11 · 1 upvotes · similarity 0.47
- Sort By Cravings · ph · 2026-09-24 · 1 upvotes · similarity 0.46
- I used Claude Code to discover connections between 100 books · hn · 2026-01-10 · 524 upvotes · similarity 0.46
- MajinBook = Anna's Archive and Goodreads · hn · 2026-08-25 · 7 upvotes · similarity 0.44
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a analytics & bi tool for Legal yet.