Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

I scraped 3B Goodreads reviews to train a better recommendation model

Details

External ID
45825733
Source
HN
Company
—
Product
I scraped 3B Goodreads reviews to train a better recommendation model
Website domain
book.sv
Launched
Nov. 5, 2025
Cohort
—
Upvotes
606
Upvotes percentile
0.9934497816593887
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Hi everyone,For the past couple months I've been working on a website with two main features:- https://book.sv - put in a list of books and get recommendations on what to read next from a model trained on over a billion reviews- https://book.sv/intersect - put in a list of books and find the users on Goodreads who have read them all (if you don't want to be included in these results, you can opt-out here: https://book.sv/remove-my-data)Technical info available here: https://book.sv/how-it-worksNote 1: If you only provide one or two books, the model doesn't have a lot to work with and may include a handful of somewhat unrelated popular books in the results. If you want recommendations based on just one book, click the "Similar" button next to the book after adding it to the input book list on the recommendations page.Note 2: This is uncommon, but if you get an unexpected non-English titled book in the results, it is probably not a mistake and it very likely has an English edition. The "canonical" edition of a book I use for display is whatever one is the most popular, which is usually the English version, but this is not the case for all books, especially those by famous French or Russian authors.

Enrichment

Theme
media discovery and streaming tools
Vertical
Media & entertainment
Function
Analytics & BI
Audience
B2C
AI stance
AI feature
Project type
Hobby / open-source project
Normalized one-liner
book recommendation model trained on goodreads reviews
Manually corrected
False

Could you build this?

Partial The web interface and vector search endpoints are easily vibe-coded, but scraping 3 billion Goodreads reviews and training a collaborative filtering or embedding model at that scale requires massive data pipelines.

What it would actually take: Building this requires a distributed web scraping cluster running across thousands of residential proxies to bypass Goodreads/Amazon rate limits and Cloudflare protections for billions of pages. Ingesting and cleaning terabytes of raw text demands distributed data processing frameworks like Apache Spark or Ray. Finally, training modern matrix factorization or graph neural network recommendation models (e.g., LightGCN, two-tower models) on billions of interactions requires specialized ML systems engineering and high-memory GPU/TPU instances.

Discussion

20 comments analyzed.

Competitors mentioned: Netflix Cinematch (SVD-based recommendation algorithm), Netflix Prize winning algorithm

Concerns raised: Missing negative feedback/low ratings in model training, Firefox AbortError with concurrency timeout issues, Too many repetitive recommendations from same series/author, Profile access issues with large shelves (5000+ books), User confusion about input format (URL vs numeric ID)

Feature requests: Filter recommendations to exclude already-read authors, Click-through recommendations (chain recommendations from output books), Include negative ratings (1-2 stars) as model features, Reduce author/series repetition in results

Competitors

Other products that read as similar to this one — 100 launches clear the similarity bar, closest 8 shown.

Attention rank: #3 of 101 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Looks like the first mover among its competitors.

Other launches for this product

Same idea, different domain

Nobody's really built a analytics & bi tool for Legal yet.