Stop AI scrapers from hammering your self-hosted blog (using porn)
Details
- External ID
- 46294144
- Source
- HN
- Company
- —
- Product
- Stop AI scrapers from hammering your self-hosted blog (using porn)
- Website domain
- github.com
- Launched
- Dec. 16, 2025
- Cohort
- —
- Upvotes
- 373
- Upvotes percentile
- 0.9732824427480916
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Alright so if you run a self-hosted blog, you've probably noticed AI companies scraping it for training data. And not just a little (RIP to your server bill).There isn't much you can do about it without cloudflare. These companies ignore robots.txt, and you're competing with teams with more resources than you. It's you vs the MJs of programming, you're not going to win.But there is a solution. Now I'm not going to say it's a great solution...but a solution is a solution. If your website contains content that will trigger their scraper's safeguards, it will get dropped from their data pipelines.So here's what fuzzycanary does: it injects hundreds of invisible links to porn websites in your HTML. The links are hidden from users but present in the DOM so that scrapers can ingest them and say "nope we won't scrape there again in the future".The problem with that approach is that it will absolutely nuke your website's SEO. So fuzzycanary also checks user agents and won't show the links to legitimate search engines, so Google and Bing won't see them.One caveat: if you're using a static site generator it will bake the links into your HTML for everyone, including googlebot. Does anyone have a work-around for this that doesn't involve using a proxy?Please try it out! Setup is one component or one import.(And don't tell me it's a terrible idea because I already know it is)package: https://www.npmjs.com/package/@fuzzycanary/core gh: https://github.com/vivienhenz24/fuzzy-canary
Enrichment
- Theme
- browser automation and scraping for AI
- Vertical
- Horizontal
- Function
- —
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- block ai scrapers from self-hosted content
- Manually corrected
- False
Could you build this?
Yes This is a simple HTTP reverse proxy or middleware script that checks user-agent/traffic signatures and serves alternate adult/decoy content to suspected scraper IPs.
Discussion
20 comments analyzed.
Competitors mentioned: reCAPTCHA, Cloudflare, PiHole
Concerns raised: SEO impact from blocking scrapers, IP addresses too ephemeral for effective filtering, Ongoing arms race between scrapers and detection methods, Default CAPTCHA images inappropriate/out of context
Feature requests: Decentralized trust algorithm for traffic filtering, Decay mechanism for IP reputation scoring
Competitors
Other products that read as similar to this one — 68 launches clear the similarity bar, closest 8 shown.
Attention rank: #2 of 69 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 44 days after the earliest competitor.
- I built an SDK that scrambles HTML so scrapers get garbage · hn · 2026-03-12 · 16 upvotes · similarity 0.47
- Trying to fix the web scraping industry's benchmark problem · hn · 2026-07-16 · 18 upvotes · similarity 0.45
- Bravo Crawler Hunter · ph · 2026-09-12 · 3 upvotes · similarity 0.42
- Snitchmd · hn · 2026-04-29 · 8 upvotes · similarity 0.42
- Roast My Website · ph · 2026-09-26 · 2 upvotes · similarity 0.41
- Self-host Reddit · hn · 2026-01-13 · 286 upvotes · similarity 0.40
- TafScraper · ph · 2026-09-28 · 3 upvotes · similarity 0.40
- BrowserAct Cloud · ph · 2026-08-14 · 287 upvotes · similarity 0.39
Other launches for this product
- No other launches for this product.