Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Stop AI scrapers from hammering your self-hosted blog (using porn)

Details

External ID
46294144
Source
HN
Company
—
Product
Stop AI scrapers from hammering your self-hosted blog (using porn)
Website domain
github.com
Launched
Dec. 16, 2025
Cohort
—
Upvotes
373
Upvotes percentile
0.9732824427480916
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Alright so if you run a self-hosted blog, you've probably noticed AI companies scraping it for training data. And not just a little (RIP to your server bill).There isn't much you can do about it without cloudflare. These companies ignore robots.txt, and you're competing with teams with more resources than you. It's you vs the MJs of programming, you're not going to win.But there is a solution. Now I'm not going to say it's a great solution...but a solution is a solution. If your website contains content that will trigger their scraper's safeguards, it will get dropped from their data pipelines.So here's what fuzzycanary does: it injects hundreds of invisible links to porn websites in your HTML. The links are hidden from users but present in the DOM so that scrapers can ingest them and say "nope we won't scrape there again in the future".The problem with that approach is that it will absolutely nuke your website's SEO. So fuzzycanary also checks user agents and won't show the links to legitimate search engines, so Google and Bing won't see them.One caveat: if you're using a static site generator it will bake the links into your HTML for everyone, including googlebot. Does anyone have a work-around for this that doesn't involve using a proxy?Please try it out! Setup is one component or one import.(And don't tell me it's a terrible idea because I already know it is)package: https://www.npmjs.com/package/@fuzzycanary/core gh: https://github.com/vivienhenz24/fuzzy-canary

Enrichment

Theme
browser automation and scraping for AI
Vertical
Horizontal
Function
—
Audience
Developer
AI stance
Not AI
Project type
Hobby / open-source project
Normalized one-liner
block ai scrapers from self-hosted content
Manually corrected
False

Could you build this?

Yes This is a simple HTTP reverse proxy or middleware script that checks user-agent/traffic signatures and serves alternate adult/decoy content to suspected scraper IPs.

Discussion

20 comments analyzed.

Competitors mentioned: reCAPTCHA, Cloudflare, PiHole

Concerns raised: SEO impact from blocking scrapers, IP addresses too ephemeral for effective filtering, Ongoing arms race between scrapers and detection methods, Default CAPTCHA images inappropriate/out of context

Feature requests: Decentralized trust algorithm for traffic filtering, Decay mechanism for IP reputation scoring

Competitors

Other products that read as similar to this one — 68 launches clear the similarity bar, closest 8 shown.

Attention rank: #2 of 69 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 44 days after the earliest competitor.

Other launches for this product