Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Self-host Reddit

2.38B posts, works offline, yours forever

Details

External ID
46602324
Source
HN
Company
—
Product
Self-host Reddit
Website domain
github.com
Launched
Jan. 13, 2026
Cohort
—
Upvotes
286
Upvotes percentile
0.9789196310935442
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Reddit's API is effectively dead for archival. Third-party apps are gone. Reddit has threatened to cut off access to the Pushshift dataset multiple times. But 3.28TB of Reddit history exists as a torrent right now, and I built a tool to turn it into something you can browse on your own hardware.The key point: This doesn't touch Reddit's servers. Ever. Download the Pushshift dataset, run my tool locally, get a fully browsable archive. Works on an air-gapped machine. Works on a Raspberry Pi serving your LAN. Works on a USB drive you hand to someone.What it does: Takes compressed data dumps from Reddit (.zst), Voat (SQL), and Ruqqus (.7z) and generates static HTML. No JavaScript, no external requests, no tracking. Open index.html and browse. Want search? Run the optional Docker stack with PostgreSQL – still entirely on your machine.API & AI Integration: Full REST API with 30+ endpoints – posts, comments, users, subreddits, full-text search, aggregations. Also ships with an MCP server (29 tools) so you can query your archive directly from AI tools.Self-hosting options: - USB drive / local folder (just open the HTML files) - Home server on your LAN - Tor hidden service (2 commands, no port forwarding needed) - VPS with HTTPS - GitHub Pages for small archivesWhy this matters: Once you have the data, you own it. No API keys, no rate limits, no ToS changes can take it away.Scale: Tens of millions of posts per instance. PostgreSQL backend keeps memory constant regardless of dataset size. For the full 2.38B post dataset, run multiple instances by topic.How I built it: Python, PostgreSQL, Jinja2 templates, Docker. Used Claude Code throughout as an experiment in AI-assisted development. Learned that the workflow is "trust but verify" – it accelerates the boring parts but you still own the architecture.Live demo: https://online-archives.github.io/redd-archiver-example/GitHub: https://github.com/19-84/redd-archiver (Public Domain)Pushshift torrent: https://academictorrents.com/details/1614740ac8c94505e4ecb9d...

Enrichment

Theme
database infrastructure and developer tools
Vertical
Horizontal
Function
Vertical SaaS
Audience
B2C
AI stance
Not AI
Project type
Hobby / open-source project
Normalized one-liner
self-hosted reddit alternative
Manually corrected
False

Could you build this?

Partial While the UI and search wrapper are straightforward, streaming, decompressing, and indexing a 3.28TB / 2.38-billion-post dataset to run efficiently on consumer hardware is a heavy data engineering problem.

What it would actually take: A workable stack uses Rust or Go to stream zstd-compressed archives into an embedded columnar or full-text engine like DuckDB or Tantivy, paired with a local SQLite index for comment hierarchy. The primary challenge is memory-constrained decompression, efficient disk layout, and fast tree reconstruction for billions of threaded comments without crashing consumer machines.

Discussion

20 comments analyzed.

Competitors mentioned: Pushshift Reddit data on Hugging Face Datasets, Voat, Lemmy instances, Archive.org YouTube metadata

Concerns raised: Reddit astroturfing and bot campaigns, Locked content unavailable on Internet Archive, NSFW content concerns, Reddit data scraping legality and ToS compliance, Moderation and extreme content on platforms

Feature requests: Docker compose configuration, YouTube channels dataset with account/creator details, Data migration to Lemmy instances

Competitors

Other products that read as similar to this one — 112 launches clear the similarity bar, closest 8 shown.

Attention rank: #5 of 113 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 72 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a vertical saas tool for Insurance yet.