Self-host Reddit
2.38B posts, works offline, yours forever
Details
- External ID
- 46602324
- Source
- HN
- Company
- —
- Product
- Self-host Reddit
- Website domain
- github.com
- Launched
- Jan. 13, 2026
- Cohort
- —
- Upvotes
- 286
- Upvotes percentile
- 0.9789196310935442
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Reddit's API is effectively dead for archival. Third-party apps are gone. Reddit has threatened to cut off access to the Pushshift dataset multiple times. But 3.28TB of Reddit history exists as a torrent right now, and I built a tool to turn it into something you can browse on your own hardware.The key point: This doesn't touch Reddit's servers. Ever. Download the Pushshift dataset, run my tool locally, get a fully browsable archive. Works on an air-gapped machine. Works on a Raspberry Pi serving your LAN. Works on a USB drive you hand to someone.What it does: Takes compressed data dumps from Reddit (.zst), Voat (SQL), and Ruqqus (.7z) and generates static HTML. No JavaScript, no external requests, no tracking. Open index.html and browse. Want search? Run the optional Docker stack with PostgreSQL – still entirely on your machine.API & AI Integration: Full REST API with 30+ endpoints – posts, comments, users, subreddits, full-text search, aggregations. Also ships with an MCP server (29 tools) so you can query your archive directly from AI tools.Self-hosting options: - USB drive / local folder (just open the HTML files) - Home server on your LAN - Tor hidden service (2 commands, no port forwarding needed) - VPS with HTTPS - GitHub Pages for small archivesWhy this matters: Once you have the data, you own it. No API keys, no rate limits, no ToS changes can take it away.Scale: Tens of millions of posts per instance. PostgreSQL backend keeps memory constant regardless of dataset size. For the full 2.38B post dataset, run multiple instances by topic.How I built it: Python, PostgreSQL, Jinja2 templates, Docker. Used Claude Code throughout as an experiment in AI-assisted development. Learned that the workflow is "trust but verify" – it accelerates the boring parts but you still own the architecture.Live demo: https://online-archives.github.io/redd-archiver-example/GitHub: https://github.com/19-84/redd-archiver (Public Domain)Pushshift torrent: https://academictorrents.com/details/1614740ac8c94505e4ecb9d...
Enrichment
- Theme
- database infrastructure and developer tools
- Vertical
- Horizontal
- Function
- Vertical SaaS
- Audience
- B2C
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- self-hosted reddit alternative
- Manually corrected
- False
Could you build this?
Partial While the UI and search wrapper are straightforward, streaming, decompressing, and indexing a 3.28TB / 2.38-billion-post dataset to run efficiently on consumer hardware is a heavy data engineering problem.
What it would actually take: A workable stack uses Rust or Go to stream zstd-compressed archives into an embedded columnar or full-text engine like DuckDB or Tantivy, paired with a local SQLite index for comment hierarchy. The primary challenge is memory-constrained decompression, efficient disk layout, and fast tree reconstruction for billions of threaded comments without crashing consumer machines.
Discussion
20 comments analyzed.
Competitors mentioned: Pushshift Reddit data on Hugging Face Datasets, Voat, Lemmy instances, Archive.org YouTube metadata
Concerns raised: Reddit astroturfing and bot campaigns, Locked content unavailable on Internet Archive, NSFW content concerns, Reddit data scraping legality and ToS compliance, Moderation and extreme content on platforms
Feature requests: Docker compose configuration, YouTube channels dataset with account/creator details, Data migration to Lemmy instances
Competitors
Other products that read as similar to this one — 112 launches clear the similarity bar, closest 8 shown.
Attention rank: #5 of 113 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 72 days after the earliest competitor.
- View Any Reddit Profile Instantly · ph · 2026-09-27 · 1 upvotes · similarity 0.46
- Snapbyte · hn · 2026-01-21 · 5 upvotes · similarity 0.43
- Pbnj · hn · 2025-12-05 · 69 upvotes · similarity 0.43
- I built a smart blocker after destroying my dopamine baseline · hn · 2025-11-02 · 20 upvotes · similarity 0.42
- Reddit GDPR Export Viewer · hn · 2026-01-17 · 5 upvotes · similarity 0.40
- Trying to fix the web scraping industry's benchmark problem · hn · 2026-07-16 · 18 upvotes · similarity 0.40
- ObsessionDB · hn · 2026-01-23 · 12 upvotes · similarity 0.40
- Posthorn, self-hosted mail gateway · hn · 2026-05-27 · 83 upvotes · similarity 0.40
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a vertical saas tool for Insurance yet.