Open database of link metadata for large-scale analysis
Details
- External ID
- 46478815
- Source
- HN
- Company
- —
- Product
- database
- Website domain
- github.com
- Launched
- Jan. 3, 2026
- Cohort
- —
- Upvotes
- 15
- Upvotes percentile
- 0.6106719367588933
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
I would like to share an open database focused on link-level metadata extraction and aggregation, which may be of interest to researchers.The project maintains a structured dataset of links enriched with metadata such as:- page title- description / summary- publication date (when available)- thumbnail / preview image- etc.The goal is to provide a reusable, inspectable set of link metadata that can be used for experiments in areas such as:- RSS and feed analysis- news analysis- link rot analysis?The database is publicly available here:https://github.com/rumca-js/RSS-Link-Database-2025There are also databases for previous years
Enrichment
- Theme
- Vertical
- Horizontal
- Function
- Data infrastructure
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Commercial product
- Normalized one-liner
- database of link metadata for analysis
- Manually corrected
- False
Could you build this?
Partial Extracting metadata and exposing a database or API is straightforward, but maintaining an open, large-scale, continuously updated crawl dataset requires heavy data pipeline infrastructure and anti-scraping management.
What it would actually take: The architecture needs a distributed web crawler (e.g., Scrapy, Playwright clusters, or Go-based workers) ingesting millions of URLs, extracting OpenGraph/schema metadata, and storing petabyte-scale outputs in ClickHouse or Parquet on object storage. The difficult parts are rate-limiting, IP proxy rotation, dynamic JavaScript rendering at scale, and high-throughput data deduplication.
Discussion
1 comment analyzed.
Concerns raised: Feed structure changes (fields added/removed, summaries truncated), Data normalization strategy unclear, Longitudinal dataset consistency
Feature requests: Handle RSS source schema evolution, Store raw payload alongside normalized data
Competitors
Other products that read as similar to this one — 54 launches clear the similarity bar, closest 8 shown.
Attention rank: #26 of 55 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 54 days after the earliest competitor.
- PhantomCollect · hn · 2025-11-11 · 7 upvotes · similarity 0.43
- HelixDB · hn · 2026-06-10 · 159 upvotes · similarity 0.42
- SEO Meta Extractor · ph · 2026-09-29 · 2 upvotes · similarity 0.42
- Image URL Dataset · ph · 2026-09-11 · 1 upvotes · similarity 0.39
- OpenFable · hn · 2026-04-08 · 5 upvotes · similarity 0.38
- Ex Situ · hn · 2026-07-21 · 51 upvotes · similarity 0.37
- YaraDB · hn · 2025-11-12 · 10 upvotes · similarity 0.37
- RSS.Style · hn · 2026-01-12 · 6 upvotes · similarity 0.37
Other launches for this product
- A memory database that forgets, consolidates, and detects contradiction
- Open database of 2k IP camera specs (JSON/CSV, CC0)
- KV and wide-column database with CDN-scale replication
- Doberman: The AI watchdog that stops Claude from deleting your database
- database
- Connect DuckDB to any database that has an ADBC driver
- Daily-updated database of malicious browser extensions
- I built a database for AI agents
- Open-source Agent in Rust that can't delete your database
Same idea, different domain
Nobody's really built a data infrastructure tool for Media & entertainment yet.