Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Open database of link metadata for large-scale analysis

Details

External ID
46478815
Source
HN
Company
—
Product
database
Website domain
github.com
Launched
Jan. 3, 2026
Cohort
—
Upvotes
15
Upvotes percentile
0.6106719367588933
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

I would like to share an open database focused on link-level metadata extraction and aggregation, which may be of interest to researchers.The project maintains a structured dataset of links enriched with metadata such as:- page title- description / summary- publication date (when available)- thumbnail / preview image- etc.The goal is to provide a reusable, inspectable set of link metadata that can be used for experiments in areas such as:- RSS and feed analysis- news analysis- link rot analysis?The database is publicly available here:https://github.com/rumca-js/RSS-Link-Database-2025There are also databases for previous years

Enrichment

Theme
Vertical
Horizontal
Function
Data infrastructure
Audience
Developer
AI stance
Not AI
Project type
Commercial product
Normalized one-liner
database of link metadata for analysis
Manually corrected
False

Could you build this?

Partial Extracting metadata and exposing a database or API is straightforward, but maintaining an open, large-scale, continuously updated crawl dataset requires heavy data pipeline infrastructure and anti-scraping management.

What it would actually take: The architecture needs a distributed web crawler (e.g., Scrapy, Playwright clusters, or Go-based workers) ingesting millions of URLs, extracting OpenGraph/schema metadata, and storing petabyte-scale outputs in ClickHouse or Parquet on object storage. The difficult parts are rate-limiting, IP proxy rotation, dynamic JavaScript rendering at scale, and high-throughput data deduplication.

Discussion

1 comment analyzed.

Concerns raised: Feed structure changes (fields added/removed, summaries truncated), Data normalization strategy unclear, Longitudinal dataset consistency

Feature requests: Handle RSS source schema evolution, Store raw payload alongside normalized data

Competitors

Other products that read as similar to this one — 54 launches clear the similarity bar, closest 8 shown.

Attention rank: #26 of 55 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 54 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a data infrastructure tool for Media & entertainment yet.