minhash-dedup
Bucket-feasible MinHash/LSH deduplication for text corpora: keeps one document per duplicate bucket, not one per connected component
Details
- External ID
- 1384136380
- Source
- GITHUB
- Company
- —
- Product
- minhash-dedup
- Website domain
- github.com
- Launched
- Sept. 23, 2026
- Cohort
- —
- Upvotes
- 16
- Upvotes percentile
- 0.5194722008711248
- Tags
- —
- Fetched at
- Sept. 27, 2026, 5:02 p.m.
- Updated at
- Sept. 27, 2026, 5:02 p.m.
Enrichment
- Theme
- low-level systems and developer tools
- Vertical
- Horizontal
- Function
- Data infrastructure
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- minhash text deduplication for text corpora
- Manually corrected
- False
Could you build this?
Yes It is a self-contained text processing utility implementing MinHash/LSH algorithms with custom clustering logic, easily implemented with standard Python or Rust libraries.
Competitors
Other products that read as similar to this one — 124 launches clear the similarity bar, closest 8 shown.
Attention rank: #50 of 125 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 320 days after the earliest competitor.
- genpark-agentic-cache-semantic-deduplicator-skill · github · 2026-09-14 · 8 upvotes · similarity 0.43
- genpark-agentic-cache-semantic-deduplicator-skill · github · 2026-09-14 · 7 upvotes · similarity 0.43
- Agfs · hn · 2025-11-18 · 9 upvotes · similarity 0.42
- udoc. Dependency-free document extraction in Rust · hn · 2026-05-20 · 5 upvotes · similarity 0.41
- hop · github · 2026-09-11 · 7 upvotes · similarity 0.41
- genpark-consistent-hashing-virtual-nodes-skill · github · 2026-09-10 · 7 upvotes · similarity 0.40
- genpark-locality-sensitive-hashing-lsh-cosine-index-skill · github · 2026-09-29 · 7 upvotes · similarity 0.39
- Marmot · hn · 2025-12-02 · 103 upvotes · similarity 0.39
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a data infrastructure tool for Media & entertainment yet.