We are building Git for data
Details
- External ID
- 46785065
- Source
- HN
- Company
- —
- Product
- —
- Website domain
- —
- Launched
- Jan. 27, 2026
- Cohort
- —
- Upvotes
- 9
- Upvotes percentile
- 0.461133069828722
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Today you can easily adopt AI coding tools because you have git for branching and rolling back if AI writes bad code. We haven't seen this same capability for data and decided to build it ourselves.Nile is a new kind of data lake, purpose built for using with AI. It can act as your data engineer or data analyst creating new tables and rolling back bad changes in seconds. We support real versions for data, schema, and ETL.We'd love your feedback on any part of what we are building - https://getnile.ai/What do you think?
Enrichment
- Theme
- database infrastructure and developer tools
- Vertical
- Horizontal
- Function
- Data infrastructure
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Commercial product
- Normalized one-liner
- version control for data
- Manually corrected
- False
Could you build this?
No Building a version-controlled data lake ('Git for data') requires deep database systems engineering, distributed systems storage primitives, and complex copy-on-write transactional architectures.
What it would actually take: A production implementation requires building or extending distributed storage formats (like Apache Iceberg, Delta Lake, or Nessie/lakeFS) with multi-table ACID transactions, zero-copy branching, snapshot isolation, and metadata catalog management. The hardest parts are handling petabyte-scale metadata commits, distributed conflict resolution, and garbage collection without performance degradation. This demands deep systems architecture and distributed database engineering expertise.
Discussion
2 comments analyzed.
Competitors mentioned: Iceberg, Delta Lake
Concerns raised: How branching/merging scales, Cost predictability, Trust and observability
Feature requests: Cost predictability features, Observability improvements, Trust/reliability guarantees
Competitors
Other products that read as similar to this one — 86 launches clear the similarity bar, closest 8 shown.
Attention rank: #47 of 87 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 87 days after the earliest competitor.
- Tracking AI Code with Git AI · hn · 2025-11-10 · 6 upvotes · similarity 0.49
- Git for AI Agents · hn · 2026-05-08 · 129 upvotes · similarity 0.45
- Git for AI Agents · hn · 2026-05-07 · 5 upvotes · similarity 0.45
- CleanCode · ph · 2026-09-14 · 1 upvotes · similarity 0.42
- I built a local data lake for AI powered data engineering and analytics · hn · 2026-04-08 · 14 upvotes · similarity 0.42
- A Git structure for AI to orchestrate across code, docs, and ops · hn · 2026-01-06 · 5 upvotes · similarity 0.40
- briefd · ph · 2026-09-23 · 5 upvotes · similarity 0.39
- Mljar Studio · hn · 2026-05-02 · 73 upvotes · similarity 0.39
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a data infrastructure tool for Media & entertainment yet.