Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

DDL to Data

Generate realistic test data from SQL schemas

Details

External ID
46511578
Source
HN
Company
—
Product
—
Website domain
—
Launched
Jan. 6, 2026
Cohort
—
Upvotes
55
Upvotes percentile
0.8155467720685112
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

I built DDL to Data after repeatedly pushing back on "just use production data and mask it" requests. Teams needed populated databases for testing, but pulling prod meant security reviews, PII scrubbing, and DevOps tickets. Hand-written seed scripts were the alternative slow, fragile, and out of sync the moment schemas changed.Paste your CREATE TABLE statements, get realistic test data back. It parses your schema, preserves foreign key relationships, and generates data that looks real, emails look like emails, timestamps are reasonable, uniqueness constraints are honored.No setup, no config. Works with PostgreSQL and MySQL.https://ddltodata.comWould love feedback from anyone who deals with test data or staging environments. What's missing?

Enrichment

Theme
database infrastructure and developer tools
Vertical
Horizontal
Function
Dev tools
Audience
Developer
AI stance
Not AI
Project type
Hobby / open-source project
Normalized one-liner
generate realistic test data from sql schemas
Manually corrected
False

Could you build this?

Yes Parsing SQL DDL schemas and generating realistic relational mock data by passing schemas to an LLM or using Faker-like generation rules is a standard vibe-codeable project.

Discussion

20 comments analyzed.

Competitors mentioned: Seedfast - CLI tool with CI/CD pipeline integration, Postgre-data-generator - LLM-based open source tool, Tabulify - free data generation tool, fakemydb - CSV/SQL insert statement generator, Snaplet/seed/copycat/snapshot - Supabase community tools

Concerns raised: Web-based UI requires copying schemas every time vs CLI-native approach, Service call costs not free compared to local tools, Faker generates overly random data, doesn't match production distributions, Test data doesn't accurately represent production cardinality and long-tail distributions, SaaS data generation business model viability questioned by users

Feature requests: CLI tool for local extraction and integration with CI/CD pipelines, SQL Server support beyond Postgres, Control over data distribution to match production patterns (skewed/long-tail), Extract statistical profiles from production without touching actual database, Realistic cardinality and one-to-many relationship ratios in generated data

Competitors

Other products that read as similar to this one — 107 launches clear the similarity bar, closest 8 shown.

Attention rank: #17 of 108 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 69 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a dev tools tool for Sales yet.