Rocky
Rust SQL engine with branches, replay, column lineage
Details
- External ID
- 47935246
- Source
- HN
- Company
- —
- Product
- Rocky and Caveman Speak in Claurst CLI Save Big Token Amaze Amaze Amaze
- Website domain
- github.com
- Launched
- April 28, 2026
- Cohort
- —
- Upvotes
- 122
- Upvotes percentile
- 0.9138817480719794
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Hi HN, I'm Hugo. I've been building Rocky over the past month, shipping fast in the open. The binary is on GitHub Releases, `dagster-rocky` on PyPI, and the VS Code extension on the Marketplace. I held off on a broader announcement until the trust-system surface was coherent enough to talk about as one thing. The governance waveplan — column classification, per-env masking, 8-field audit trail on every run, `rocky compliance` rollup, role-graph reconciliation, retention policies — landed end-to-end last week in engine-v1.16.0 and rounded out in v1.17.4 (tagged 2026-04-26). That's the milestone I'd been waiting for.The pitch: keep Databricks or Snowflake. Bring Rocky for the DAG. Rocky is a Rust-based control plane for warehouse pipelines. Storage and compute stay with your warehouse. Rocky owns the graph — dependencies, compile-time types, drift, incremental logic, cost, lineage, governance. The things your current stack can't give you because it doesn't own the DAG.A few things I think are interesting:- Branches + replay. `rocky branch create stg` gives you a logical copy of a pipeline's tables (schema-prefix today; native Delta SHALLOW CLONE and Snowflake zero-copy are next). `rocky replay <run_id>` reconstructs which SQL ran against which inputs. Git-grade workflow on a warehouse.- Column-level lineage from the compiler, not a post-hoc graph crawl. The type checker traces columns through joins, CTEs, and windows. VS Code surfaces it inline via LSP.- Governance as a first-class surface. Column classification tags plus per-env masking policies, applied to the warehouse via Unity Catalog (Databricks) or masking policies (Snowflake). 8-field audit trail on every run. `rocky compliance` rollup that CI can gate on. Role-graph reconciliation via SCIM + per-catalog GRANT. Retention policies with a warehouse-side drift probe.- Cost attribution. Every run produces per-model cost (bytes, duration). `[budget]` blocks in `rocky.toml`; breaches fire a `budget_breach` hook event.- Compile-time portability + blast radius. Dialect-divergence lint across Databricks / Snowflake / BigQuery / DuckDB (12 constructs). `SELECT *` downstream-impact lint.- Schema-grounded AI. Generated SQL goes through the compiler — AI suggestions type-check before they can land.What Rocky isn't:- Not a warehouse — it's the control plane on top.- Not a Fivetran replacement. `rocky load` handles files (CSV/Parquet/JSONL); for SaaS sources use Fivetran, Airbyte, or warehouse-native CDC.- Not dbt Cloud — no hosted UI, no managed scheduler. First-class Dagster integration if you need orchestration.Adapters: Databricks (GA), Snowflake (Beta), BigQuery (Beta), DuckDB (local dev / playground). Apache 2.0.I'd love feedback on the trust-system framing, the governance surface (particularly classification-to-masking resolution in `rocky compile` and the `rocky compliance` CI gate), the branches/replay design, the cost-attribution primitives, or anything else that catches your eye. Happy to go deep in the thread.
Enrichment
- Theme
- database infrastructure and developer tools
- Vertical
- Horizontal
- Function
- Data infrastructure
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Commercial product
- Normalized one-liner
- sql engine with branches and replay
- Manually corrected
- False
Could you build this?
No Creating a transactional SQL database engine written in Rust featuring custom branchable storage, deterministic query replay, and column-level lineage requires deep database internals, compiler theory, and systems programming expertise.
What it would actually take: The architecture requires a custom SQL parser/analyzer, a cost-based query optimizer, an execution engine (or DataFusion extension), and an append-only, copy-on-write storage engine capable of snapshot branching and deterministic transaction replaying. Building accurate column-level lineage also demands static code analysis across arbitrary nested SQL subqueries and window functions.
Discussion
20 comments analyzed.
Competitors mentioned: dbt (data build tool), Dagster, Databricks, Iceberg
Concerns raised: Name 'Rocky' conflicts with existing popular Linux distribution, Missing proper attribution to prior work on branches and lineage, README doesn't clearly connect dots on benefits for data engineers, Limited to SQL-first approach, no Python support mentioned
Feature requests: Lineage diff command (compare lineage between branches), Nested branches support, Python language support beyond SQL, Cross-system merge semantics for branch operations
Competitors
Other products that read as similar to this one — 113 launches clear the similarity bar, closest 8 shown.
Attention rank: #15 of 114 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 181 days after the earliest competitor.
- Altimate Code · hn · 2026-03-19 · 20 upvotes · similarity 0.43
- Turbolite · hn · 2026-03-26 · 185 upvotes · similarity 0.42
- Pgrust, Postgres in Rust (passing 100% of Postgres regression tests) · hn · 2026-06-25 · 6 upvotes · similarity 0.42
- HoundDog.ai · hn · 2026-02-02 · 16 upvotes · similarity 0.41
- Tangent · hn · 2025-11-20 · 28 upvotes · similarity 0.40
- Bloodhound · hn · 2025-12-09 · 5 upvotes · similarity 0.39
- I implemented a neural network in SQL · hn · 2026-07-13 · 121 upvotes · similarity 0.39
- Typol · hn · 2026-06-07 · 5 upvotes · similarity 0.39
Other launches for this product
Same idea, different domain
Nobody's really built a data infrastructure tool for Media & entertainment yet.