Dux, distributed DuckDB-backed dataframes on the Beam
Details
- External ID
- 47594412
- Source
- HN
- Company
- —
- Product
- Dux, distributed DuckDB-backed dataframes on the Beam
- Website domain
- github.com
- Launched
- March 31, 2026
- Cohort
- —
- Upvotes
- 7
- Upvotes percentile
- 0.4108241082410824
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Hey all! I wrote Explorer[1] a good few years ago now with the dream of fast dataframes with a dplyr-like API in a really powerful, ergonomic language (Elixir). It's proved pretty successful. Explorer is used in production at my company, and it's my go-to for quick data analysis.But maintaining it became a true albatross. Polars is an amazing project, but the development process is fast and a lot is very focused on the Python lib. We found that trying to maintain Explorer against Polars was a maintenance nightmare and eventually hit points where we had to give up features and found it extremely difficult to update to the latest.We also tried distributing Explorer and only got so far. A reasonable alternative to Spark was always what I wanted, and I could (tantalisingly, frustratingly) see the pieces there in dataframes and the BEAM, but couldn't make it happen.We also always knew that the right direction was to be 'lazy by default', accumulating ops and only executing when the dataframe needs to be realised. But this was very difficult with Polars's Series API and eager/lazy split.Enter DuckDB. A few weeks ago, I made a duckdb backend for Explorer. But in doing so I saw that DuckDB would allow us to realise the lazy-by-default and distributed vision. So I went for it.And here we are. Dux as in ducks as in multiple ducks. Plus an 'x' in the name because, you know, it's Elixir.It's faster than Explorer on a single node. It has a simple, dataframe only API. It distributes arbitrarily on Erlang clusters on the BEAM. Startup is faster than Spark and for many use cases it's simpler and faster. DuckDB functions are all transparently available, as are custom SQL macros. We have a full graph API, as in GraphX/NetworkX. You can install and use any duckdb extensions, including in distribution. And on the maintenance side, it doesn't use a NIF (it depends on the ADBC library[2] and a DuckDB driver) -- the API is primarily about compiling to SQL.DuckDB is incredible for OLAP on out of memory data. Distribution enables fast exploration of in-memory data and real-time applications. The BEAM gives us battle-hardened distribution almost for free.Give it a shot! I'd love feedback and of course PRs are welcome. Oh, I also made a webpage for it[3].[1] https://github.com/elixir-explorer/explorer[2] https://github.com/livebook-dev/adbc[3] https://dux.now
Enrichment
- Theme
- database infrastructure and developer tools
- Vertical
- Horizontal
- Function
- Data infrastructure
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Commercial product
- Normalized one-liner
- distributed duckdb-backed dataframes
- Manually corrected
- False
Could you build this?
No Writing a distributed dataframe framework in Elixir over DuckDB requires deep systems engineering, distributed query planning, and native C/Rust interop via NIFs.
What it would actually take: The architecture requires Erlang/Elixir BEAM actor systems coordinating native DuckDB C++ instances across nodes, serialization protocols (like Apache Arrow), and distributed query pushdown semantics. This demands specialized domain knowledge in database internals, systems programming, and distributed systems consensus.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 104 launches clear the similarity bar, closest 8 shown.
Attention rank: #61 of 105 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 153 days after the earliest competitor.
- DataKit, your all in browser data studio is open source now · hn · 2025-12-08 · 6 upvotes · similarity 0.50
- Connect DuckDB to any database that has an ADBC driver · hn · 2026-07-08 · 7 upvotes · similarity 0.49
- Data Studio · hn · 2026-02-17 · 28 upvotes · similarity 0.47
- repere · hn · 2026-01-20 · 5 upvotes · similarity 0.47
- Dbxlite · hn · 2025-12-12 · 5 upvotes · similarity 0.46
- Ducklang: Achieving 100x more requests per second than NextJS · hn · 2026-01-01 · 9 upvotes · similarity 0.45
- GEDB · hn · 2026-02-16 · 9 upvotes · similarity 0.43
- Turbolite · hn · 2026-03-26 · 185 upvotes · similarity 0.43
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a data infrastructure tool for Media & entertainment yet.