DataKit, your all in browser data studio is open source now
Details
- External ID
- 46197050
- Source
- HN
- Company
- —
- Product
- DataKit, your all in browser data studio is open source now
- Website domain
- github.com
- Launched
- Dec. 8, 2025
- Cohort
- —
- Upvotes
- 6
- Upvotes percentile
- 0.2652671755725191
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Hey HN! I'm open-sourcing DataKit today.GitHub: https://github.com/datakitpage/datakit Live demo: https://datakit.pageDataKit is a browser-based data analysis platform that processes multi-gigabyte files (CSV, Parquet, JSON, Excel) entirely client-side using DuckDB-WASM. Your data never leaves your browser.What it does: • Process large files (tested up to 20GB) without any server • Full SQL interface powered by DuckDB compiled to WebAssembly • Python notebooks via Pyodide for data science workflows • Connect to remote sources (PostgreSQL, MotherDuck, S3) with optional proxy • AI assistant that only sees column schemas, not actual dataI was done with having to choose between cloud tools and heavy local installations. I wanted something that just works in a browser tab but has real power.It's AGPL licensed with commercial licenses available for enterprises.I've been building this solo as a side project for the past few months. Would love your feedback on: - Performance bottlenecks you encounter - Features you'd need for your workflows - The architecture decisions (all client-side vs hybrid)
Enrichment
- Theme
- database infrastructure and developer tools
- Vertical
- Horizontal
- Function
- Analytics & BI
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- browser-based data analysis tool
- Manually corrected
- False
Could you build this?
Partial While the UI is standard React/web tooling, orchestrating in-browser multi-gigabyte data processing via DuckDB-WASM, web workers, and memory-efficient streaming across various formats requires non-trivial browser systems engineering.
What it would actually take: The stack uses Next.js/React, DuckDB-WASM, Apache Arrow, and Web Workers. The hard part is managing client-side memory constraints and WebAssembly limits when streaming and querying multi-gigabyte Parquet, Excel, and CSV files in the browser without freezing the UI thread or running out of memory (OOM).
Discussion
1 comment analyzed.
Concerns raised: Parquet files not displaying in browser
Competitors
Other products that read as similar to this one — 124 launches clear the similarity bar, closest 8 shown.
Attention rank: #107 of 125 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 40 days after the earliest competitor.
- Data Studio · hn · 2026-02-17 · 28 upvotes · similarity 0.60
- Rawkit · hn · 2026-02-12 · 8 upvotes · similarity 0.53
- Dbxlite · hn · 2025-12-12 · 5 upvotes · similarity 0.52
- SqlKit · ph · 2026-09-15 · 1 upvotes · similarity 0.51
- Dux, distributed DuckDB-backed dataframes on the Beam · hn · 2026-03-31 · 7 upvotes · similarity 0.50
- repere · hn · 2026-01-20 · 5 upvotes · similarity 0.45
- GetDataFlux Analytics · ph · 2026-09-15 · 2 upvotes · similarity 0.45
- VisuaLeaf · hn · 2026-04-29 · 9 upvotes · similarity 0.44
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a analytics & bi tool for Legal yet.