I built a database for AI agents
Details
- External ID
- 47678048
- Source
- HN
- Company
- —
- Product
- database
- Website domain
- github.com
- Launched
- April 7, 2026
- Cohort
- —
- Upvotes
- 12
- Upvotes percentile
- 0.6446015424164524
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Hey HN,I just spent the last few weeks building a database for agents.Over the last year I built PostHog AI, the company's business analyst agent, where we experimented on giving raw SQL access to PostHog databases vs. exposing tools/MCPs. Needless to say, SQL wins.I left PostHog 3 weeks ago to work on side-projects. I wanted to experiment more with SQL+agents.I built an MVP exposing business data through DuckDB + annotated schemas, and ran a benchmark with 11 LLMs (from Kimi 2.5 to Claude Opus 4.6) answering business questions with either 1) per-source MCP access (e.g. one Stripe MCP, one Hubspot MCP) or 2) my annotated SQL layer.My solution consistently reached 2-3x accuracy (correct vs. incorrect answers), using 16-22x less tokens per correct answer, and being 2-3x faster. Benchmark in the repo!The insight is that tool calls/MCPs/raw APIs force the agent to join information in-context. SQL does that natively.What I have today: - 101 connectors (SaaS APIs, databases, file storages) sync to Parquet via dlt, locally or in your S3/GCS/Azure bucket - DuckDB is the query engine — cross-source JOINs across sources work natively, plus guardrails for safe mutations / reverse ETL - After each sync a Claude agent annotates the schema: table descriptions, column docs, PII flags, relationship mapsIt works with all major agent frameworks (LangChain, CrewAI, LlamaIndex, Pydantic AI, Mastra), and local agents like Claude Code, Cursor, Codex and OpenClaw.I love dinosaurs and the domain was available, so it's called Dinobase.It's not bug free and I'm here to ask for feedback or major holes in the project I can't see, because the results seem almost too good. Thanks!
Enrichment
- Theme
- database infrastructure and developer tools
- Vertical
- Horizontal
- Function
- Data infrastructure
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- database for ai agents
- Manually corrected
- False
Could you build this?
Partial Wrapping DuckDB or SQLite with an LLM-friendly API is straightforward, but creating a production-grade specialized database with agentic sandboxing, schema reflection, and transactional safety requires non-trivial database systems work.
What it would actually take: The system requires an embedded OLAP/OLTP engine (like DuckDB, SQLite, or custom storage engine) exposed via a strict sandboxed protocol (e.g., eBPF/Wasm or isolated microVMs) with automated rollbacks and fine-grained permissions. The hard engineering problems include sub-millisecond query execution, strict resource governance to prevent runaway agent queries, and dynamic automated indexing based on agent query patterns. This requires database internals expertise and systems programming.
Discussion
10 comments analyzed.
Competitors mentioned: Spider 2.0 (text-to-SQL benchmark), MCPs (Model Context Protocols) for individual data sources, API-based approaches requiring pagination
Concerns raised: Schema drift handling - how breaking changes from SaaS vendors are detected and marked, Stale columns not being dropped after schema changes, Edge cases with SaaS platforms like Salesforce custom objects, Multi-hop query performance (hallucinations and partial answers)
Feature requests: Automated schema migrations to reconcile and drop old columns, Explicit marking of breaking changes when schema drift occurs, Benchmark breakdown by question type (aggregations vs multi-hop vs lookups)
Competitors
Other products that read as similar to this one — 338 launches clear the similarity bar, closest 8 shown.
Attention rank: #148 of 339 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 160 days after the earliest competitor.
- DeepSQL · hn · 2026-07-20 · 52 upvotes · similarity 0.56
- SqlKit · ph · 2026-09-15 · 1 upvotes · similarity 0.54
- Declarative open-source framework for MCPs with search and execute · hn · 2026-02-24 · 11 upvotes · similarity 0.53
- Altimate Code · hn · 2026-03-19 · 20 upvotes · similarity 0.50
- Risk Analysis Database of Every MCP Server · hn · 2026-02-05 · 22 upvotes · similarity 0.50
- Pglens · hn · 2026-03-29 · 13 upvotes · similarity 0.48
- API to MCP · ph · 2026-06-19 · 195 upvotes · similarity 0.48
- MCPfinder · hn · 2026-04-20 · 9 upvotes · similarity 0.47
Other launches for this product
- A memory database that forgets, consolidates, and detects contradiction
- Open database of 2k IP camera specs (JSON/CSV, CC0)
- KV and wide-column database with CDN-scale replication
- Doberman: The AI watchdog that stops Claude from deleting your database
- database
- Connect DuckDB to any database that has an ADBC driver
- Open database of link metadata for large-scale analysis
- Daily-updated database of malicious browser extensions
- Open-source Agent in Rust that can't delete your database
Same idea, different domain
Nobody's really built a data infrastructure tool for Media & entertainment yet.