Ktx
Open-source executable context layer for data agents
Details
- External ID
- 48309986
- Source
- HN
- Company
- —
- Product
- Kctx
- Website domain
- github.com
- Launched
- May 28, 2026
- Cohort
- —
- Upvotes
- 93
- Upvotes percentile
- 0.8990306946688207
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Hi HN, we’re open-sourcing ktx. It’s an executable context layer that makes agents reliable on your data stack.We built it after going through the experience of building production-grade data agents for dozens of companies. If you’ve also tried building them, or simply tried using Claude Code or Codex on your data warehouse, you’ll know that accuracy is the #1 issue. Agents are great at generating valid SQL, but it’s not always correct SQL.To cite a few examples of “agents gone wrong”:- Stale column + hidden business rule: when preparing a board report, a finance analyst asks Claude Code for “ARR by customer segment”, it derives ARR from multiple tables (subscriptions, plans, accounts), then groups by accounts.industry. But CC doesn’t know that this industry column was deprecated a few months prior, or that past board reports excluded paused subscriptions from the ARR calculation- Join fanout: a data analyst at a retailer uses their company’s internal agent to prep a product revenue deck for a QBR. The agent joins orders to order_items, then sums orders.total_amount_cents grouped by order_items.product_id. The SQL runs fine, but each order’s revenue is repeated once per line item, which most people will miss if most orders only have 1 item- Missing attribution logic: a marketing analyst asks Codex “Which campaigns drove the most revenue?” Codex joins marketing_touches to users to orders and groups by utm_campaign. But since each order can have multiple touches before purchase, the same order can be credited to first touch, last touch, every touch, or every campaign the user clicked before buying. If the agent chooses the method that doesn’t match the team’s attribution logic, they’ll make suboptimal decisionsTo solve this at first we gave the agent more context through skills + a wiki-style knowledge base. That gives it some useful extra context but still relies on it writing the SQL without incorrect assumptions.The next solution we explored was implementing a classic semantic layer. That solves the executable part, but they’re such a pain to build and maintain since they were made for legacy BI tools. Plus as a standalone tool, they lack all the useful context from unstructured data sources like internal docs.So we built ktx and split it into 2 parts:1. Business context goes in Markdown wiki pages that are auto-ingested and auto-populated2. Queryable definitions go into YAML files that define tables, row grain, joins, measures, dimensions, filters, and filter groupsThat way, when an agent needs a metric, it asks ktx for a measure, dimensions, filters, and filter groups instead of writing the whole query itself. ktx’s planner chooses the join path, uses grain and relationship metadata, catches issues like join fanout and chasm joins, and compiles the warehouse SQL, while utilizing the extra unstructured knowledge it has access to.ktx is Apache 2.0. It can ingest from most warehouses (BigQuery, Snowflake, Postgres & others), modeling tools (dbt, MetricFlow, LookML), BI tools (Looker, Metabase), doc tools like Notion, and corrections from user interactions.Install manually:npm install -g @kaelio/ktxktx setupOr give this prompt to your agent:Run npx skills add Kaelio/ktx --skill ktx and use ktx skill to install and configure ktxWe’d especially like feedback from people who’ve tried using Claude Code, Codex, or building custom agents on analytics warehouses. Where did they fail? And what did you try to make the answers more reliable?
Enrichment
- Theme
- AI agent frameworks and developer tools
- Vertical
- Horizontal
- Function
- Data infrastructure
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- context layer for data agents
- Manually corrected
- False
Could you build this?
Partial While wrapping data tools and context formatting for LLMs is straightforward, building an executable context layer that reliably interfaces with heterogeneous data warehouses, enforces fine-grained schema security, and coordinates agent tool executions requires deep data engineering knowledge.
What it would actually take: A complete system requires connectors across disparate OLAP and transactional databases (Snowflake, BigQuery, Postgres), an AST-level query validator/sandboxer for safe execution, and deterministic context retrieval primitives. The hardest part is handling dynamic schema discovery, query optimization, and sandboxed rollback mechanisms across diverse dialects. It requires senior data platform and systems engineering experience to build robust database execution layers.
Discussion
20 comments analyzed.
Competitors mentioned: Graphene (code-first BI tool for query->viz with GSQL), GrapheneData, McKinsey semantic layer solutions
Concerns raised: Setup performance slow (~19 minutes for 20 tables), Schema drift - agents run old queries after DB updates, Context grows too large over time, pg_description data not being ingested into wiki
Feature requests: Support for OpenAI API / Codex backend (now released), Auto memory/context compression, Agent self-correction mechanism for query errors, Manual context compression feature
Competitors
Other products that read as similar to this one — 200 launches clear the similarity bar, closest 8 shown.
Attention rank: #22 of 201 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 205 days after the earliest competitor.
- Altimate Code · hn · 2026-03-19 · 20 upvotes · similarity 0.53
- Athenic · hn · 2026-06-10 · 5 upvotes · similarity 0.48
- Data Formulator · hn · 2025-11-11 · 38 upvotes · similarity 0.48
- LangAlpha · hn · 2026-04-14 · 148 upvotes · similarity 0.47
- FactIQ · hn · 2026-07-08 · 7 upvotes · similarity 0.46
- I built a database for AI agents · hn · 2026-04-07 · 12 upvotes · similarity 0.46
- Kilroy · hn · 2026-04-16 · 5 upvotes · similarity 0.46
- SLayer, a semantic layer maintained by your agent · hn · 2026-05-11 · 12 upvotes · similarity 0.44
Other launches for this product
Same idea, different domain
Nobody's really built a data infrastructure tool for Media & entertainment yet.