Keep large tool output out of LLM context: 3x accuracy 95% fewer tokens
Details
- External ID
- 47261565
- Source
- HN
- Company
- —
- Product
- Keep large tool output out of LLM context: 3x accuracy 95% fewer tokens
- Website domain
- github.com
- Launched
- March 5, 2026
- Cohort
- —
- Upvotes
- 10
- Upvotes percentile
- 0.5781057810578106
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
LLM agents often place raw JSON tool outputs directly in the prompt. After a few tool calls, earlier results get compacted or truncated and answers become incorrect or inconsistent.I built Sift, a drop-in MCP gateway that stores tool outputs as local artifacts (filesystem blobs indexed in SQLite) and returns an `artifact_id` plus compact schema hints when responses are large or paginated.Instead of reasoning over full JSON in the prompt, the model runs a small Python query: def run(data, schema, params): return max(data, key=lambda x: x["magnitude"])["place"] Query code runs in a constrained subprocess (AST/import guards + timeout/memory caps). Only the computed result is returned to the model.Benchmark (Claude Sonnet 4.6, 103 questions across 12 datasets):- Baseline (raw JSON in prompt): 34/103 (33%), 10.7M input tokens- Sift (artifact + code query): 102/103 (99%), 489K input tokensOpen benchmark + MIT code: https://github.com/lourencomaciel/sift-gatewayInstall: pipx install sift-gateway sift-gateway init --from claude Works with Claude Code, Cursor, Windsurf, Zed, and VS Code. Existing MCP servers and tools require no changes.
Enrichment
- Theme
- niche developer utilities and toolchains
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI feature
- Project type
- —
- Normalized one-liner
- reduce llm context with tool output management
- Manually corrected
- False
Could you build this?
Yes It is a lightweight MCP proxy server that intercepts JSON payloads, writes them to local disk/SQLite, and returns reference IDs, which is well-suited for AI-assisted development.
Discussion
1 comment analyzed.
Concerns raised: Context compaction issues with tool-heavy agents
Competitors
Other products that read as similar to this one — 59 launches clear the similarity bar, closest 8 shown.
Attention rank: #28 of 60 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 124 days after the earliest competitor.
- Calling tools w/ Python improves LLM perf. vs MCP (77.1% on BrowseComp) · hn · 2025-12-16 · 10 upvotes · similarity 0.46
- A new benchmark for testing LLMs for deterministic outputs · hn · 2026-04-29 · 60 upvotes · similarity 0.43
- Mcp2cli · hn · 2026-03-09 · 146 upvotes · similarity 0.43
- Rust-split · hn · 2026-09-12 · 5 upvotes · similarity 0.40
- Reducing LLM input tokens by 70% · hn · 2026-05-12 · 56 upvotes · similarity 0.38
- Model Tools Protocol (MTP) · hn · 2026-02-10 · 9 upvotes · similarity 0.38
- Context Gateway · hn · 2026-03-13 · 97 upvotes · similarity 0.37
- Prompt-refiner · hn · 2025-12-17 · 7 upvotes · similarity 0.37
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.