An LLM response cache that's aware of dynamic data
Details
- External ID
- 46532755
- Source
- HN
- Company
- —
- Product
- An LLM response cache that's aware of dynamic data
- Website domain
- butter.dev
- Launched
- Jan. 7, 2026
- Cohort
- —
- Upvotes
- 17
- Upvotes percentile
- 0.6324110671936759
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Raymond here from Butter.dev, an LLM response cache built as a chat-completions proxy. Today we're launching a key feature for the platform: the ability to generalize on dynamic, templated inputs.Caching at the HTTP request level has the obvious problem of generalizability. Nearly no request is identical, due to templated variables (like names) and metadata (like timestamps), so exact-match cache lookups rarely hit. We solve this at Butter by using LLMs to detect dynamic content in requests and derive their inter-relationships, allowing the cache entry to be stored as a template + variables + deterministic code. This allows future requests to contain different variable data, yet still serve from cache.We've found this approach greatly improves cache hit rate, and believe it could be useful for agents performing repetitive back-office tasks, computer use, or data transformations where input data is frequently of the same shape.- You can see a demo of learning patterns here: https://www.youtube.com/watch?v=ORDfPnk9rCA- We wrote more about the technical approach here: https://blog.butter.dev/on-automatic-template-induction-for-...- It's free to try out here: https://butter.dev/auth
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- llm response cache for dynamic data
- Manually corrected
- False
Could you build this?
Partial Creating a simple HTTP reverse proxy for LLM caching is trivial, but automatic template and variable induction from arbitrary natural language queries without generating semantic false positives requires non-trivial algorithmic design.
What it would actually take: The system requires a high-performance reverse proxy (e.g., in Rust or Go) paired with an automated grammar induction/tree-sitter AST parser that discovers structural templates and variable slots from unstructured prompt streams. It demands specialized algorithms for generalized pattern matching and low-latency cache-tree indexing to keep overhead under a few milliseconds.
Discussion
1 comment analyzed.
Competitors
Other products that read as similar to this one — 48 launches clear the similarity bar, closest 8 shown.
Attention rank: #17 of 49 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 56 days after the earliest competitor.
- Llmbuffer · hn · 2026-06-10 · 7 upvotes · similarity 0.48
- Cachet · hn · 2026-06-23 · 5 upvotes · similarity 0.46
- Continual Learning with .md · hn · 2026-04-13 · 34 upvotes · similarity 0.40
- Llm.sql · hn · 2026-04-24 · 8 upvotes · similarity 0.40
- An agent that tunes its own cache · hn · 2026-05-08 · 7 upvotes · similarity 0.39
- Askfeather.ai · hn · 2026-02-03 · 7 upvotes · similarity 0.39
- genpark-agentic-cache-semantic-deduplicator-skill · github · 2026-09-14 · 8 upvotes · similarity 0.39
- genpark-agentic-cache-semantic-deduplicator-skill · github · 2026-09-14 · 7 upvotes · similarity 0.39
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.