Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Librarian

Cut token costs by up to 85% for LangGraph and OpenClaw

Details

External ID
47169742
Source
HN
Company
—
Product
Librarian
Website domain
uselibrarian.dev
Launched
Feb. 26, 2026
Cohort
—
Upvotes
8
Upvotes percentile
0.4393530997304582
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Hi HN,I'm building Librarian (https://uselibrarian.dev/), an open-source (MIT) context management tool that stops AI agents from burning tokens by blindly re-reading their entire conversation history on every turn.The Problem: If you're building agentic loops in frameworks like LangGraph or OpenClaw, you hit two walls fast:Financial Cost: Token usage scales quadratically over long conversations. Passing the whole history every time gets incredibly expensive.Context Rot: As the context window fills up, the LLM suffers from the "Lost in the Middle" effect. Response latency spikes, and reasoning accuracy drops.The standard workaround is vector search (RAG) over past messages, but that completely loses temporal logic and conversational dependencies.How Librarian Fixes This: We replaced brute-force context windowing with a lightweight reasoning pipeline:Index: After a message, a smaller model asynchronously creates a compressed summary (~100 tokens), building an index of the conversation.Select: When a new prompt arrives, Librarian reads the summary index and reasons about which specific historical messages are actually relevant to the current turn.Hydrate: It fetches only those selected messages and passes them to the responder.The Results: Instead of passing 2,000+ tokens of noise, you pass a highly curated context of ~800 tokens. In our 50-turn benchmarks, this reduces token costs by up to 85% while actually increasing answer accuracy (82% vs 78% for brute-force) because the distracting noise is removed. It currently works as a drop-in integration for LangGraph and OpenClaw.I'd love for you to check out the benchmark suite, try the integrations, and tear the methodology apart. I'll be hanging out in the comments to answer questions, debug, or hear why this approach is terrible. Thanks!

Enrichment

Theme
browser automation and scraping for AI
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI feature
Project type
Commercial product
Normalized one-liner
reduce token costs for ai agents
Manually corrected
False

Could you build this?

Yes Librarian is an agent context management pipeline that summarizes previous turns asynchronously and selectively hydrates context, which is standard orchestration logic with LLM calls.

Discussion

7 comments analyzed.

Competitors mentioned: Claude prompt caching, Standard append-only chat history, Llama 3, Mistral

Concerns raised: Disables prompt caching mechanism for conversation history, Caching still costs money despite optimization claims, Latency penalty for shorter conversations due to LLM inference step, Context rot issue not fully solved, Prompt caching only helps for 'hot' conversations, not long-term daily usage

Feature requests: Optimize selection latency below 1 second for shorter conversations, Fine-tuned lightweight model for context routing task, Handle conversation cool-off periods with optimized system message caching

Competitors

Other products that read as similar to this one — 125 launches clear the similarity bar, closest 8 shown.

Attention rank: #72 of 126 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 114 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.