Token compression CLI to save Codex/Astra costs
Details
- External ID
- 49911910
- Source
- HN
- Company
- —
- Product
- —
- Website domain
- —
- Launched
- Sept. 30, 2026
- Cohort
- —
- Upvotes
- 9
- Upvotes percentile
- 0.5462249614791987
- Tags
- —
- Fetched at
- Oct. 1, 2026, 5:01 p.m.
- Updated at
- Oct. 1, 2026, 5:01 p.m.
Description
Hey HN! Yolanda and Spencer here - wanted to share a token compression tool that we’ve built for ourselves to save 30% costs on codex!After maxing out sub and burning $700/day per person on api, we fine tuned a compression model to trim codex's tool call output to reduce input token + cache. It cut down tokens by 29.6% and now I just leave it on by default in Codex.To avoid messing up w/ cache, we use proxy + fine tuned qwen model trained on preserving agent trajectory to remove tool call results before they go back to the model, leaving kv cache untouched.The cli is free for everyone to use (https://github.com/spenmcke/compress). Just lmk ur feedback and hacks to shave even more costs on astra! If you want to integrate it into your product to offer the best models at low cost, I can set you up with an sdk and api keysPS: It’s built for coding agents, not conversational agents. I optimized it for file retrieval accuracy, trajectory preservation, and quality to get up to 30% cost reduction depending on how context-heavy the task is.On security and privacy side, it's a proxy wrapping your local codex and ZDR so it doesn't retain any queries. It’s on by default in codex and when you don’t want compression, you can use `codex --uncompress` to disable it.Give it a try: code is in https://github.com/spenmcke/compressYou can install the cli using`curl -fsSL https://install.everestagi.com/install.sh | sh && source ~/.config/everest/shell.sh`Love to hear any feedback and learn your hacky ways to save token costs too!
Enrichment
- Theme
- memory systems for ai agents
- Vertical
- Horizontal
- Function
- Dev tools
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- token compression cli for ai coding assistants
- Manually corrected
- False
Could you build this?
Partial Writing a CLI proxy is easy, but fine-tuning a custom compression model to reliably strip tool call tokens without breaking code syntax or downstream context requires ML training expertise and data curation.
What it would actually take: The system requires an MITM proxy or CLI wrapper (Python/Go) that intercepts LLM agent API calls, paired with a specialized lightweight language model (like a fine-tuned LLaMA or Qwen) trained specifically on agent scratchpads and AST diffs to compress outputs while preserving semantic fidelity. The hard part is generating high-quality training pairs and fine-tuning an ultra-fast, low-latency inference endpoint that achieves a net cost and latency reduction.
Discussion
4 comments analyzed.
Concerns raised: How 30% savings are measured
Competitors
Other products that read as similar to this one — 152 launches clear the similarity bar, closest 8 shown.
Attention rank: #72 of 153 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 322 days after the earliest competitor.
- Edgee Codex Compressor V2 · ph · 2026-09-18 · 90 upvotes · similarity 0.59
- Edgee Codex Compressor · ph · 2026-04-12 · 166 upvotes · similarity 0.58
- I nerfed our coding agents on purpose · hn · 2026-06-05 · 27 upvotes · similarity 0.54
- Ctx, save tokens by loading only the relevant tools · hn · 2026-06-16 · 8 upvotes · similarity 0.54
- Librarian · hn · 2026-02-26 · 8 upvotes · similarity 0.53
- Edgee Claude Code Compressor V2 · ph · 2026-07-06 · 179 upvotes · similarity 0.50
- Hydra · hn · 2026-04-21 · 9 upvotes · similarity 0.49
- compress · github · 2026-09-29 · 69 upvotes · similarity 0.48
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a dev tools tool for Sales yet.