Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Token compression CLI to save Codex/Astra costs

Details

External ID
49911910
Source
HN
Company
—
Product
—
Website domain
—
Launched
Sept. 30, 2026
Cohort
—
Upvotes
9
Upvotes percentile
0.5462249614791987
Tags
—
Fetched at
Oct. 1, 2026, 5:01 p.m.
Updated at
Oct. 1, 2026, 5:01 p.m.

Description

Hey HN! Yolanda and Spencer here - wanted to share a token compression tool that we’ve built for ourselves to save 30% costs on codex!After maxing out sub and burning $700/day per person on api, we fine tuned a compression model to trim codex's tool call output to reduce input token + cache. It cut down tokens by 29.6% and now I just leave it on by default in Codex.To avoid messing up w/ cache, we use proxy + fine tuned qwen model trained on preserving agent trajectory to remove tool call results before they go back to the model, leaving kv cache untouched.The cli is free for everyone to use (https://github.com/spenmcke/compress). Just lmk ur feedback and hacks to shave even more costs on astra! If you want to integrate it into your product to offer the best models at low cost, I can set you up with an sdk and api keysPS: It’s built for coding agents, not conversational agents. I optimized it for file retrieval accuracy, trajectory preservation, and quality to get up to 30% cost reduction depending on how context-heavy the task is.On security and privacy side, it's a proxy wrapping your local codex and ZDR so it doesn't retain any queries. It’s on by default in codex and when you don’t want compression, you can use `codex --uncompress` to disable it.Give it a try: code is in https://github.com/spenmcke/compressYou can install the cli using`curl -fsSL https://install.everestagi.com/install.sh | sh && source ~/.config/everest/shell.sh`Love to hear any feedback and learn your hacky ways to save token costs too!

Enrichment

Theme
memory systems for ai agents
Vertical
Horizontal
Function
Dev tools
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
token compression cli for ai coding assistants
Manually corrected
False

Could you build this?

Partial Writing a CLI proxy is easy, but fine-tuning a custom compression model to reliably strip tool call tokens without breaking code syntax or downstream context requires ML training expertise and data curation.

What it would actually take: The system requires an MITM proxy or CLI wrapper (Python/Go) that intercepts LLM agent API calls, paired with a specialized lightweight language model (like a fine-tuned LLaMA or Qwen) trained specifically on agent scratchpads and AST diffs to compress outputs while preserving semantic fidelity. The hard part is generating high-quality training pairs and fine-tuning an ultra-fast, low-latency inference endpoint that achieves a net cost and latency reduction.

Discussion

4 comments analyzed.

Concerns raised: How 30% savings are measured

Competitors

Other products that read as similar to this one — 152 launches clear the similarity bar, closest 8 shown.

Attention rank: #72 of 153 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 322 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a dev tools tool for Sales yet.