Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Codex context bloat? 87% avg reduction on SWE-bench Verified traces

Details

External ID
47896087
Source
HN
Company
—
Product
Codex context bloat? 87% avg reduction on SWE-bench Verified traces
Website domain
npmjs.com
Launched
April 24, 2026
Cohort
—
Upvotes
10
Upvotes percentile
0.5919023136246787
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

If you had to build a context window manager in 24h, would you stick to the existing model or come up with something better?Here's what I did:1. Built a proxy that intercepts Codex's calls to OpenAI and rewrites them on the fly.2. Replayed 3,807 rounds of SWE-bench Verified traces through it: avg prompt 44k → 6k tokens (-87%).3. Posted it to HN to get the next reduction applied to my confidence interval — starting with the inevitable "How about accuracy?"npx -y pando-proxy · github.com/human-software-us/pando-proxy

Enrichment

Theme
Codex monitoring and optimization tools
Vertical
Horizontal
Function
Dev tools
Audience
Developer
AI stance
AI feature
Project type
Commercial product
Normalized one-liner
context reduction for ai coding
Manually corrected
False

Could you build this?

Partial The proxy server itself is simple, but achieving an 87% context reduction without degrading code benchmark performance requires sophisticated prompt deduplication, chunking heuristics, and active memory management.

What it would actually take: An HTTP/reverse-proxy built in Node.js or Go that intercepts OpenAI Responses/Completions payloads, tokenizes request bodies, tracks conversation DAGs, and applies stateful diffing and retrieval functions. The hardest part is the semantic chunking and active memory lifecycle algorithm that decides what context to drop or recall without hallucination or broken dependencies. Requires deep expertise in LLM context manipulation, SWE-bench evaluation methodology, and stateful proxy infrastructure.

Discussion

2 comments analyzed.

Concerns raised: Proxy overhead makes short sessions more expensive than direct API calls, Two extra model calls per round (chunker + working_memory_update) add hidden costs, Unclear total cost comparison when factoring in multiple model invocations

Competitors

Other products that read as similar to this one — 54 launches clear the similarity bar, closest 8 shown.

Attention rank: #27 of 55 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 170 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a dev tools tool for Sales yet.