Optimize and serve models with Fable quality at half the cost
Details
- External ID
- 49063454
- Source
- HN
- Company
- —
- Product
- Optimize and serve models with Fable quality at half the cost
- Website domain
- github.com
- Launched
- July 26, 2026
- Cohort
- —
- Upvotes
- 71
- Upvotes percentile
- 0.8823178016726404
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Hi HN, we built world-model-optimizer, an open source tool to continually improve a specialized model for an agent.It does this by simulating production tool responses through text world modeling (similar to QwenAgentWorld, summary here https://x.com/silennai/status/2073887455884058814).We can then use this to train a router for frontier, OS, and local models (use defaults or pick which ones to optimize against).wmo ingests agent traces, builds the simulation, embeds the traces, runs different models you choose against the simulation scenarios, and then uses a KNN for model selection (similar to https://arxiv.org/abs/2505.19797).- Cache aware: cache is taken into account for the effective price in routing.- Confidence gated: we don't deviate from the best fit model when paired evidence over retrieved neighbors is below 0.5 standard errors or on queries unlike anything in the fit set.- Optimize for cost or quality: train a balanced, cost max, or quality max router.Usage`wmo build` creates the simulation (or add your own benchmark)`wmo optimize` tunes the router`wmo serve` starts the server and can run everything fully locally. The simulation and router can update over time as more agent traces are gathered and new models are added.Router results vs Fable- RouterBench: -66.5% cost, -1.7% performance, -24.7% latency p50. 77.5% of traffic to Sonnet 5, 16.1% Fable 5.- TauBench: -44.5% cost, +6.3% performance, -20% latency. 83% to Opus 5, 17% to Kimi-K2.6 (over K3).- Terminal Bench 2: -64% cost, +8% performance, -50.6% latency. Sonnet 5 is fully along the pareto front. Training a specialized router per task isn't cheap. In sparse data regimes the value can be "here's the best model".We're working on sample effiient continual learning for agent specific models at experientiallabs.ai"
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- B2B
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- optimize and serve ai models at lower cost
- Manually corrected
- False
Could you build this?
No Building an LLM world-model optimizer that simulates dynamic production tool environments to fine-tune and distill agent models involves cutting-edge AI research and distributed model training.
What it would actually take: The architecture requires PyTorch, vLLM/DeepSpeed, synthetic data generation pipelines, and reinforcement learning / rejection sampling fine-tuning frameworks (DPO/PPO). The hard part is training an accurate text world-model that simulates complex multi-step tool call responses realistically and using that environment to distill or align agent models without catastrophic drift. This demands machine learning research scientists with deep experience in agent trajectories, synthetic data curation, and GPU cluster training.
Discussion
20 comments analyzed.
Competitors mentioned: Frontier models (GPT-4, Claude, etc.) for task routing, Other distillation approaches
Concerns raised: Simulator exploitation/reward hacking in isolated environments, High training costs (thousands of dollars), Unclear cold-start performance with limited traces, Reconstruction fidelity vs. downstream performance disagreement, Results not yet solidified/mature enough
Feature requests: Pre-trained routers in app for faster onboarding, Cost reduction for training and inference, Better documentation on traces needed for reliable routing
Competitors
Other products that read as similar to this one — 67 launches clear the similarity bar, closest 8 shown.
Attention rank: #10 of 68 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 270 days after the earliest competitor.
- Smart model routing directly in Claude, Codex and Cursor · hn · 2026-06-26 · 216 upvotes · similarity 0.45
- Echo · hn · 2026-07-23 · 484 upvotes · similarity 0.43
- InstinctFlash · hn · 2026-09-22 · 27 upvotes · similarity 0.43
- We built open OpenRouter that turns usage into a better model · hn · 2026-08-27 · 222 upvotes · similarity 0.41
- Zenith: sota harness for normal models to beat Fable on FrontierSWE · hn · 2026-06-29 · 8 upvotes · similarity 0.39
- Telem · hn · 2026-08-27 · 8 upvotes · similarity 0.39
- Forge · hn · 2026-05-19 · 687 upvotes · similarity 0.39
- Throttle · ph · 2026-09-25 · 1 upvotes · similarity 0.37
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.