Echo
Fable-level results at 1/3 the cost using open-weight models
Details
- External ID
- 49026810
- Source
- HN
- Company
- —
- Product
- —
- Website domain
- —
- Launched
- July 23, 2026
- Cohort
- —
- Upvotes
- 484
- Upvotes percentile
- 0.992831541218638
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
I’ve been building Echo (https://echo.tracerml.ai/), an experiment in making one AI system out of a pool of open-weight models rather than choosing a single model and using it for every task.It started with a simple experiment. I took a group of models, including GLM-5.2, Kimi K2.7 and others, and ran them on the same evaluations. Then I measured what would happen if, for each problem, you somehow knew in advance which models would be useful and how their outputs should be combined.That hypothetical system performed substantially better than any individual model in the pool. Of course, it is not something you can actually deploy because it relies on knowing which decisions were good after seeing the result. Echo is my attempt to recover some of that advantage without having that information in advance.For each request, Echo decides how much computation to allocate, which models should participate, and how their work should be combined. Some prompts may only need a relatively small amount of inference, while others benefit from multiple models working on different parts of the problem.One thing that surprised me while building it was how complementary the models are. A model that is clearly weaker overall can still be extremely useful on particular problems or as part of a combination.On my first evaluation mix, Echo consistently performed better than the best individual model in its pool. It also reached roughly the same aggregate result as Fable, which I used as one of the stronger comparison systems, at around one third of the inference cost.There are still some cases where Echo makes the wrong allocation or combination decision. I’m currently spending a lot of time understanding those failures, as well as testing whether the same approach holds up on coding and agentic tasks where measuring the quality of each decision becomes much harder.I built a chat interface (echo.tracerml.ai) and an OpenAI-compatible API (https://echo.tracerml.ai/docs/api) so the system can be tested outside the evaluation setup.Here is a short/high level video on how it works: https://www.youtube.com/watch?v=lJFJSvOdXhgI wrote up the evaluation methodology, individual model results, costs and current limitations here: https://echo.tracerml.ai/evalI would love for you to try it! Especially if you hit any weird failure cases or places where the allocation looks unintuitive.
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- text-to-speech results using open-weight models
- Manually corrected
- False
Could you build this?
Partial Creating a model routing or ensemble wrapper across several open-source LLMs is doable, but beating state-of-the-art models at 1/3 cost requires non-trivial routing algorithms, benchmarking, and multi-model inference orchestration.
What it would actually take: The system requires a distributed inference pipeline (vLLM, SGLang, or TensorRT-LLM) coupled with an intelligent router (e.g., classifier or reward model) that determines optimal model allocation per query difficulty. The hard parts are training a reliable routing classifier, maintaining high-throughput multi-model GPU clusters, and speculative or ensemble verification routines to achieve target benchmark performance. This requires ML infra engineering and access to GPU server clusters.
Discussion
20 comments analyzed.
Competitors mentioned: Claude, Qwen
Concerns raised: Qwen's output can anchor solution space to bad approaches when passed to Claude, Poor performance on UI-related tasks compared to Claude, Token efficiency unclear - sometimes costs more tokens than having stronger model do full implementation, Ensemble approach effectiveness depends on model correlation and task granularity
Feature requests: Better UI handling capabilities, Clearer guidance on when to use ensemble vs single strong model
Competitors
Other products that read as similar to this one — 76 launches clear the similarity bar, closest 8 shown.
Attention rank: #2 of 77 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 254 days after the earliest competitor.
- Tracer: Fable-level AI at 1/3 the cost using open-weight models · yc · 2026-07-31 · 18 upvotes · similarity 0.45
- Optimize and serve models with Fable quality at half the cost · hn · 2026-07-26 · 71 upvotes · similarity 0.43
- Open-source simulation testing infra for voice agents · hn · 2026-09-10 · 16 upvotes · similarity 0.43
- Humor Arena · hn · 2026-08-05 · 10 upvotes · similarity 0.42
- Run open-weight OCR, VLM and vision models behind one API · hn · 2026-09-04 · 5 upvotes · similarity 0.41
- Sup AI, a confidence-weighted ensemble (52.15% on Humanity's Last Exam) · hn · 2026-03-26 · 26 upvotes · similarity 0.40
- Sparrow-2 · hn · 2026-09-08 · 11 upvotes · similarity 0.38
- Frugon · hn · 2026-07-07 · 67 upvotes · similarity 0.38
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.