Sup AI, a confidence-weighted ensemble (52.15% on Humanity's Last Exam)
Details
- External ID
- 47531922
- Source
- HN
- Company
- —
- Product
- Sup AI, a confidence-weighted ensemble (52.15% on Humanity's Last Exam)
- Website domain
- sup.ai
- Launched
- March 26, 2026
- Cohort
- —
- Upvotes
- 26
- Upvotes percentile
- 0.7755227552275523
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Hi HN. I'm Ken, a 20-year-old Stanford CS student. I built Sup AI.I started working on this because no single AI model is right all the time, but their errors don’t strongly correlate. In other words, models often make unique mistakes relative to other models. So I run multiple models in parallel and synthesize the outputs by weighting segments based on confidence. Low entropy in the output token probability distributions correlates with accuracy. High entropy is often where hallucinations begin.My dad Scott (AI Research Scientist at TRI) is my research partner on this. He sends me papers at all hours, we argue about whether they actually apply and what modifications make sense, and then I build and test things. The entropy-weighting approach came out of one of those conversations.In our eval on Humanity's Last Exam, Sup scored 52.15%. The best individual model in the same evaluation run got 44.74%. The relative gap is statistically significant (p < 0.001).Methodology, eval code, data, and raw results:- https://sup.ai/research/hle-white-paper-jan-9-2026- https://github.com/supaihq/hleLimitations:- We evaluated 1,369 of the 2,500 HLE questions (details in the above links)- Not all APIs expose token logprobs; we use several methods to estimate confidence when they don'tWe tried offering free access and it got abused so badly it nearly killed us. Right now the sustainable option is a $5 starter credit with card verification (no auto-charge). If you don't want to sign up, drop a prompt in the comments and I'll run it myself and post the result.Try it at https://sup.ai. My dad Scott (@scottmu) is in the thread too. Would love blunt feedback, especially where this really works for you and where it falls short.Here's a short demo video: https://www.youtube.com/watch?v=DRcns0rRhsg
Enrichment
- Theme
- specialized AI models and agent reasoning tools
- Vertical
- Horizontal
- Function
- Agent / copilot
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- confidence-weighted ai ensemble
- Manually corrected
- False
Could you build this?
Partial Querying multiple LLM APIs in parallel and writing a weighted majority voting script is simple, but achieving a state-of-the-art score (52%+) on Humanity's Last Exam requires sophisticated calibration, learned error-correlation matrices, and non-trivial confidence estimation algorithms.
What it would actually take: The architecture involves an orchestration pipeline (e.g., Python/FastAPI) that dispatches queries to a diverse suite of frontier models (Claude, GPT-4, Gemini, reasoning models) and collects output logits/confidence scores. The hard part is training an offline meta-model or Bayesian aggregator that evaluates token-level confidence, models the joint error distribution among LLMs across distinct subjects, and performs optimal consensus decoding. It requires machine learning research expertise in ensemble methods, uncertainty estimation, and evaluation benchmarks.
Discussion
20 comments analyzed.
Competitors mentioned: ChatGPT, Gemini, Claude, Grok (individual model usage), Basic majority voting ensembles
Concerns raised: 7% HLE gain doesn't justify ensemble cost vs single model, Ensemble accuracy depends on independence assumption between responses, Lack of ablation studies on ensemble mechanisms, Limited benchmark data beyond HLE; need coding/domain-specific results, Hallucination detection still imperfect despite entropy weighting
Feature requests: User-configurable accuracy/speed tradeoff levers, Dynamic timeout based on confidence levels, Results across more benchmarks and domains, Ablation study comparing different ensemble voting mechanisms
Competitors
Other products that read as similar to this one — 280 launches clear the similarity bar, closest 8 shown.
Attention rank: #64 of 281 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 146 days after the earliest competitor.
- Slop or not · hn · 2026-03-12 · 19 upvotes · similarity 0.47
- Token Economics Calculator for AI inference hardware · hn · 2025-11-19 · 13 upvotes · similarity 0.46
- I built a 2-min quiz that shows you how bad you are at estimating · hn · 2026-04-06 · 20 upvotes · similarity 0.44
- Try Archetype 360 · hn · 2026-03-02 · 10 upvotes · similarity 0.44
- jj-benchmark · hn · 2026-03-12 · 5 upvotes · similarity 0.43
- FutureSearch, AI forecasting you can verify · hn · 2026-08-03 · 11 upvotes · similarity 0.43
- Benchmark your team's AI coding security posture · hn · 2025-11-05 · 5 upvotes · similarity 0.43
- LLMadness · hn · 2026-03-19 · 5 upvotes · similarity 0.43
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a agent / copilot tool for Agriculture yet.