LLMadness
March Madness Model Evals
Details
- External ID
- 47437420
- Source
- HN
- Company
- —
- Product
- LLMadness
- Website domain
- llmadness.com
- Launched
- March 19, 2026
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.1070110701107011
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
I wanted to play around with the non-coding agentic capabilities of the top LLMs so I built a model eval predicting the March Madness bracket.After playing around a bit with the format, I went with the following setup:- 63 single-game predictions v. full one-shot bracket- Maxed out at 10 tool calls per game- Upset-specific instruction in the system prompt- Exponential scoring by round (1, 2, 4, 8, 16, 32)There were some interesting learnings:- Unsurprisingly, most brackets are close to chalk. Very few significant upsets were predicted.- There was a HUGE cost and token disparity with the exact same setup and constraints. Both Claude models spent over $40 to fill in the bracket while MiMo-V2-Flash spent $0.39. I spent a total of $138.69 on all 15 model runs.- There was also a big disparity in speed. Claude Opus 4.6 took almost 2 full days to finish the 2 play-ins and 63 bracket games. Qwen 3.5 Flash took under 10 minutes.- Even when given the tournament year (2026), multiple models pulled in information from previous years. Claude seemed to be the biggest offender, really wanting Cooper Flagg to be on this year's Duke team.This was a really fun way to combine two of my interests and I'm excited to see how the models perform over the coming weeks. You can click into each bracket node to see the full model trace and rationale behind the picks.The stack is Typescript, Next.js, React, and raw CSS. No DB, everything stored in static JSON files. After each game, I update the actual results and re-deploy via GitHub Pages.I wanted to work as fast as possible since the brackets lock today so almost all of the code was AI-generated (shocker).Hope you enjoy checking it out!
Enrichment
- Theme
- indie hacker passion projects
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI feature
- Project type
- Hobby / open-source project
- Normalized one-liner
- llm model evaluation benchmark
- Manually corrected
- False
Could you build this?
Yes It is a Next.js tournament leaderboard comparing LLM benchmark responses for NCAA bracket predictions generated via API tool-calling scripts.
Discussion
2 comments analyzed.
Concerns raised: AI picks similar to random/simple heuristics (just higher seeds), Unclear value proposition at $50 price point
Competitors
Other products that read as similar to this one — 109 launches clear the similarity bar, closest 8 shown.
Attention rank: #101 of 110 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 139 days after the earliest competitor.
- March Madness Bracket Challenge for AI Agents Only · hn · 2026-03-17 · 67 upvotes · similarity 0.56
- Pencil Puzzle Bench · hn · 2026-03-03 · 5 upvotes · similarity 0.48
- An unmetered LLM API–$6/month, no token tracking, no limits · hn · 2026-07-06 · 12 upvotes · similarity 0.45
- Sup AI, a confidence-weighted ensemble (52.15% on Humanity's Last Exam) · hn · 2026-03-26 · 26 upvotes · similarity 0.43
- I was laid off, so I built a NerdWallet for startup equity liquidity · hn · 2026-03-16 · 9 upvotes · similarity 0.41
- Group of Death · hn · 2026-06-05 · 6 upvotes · similarity 0.41
- Friendly prediction markets to turn trips into a running tournament · hn · 2026-04-27 · 5 upvotes · similarity 0.40
- New eval from SWE-bench team evalutes LMs based on goals not tickets · hn · 2025-11-05 · 5 upvotes · similarity 0.40
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.