Watch LLMs play 21,000 hands of Poker
Details
- External ID
- 46540794
- Source
- HN
- Company
- —
- Product
- Watch LLMs play 21,000 hands of Poker
- Website domain
- adfontes.io
- Launched
- Jan. 8, 2026
- Cohort
- —
- Upvotes
- 36
- Upvotes percentile
- 0.7424242424242424
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
PokerBench is my attempt at a new LLM benchmark wherein frontier models play Texas Hold'em in an arena setting. It also features a simulator to view individual games and observe how the different models reason about poker strategy. Opus/Haiku, Gemini Pro/Flash, GPT-5.2/5 mini, and Grok 4.1 Fast Reasoning have all been included.All code -> https://github.com/JoeAzar/pokerbench
Enrichment
- Theme
- AI agent games and chess tools
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI feature
- Project type
- Hobby / open-source project
- Normalized one-liner
- watch llms play poker
- Manually corrected
- False
Could you build this?
Yes A simulation harness and web visualizer where an open-source poker game engine drives turns, sends prompts to various LLM provider APIs, and logs the game state for playback.
Discussion
19 comments analyzed.
Competitors mentioned: Deterministic bot with probability tables, Trading benchmarks (Deepseek), Open source models
Concerns raised: Very expensive to run games ($30/game for 6-handed, $6/game for 4-handed), Sample size may be too small (160 games, ~21k hands), Results may be random walk, need reinitialization to test robustness, Win rate vs. profit inconsistency (high win rate models losing money)
Feature requests: Incorporate open source/open weights models, Add tool use incorporation, Explain why smaller models outperform larger ones, Run multiple simultaneous trials for robustness
Competitors
Other products that read as similar to this one — 57 launches clear the similarity bar, closest 8 shown.
Attention rank: #14 of 58 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 57 days after the earliest competitor.
- Play poker with LLMs, or watch them play against each other · hn · 2026-01-10 · 163 upvotes · similarity 0.65
- I taught LLMs to play Magic: The Gathering against each other · hn · 2026-02-17 · 117 upvotes · similarity 0.51
- JevPokerBench · github · 2026-09-21 · 11 upvotes · similarity 0.49
- Can you beat frontier LLMs at social strategy games? | Multi-Agent Arena by Olam Labs, evaluating through multi-agent simulations · yc · 2026-08-05 · 9 upvotes · similarity 0.48
- PokerGame · ph · 2026-09-18 · 3 upvotes · similarity 0.48
- PokerRiskEngine · ph · 2026-09-30 · 1 upvotes · similarity 0.46
- Pencil Puzzle Bench · hn · 2026-03-03 · 5 upvotes · similarity 0.45
- BluffKing · ph · 2026-09-18 · 2 upvotes · similarity 0.42
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.