Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Watch LLMs play 21,000 hands of Poker

Details

External ID
46540794
Source
HN
Company
—
Product
Watch LLMs play 21,000 hands of Poker
Website domain
adfontes.io
Launched
Jan. 8, 2026
Cohort
—
Upvotes
36
Upvotes percentile
0.7424242424242424
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

PokerBench is my attempt at a new LLM benchmark wherein frontier models play Texas Hold'em in an arena setting. It also features a simulator to view individual games and observe how the different models reason about poker strategy. Opus/Haiku, Gemini Pro/Flash, GPT-5.2/5 mini, and Grok 4.1 Fast Reasoning have all been included.All code -> https://github.com/JoeAzar/pokerbench

Enrichment

Theme
AI agent games and chess tools
Vertical
Horizontal
Function
Observability & eval
Audience
Developer
AI stance
AI feature
Project type
Hobby / open-source project
Normalized one-liner
watch llms play poker
Manually corrected
False

Could you build this?

Yes A simulation harness and web visualizer where an open-source poker game engine drives turns, sends prompts to various LLM provider APIs, and logs the game state for playback.

Discussion

19 comments analyzed.

Competitors mentioned: Deterministic bot with probability tables, Trading benchmarks (Deepseek), Open source models

Concerns raised: Very expensive to run games ($30/game for 6-handed, $6/game for 4-handed), Sample size may be too small (160 games, ~21k hands), Results may be random walk, need reinitialization to test robustness, Win rate vs. profit inconsistency (high win rate models losing money)

Feature requests: Incorporate open source/open weights models, Add tool use incorporation, Explain why smaller models outperform larger ones, Run multiple simultaneous trials for robustness

Competitors

Other products that read as similar to this one — 57 launches clear the similarity bar, closest 8 shown.

Attention rank: #14 of 58 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 57 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a observability & eval tool for Media & entertainment yet.