Can you beat frontier LLMs at social strategy games? | Multi-Agent Arena by Olam Labs, evaluating through multi-agent simulations
Play games like Risk, Poker, or Codenames against frontier AIs. We're running public long-horizon games to do evaluations on social behavior and agentic performance.
Details
- External ID
- 109432
- Source
- YC
- Company
- Olam Labs
- Product
- Can you beat frontier LLMs at social strategy games? | Multi-Agent Arena by Olam Labs, evaluating through multi-agent simulations
- Website domain
- olamlabs.ai
- Launched
- Aug. 5, 2026
- Cohort
- Summer 2026
- Upvotes
- 9
- Upvotes percentile
- 0.3160919540229885
- Tags
- Artificial Intelligence, Reinforcement Learning, Gaming, Data Engineering
- Fetched at
- Sept. 30, 2026, 5 p.m.
- Updated at
- Sept. 30, 2026, 5 p.m.
Enrichment
- Theme
- AI agent games and chess tools
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- multi-agent simulation evaluation platform
- Manually corrected
- False
Could you build this?
Yes A web-based multiplayer game arena connecting human players with LLM agents via standard game state engines and commercial LLM APIs can easily be built with vibe coding.
Competitors
Other products that read as similar to this one — 68 launches clear the similarity bar, closest 8 shown.
Attention rank: #47 of 69 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 273 days after the earliest competitor.
- Multi-Agent Arena · ph · 2026-09-27 · 2 upvotes · similarity 0.78
- A real-time strategy game that AI agents can play · hn · 2026-02-25 · 220 upvotes · similarity 0.55
- Watch LLMs play 21,000 hands of Poker · hn · 2026-01-08 · 36 upvotes · similarity 0.48
- llm-multi-agent-game-evaluation · github · 2026-09-23 · 19 upvotes · similarity 0.47
- Play poker with LLMs, or watch them play against each other · hn · 2026-01-10 · 163 upvotes · similarity 0.45
- LLM Skirmish · hn · 2026-02-04 · 5 upvotes · similarity 0.43
- A business SIM where humans beat GPT-5 by 9.8 X · hn · 2025-11-19 · 23 upvotes · similarity 0.42
- CivBench a long-horizon AI benchmark for multi-agent games · hn · 2026-02-25 · 12 upvotes · similarity 0.42
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.