llm-multi-agent-game-evaluation
A dynamic multi-agent game framework for evaluating LLM strategic reasoning, communication-driven decision-making, deception, and social-norm behavior across RPS, Hawk-Dove, and Liars Bar environments.
Details
- External ID
- 1383099972
- Source
- GITHUB
- Company
- —
- Product
- llm-multi-agent-game-evaluation
- Website domain
- github.com
- Launched
- Sept. 23, 2026
- Cohort
- —
- Upvotes
- 19
- Upvotes percentile
- 0.5840379195490648
- Tags
- —
- Fetched at
- Sept. 27, 2026, 5:02 p.m.
- Updated at
- Sept. 27, 2026, 5:02 p.m.
Enrichment
- Theme
- autonomous agent research and evaluation
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- game-based evaluation framework for multi-agent llms
- Manually corrected
- False
Could you build this?
Yes This is a prompt-and-environment benchmarking harness orchestrating turn-based games (RPS, Hawk-Dove, Liars Bar) via standard LLM API calls and tracking state.
Competitors
Other products that read as similar to this one — 1170 launches clear the similarity bar, closest 8 shown.
Attention rank: #420 of 1171 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 328 days after the earliest competitor.
- agent-collusion · github · 2026-09-21 · 8 upvotes · similarity 0.67
- genpark-agent-self-consistency-consensus-scorer-skill · github · 2026-09-28 · 7 upvotes · similarity 0.61
- genpark-shapley-value-cooperative-game-skill · github · 2026-09-10 · 7 upvotes · similarity 0.59
- Galdor · hn · 2026-06-13 · 7 upvotes · similarity 0.58
- Multi-Agent-Game-Localizer · github · 2026-09-26 · 26 upvotes · similarity 0.56
- aa-agentperf-local · github · 2026-09-26 · 48 upvotes · similarity 0.55
- decision-playground · github · 2026-09-19 · 15 upvotes · similarity 0.54
- Strategy-RSI · github · 2026-09-11 · 10 upvotes · similarity 0.54
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.