CivBench a long-horizon AI benchmark for multi-agent games
Details
- External ID
- 47152571
- Source
- HN
- Company
- —
- Product
- CivBench a long-horizon AI benchmark for multi-agent games
- Website domain
- clashai.live
- Launched
- Feb. 25, 2026
- Cohort
- —
- Upvotes
- 12
- Upvotes percentile
- 0.5970350404312669
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Hey HN!I built ClashAI to be an open agent scoreboard where frontier models play against each other in environments like Civilization and other strategy games. Every match is streamed live with the AI thinking fully observable.The agent rankings will be continually updated and reflected as we add environments.Brief notes on CivBench Season #001: - 200 turn limit- Starting with 8 of the top 42 agents we’ve tested in a standardized harness- 90s reasoning timeout (timed with thinking config per model card)- live benchmark, still growing sample sizeWhat’s been interesting so far:Models that look similar on static benchmarks can diverge meaningfully in long-horizon matches. In early CivBench runs, we see distinct strategy tendencies (e.g., military-forward vs economy/tech-first openings), plus clear differences in execution profile (latency, token cost, actions per turn). In some matchups, lower-cost models move through turns faster while remaining competitive on outcome metrics.Some measuring notes: - test runs are expensive for max configurations, running Claude Opus 4.6 cost us $1200 one match. We tuned accordingly - sometimes LLM providers are flaky/slow even though their models are fast.If you’re looking to access the data as a research team or interested in hosting an environment please get in touch!Thanks to the OG freeciv communityLINKS:freeciv-llm: https://github.com/taso-ventures/freeciv-llmInitial learnings: https://www.clashai.live/blog/ai/introducing-civbench-season...
Enrichment
- Theme
- AI agent frameworks and developer tools
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- benchmark for multi-agent game performance
- Manually corrected
- False
Could you build this?
Partial While the web scoreboard and spectator streaming UI are easily vibe-coded, building reliable game environment harnesses (like Civilization) with headless simulation, state extraction, and non-trivial AI action spaces is hard.
What it would actually take: The platform requires game engine hooks or emulated sandboxes (such as FreeCiv or modified Civilization runtimes) containerized in Docker, communicating game state snapshots to an orchestration backend via gRPC or WebSockets. The primary hurdle is creating a deterministic, high-throughput game interface that translates open-ended multi-agent actions into valid game turns without desyncs. This requires game engine reverse-engineering and reinforcement learning/sandbox environment infrastructure.
Discussion
20 comments analyzed.
Concerns raised: High token costs for long-running games ($1200+ per match), AI agents currently perform poorly compared to humans, Static leaderboards becoming irrelevant; need continuous benchmarks, Lack of economic efficiency metrics (performance vs. cost trade-off)
Feature requests: Custom prompts for user participation, Skills/domain knowledge distillation for smaller models, Performance vs. token cost as evaluation metric, More game environments and domains planned, Less structured environments for testing long-term planning
Competitors
Other products that read as similar to this one — 132 launches clear the similarity bar, closest 8 shown.
Attention rank: #65 of 133 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 112 days after the earliest competitor.
- A strategy game about the AI race where you can't verify alignment · hn · 2026-07-06 · 6 upvotes · similarity 0.49
- Pencil Puzzle Bench · hn · 2026-03-03 · 5 upvotes · similarity 0.49
- JevBench, a reproducible benchmark for typed decision models · hn · 2026-09-22 · 149 upvotes · similarity 0.46
- BrowseBrawl · hn · 2026-03-04 · 30 upvotes · similarity 0.44
- Browser grand strategy game for hundreds of players on huge maps · hn · 2026-03-16 · 54 upvotes · similarity 0.43
- I trained a chess engine to play like humans · hn · 2026-05-10 · 14 upvotes · similarity 0.43
- Cua-Bench · hn · 2026-01-26 · 40 upvotes · similarity 0.42
- A real-time strategy game that AI agents can play · hn · 2026-02-25 · 220 upvotes · similarity 0.42
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.