Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

CivBench a long-horizon AI benchmark for multi-agent games

Details

External ID
47152571
Source
HN
Company
—
Product
CivBench a long-horizon AI benchmark for multi-agent games
Website domain
clashai.live
Launched
Feb. 25, 2026
Cohort
—
Upvotes
12
Upvotes percentile
0.5970350404312669
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Hey HN!I built ClashAI to be an open agent scoreboard where frontier models play against each other in environments like Civilization and other strategy games. Every match is streamed live with the AI thinking fully observable.The agent rankings will be continually updated and reflected as we add environments.Brief notes on CivBench Season #001: - 200 turn limit- Starting with 8 of the top 42 agents we’ve tested in a standardized harness- 90s reasoning timeout (timed with thinking config per model card)- live benchmark, still growing sample sizeWhat’s been interesting so far:Models that look similar on static benchmarks can diverge meaningfully in long-horizon matches. In early CivBench runs, we see distinct strategy tendencies (e.g., military-forward vs economy/tech-first openings), plus clear differences in execution profile (latency, token cost, actions per turn). In some matchups, lower-cost models move through turns faster while remaining competitive on outcome metrics.Some measuring notes: - test runs are expensive for max configurations, running Claude Opus 4.6 cost us $1200 one match. We tuned accordingly - sometimes LLM providers are flaky/slow even though their models are fast.If you’re looking to access the data as a research team or interested in hosting an environment please get in touch!Thanks to the OG freeciv communityLINKS:freeciv-llm: https://github.com/taso-ventures/freeciv-llmInitial learnings: https://www.clashai.live/blog/ai/introducing-civbench-season...

Enrichment

Theme
AI agent frameworks and developer tools
Vertical
Horizontal
Function
Observability & eval
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
benchmark for multi-agent game performance
Manually corrected
False

Could you build this?

Partial While the web scoreboard and spectator streaming UI are easily vibe-coded, building reliable game environment harnesses (like Civilization) with headless simulation, state extraction, and non-trivial AI action spaces is hard.

What it would actually take: The platform requires game engine hooks or emulated sandboxes (such as FreeCiv or modified Civilization runtimes) containerized in Docker, communicating game state snapshots to an orchestration backend via gRPC or WebSockets. The primary hurdle is creating a deterministic, high-throughput game interface that translates open-ended multi-agent actions into valid game turns without desyncs. This requires game engine reverse-engineering and reinforcement learning/sandbox environment infrastructure.

Discussion

20 comments analyzed.

Concerns raised: High token costs for long-running games ($1200+ per match), AI agents currently perform poorly compared to humans, Static leaderboards becoming irrelevant; need continuous benchmarks, Lack of economic efficiency metrics (performance vs. cost trade-off)

Feature requests: Custom prompts for user participation, Skills/domain knowledge distillation for smaller models, Performance vs. token cost as evaluation metric, More game environments and domains planned, Less structured environments for testing long-term planning

Competitors

Other products that read as similar to this one — 132 launches clear the similarity bar, closest 8 shown.

Attention rank: #65 of 133 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 112 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a observability & eval tool for Media & entertainment yet.