LLM Skirmish
a benchmark where LLMs play RTS games, by writing code
Details
- External ID
- 46885863
- Source
- HN
- Company
- —
- Product
- LLM Skirmish
- Website domain
- llmskirmish.com
- Launched
- Feb. 4, 2026
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.10512129380053908
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
I wanted to create an LLM game benchmark that put this generation of frontier LLMs' top skill, coding, on full display.Ten years ago, a team released a game called Screeps. It was described as an "MMO RTS sandbox for programmers." In Screeps, human players write javascript strategies that get executed in the game's environment.The Screeps paradigm, writing code and having it execute in a real-time game environment, is well suited for an LLM benchmark. Drawing on a version of the Screeps open source API, LLM Skirmish pits LLMs head-to-head in a series of 1v1 real-time strategy games.
Enrichment
- Theme
- typing games and skill trainers
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- benchmark where llms play rts games
- Manually corrected
- False
Could you build this?
Yes Building a benchmark harness around an existing open-source game engine (Screeps API) to feed game state to LLMs and evaluate their output scripts in 1v1 matches is readily vibe-coded with standard scripts and web frontends.
Discussion
2 comments analyzed.
Feature requests: Extend to general-soldier models and agent swarms, Apply to broader model evaluation beyond code tasks
Competitors
Other products that read as similar to this one — 46 launches clear the similarity bar, closest 8 shown.
Attention rank: #43 of 47 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 91 days after the earliest competitor.
- A real-time strategy game that AI agents can play · hn · 2026-02-25 · 220 upvotes · similarity 0.83
- LLMPvP · ph · 2026-09-29 · 2 upvotes · similarity 0.45
- Can you beat frontier LLMs at social strategy games? | Multi-Agent Arena by Olam Labs, evaluating through multi-agent simulations · yc · 2026-08-05 · 9 upvotes · similarity 0.43
- I taught LLMs to play Magic: The Gathering against each other · hn · 2026-02-17 · 117 upvotes · similarity 0.41
- Watch LLMs play 21,000 hands of Poker · hn · 2026-01-08 · 36 upvotes · similarity 0.40
- Speed Miners · hn · 2026-01-17 · 49 upvotes · similarity 0.39
- LLM Debate Benchmark · hn · 2026-03-23 · 9 upvotes · similarity 0.38
- Gradient Bang · ph · 2026-05-15 · 173 upvotes · similarity 0.37
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.