A business SIM where humans beat GPT-5 by 9.8 X
Details
- External ID
- 45982346
- Source
- HN
- Company
- —
- Product
- —
- Website domain
- —
- Launched
- Nov. 19, 2025
- Cohort
- —
- Upvotes
- 23
- Upvotes percentile
- 0.7216157205240175
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Hi HN,Can current AI systems actually run a business?There’s a growing belief that LLM agents can already manage entire teams, replace the entire software stack or even act as an AI CEO.So we built a controlled, measurable environment to evaluate this premise.Why did we build this benchmark?A modern enterprise operates in a dynamic environment with high uncertainty and incomplete information. The CEO has to deal with delayed consequences, staffing/resource tradeoffs and death by a thousand cuts of failure modes.If we ever want AI systems that can meaningfully make operational or strategic decisions, say an AI CEO, then they must be able to handle these dynamics.So we made one.What did we build?Mini Amusement Parks (MAPs) is a RollerCoaster Tycoon style business simulator with: - Stochastic events - Incomplete information - Staffing, restocking, maintenance - Long horizon planning - Compounding operational failures - Resource constraints - Spatial layout affecting outcomesYou can play it & make it to the leaderboard here: https://maps.skyfall.ai/play (it’s fun)It looks like a simple game. But underneath, it’s a benchmark designed to answer one question:Can an agent operate a business coherently over time?What we testedWe evaluated: - Humans (internal and external testers) - Multiple GPT-5 agents - Variants with additional tools, documents, practice mode, planning scaffolds, etc.We intentionally stacked in favour of the models - full documentation, step by step action interfaces, sandbox exploration mode, extra observations, multiple prompting strategies, etc.What happened?Humans destroyed the agents by FAR. Even the strongest model, with documentation, tool use, and sandbox “practice”, reached <10% of human performance. The failure modes were consistent: - chasing flashy upgrades instead of profitable ones - ignoring maintenance, staffing, restocking - overreacting to noise - zero long-term plan - sandbox training often made things worseIt became clear: LLMs can use tools, but they cannot run systems. They break when randomness, time, and spatial constraints matter.Why does this matter?There’s a growing narrative that: - LLMs will run entire companies - LLMs will take over the jobs of CEOs - LLMs can be autonomous agents - LLMs can manage workflows end-to-endMAPs show the complete opposite.Operating a business requires: foresight, risk modeling, temporal reasoning, causal understanding, prioritization under uncertainty, adaptive planning. These are the basics of what a functional and real AI CEO would need and this is exactly where the current models break.If an LLM can’t run a toy business, how can you trust it with a real business?This benchmark is our first step toward understanding what an AI system would actually need in order to exhibit enterprise level decision making and the basics of the AI CEO. AI CEO is not a chatbot, not chain of thought, definitely not an agent wrapper but a true demonstration of operational intelligence.We’re sharing this because: - we want the community to try to beat the models - we want criticism of the benchmark - most importantly, we want an honest discussion about what “AI CEO” is and should do (surely it’s not LLMs)If you want to try beating the agents (it’s fun!): https://maps.skyfall.ai/playIf you want the read more about it, you can do so here: https://skyfall.ai/blog/building-the-foundations-of-an-ai-ce...Check our the launch video here: https://www.youtube.com/watch?v=7oqVAWw5Ii8Happy to answer questions in the thread.
Enrichment
- Theme
- AI agents for business operations
- Vertical
- Horizontal
- Function
- Analytics & BI
- Audience
- B2B
- AI stance
- Not AI
- Project type
- Commercial product
- Normalized one-liner
- business simulation game
- Manually corrected
- False
Could you build this?
Partial A simple simulated business game can be vibe coded, but creating a statistically valid, multi-agent AI benchmark environment with complex economic dynamics requires rigorous benchmark design and econometric modeling.
What it would actually take: The system requires a deterministic simulation engine modeling company financials, market dynamics, employee workflows, and customer behavior. It needs an evaluation harness to interface with LLM agent APIs (function calling, planning loops), manage context windows across multi-day turn-based simulations, and ensure non-trivial, non-memorizable economic challenges. Requires expertise in game theory, business simulations, and LLM evaluation benchmarks.
Discussion
14 comments analyzed.
Competitors mentioned: WorkArena++ (ServiceNow), OpenAI Atlas, VLMs (Vision Language Models), AI web browsers
Concerns raised: LLMs struggling with composite tasks and low accuracy, Profitability uncertain for self-publishing children's books, More variables than appear (inventory, timing, shipping, marketing), Far from running autonomous business operations
Feature requests: Benchmark/gym for minimal-scope digital businesses, VLM integration for spatial reasoning, Better handling of multi-variable business constraints, Domain-specific improvements for planning and execution
Competitors
Other products that read as similar to this one — 301 launches clear the similarity bar, closest 8 shown.
Attention rank: #96 of 302 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 15 days after the earliest competitor.
- I built a text-based business simulator to replace video courses · hn · 2026-01-16 · 90 upvotes · similarity 0.48
- Maingen: Simulations of real industrial operations, so agents can run the physical world · yc · 2026-08-07 · 5 upvotes · similarity 0.46
- Playbook — AI agents that run wealth management operations · yc · 2026-08-13 · 5 upvotes · similarity 0.45
- AI agents play SimCity through a REST API · hn · 2026-02-09 · 216 upvotes · similarity 0.44
- Poor 2 Power: Tycoon Simulator · ph · 2026-09-20 · 2 upvotes · similarity 0.44
- Async - Specialized AI agents, that automate costly work for small businesses. · yc · 2026-08-10 · 8 upvotes · similarity 0.43
- A game/benchmark where AI bots hunt each other · hn · 2026-01-08 · 5 upvotes · similarity 0.43
- Can you beat frontier LLMs at social strategy games? | Multi-Agent Arena by Olam Labs, evaluating through multi-agent simulations · yc · 2026-08-05 · 9 upvotes · similarity 0.42
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a analytics & bi tool for Legal yet.