Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

New eval from SWE-bench team evalutes LMs based on goals not tickets

Details

External ID
45824582
Source
HN
Company
—
Product
New eval from SWE-bench team evalutes LMs based on goals not tickets
Website domain
codeclash.ai
Launched
Nov. 5, 2025
Cohort
—
Upvotes
5
Upvotes percentile
0.0982532751091703
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Current evals test LMs on tasks: "fix this bug," "write a test"But we code to achieve goals: maximize revenue, cut costs, win usersMeet CodeClash: LMs compete via their codebases across multi-round tournaments to achieve high-level goals.Because real software dev isn’t about following instructions. It’s about achieving outcomes.Here's how it works:Two LMs enter a tournament. Each maintains its own codebase.Every round:1. Edit Phase: LMs modify their codebases however they like 2. Competition phase: Codebases battle in an arena. 3. RepeatThe LM that wins the majority of rounds is declared winner.Arenas can be anything like games, trading sims, cybersec envs. We currently have 6 arenas implemented and support for 8 different programming languages.This has been one of our biggest projects in terms of scale to date. Over the past few months, we've completed 1.5k tournaments, totalling more than 50,400 agent runs. And you can look at all of these runs right now from your browser (links below!)You can find the rankings on our website (spoiler: Sonnet 4.5 tops the list), but almost more interesting: Humans are still way ahead! In one of our arena, even the worst solution from the human leaderboard is miles ahead of the best LM!And we're not surprised: LMs consistently fail to properly adapt to outcomes, hallucinate about reasons for failure, and produce ever messier codebases with every round.More information:https://codeclash.ai/ https://arxiv.org/pdf/2511.00839 https://github.com/codeclash-ai/codeclash

Enrichment

Theme
AI agent frameworks and developer tools
Vertical
—
Function
Observability & eval
Audience
Developer
AI stance
AI feature
Project type
Commercial product
Normalized one-liner
llm evaluation benchmark for goal-based tasks
Manually corrected
False

Could you build this?

No CodeClash is an advanced AI evaluation platform running multi-round competitive tournaments across complex environments like Halite and BattleSnake with secure code sandboxing and statistical Elo benchmarking. Building and scientifically validating a frontier LLM benchmark requires significant AI research expertise and dedicated execution infrastructure.

What it would actually take: A full version requires containerized, sandboxed multi-agent execution harnesses (Docker/Firecracker) capable of compiling and running arbitrary untrusted code safely at scale. It needs game engine integration, deterministic round-robin or Swiss-tournament matchmaking, dynamic LLM prompt/agent scaffolding, and robust statistical Elo rating algorithms. Building this requires deep ML research expertise in benchmarking methodology, rigorous sandboxing security, and substantial compute budget to run thousands of model iterations.

Discussion

1 comment analyzed.

Feature requests: Integrate reinforcement learning into competitive framework

Competitors

Other products that read as similar to this one — 42 launches clear the similarity bar, closest 8 shown.

Attention rank: #41 of 43 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 4 days after the earliest competitor.

Other launches for this product