Cua-Bench
a benchmark for AI agents in GUI environments
Details
- External ID
- 46768906
- Source
- HN
- Company
- —
- Product
- Cua-Bench
- Website domain
- github.com
- Launched
- Jan. 26, 2026
- Cohort
- —
- Upvotes
- 40
- Upvotes percentile
- 0.761528326745718
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Hey HN, we're excited to share Cua-Bench ( https://github.com/trycua/cua ), an open-source framework for evaluating and training computer-use agents across different environments.Computer-use agents show massive performance variance across different UIs—an agent with 90% success on Windows 11 might drop to 9% on Windows XP for the same task. The problem is OS themes, browser versions, and UI variations that existing benchmarks don't capture.The existing benchmarks (OSWorld, Windows Agent Arena, AndroidWorld) were great but operated in silos—different harnesses, different formats, no standardized way to test the same agent across platforms. More importantly, they were evaluation-only. We needed environments that could generate training data and run RL loops, not just measure performance. Cua-Bench takes a different approach: it's a unified framework that standardizes environments across platforms and supports the full agent development lifecycle—benchmark, train, deploy.With Cua-Bench, you can:- Evaluate agents across multiple benchmarks with one CLI (native tasks + OSWorld + Windows Agent Arena adapters)- Test the same agent on different OS variations (Windows 11/XP/Vista, macOS themes, Linux, Android via QEMU)- Generate new tasks from natural language prompts- Create simulated environments for RL training (shell apps like Spotify, Slack with programmatic rewards)- Run oracle validations to verify environments before agent evaluation- Monitor agent runs in real-time with traces and screenshotsAll of this works on macOS, Linux, Windows, and Android, and is self-hostable.To get started:Install cua-bench:% pip install cua-benchRun a basic evaluation:% cb run dataset datasets/cua-bench-basic --agent demoOpen the monitoring dashboard:% cb run watch <run_id>For parallelized evaluations across multiple workers:% cb run dataset datasets/cua-bench-basic --agent your-agent --max-parallel 8Want to test across different OS variations? Just specify the environment:% cb run task slack_message --agent your-agent --env windows_xp% cb run task slack_message --agent your-agent --env macos_sonomaGenerate new tasks from prompts:% cb task generate "book a flight on kayak.com"Validate environments with oracle implementations:% cb run dataset datasets/cua-bench-basic --oracleThe simulated environments are particularly useful for RL training—they're HTML/JS apps that render across 10+ OS themes with programmatic reward verification. No need to spin up actual VMs for training loops.We're seeing teams use Cua-Bench for:- Training computer-use models on mobile and desktop environments- Generating large-scale training datasets (working with labs on millions of screenshots across OS variations)- RL fine-tuning with shell app simulators- Systematic evaluation across OS themes and browser versions- Building task registries (collaborating with Snorkel AI on task design and data curation, similar to their Terminal-Bench work)Cua-Bench is 100% open-source under the MIT license. We're actively developing it as part of Cua (https://github.com/trycua/cua), our Computer Use Agent SDK, and we'd love your feedback, bug reports, or feature ideas.GitHub: https://github.com/trycua/cuaDocs: https://cua.ai/docs/cuabenchTechnical Report: https://cuabench.aiWe'll be here to answer any technical questions and look forward to your comments!
Enrichment
- Theme
- developer tools for AI agents
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- benchmark for ai agents in gui environments
- Manually corrected
- False
Could you build this?
Partial The benchmarking orchestration scripts can be vibe-coded, but setting up multi-OS sandboxed virtual machine orchestration with deterministic action evaluation requires deep systems and virtualization engineering.
What it would actually take: The system requires an orchestration backend spinning up disposable VM environments (QEMU/KVM, cloud VMs, or Docker containers running GUI desktops like XFCE/Windows) paired with hypervisor-level guest automation agents. The hard part is building deterministic verification harnesses that can programmatically assert whether complex GUI tasks succeeded across heterogeneous OS versions and application states without manual intervention. This demands deep infrastructure expertise in virtualization, OS-level automation, and benchmark harness design.
Discussion
8 comments analyzed.
Competitors mentioned: rtrvr.ai (DOM-only web agent), OSWorld/Windows Agent Arena (benchmarks), UiPath Enterprise Benchmark (CUA benchmark), Clawd/Moltbot (AI sandboxing), Lume CLI (Claude sandboxing)
Concerns raised: Vision/GUI agents struggle with popups and overlays, Requires large vision models and CDP browser integration, High infrastructure costs, Non-determinism in real GUIs (animations, loading states, timing issues), Windows edge cases (UAC prompts, driver dialogs, update notifications)
Feature requests: Support for adversarial/hard-mode scenarios (UAC, driver dialogs), More realistic benchmark tasks curated with actual workers, Integration with other CUA benchmarks
Competitors
Other products that read as similar to this one — 316 launches clear the similarity bar, closest 8 shown.
Attention rank: #73 of 317 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 83 days after the earliest competitor.
- Τ³-Bench is out · hn · 2026-03-25 · 12 upvotes · similarity 0.51
- Agentic CUDA Kernel Optimizer · hn · 2026-09-25 · 37 upvotes · similarity 0.50
- Compute:Arena · ph · 2026-09-17 · 78 upvotes · similarity 0.49
- Cheddar-bench · hn · 2026-02-22 · 9 upvotes · similarity 0.47
- CUA-Sandbox-Efficient-Environments-for-Computer-Use-Reinforcement-Learning · github · 2026-09-26 · 9 upvotes · similarity 0.46
- agentic-cuda-optimizer · github · 2026-09-24 · 35 upvotes · similarity 0.46
- Compute:Arena · hn · 2026-09-17 · 5 upvotes · similarity 0.45
- OpenBenchmarks · hn · 2026-07-11 · 6 upvotes · similarity 0.44
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.