TinyFish Web Agent (82% on hard tasks vs. Operator's 43%)
Details
- External ID
- 46991520
- Source
- HN
- Company
- —
- Product
- TinyFish Web Agent (82% on hard tasks vs. Operator's 43%)
- Website domain
- tinyfish.ai
- Launched
- Feb. 12, 2026
- Cohort
- —
- Upvotes
- 17
- Upvotes percentile
- 0.6610512129380054
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Enterprises need ~90% accuracy to deploy web agents. Until now, no agent has come close on real-world tasks. TinyFish is the first production-ready web agent. Here's the evidence.Results of hard task scores on Online-Mind2Web (300 tasks, 136 live websites, human-correlated judge):- TinyFish: 81.9% - OpenAI Operator: 43.2% - Claude Computer Use: 32.4% - Browser Use: 8.1%Why not WebVoyager like everyone else?Because it's broken. Easy tasks, Google Search shortcuts, and a judge that agrees with humans only 62% of the time. Browser Use self-reported 89% on WebVoyager — then scored 8.1% on hard tasks here.We evaluated TinyFish against Online-Mind2Web instead — 300 real tasks, 136 live websites, three difficulty levels, and a judge that agrees with humans 85% of the time. No shortcuts. No easy mode.The cookbook repo is open source: https://github.com/tinyfish-io/tinyfish-cookbookYou can see all failure task runs form here: https://tinyurl.com/tinyfish-mind2webHappy to answer questions about the architecture, the benchmark methodology, or why we think WebVoyager scores are misleading.
Enrichment
- Theme
- browser automation and scraping for AI
- Vertical
- Horizontal
- Function
- Agent / copilot
- Audience
- B2B
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- web automation agent for complex tasks
- Manually corrected
- False
Could you build this?
No Achieving state-of-the-art 82% task completion on complex web benchmarks requires novel reinforcement learning/agentic training architectures, deep visual DOM parsing, and robust enterprise browser infrastructure that far exceeds prompt engineering.
What it would actually take: Building a high-accuracy web agent requires fine-tuning multimodal models with extensive trajectory data using imitation and reinforcement learning (e.g., PPO/DPO on browser traces). The architecture involves a distributed serverless headless browser fleet (Playwright/Chromium clusters) coupled with specialized DOM tree pruning, visual grounding, and anti-bot mitigation. You need seasoned ML researchers experienced in agentic workflows and distributed systems engineers.
Discussion
12 comments analyzed.
Competitors mentioned: OpenAI Operator, Claude Computer Use, Browser Use, WebVoyager
Concerns raised: Evaluation methodology doesn't control for website non-determinism, A/B tests, and cookie consent modals across sessions, Large performance drop on hard tasks suggests limited real-world applicability despite strong easy-task numbers, WebVoyager benchmark only agrees with humans 62% of the time, undermining validity of reported metrics, Ethical concerns about using own product to hype up the post, Latency/wall-clock time for hard task execution not disclosed
Feature requests: Structured DOM extraction before passing to model instead of pure vision, Better error recovery when actions don't produce expected state changes, Breakdown of failure modes and variance by difficulty level
Competitors
Other products that read as similar to this one — 23 launches clear the similarity bar, closest 8 shown.
Attention rank: #12 of 24 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 42 days after the earliest competitor.
- Agent Arena · hn · 2026-02-06 · 47 upvotes · similarity 0.36
- UserAgent-list · github · 2026-09-18 · 171 upvotes · similarity 0.34
- Webhook Skills · hn · 2026-02-04 · 9 upvotes · similarity 0.34
- The new Firecrawl /search · ph · 2026-07-24 · 259 upvotes · similarity 0.33
- Llama 3.2 3B and Keiro Research achieves 85% on SimpleQA · hn · 2026-03-07 · 6 upvotes · similarity 0.33
- BrowserAct · ph · 2026-06-25 · 561 upvotes · similarity 0.33
- Web Search Agents by Nimble · ph · 2026-09-14 · 306 upvotes · similarity 0.32
- SanityCheck · ph · 2026-09-10 · 7 upvotes · similarity 0.32
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a agent / copilot tool for Agriculture yet.