Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

TinyFish Web Agent (82% on hard tasks vs. Operator's 43%)

Details

External ID
46991520
Source
HN
Company
—
Product
TinyFish Web Agent (82% on hard tasks vs. Operator's 43%)
Website domain
tinyfish.ai
Launched
Feb. 12, 2026
Cohort
—
Upvotes
17
Upvotes percentile
0.6610512129380054
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Enterprises need ~90% accuracy to deploy web agents. Until now, no agent has come close on real-world tasks. TinyFish is the first production-ready web agent. Here's the evidence.Results of hard task scores on Online-Mind2Web (300 tasks, 136 live websites, human-correlated judge):- TinyFish: 81.9% - OpenAI Operator: 43.2% - Claude Computer Use: 32.4% - Browser Use: 8.1%Why not WebVoyager like everyone else?Because it's broken. Easy tasks, Google Search shortcuts, and a judge that agrees with humans only 62% of the time. Browser Use self-reported 89% on WebVoyager — then scored 8.1% on hard tasks here.We evaluated TinyFish against Online-Mind2Web instead — 300 real tasks, 136 live websites, three difficulty levels, and a judge that agrees with humans 85% of the time. No shortcuts. No easy mode.The cookbook repo is open source: https://github.com/tinyfish-io/tinyfish-cookbookYou can see all failure task runs form here: https://tinyurl.com/tinyfish-mind2webHappy to answer questions about the architecture, the benchmark methodology, or why we think WebVoyager scores are misleading.

Enrichment

Theme
browser automation and scraping for AI
Vertical
Horizontal
Function
Agent / copilot
Audience
B2B
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
web automation agent for complex tasks
Manually corrected
False

Could you build this?

No Achieving state-of-the-art 82% task completion on complex web benchmarks requires novel reinforcement learning/agentic training architectures, deep visual DOM parsing, and robust enterprise browser infrastructure that far exceeds prompt engineering.

What it would actually take: Building a high-accuracy web agent requires fine-tuning multimodal models with extensive trajectory data using imitation and reinforcement learning (e.g., PPO/DPO on browser traces). The architecture involves a distributed serverless headless browser fleet (Playwright/Chromium clusters) coupled with specialized DOM tree pruning, visual grounding, and anti-bot mitigation. You need seasoned ML researchers experienced in agentic workflows and distributed systems engineers.

Discussion

12 comments analyzed.

Competitors mentioned: OpenAI Operator, Claude Computer Use, Browser Use, WebVoyager

Concerns raised: Evaluation methodology doesn't control for website non-determinism, A/B tests, and cookie consent modals across sessions, Large performance drop on hard tasks suggests limited real-world applicability despite strong easy-task numbers, WebVoyager benchmark only agrees with humans 62% of the time, undermining validity of reported metrics, Ethical concerns about using own product to hype up the post, Latency/wall-clock time for hard task execution not disclosed

Feature requests: Structured DOM extraction before passing to model instead of pure vision, Better error recovery when actions don't produce expected state changes, Breakdown of failure modes and variance by difficulty level

Competitors

Other products that read as similar to this one — 23 launches clear the similarity bar, closest 8 shown.

Attention rank: #12 of 24 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 42 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a agent / copilot tool for Agriculture yet.