OpenBenchmarks
Helping agents discover and pick the right SaaS APIs
Details
- External ID
- 48875730
- Source
- HN
- Company
- —
- Product
- OpenBenchmarks
- Website domain
- openbenchmarks.com
- Launched
- July 11, 2026
- Cohort
- —
- Upvotes
- 6
- Upvotes percentile
- 0.2873357228195938
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
I'm Fenil, co-founder/CEO of OpenFunnel (YC F24), building this with my co-founder/CTO Aditya. We're launching OpenBenchmarks (https://openbenchmarks.com), open-source, reproducible benchmarks for SaaS APIs, starting with the category we know best: GTM APIs.## Why we built thisMore and more B2B software evaluation will/already runs through reasoning models inside agentic workflows rather than through people. And buyers increasingly pick vendors that are API-first and ship MCPs, so they can wire them into internal workflows. Strong reasoning models are skeptical of marketing. When a genuinely neutral benchmark is available, they discount SEO and self-published benchmarks, and they lose trust the moment something looks like a marketing claim.The thing that survives that skepticism is an independent build-first benchmark the agent can reproduce itself and trust.Reproducible by default: every cell ships the literal HTTP request/response plus the judge prompt/responseAll benchmarks live at https://github.com/openbenchmarks-labs. We start with GTM APIs, with more on the way.## On Benching ourselvesWe benched our own product (OpenFunnel) as a vendor in the lookalikes benchmark, on purpose.We increasingly saw buyers asking for benchmarks on sales calls and were also curious to see if agents also made decisions the same way.We wanted to see whether an agent could run the whole loop end to end: discover the benchmark while researching a user's query, weigh it as the deciding factor in picking a winner for that user's category, then sign up and auth through to the chosen vendor to complete the task. That last stretch needs live vendors with agent-auth wired up: us and a few others.We came out #1 on the current seed (89% vs 74%). Because it's open, we don't win everything: depending on seed and metric (precision@10/@50/@100), we win some and lose others.## DogfoodingWe dogfooded it and ran 200 incognito simulated buyer flows through Claude Code with live web search, from query to discovery to sign-up, and watched which sources the model fetched and which ones made it into the final decision. The benchmark got opened in a large majority of runs, across the range of queries an actual buyer would ask in Claude Code, from building a lookalikes workflow to picking a lookalike API provider. And when it was opened, it usually drove the final call, over self-published GEO pages and vendor benchmarks with years of domain authority.We've studied the effect and potential ROI of being benched, we're taking OpenFunnel off the benchmark.## What's nextMore GTM benchmarks, then beyond: devtools and infrastructure.We're are also working with traditional SaaS companies thinking about going API-first and opening up to a new class of customer: agentsTry it, or hand it to your agent: https://openbenchmarks.com
Enrichment
- Theme
- browser automation and scraping for AI
- Vertical
- Horizontal
- Function
- Agent / copilot
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- saas api discovery for agents
- Manually corrected
- False
Could you build this?
Yes Benchmarking SaaS APIs involves scheduled cron workers sending synthetic API requests, evaluating responses/latency against a ground truth dataset, and displaying the results on a dashboard.
Discussion
2 comments analyzed.
Competitors mentioned: Brave
Concerns raised: Hidden variables in search results based on location, Unclear how Claude search actually works
Competitors
Other products that read as similar to this one — 151 launches clear the similarity bar, closest 8 shown.
Attention rank: #113 of 152 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 249 days after the earliest competitor.
- Open Benchmarks Grants– a $3M commitment to close the AI eval gap · hn · 2026-02-11 · 6 upvotes · similarity 0.52
- Cheddar-bench · hn · 2026-02-22 · 9 upvotes · similarity 0.47
- OpenTrustBench · ph · 2026-09-16 · 2 upvotes · similarity 0.46
- Trying to fix the web scraping industry's benchmark problem · hn · 2026-07-16 · 18 upvotes · similarity 0.44
- Cua-Bench · hn · 2026-01-26 · 40 upvotes · similarity 0.44
- Axiomeer · hn · 2026-02-03 · 13 upvotes · similarity 0.43
- OpenGem · hn · 2026-02-22 · 7 upvotes · similarity 0.41
- Τ³-Bench is out · hn · 2026-03-25 · 12 upvotes · similarity 0.41
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a agent / copilot tool for Agriculture yet.