Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Trunchbull, run real models against any benchmark in your browser

Details

External ID
49273695
Source
HN
Company
—
Product
Trunchbull, run real models against any benchmark in your browser
Website domain
trunchbull.dev
Launched
Aug. 12, 2026
Cohort
—
Upvotes
6
Upvotes percentile
0.3125
Tags
—
Fetched at
Sept. 10, 2026, 5:32 a.m.
Updated at
Sept. 10, 2026, 5:32 a.m.

Description

Hi HN,Today I'm showcasing Trunchbull, a benchmarking platform designed for authoring benchmarks and running them against different models. We have direct support for benchmarks that use the harbor authoring system, custom tool authoring via the vercel ai sdk and configuration limits.We've also already imported terminalbench 2.0, as a sort of proof of concept that our harbor task orchestrator works, although you currently need a paid account as we are provisioning sandbox environments.I've made several popular benchmarks publicly available for testing. You dont need an account or your credit card information to access these:[Public benchmark demos](https://trunchbull.dev/sandboxes)[GSM8K math reasoning](https://trunchbull.dev/try/gsm8k)[SkateBench](https://trunchbull.dev/try/skatebench)[ARC-Challenge](https://trunchbull.dev/try/arc-challenge)[TruthfulQA](https://trunchbull.dev/try/truthfulqa-mc1)[Medical AI Failure Atlas](https://trunchbull.dev/try/medical-ai-failure-atlas)These demos let you pick from a preselected list of models, and will systematically test the selected models against their case scenarios.Any and all feedback is welcome, but i'm particularly interested in: - knowing what kind of benchmark evidence youd like to expect - any improvements on our benchmark run page, anything that can provide clarity or better understanding of the benchmark u just ran. - what you'd prefer to see on the overview page. - improvements on our documentation - better configuration and spend limits.

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Dev tools
Audience
Developer
AI stance
Not AI
Project type
Commercial product
Normalized one-liner
run models against benchmarks in browser
Manually corrected
False

Could you build this?

Partial Benchmarking models against standard datasets is approachable, but Trunchbull deploys arbitrary GitHub repos into isolated edge workers (Cloudflare/Docker containers) with hard resource limits and interactive terminal tracing.

What it would actually take: The architecture combines a web frontend, Cloudflare AI Gateway/orchestration, and an isolated sandbox execution layer (like Firecracker microVMs or gVisor-based edge runtimes). The hard part is building safe multi-tenant ephemeral execution environments that support process inspection, network gating, hard token/dollar budget enforcement, and deep step-by-step tracing without opening security vulnerabilities. This requires specialized infrastructure engineering in container virtualization and distributed job scheduling.

Discussion

No comments on this launch.

Competitors

Other products that read as similar to this one — 24 launches clear the similarity bar, closest 8 shown.

Attention rank: #17 of 25 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 265 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a dev tools tool for Sales yet.