JevBench, a reproducible benchmark for typed decision models
Details
- External ID
- 49800574
- Source
- HN
- Company
- —
- Product
- JevBench, a reproducible benchmark for typed decision models
- Website domain
- benchmarkheaven.com
- Launched
- Sept. 22, 2026
- Cohort
- —
- Upvotes
- 149
- Upvotes percentile
- 0.9497607655502392
- Tags
- —
- Fetched at
- Sept. 26, 2026, 10:53 p.m.
- Updated at
- Sept. 26, 2026, 10:53 p.m.
Description
Hi HN! I built JevBench because Jev kicks ass, and the world deserves to know how the serious open source and fake lookalike projects really perform in comparison.Jev-class models return bounded choices and probabilities instead of text, and are disruptively faster and cheaper than LLMs, while being similarly intelligent on the text input they operate on.JevBench allows looking at accuracy, latency and price all at once, in a weighted way - you can even configure the weighting.A full run asks 534 English decisions. The v1.3 score combines chance-corrected Intelligence, Calibration, Speed and Cost.Leaderboard right now: #1 - Jev 74.4 #2 - SemIf 73.1 #3 - djev 73.0 #4 - Winnow-12B Q8 71.2 #5 reflex 4B 70.3. MIT harness, public items, frozen artifacts, scoring code and public per-task outcomes:https://github.com/fstandhartinger/jevbenchTwo no-signup demos:https://who-is-right.app.mintapis.comhttps://is-it-ai-slop.app.mintapis.comLimitations: English-only; latency from one German server; local/demo latency gets a disclosed ×2 adjustment (+150 ms on my servers) which is an informed assumption; held-out prompts still reach evaluated services; ~1-point gaps can be noise.Wdyt?
Enrichment
- Theme
- decision model runtimes and tools
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- benchmark suite for typed decision models
- Manually corrected
- False
Could you build this?
Partial Building the benchmark web leaderboard UI is straightforward, but defining, evaluating, and reliably scoring 50+ typed decision systems across 500+ specialized decision problems requires deep domain test harness design.
What it would actually take: The core platform requires a standardized execution harness that interfaces with various model providers and typed decision architectures, verifying strictly formatted JSON/probabilistic outputs against ground-truth decision distributions. It requires rigorous benchmarking methodologies, cost/latency profiling infra across international data residency zones, and a static or dynamic frontend (like Next.js with automated CI/CD benchmark runner pipelines).
Discussion
15 comments analyzed.
Competitors mentioned: SemIf, Qwen, ChatGPT
Concerns raised: Slop detector easily fooled by AI text, High false positive rate on human text and keysmash, Jev performs on-par with free alternatives despite high funding, Jev is expensive and closed-source, Vibecoded UI with unnecessary explanatory LLM text
Feature requests: Transparent source and methodology, Flagging AI written email
Competitors
Other products that read as similar to this one — 349 launches clear the similarity bar, closest 8 shown.
Attention rank: #53 of 350 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 322 days after the earliest competitor.
- jevbench · github · 2026-09-19 · 101 upvotes · similarity 0.79
- jevk5 · github · 2026-09-22 · 111 upvotes · similarity 0.64
- jev-benchmarks · github · 2026-09-17 · 15 upvotes · similarity 0.63
- open-alternative-jev · github · 2026-09-18 · 50 upvotes · similarity 0.60
- AnyDecisionModel · github · 2026-09-21 · 9 upvotes · similarity 0.59
- open-jev-typed-decision-engine · github · 2026-09-19 · 43 upvotes · similarity 0.59
- jev-forge · github · 2026-09-19 · 24 upvotes · similarity 0.59
- openJev-verdict-2.0 · github · 2026-09-19 · 276 upvotes · similarity 0.58
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.