A new benchmark for testing LLMs for deterministic outputs
Details
- External ID
- 47950283
- Source
- HN
- Company
- —
- Product
- A new benchmark for testing LLMs for deterministic outputs
- Website domain
- interfaze.ai
- Launched
- April 29, 2026
- Cohort
- —
- Upvotes
- 60
- Upvotes percentile
- 0.8508997429305912
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
When building workflows that rely on LLMs, we commonly use structured output for programmatic use cases like converting an invoice into rows or meeting transcripts into tickets or even complex PDFs into database entries.The model may return the schema you want, but with hallucinated values like `invoice_date` being off by 2 months or the transcript array ordered wrongly. The JSON is valid, but the values are not.Structured output today is a big part of using LLMs, especially when building deterministic workflows.Current structured output benchmarks (e.g., JSONSchemaBench) only validate the pass rate for JSON schema and types, and not the actual values within the produced JSON.So we designed the Structured Output Benchmark (SOB) that fixes this by measuring both the JSON schema pass rate, types, and the value accuracy across all three modalities, text, image, and audio.For our test set, every record is paired with a JSON Schema and a ground-truth answer that was verified against the source context manually by a human and an LLM cross-check, so a missing or hallucinated value will be considered to be wrong.Open source is doing pretty well with GLM 4.7 coming in number 2 right after GPT 5.4.We noticed the rankings shift across modalities: GLM-4.7 leads text, Gemma-4-31B leads images, Gemini-2.5-Flash leads audio.For example, GPT-5.4 ranks 3rd on text but 9th on images.Model size is not a predictor, either: Qwen3.5-35B and GLM-4.7 beat GPT-5 and Claude-Sonnet-4.6 on Value Accuracy. Phi-4 (14B) beats GPT-5 and GPT-5-mini on text.Structured hallucinations are the hardest bug. Such values are type-correct, schema-valid, and plausible, so they slip through most guardrails. For example, in one audio record, the ground truth is "target_market_age": "15 to 35 years", and a model returns "25 to 35". This is invisible without field-level checks.Our goal is to be the best general model for deterministic tasks, and a key aspect of determinism is a controllable and consistent output structure. The first step to making structured output better is to measure it and hold ourselves against the best.
Enrichment
- Theme
- niche developer utilities and toolchains
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- benchmark for llm deterministic output testing
- Manually corrected
- False
Could you build this?
Partial The benchmarking runner and leaderboard site are simple to assemble, but curating a multi-modal gold-standard evaluation dataset across complex real-world text, images, and audio with statistically rigorous grading metrics is labor- and domain-intensive.
What it would actually take: Requires building automated test harness pipelines against multiple model vendor APIs, managing diverse multimodal corpora (annotated receipts, invoices, audio transcripts), and designing field-level semantic comparison metrics (Levenshtein, date parsing, numerical tolerances) alongside strict validation logic.
Discussion
20 comments analyzed.
Competitors mentioned: BAML for structured output tasks, Qwen models (3.5-35B noted as cost-effective), Reasoning models (Claude, GPT) for two-pass approach, Batch invariant kernel implementations (DeepSeek v4)
Concerns raised: Interfaze-Beta on own leaderboard raises neutrality questions, Single-pass JSON generation fragile; two-pass approach more reliable, Valid but incorrect JSON data harder to catch than parse errors, Benchmark doesn't show cost-per-performance or peak reasoning performance, LLM determinism broken by batching implementations
Feature requests: Add per-dollar performance metrics to benchmark, Extend benchmark to large reasoning models, Include Pareto frontier analysis for cost-vs-performance, Add Gemini 3.1 Flash Lite and other latest model versions, Show intermediary representation results (reasoning passes)
Competitors
Other products that read as similar to this one — 107 launches clear the similarity bar, closest 8 shown.
Attention rank: #13 of 108 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 179 days after the earliest competitor.
- Smelt · hn · 2026-03-07 · 6 upvotes · similarity 0.55
- Tabstack Structured Extraction · ph · 2026-06-11 · 199 upvotes · similarity 0.46
- Misata · hn · 2025-12-16 · 24 upvotes · similarity 0.44
- ISON · hn · 2025-12-26 · 7 upvotes · similarity 0.44
- Keep large tool output out of LLM context: 3x accuracy 95% fewer tokens · hn · 2026-03-05 · 10 upvotes · similarity 0.43
- jev-bench · ph · 2026-09-28 · 2 upvotes · similarity 0.42
- open-jev · github · 2026-09-16 · 19 upvotes · similarity 0.41
- genpark-json-schema-to-regex-converter-skill · github · 2026-09-29 · 7 upvotes · similarity 0.41
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.