Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

A new benchmark for testing LLMs for deterministic outputs

Details

External ID
47950283
Source
HN
Company
—
Product
A new benchmark for testing LLMs for deterministic outputs
Website domain
interfaze.ai
Launched
April 29, 2026
Cohort
—
Upvotes
60
Upvotes percentile
0.8508997429305912
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

When building workflows that rely on LLMs, we commonly use structured output for programmatic use cases like converting an invoice into rows or meeting transcripts into tickets or even complex PDFs into database entries.The model may return the schema you want, but with hallucinated values like `invoice_date` being off by 2 months or the transcript array ordered wrongly. The JSON is valid, but the values are not.Structured output today is a big part of using LLMs, especially when building deterministic workflows.Current structured output benchmarks (e.g., JSONSchemaBench) only validate the pass rate for JSON schema and types, and not the actual values within the produced JSON.So we designed the Structured Output Benchmark (SOB) that fixes this by measuring both the JSON schema pass rate, types, and the value accuracy across all three modalities, text, image, and audio.For our test set, every record is paired with a JSON Schema and a ground-truth answer that was verified against the source context manually by a human and an LLM cross-check, so a missing or hallucinated value will be considered to be wrong.Open source is doing pretty well with GLM 4.7 coming in number 2 right after GPT 5.4.We noticed the rankings shift across modalities: GLM-4.7 leads text, Gemma-4-31B leads images, Gemini-2.5-Flash leads audio.For example, GPT-5.4 ranks 3rd on text but 9th on images.Model size is not a predictor, either: Qwen3.5-35B and GLM-4.7 beat GPT-5 and Claude-Sonnet-4.6 on Value Accuracy. Phi-4 (14B) beats GPT-5 and GPT-5-mini on text.Structured hallucinations are the hardest bug. Such values are type-correct, schema-valid, and plausible, so they slip through most guardrails. For example, in one audio record, the ground truth is "target_market_age": "15 to 35 years", and a model returns "25 to 35". This is invisible without field-level checks.Our goal is to be the best general model for deterministic tasks, and a key aspect of determinism is a controllable and consistent output structure. The first step to making structured output better is to measure it and hold ourselves against the best.

Enrichment

Theme
niche developer utilities and toolchains
Vertical
Horizontal
Function
Observability & eval
Audience
Developer
AI stance
Not AI
Project type
Hobby / open-source project
Normalized one-liner
benchmark for llm deterministic output testing
Manually corrected
False

Could you build this?

Partial The benchmarking runner and leaderboard site are simple to assemble, but curating a multi-modal gold-standard evaluation dataset across complex real-world text, images, and audio with statistically rigorous grading metrics is labor- and domain-intensive.

What it would actually take: Requires building automated test harness pipelines against multiple model vendor APIs, managing diverse multimodal corpora (annotated receipts, invoices, audio transcripts), and designing field-level semantic comparison metrics (Levenshtein, date parsing, numerical tolerances) alongside strict validation logic.

Discussion

20 comments analyzed.

Competitors mentioned: BAML for structured output tasks, Qwen models (3.5-35B noted as cost-effective), Reasoning models (Claude, GPT) for two-pass approach, Batch invariant kernel implementations (DeepSeek v4)

Concerns raised: Interfaze-Beta on own leaderboard raises neutrality questions, Single-pass JSON generation fragile; two-pass approach more reliable, Valid but incorrect JSON data harder to catch than parse errors, Benchmark doesn't show cost-per-performance or peak reasoning performance, LLM determinism broken by batching implementations

Feature requests: Add per-dollar performance metrics to benchmark, Extend benchmark to large reasoning models, Include Pareto frontier analysis for cost-vs-performance, Add Gemini 3.1 Flash Lite and other latest model versions, Show intermediary representation results (reasoning passes)

Competitors

Other products that read as similar to this one — 107 launches clear the similarity bar, closest 8 shown.

Attention rank: #13 of 108 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 179 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a observability & eval tool for Media & entertainment yet.