Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Pencil Puzzle Bench

LLM Benchmark for Multi-Step Verifiable Reasoning

Details

External ID
47235084
Source
HN
Company
—
Product
Pencil Puzzle Bench
Website domain
ppbench.com
Launched
March 3, 2026
Cohort
—
Upvotes
5
Upvotes percentile
0.1070110701107011
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

I've been working on applying LLMs to long-context, verifiable problems over the past year, and today I'm releasing a benchmark of 62,000 pencil puzzles across 94 types (sudoku, nonori, slitherlink, etc.). The benchmark also allows for intermediate checks /rule breaks for all varieties at any step.I tested 51 models against a subset (300 puzzles) in two modes: single-shot (output the full solution) and agentic (iterate with verifier feedback).Some results:- Best model (GPT 5.2@xhigh) solves 56%. (~ half the puzzles are unsolved by any model)- Agentic solves average 29 turns. The longest attempt took ~1,200 turns over 14 hours.- Cost per success varies wildly (cheapest: $0.00033 — Grok 4.1 Fast Reasoning, most expensive: $238.16 — Claude Sonnet 4.6 (1M context))- Reasoning depth (eg. @medium, @high, @xhigh) dramatically improves capability (up to repeated infrastructure failure for @xhigh)- Stark difference between US closed models (3 at >33%) and Chinese open models (top: 6%)Made the website to show off the dataset + play every puzzle, and even every replay AI agent solves step-by-step (fun to watch how it gets to solutions).Also here's the paper: https://arxiv.org/abs/2603.02119I didn't test human ability to solve, but it seems these puzzles are pretty difficult. I'd be curious how HN audience fares on the puzzles.

Enrichment

Theme
indie puzzle games and learning toys
Vertical
Horizontal
Function
Observability & eval
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
reasoning benchmark for llms
Manually corrected
False

Could you build this?

Partial A basic evaluation script and web dashboard can be vibe-coded, but building deterministic generators and step-by-step rule verification engines for 94 distinct combinatorial puzzle types requires substantial algorithmic implementation.

What it would actually take: The backend requires implementing or integrating high-performance SAT/SMT solvers (like Z3) or custom constraint satisfaction engines for dozens of logic puzzle variants (Sudoku, Slitherlink, Nonograms, Nurikabe, etc.). The hardest engineering challenge is constructing intermediate verification oracles that validate partial board states, detect non-trivial rule violations dynamically, and guarantee unique solutions across 62,000 generated puzzle instances. This necessitates strong expertise in discrete mathematics, constraint satisfaction problems (CSP), and formal verification.

Discussion

No comments on this launch.

Competitors

Other products that read as similar to this one — 251 launches clear the similarity bar, closest 8 shown.

Attention rank: #224 of 252 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 125 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a observability & eval tool for Media & entertainment yet.