Part_HackBench
A stress-testing benchmark for partial-completion evaluators in stateful multi-turn agents, focusing on temporary progress, rollback, and attribution failures.
Details
- External ID
- 1363881689
- Source
- GITHUB
- Company
- —
- Product
- Part_HackBench
- Website domain
- github.com
- Launched
- Sept. 10, 2026
- Cohort
- —
- Upvotes
- 21
- Upvotes percentile
- 0.6211247758134768
- Tags
- —
- Fetched at
- Sept. 14, 2026, 5:28 p.m.
- Updated at
- Sept. 14, 2026, 5:28 p.m.
Enrichment
- Theme
- ai agent infrastructure and tooling
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- stress-testing benchmark for multi-turn agent evaluators
- Manually corrected
- False
Could you build this?
Yes Building an LLM benchmark harness with multi-turn state tracking and rollback testing is straightforward software engineering that can be vibe-coded with standard Python testing libraries.
Competitors
Other products that read as similar to this one — 1720 launches clear the similarity bar, closest 8 shown.
Attention rank: #550 of 1721 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 316 days after the earliest competitor.
- genpark-agent-multi-turn-eval-scorer-skill · github · 2026-09-26 · 7 upvotes · similarity 0.63
- genpark-agent-multi-turn-eval-scorer-skill · github · 2026-09-26 · 7 upvotes · similarity 0.63
- genpark-agent-trajectory-pass-fail-evaluator-skill · github · 2026-09-29 · 7 upvotes · similarity 0.62
- genpark-agent-trajectory-pass-fail-evaluator-skill · github · 2026-09-29 · 7 upvotes · similarity 0.62
- Polymath · yc · 2026-02-26 · 43 upvotes · similarity 0.60
- Teleport-env · hn · 2026-05-28 · 9 upvotes · similarity 0.59
- aa-agentperf-local · github · 2026-09-26 · 48 upvotes · similarity 0.59
- dolphinbench · github · 2026-09-22 · 29 upvotes · similarity 0.59
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.