genpark-agent-trajectory-pass-fail-evaluator-skill
Step-by-step agent trajectory evaluator comparing execution traces against golden tool call sequences
Details
- External ID
- 1394935113
- Source
- GITHUB
- Company
- —
- Product
- genpark-agent-trajectory-pass-fail-evaluator-skill
- Website domain
- github.com
- Launched
- Sept. 29, 2026
- Cohort
- —
- Upvotes
- 7
- Upvotes percentile
- 0.05976172175249808
- Tags
- agent-benchmarks, agent-reliability, agent-skills, developer-tools, evaluation, mcp, python-standard-library, quality-assurance, testing, tool-call-eval, trajectory-evaluator
- Fetched at
- Sept. 30, 2026, 1:02 a.m.
- Updated at
- Sept. 30, 2026, 1:02 a.m.
Enrichment
- Theme
- autonomous agent research and evaluation
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- trajectory evaluator for ai agents
- Manually corrected
- False
Could you build this?
Yes This is a focused utility script/skill that compares JSON/array execution traces of agent tool calls against golden benchmarks.
Competitors
Other products that read as similar to this one — 1760 launches clear the similarity bar, closest 8 shown.
Attention rank: #1543 of 1761 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 335 days after the earliest competitor.
- genpark-agent-multi-turn-eval-scorer-skill · github · 2026-09-26 · 7 upvotes · similarity 0.75
- genpark-agent-multi-turn-eval-scorer-skill · github · 2026-09-26 · 7 upvotes · similarity 0.75
- Part_HackBench · github · 2026-09-10 · 21 upvotes · similarity 0.62
- Buildbox: Agent analytics for real user outcomes · yc · 2026-08-03 · 10 upvotes · similarity 0.60
- genpark-continual-fewshot-exemplar-distiller-skill · github · 2026-09-26 · 7 upvotes · similarity 0.60
- aa-agentperf-local · github · 2026-09-26 · 48 upvotes · similarity 0.59
- Agent-skills-eval · hn · 2026-05-07 · 79 upvotes · similarity 0.59
- Skill-up · hn · 2026-07-30 · 5 upvotes · similarity 0.57
Other launches for this product
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.