aa-agentperf-local
Benchmark local LLM serving by replaying real agent trajectories
Details
- External ID
- 1388431549
- Source
- GITHUB
- Company
- —
- Product
- aa-agentperf-local
- Website domain
- artificialanalysis.ai
- Launched
- Sept. 26, 2026
- Cohort
- —
- Upvotes
- 48
- Upvotes percentile
- 0.813412759415834
- Tags
- ai-agents, artificial-analysis, benchmark, inference, llm, local-llm
- Fetched at
- Sept. 30, 2026, 5:02 p.m.
- Updated at
- Sept. 30, 2026, 5:02 p.m.
Enrichment
- Theme
- developer tools for ai agents
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- local llm benchmark for agent trajectories
- Manually corrected
- False
Could you build this?
Partial Building a CLI benchmark harness that replays complex multi-step agent trajectories against local LLM serving engines (vLLM, Ollama) requires precise state tracking and deterministic execution environments.
What it would actually take: The tool requires a Python or Go CLI that parses trajectory logs (e.g., SWE-bench or custom action logs), mocks or sandboxes environmental execution (Docker/microVMs), and measures TTFT, TPOT, and concurrency metrics across local inference runtimes like vLLM, llama.cpp, or SGLang. The hard part is managing realistic sandboxed environments to replay multi-step tool calls without side-effect leaks.
Competitors
Other products that read as similar to this one — 2029 launches clear the similarity bar, closest 8 shown.
Attention rank: #342 of 2030 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 332 days after the earliest competitor.
- Axe · hn · 2026-03-03 · 6 upvotes · similarity 0.72
- Galdor · hn · 2026-06-13 · 7 upvotes · similarity 0.72
- Lore · hn · 2026-06-09 · 6 upvotes · similarity 0.65
- AgentNexus · hn · 2026-06-13 · 6 upvotes · similarity 0.65
- Cloclo · hn · 2026-04-06 · 5 upvotes · similarity 0.65
- Local Coding Agent with LLMs to Delegate Tool Calls to Small AI Models · hn · 2026-05-28 · 9 upvotes · similarity 0.65
- BLINDSPOT · github · 2026-09-14 · 15 upvotes · similarity 0.65
- Wasmer SDK · hn · 2026-09-01 · 5 upvotes · similarity 0.65
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.