Automated Testing for AI Agents
Details
- External ID
- 47270208
- Source
- HN
- Company
- —
- Product
- Automated Testing for AI Agents
- Website domain
- zalor.ai
- Launched
- March 6, 2026
- Cohort
- —
- Upvotes
- 8
- Upvotes percentile
- 0.491389913899139
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Hi Hacker News! We're launching Zalor, an agent testing platform.Agents often break when you tweak system prompts, swap models, or add tools. Zalor automatically generates test scenarios and evaluates your agent so you know it's reliable before deploying to production.We currently support the OpenAI Agents SDK and are onboarding other frameworks. A GitHub integration is coming so you can get feedback on every update.Looking forward to hearing feedback from people building agents.
Enrichment
- Theme
- developer tools for AI agents
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI feature
- Project type
- Commercial product
- Normalized one-liner
- testing framework for ai agents
- Manually corrected
- False
Could you build this?
Partial Generating test scenarios and evaluating agent calls with an LLM-as-a-judge can be started via vibe coding, but reliable non-deterministic evaluation and rigorous synthetic test generation require specialized agent-benchmarking expertise.
What it would actually take: The platform requires a mock environment/proxy that intercepts LLM and tool calls, a stateful graph-based scenario simulator, and an evaluation pipeline using LLM judges, semantic diffing, and behavioral property assertions. The hard technical problem is reliably generating diverse edge-case scenarios that faithfully replicate production failures without hallucinating false positives, alongside managing non-deterministic multi-turn rollouts.
Discussion
6 comments analyzed.
Feature requests: GitHub integration
Competitors
Other products that read as similar to this one — 438 launches clear the similarity bar, closest 8 shown.
Attention rank: #227 of 439 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 123 days after the earliest competitor.
- SanityCheck · ph · 2026-09-10 · 7 upvotes · similarity 0.50
- AutonixQ · ph · 2026-09-09 · 1 upvotes · similarity 0.48
- QA Agent Builder · ph · 2026-09-14 · 1 upvotes · similarity 0.48
- TestSprite 3.0 · ph · 2026-05-22 · 471 upvotes · similarity 0.47
- Agent Reliability Toolkit · ph · 2026-09-10 · 1 upvotes · similarity 0.47
- Flakestorm · hn · 2026-01-05 · 6 upvotes · similarity 0.47
- Spec27 · hn · 2026-04-30 · 13 upvotes · similarity 0.46
- Cekura Red Teaming: Stress-test your AI agents for Jailbreaks, Bias, Toxicity and more · yc · 2026-01-16 · 17 upvotes · similarity 0.46
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.