Spec27
Spec-driven validation for AI agents
Details
- External ID
- 47959984
- Source
- HN
- Company
- —
- Product
- Spec27
- Website domain
- spec27.ai
- Launched
- April 30, 2026
- Cohort
- —
- Upvotes
- 13
- Upvotes percentile
- 0.6658097686375322
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Hi HN! We’re a team of ML validation specialists and we’ve been building /Spec27, a tool for testing whether AI agents still do their job safely and reliably as models, prompts, tools, and surrounding systems change.We started working on this because a lot of current LLM evaluation work seems aimed at scoring general model behavior, while many teams are deploying systems that have a specific mission to fulfill. Many of the tools also assume you have full access to the agent stack and traces so you can place SDKs and Gateways, but a lot of agents are being created on vendor platforms where this isn’t possible.As a result, we approaches it from the outside in: all tests just run to the primary interfaces of an Agent and don’t assume anything about internals. The other important things about the approach is spec-driven. Instead of treating testing as a one-off benchmark or static eval set, we let teams define reusable specifications for the behavior they want from an agent, then generate tests against those specs. With this you can automatically generate adversarial and robustness checks, so you can see what an agent is sensitive to and what kinds of changes cause it to fail.We’ve worked on validation for other AI systems before, including vision and tabular workflows, and /Spec27 is our new product for language-model-based agents. Currently in early access, so we’d love feedback! The current version is strongest for single-turn agent and application validation. We do not fully support multi-turn interactions yet, and better telemetry/tool-call integration is still on our roadmap.We’ve made the product open to try for HN readers, with a sample flow so it’s easy to poke around without much setup. We’d especially love feedback from people deploying internal agents, vendor agents, or other AI systems where reliability matters more than benchmark scores.
Enrichment
- Theme
- AI agent frameworks and developer tools
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- validation framework for ai agents
- Manually corrected
- False
Could you build this?
Partial While test harnesses and LLM evaluation UI can be quickly bootstrapped, creating robust, deterministic semantic assertions and synthetic multi-step agent simulation environments requires deep evaluation methodology.
What it would actually take: A production version requires an agent sandbox orchestration system (Docker/Wasm runners), dynamic mocking of external tools/APIs, and specialized statistical drift detection algorithms for non-deterministic model outputs. The team needs ML test engineering experience and distributed trace-replay infrastructure to reliably benchmark multi-turn agent behavior.
Discussion
9 comments analyzed.
Competitors mentioned: Github CLI
Concerns raised: Hallucination detection and prevention, Agent safety/validation practices undercooked, Scaling agent workflows with growing codebase
Feature requests: Full flow examples/walkthroughs, Multi-turn extension support
Competitors
Other products that read as similar to this one — 997 launches clear the similarity bar, closest 8 shown.
Attention rank: #329 of 998 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 181 days after the earliest competitor.
- First autonomous ML and AI engineering Agent · hn · 2026-01-22 · 5 upvotes · similarity 0.54
- Autofix Bot · hn · 2025-12-11 · 37 upvotes · similarity 0.54
- Multi-agent autoresearch for ANE inference beats Apple's CoreML by 6× · hn · 2026-03-31 · 6 upvotes · similarity 0.50
- Statewright · hn · 2026-05-12 · 126 upvotes · similarity 0.50
- Multi-Agent Code Review · hn · 2025-11-04 · 5 upvotes · similarity 0.49
- Morph Reflexes · hn · 2026-06-30 · 20 upvotes · similarity 0.49
- Implement-spec · hn · 2026-08-04 · 12 upvotes · similarity 0.49
- AgentX · ph · 2026-06-22 · 523 upvotes · similarity 0.48
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.