NUA an agent that tests for product correctness
Details
- External ID
- 48364701
- Source
- HN
- Company
- —
- Product
- NUA an agent that tests for product correctness
- Website domain
- trynua.dev
- Launched
- June 2, 2026
- Cohort
- —
- Upvotes
- 8
- Upvotes percentile
- 0.4952185792349727
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
We’ve been using background Claude loops a lot recently, and we would wake up to PRs that didn’t solve the problem we wanted, made on assumptions that were wrong. Furthermore, the tests that the agents wrote were usually tautological, and didn’t test for intent. We wanted an agent that took all the context a company has, and writes tests that check for product correctness as well.For example, we work in reg tech, so bugs aren’t always technical. What we often see is things like insider trading alerts that should’ve fired that didn’t. We wanted an agent that turns laws and regulations into tests.For now, users can upload PDF, MD, TXT, and DOCX files, but we’re planning integrations like Slack, Notion, Linear, and Zoom in the future.We’re early on, so we would love to know what you all think!
Enrichment
- Theme
- AI agent frameworks and developer tools
- Vertical
- Horizontal
- Function
- Agent / copilot
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- ai agent for testing product correctness
- Manually corrected
- False
Could you build this?
Partial While wrapping LLM APIs with test harness prompts is straightforward, accurately verifying semantic intent and avoiding tautological test generation across arbitrary codebases requires sophisticated static/dynamic analysis and ground-truth evaluation pipelines.
What it would actually take: A functional system requires integrating headless AST parsing and symbolic execution with LLM agent loops, plus a deterministic test execution sandbox (e.g. Dockerized test runners). The hardest problem is semantic intent extraction—differentiating genuine functional behavior from tautological assertions without hallucinating expected outcomes. Building this reliably requires deep expertise in formal software verification, compiler design, and evaluation engineering.
Discussion
4 comments analyzed.
Competitors mentioned: Playwright
Concerns raised: Tests run on prod app
Feature requests: GitHub Actions integration for running tests on every PR
Competitors
Other products that read as similar to this one — 108 launches clear the similarity bar, closest 8 shown.
Attention rank: #58 of 109 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 210 days after the earliest competitor.
- Ax-check.com · hn · 2026-09-17 · 37 upvotes · similarity 0.49
- CodeLeash: framework for quality agent development, NOT an orchestrator · hn · 2026-02-27 · 12 upvotes · similarity 0.45
- Continue · hn · 2026-02-17 · 44 upvotes · similarity 0.43
- Agent Checker · ph · 2026-09-06 · 2 upvotes · similarity 0.40
- 127 PRs to Prod this wknd with 18 AI agents: metaswarm. MIT licensed · hn · 2026-02-03 · 5 upvotes · similarity 0.39
- Spec27 · hn · 2026-04-30 · 13 upvotes · similarity 0.38
- NPM registry built for AI agents (MCP-first, <100ms health scores) · hn · 2026-02-02 · 6 upvotes · similarity 0.38
- The Order of the Agents · hn · 2026-04-25 · 5 upvotes · similarity 0.37
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a agent / copilot tool for Agriculture yet.