Nyx
multi-turn, adaptive, offensive testing harness for AI agents
Details
- External ID
- 47827802
- Source
- HN
- Company
- β
- Product
- Nyx
- Website domain
- fabraix.com
- Launched
- April 19, 2026
- Cohort
- β
- Upvotes
- 20
- Upvotes percentile
- 0.7506426735218509
- Tags
- β
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
We built Nyx to solve a problem we kept hitting while building agents: AI agents break in ways traditional software doesn't. Logic bugs, reasoning failures, edge cases that manual testing and static benchmarks never explore.Nyx is an autonomous testing harness that probes your AI agents to find failure modes before users do. Itβs used to find logic bugs, instruction following failures, edge cases in agent behavior, and for red-team security testing (jailbreaks, prompt injection, tool hijacking)Technical approach: * Pure blackbox (no special access needed - test like your users interact) * Multi-turn adaptive conversations * Multi-modal testing (voice, text, images, documents, browser interactions) * Massively parallel by defaultInstead of spending time writing static evals for the key failure modes of your AI agents, point Nyx at any system and it autonomously discovers failure modes that matter. We typically find issues in under 10 minutes that manual audits take hours to surface.This is early work and we know the methodology is still going to evolve. We would love nothing more than feedback from the community as we iterate on this.
Enrichment
- Theme
- ai cybersecurity and penetration testing
- Vertical
- Security
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI feature
- Project type
- Commercial product
- Normalized one-liner
- offensive testing harness for ai agents
- Manually corrected
- False
Could you build this?
Partial While an agent testing dashboard is straightforward, building an adaptive multi-turn red-teaming harness requires complex automated reasoning, jailbreak/adversarial attack heuristics, and state tracking.
What it would actually take: A production harness requires an orchestrator running dynamic test policies using advanced LLM-as-attacker models with tree-of-thought or Monte Carlo tree search algorithms to probe edge cases. It needs sandboxed execution environments to safely trigger agent tool-calling side effects, deterministic replay capture, and formal evaluation rubrics. The core technical hurdle is generating novel, context-aware adversarial vectors across multi-turn trajectories without stalling into repetitive loops.
Discussion
8 comments analyzed.
Competitors mentioned: coverage-guided fuzzing tools
Concerns raised: added complexity vs. simpler approaches, unclear how it differs from existing fuzzing literature
Feature requests: CI/CD pipeline integration
Competitors
Other products that read as similar to this one — 157 launches clear the similarity bar, closest 8 shown.
Attention rank: #34 of 158 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 167 days after the earliest competitor.
- Xalgorix · hn · 2026-07-06 · 5 upvotes · similarity 0.46
- GAUNTLEX · ph · 2026-09-19 · 2 upvotes · similarity 0.45
- Fabraix · ph · 2026-05-08 · 196 upvotes · similarity 0.44
- AgentX · ph · 2026-06-22 · 523 upvotes · similarity 0.44
- Autofix Bot · hn · 2025-12-11 · 37 upvotes · similarity 0.44
- Canary (YC) · hn · 2026-09-24 · 8 upvotes · similarity 0.43
- Decipher AI: Agentic QA for the era of coding agents · yc · 2026-01-28 · 13 upvotes · similarity 0.42
- Agnost AI · ph · 2026-08-25 · 289 upvotes · similarity 0.42
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.