Open-source playground to red-team AI agents with exploits published
Details
- External ID
- 47392677
- Source
- HN
- Company
- —
- Product
- Open-source playground to red-team AI agents with exploits published
- Website domain
- github.com
- Launched
- March 15, 2026
- Cohort
- —
- Upvotes
- 30
- Upvotes percentile
- 0.7958179581795818
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
We build runtime security for AI agents. The playground started as an internal tool that we used to test our own guardrails. But we kept finding the same types of vulnerabilities because we think about attacks a certain way. At some point you need people who don't think like you.So we open-sourced it. Each challenge is a live agent with real tools and a published system prompt. Whenever a challenge is over, the full winning conversation transcript and guardrail logs get documented publicly.Building the general-purpose agent itself was probably the most fun part. Getting it to reliably use tools, stay in character, and follow instructions while still being useful is harder than it sounds. That alone reminded us how early we all are in understanding and deploying these systems at scale.First challenge was to get an agent to call a tool it's been told to never call.Someone got through in around 60 seconds without ever asking for the secret directly (which taught us a lot).Next challenge is focused on data exfiltration with harder defences: https://playground.fabraix.com
Enrichment
- Theme
- developer tools for AI agents
- Vertical
- Security
- Function
- Dev tools
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- red-team playground for ai agents
- Manually corrected
- False
Could you build this?
Yes The playground is an open-source evaluation interface that sends prompt injection/jailbreak test payloads against agent configurations and logs the resulting exploit traces.
Discussion
13 comments analyzed.
Concerns raised: Guardrails insufficient at preventing multi-step exploit chains, LLM-as-a-judge defense is fragile and can be socially engineered, Prompt-level guardrails fail on session-level patterns, Scoping agent permissions reduces both risk and useful capabilities, No user feedback when exploitation attempts succeed
Feature requests: Add confirmation message for successful extractions, Add leaderboard for exploit submissions, Improve stateful guardrail evaluation across multi-turn sessions, Action-boundary enforcement instead of prompt-level guardrails
Competitors
Other products that read as similar to this one — 377 launches clear the similarity bar, closest 8 shown.
Attention rank: #85 of 378 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 132 days after the earliest competitor.
- OpenAPPA · hn · 2026-09-28 · 23 upvotes · similarity 0.51
- Running AI agents across environments needs a proper solution · hn · 2026-03-24 · 8 upvotes · similarity 0.51
- Open-source sandbox for your product team · hn · 2026-07-01 · 17 upvotes · similarity 0.49
- Agent Arena · hn · 2026-02-06 · 47 upvotes · similarity 0.49
- Open-source playground to red-team AI agents against public prompts · hn · 2026-08-09 · 13 upvotes · similarity 0.49
- Playground · ph · 2026-07-13 · 198 upvotes · similarity 0.47
- Agent Arena · ph · 2026-06-26 · 354 upvotes · similarity 0.47
- I built Comrade · hn · 2026-04-20 · 5 upvotes · similarity 0.47
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a dev tools tool for Sales yet.