Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Open-source playground to red-team AI agents with exploits published

Details

External ID
47392677
Source
HN
Company
—
Product
Open-source playground to red-team AI agents with exploits published
Website domain
github.com
Launched
March 15, 2026
Cohort
—
Upvotes
30
Upvotes percentile
0.7958179581795818
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

We build runtime security for AI agents. The playground started as an internal tool that we used to test our own guardrails. But we kept finding the same types of vulnerabilities because we think about attacks a certain way. At some point you need people who don't think like you.So we open-sourced it. Each challenge is a live agent with real tools and a published system prompt. Whenever a challenge is over, the full winning conversation transcript and guardrail logs get documented publicly.Building the general-purpose agent itself was probably the most fun part. Getting it to reliably use tools, stay in character, and follow instructions while still being useful is harder than it sounds. That alone reminded us how early we all are in understanding and deploying these systems at scale.First challenge was to get an agent to call a tool it's been told to never call.Someone got through in around 60 seconds without ever asking for the secret directly (which taught us a lot).Next challenge is focused on data exfiltration with harder defences: https://playground.fabraix.com

Enrichment

Theme
developer tools for AI agents
Vertical
Security
Function
Dev tools
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
red-team playground for ai agents
Manually corrected
False

Could you build this?

Yes The playground is an open-source evaluation interface that sends prompt injection/jailbreak test payloads against agent configurations and logs the resulting exploit traces.

Discussion

13 comments analyzed.

Concerns raised: Guardrails insufficient at preventing multi-step exploit chains, LLM-as-a-judge defense is fragile and can be socially engineered, Prompt-level guardrails fail on session-level patterns, Scoping agent permissions reduces both risk and useful capabilities, No user feedback when exploitation attempts succeed

Feature requests: Add confirmation message for successful extractions, Add leaderboard for exploit submissions, Improve stateful guardrail evaluation across multi-turn sessions, Action-boundary enforcement instead of prompt-level guardrails

Competitors

Other products that read as similar to this one — 377 launches clear the similarity bar, closest 8 shown.

Attention rank: #85 of 378 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 132 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a dev tools tool for Sales yet.