Flakestorm
Chaos engineering for AI agents (local-first, open source)
Details
- External ID
- 46495434
- Source
- HN
- Company
- —
- Product
- —
- Website domain
- —
- Launched
- Jan. 5, 2026
- Cohort
- —
- Upvotes
- 6
- Upvotes percentile
- 0.2549407114624506
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Hi everyone,I’ve been working on an open-source tool called Flakestorm to test the reliability of AI agents before they hit production.Most agent testing today focuses on eval scores or happy-path prompts. In practice, agents tend to fail in more mundane ways: typos, tone shifts, long context, malformed input, or simple prompt injections — especially when running on smaller or local models. Flakestorm applies chaos-engineering ideas to agents. Instead of testing one prompt, it takes a “golden prompt”, generates adversarial mutations (semantic variations, noise, injections, encoding edge cases), runs them against your agent, and produces a robustness score plus a detailed HTML report showing what broke.Key points: Local-first (uses Ollama for mutation generation)Tested with Qwen / Gemma / other small models Works against HTTP agents, LangChain chains, or Python callables No cloud or API keys required This started as a way to debug my own agents after seeing them behave unpredictably under real user input. I’m still early and trying to understand how useful this is outside my own workflow.I’d really appreciate feedback on: Whether this overlaps with how you test agents today Failure modes you’ve seen that aren’t covered Whether “chaos testing for agents” is a useful framing, or if this should be thought of differently Repo: https://github.com/flakestorm/flakestorm Docs are admittedly long.Thanks for taking a look.
Enrichment
- Theme
- AI agent frameworks and developer tools
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI feature
- Project type
- Hobby / open-source project
- Normalized one-liner
- chaos testing for ai agents
- Manually corrected
- False
Could you build this?
Yes Flakestorm is an open-source test harness that mutates LLM inputs (adding typos, tone shifts, context limits) and runs assertions on agent responses.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 369 launches clear the similarity bar, closest 8 shown.
Attention rank: #255 of 370 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 63 days after the earliest competitor.
- KeelTest · hn · 2026-01-07 · 30 upvotes · similarity 0.50
- Fabraix · ph · 2026-05-08 · 196 upvotes · similarity 0.49
- Automated Testing for AI Agents · hn · 2026-03-06 · 8 upvotes · similarity 0.47
- We built an AI Agent to reproduce bugs · hn · 2026-04-14 · 12 upvotes · similarity 0.46
- Spec27 · hn · 2026-04-30 · 13 upvotes · similarity 0.46
- Agent Arena · hn · 2026-02-06 · 47 upvotes · similarity 0.45
- Agent Reliability Toolkit · ph · 2026-09-10 · 1 upvotes · similarity 0.44
- Fulcrum: The Agentic Debugger for AI Systems · yc · 2026-02-07 · 4 upvotes · similarity 0.43
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.