Terminal-Wrench, a dataset of 331 realistic hackable environments
Details
- External ID
- 47773298
- Source
- HN
- Company
- —
- Product
- Terminal-Wrench, a dataset of 331 realistic hackable environments
- Website domain
- github.com
- Launched
- April 15, 2026
- Cohort
- —
- Upvotes
- 6
- Upvotes percentile
- 0.2808483290488432
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
I want to share a new dataset of 331 reward-hackable environments. These are real environments used in Terminal Bench and adjacent benchmarks. I first got interested in this because, as a reviewer of Terminal Bench, I noticed a lot of our tasks were hackable. I also noticed that many contributors to the benchmark do so because it provides credibility when selling environments to labs. Hence, TBench tasks are, in my opinion, held to a higher quality standard than those being used today for RL. No one is spending hours manually reviewing the $1B in tasks being purchased by major labs. As far as I understand, while everyone knows environments are hackable, nobody has released hundreds of "realistic" environments.
Enrichment
- Theme
- self-hosted infrastructure and security tools
- Vertical
- Security
- Function
- Dev tools
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- dataset of hacking environments
- Manually corrected
- False
Could you build this?
No This is a specialized, curated research benchmark of 331 verified reward-hackable terminal environments requiring deep AI safety and reinforcement learning evaluation expertise.
What it would actually take: Building this requires designing, containerizing (Docker), and thoroughly testing hundreds of real-world command-line task environments, each instrumented with specific verifiers and exploit vectors for reward hacking. It requires deep research expertise in RLHF, agent alignment, and benchmark design, as each environment must balance realistic task completion with detectable reward specification flaws.
Discussion
2 comments analyzed.
Competitors mentioned: Berkeley 100% hack
Concerns raised: Exploits may still work on more secure harness, Unclear differentiation from existing Berkeley research
Competitors
Other products that read as similar to this one — 18 launches clear the similarity bar, closest 8 shown.
Attention rank: #15 of 19 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 133 days after the earliest competitor.
- Fresh · hn · 2025-12-03 · 187 upvotes · similarity 0.40
- RewardHackBench: Using sandboxes to stop agents from cheating · hn · 2026-06-17 · 9 upvotes · similarity 0.40
- OSS Agent I built topped the TerminalBench on Gemini-3-flash-preview · hn · 2026-04-27 · 393 upvotes · similarity 0.39
- hacking.community · ph · 2026-09-09 · 1 upvotes · similarity 0.36
- Bucket · hn · 2026-01-25 · 9 upvotes · similarity 0.35
- SnapEnv · ph · 2026-09-14 · 3 upvotes · similarity 0.35
- KeyEnv · hn · 2026-01-18 · 5 upvotes · similarity 0.35
- Graft · hn · 2026-03-15 · 5 upvotes · similarity 0.35
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a dev tools tool for Sales yet.