OSS Agent I built topped the TerminalBench on Gemini-3-flash-preview
Details
- External ID
- 47920787
- Source
- HN
- Company
- —
- Product
- OSS Agent I built topped the TerminalBench on Gemini-3-flash-preview
- Website domain
- github.com
- Launched
- April 27, 2026
- Cohort
- —
- Upvotes
- 393
- Upvotes percentile
- 0.987146529562982
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Scored 65.2% vs google's official 47.8%, and the existing top closed source model Junie CLI's 64.3%.Since there are a lot of reports of deliberate cheating on TerminalBench 2.0 lately (https://debugml.github.io/cheating-agents/), I would like to also clarify a few things1. Absolutely no {agents/skills}.md files were inserted at any point. No cheating mechanisms whatsoever2. The cli agent was run in leaderboard compliant way (no modification of resources or timeouts)3. The full terminal bench run was done using the fully open source version of the agent, no difference between what is on github and what was run.I was originally going to wait for it to land on the leaderboard, but it has been 8 days and the maintainers do not respond unfortunately (there is a large backlog of the pull requests on their HF) so I decided to post anyways.HF PR: https://huggingface.co/datasets/harborframework/terminal-ben...It is astounding how much the harness matters, based on this and other experiments I have done.
Enrichment
- Theme
- ai coding agents and tooling
- Vertical
- Horizontal
- Function
- Agent / copilot
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- agent for terminal tasks
- Manually corrected
- False
Could you build this?
Yes A high-performing terminal benchmark agent is typically a structured Python CLI script executing a loop with dynamic prompt engineering, tool definitions for bash execution, and LLM API calls.
Discussion
20 comments analyzed.
Competitors mentioned: Qwen (local model alternative), OpenRouter (API provider), Codex (Claude alternative), OpenCode (plugin system reference)
Concerns raised: Retry mechanism fails on rate limiting (429 errors), session gets stuck, Generalization beyond structured benchmarks - performance on messy files and ambiguous tasks unclear, Lacks provider subscription support like Codex CLI, Missing plugin system like OpenCode
Feature requests: Fix retry mechanism for rate-limited endpoints, Add provider subscription support (not just API keys), Implement plugin system, Support for mixed local and cloud models (e.g., GPT with local Qwen)
Competitors
Other products that read as similar to this one — 159 launches clear the similarity bar, closest 8 shown.
Attention rank: #5 of 160 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 174 days after the earliest competitor.
- Cheddar-bench · hn · 2026-02-22 · 9 upvotes · similarity 0.43
- RewardHackBench: Using sandboxes to stop agents from cheating · hn · 2026-06-17 · 9 upvotes · similarity 0.43
- Continue · hn · 2026-02-17 · 44 upvotes · similarity 0.42
- You can now run Gemini CLI in the browser · hn · 2026-04-30 · 5 upvotes · similarity 0.41
- Lessons learned from running Claude Code swarms at scale · hn · 2026-06-05 · 10 upvotes · similarity 0.41
- Agents, run any coding agent on your subscription not API costs · hn · 2026-05-31 · 6 upvotes · similarity 0.41
- AI agents run my one-person company on Gemini's free tier · hn · 2026-03-08 · 16 upvotes · similarity 0.40
- Halo · hn · 2026-07-07 · 37 upvotes · similarity 0.40
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a agent / copilot tool for Agriculture yet.