Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

OSS Agent I built topped the TerminalBench on Gemini-3-flash-preview

Details

External ID
47920787
Source
HN
Company
—
Product
OSS Agent I built topped the TerminalBench on Gemini-3-flash-preview
Website domain
github.com
Launched
April 27, 2026
Cohort
—
Upvotes
393
Upvotes percentile
0.987146529562982
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Scored 65.2% vs google's official 47.8%, and the existing top closed source model Junie CLI's 64.3%.Since there are a lot of reports of deliberate cheating on TerminalBench 2.0 lately (https://debugml.github.io/cheating-agents/), I would like to also clarify a few things1. Absolutely no {agents/skills}.md files were inserted at any point. No cheating mechanisms whatsoever2. The cli agent was run in leaderboard compliant way (no modification of resources or timeouts)3. The full terminal bench run was done using the fully open source version of the agent, no difference between what is on github and what was run.I was originally going to wait for it to land on the leaderboard, but it has been 8 days and the maintainers do not respond unfortunately (there is a large backlog of the pull requests on their HF) so I decided to post anyways.HF PR: https://huggingface.co/datasets/harborframework/terminal-ben...It is astounding how much the harness matters, based on this and other experiments I have done.

Enrichment

Theme
ai coding agents and tooling
Vertical
Horizontal
Function
Agent / copilot
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
agent for terminal tasks
Manually corrected
False

Could you build this?

Yes A high-performing terminal benchmark agent is typically a structured Python CLI script executing a loop with dynamic prompt engineering, tool definitions for bash execution, and LLM API calls.

Discussion

20 comments analyzed.

Competitors mentioned: Qwen (local model alternative), OpenRouter (API provider), Codex (Claude alternative), OpenCode (plugin system reference)

Concerns raised: Retry mechanism fails on rate limiting (429 errors), session gets stuck, Generalization beyond structured benchmarks - performance on messy files and ambiguous tasks unclear, Lacks provider subscription support like Codex CLI, Missing plugin system like OpenCode

Feature requests: Fix retry mechanism for rate-limited endpoints, Add provider subscription support (not just API keys), Implement plugin system, Support for mixed local and cloud models (e.g., GPT with local Qwen)

Competitors

Other products that read as similar to this one — 159 launches clear the similarity bar, closest 8 shown.

Attention rank: #5 of 160 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 174 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a agent / copilot tool for Agriculture yet.