Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

jj-benchmark

Evaluating AI agents on Jujutsu version control

Details

External ID
47352189
Source
HN
Company
—
Product
jj-benchmark
Website domain
github.io
Launched
March 12, 2026
Cohort
—
Upvotes
5
Upvotes percentile
0.1070110701107011
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Hi HN, Meng from TabbyML here.We decided to build this simply because we find Jujutsu (jj) really interesting, and many folks on our team have started trying it out recently. Since it introduces a very different workflow compared to traditional Git, we thought it would be a fun challenge to see how well current AI coding agents can actually use it.To build this, we created a semi-automated pipeline. We used AI to research the official Jujutsu documentation and websites, which then helped us bootstrap a dataset of 63 distinct evaluation tasks. Each task includes instructions, bootstrap scripts, and tests. We then ran the evaluations using the Harbor framework and our Pochi agent.Some interesting insights from our initial leaderboard:Claude 4.6 Sonnet is the clear winner: It achieved a 92% success rate (passing 58/63 tasks), beating out Opus and OpenAI's top models. It seems exceptionally good at parsing the novel CLI rules of jj. The Speed vs. Accuracy Trade-off: While GPT-5.4 sits at #5 with an 81% success rate, it is incredibly fast, averaging just 77.6s per task. In contrast, Gemini-3.1-pro achieved 84% but took over 3x as long (267.6s average). Open Weights / Regional Models are competitive: Models like Kimi-k2.5 (79%) put up a very respectable fight on a relatively niche tool. The benchmark isn't completely solved yet, but the fact that top models can successfully navigate a relatively new version control system by reasoning through the tasks is pretty exciting.If there are specific jj edge cases you think we should add to the dataset, feel free to open up a PR!

Enrichment

Theme
AI agent frameworks and developer tools
Vertical
Horizontal
Function
Observability & eval
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
benchmark for ai agents on jujutsu version control
Manually corrected
False

Could you build this?

Yes This is an automated evaluation benchmark and web leaderboard running test scripts against CLI tools, easily orchestratable with standard containerized runners and LLM API calls.

Discussion

2 comments analyzed.

Competitors mentioned: Claude Code, Harbor evaluation framework

Feature requests: Evaluate with jj-specific SKILL.md documentation, Support for terminal-based coding agents

Competitors

Other products that read as similar to this one — 156 launches clear the similarity bar, closest 8 shown.

Attention rank: #143 of 157 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 129 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a observability & eval tool for Media & entertainment yet.