Try Benzi
A coding harness/agent beating Claude Code itself on Sonnet
Details
- External ID
- 49226627
- Source
- HN
- Company
- —
- Product
- Try Benzi
- Website domain
- fly.dev
- Launched
- Aug. 8, 2026
- Cohort
- —
- Upvotes
- 11
- Upvotes percentile
- 0.6088709677419355
- Tags
- —
- Fetched at
- Sept. 10, 2026, 5:32 a.m.
- Updated at
- Sept. 10, 2026, 5:32 a.m.
Description
Hi y'all. Been working on something that should've been made a long time ago imo. It compiles codebases into O(1) hashmaps that the agent queries to discover the structure of your code/answer questions/write code.It also does complete static analysis checks on any writes the agent makes.Don't take my word for it though. Here are the benchmarks: https://benzi.fly.dev/benchmark. on 2/20 tests, Claude Code (mostly Sonnet on one task) regressed or timed out. Benzi didn't because of course, it has a map it can query and not get lost in the sauce. On the other 18 it is cheaper, faster, or often both.Would love to get some early adoption and criticism!(Only available on Windows for now. soz. and also keep an eye on the benchmarking page; I think I can push it far more -- no promises, still work in progress. Haven't thoroughly tested greenfielding experience either.)(another note: VS Code extension/website is running DeepSeek V4 flash. not Sonnet. Everything mostly built with CC Sonnet tho )
Enrichment
- Theme
- Claude integrations and coding agents
- Vertical
- Horizontal
- Function
- Agent / copilot
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- coding agent outperforming claude
- Manually corrected
- False
Could you build this?
Partial While an agent runner harness can be vibe coded, building a full incremental compiler that parses entire arbitrary codebases into exact O(1) queryable symbol graphs and executes deep static analysis on agent diffs requires dedicated language tooling expertise.
What it would actually take: Requires tree-sitter or language-server protocol (LSP) AST parsers across multiple languages integrated into an in-memory graph/hash index for scope, references, and definitions. It also requires integrating language-specific static analyzers/linters into an automated validation loop. A developer needs compilers/static analysis expertise and robust handling of language edge cases.
Discussion
3 comments analyzed.
Competitors mentioned: Claude Code, DeepSeek Harness, OpenCode
Concerns raised: Benchmark score parity with DeepSeek (78% vs 78.6%), SWE-bench Verified ranking (#4 on leaderboard)
Competitors
Other products that read as similar to this one — 477 launches clear the similarity bar, closest 8 shown.
Attention rank: #186 of 478 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 281 days after the earliest competitor.
- 1Code · hn · 2026-01-15 · 75 upvotes · similarity 0.52
- Cockpit for you Claude Code agents in Rust · hn · 2026-08-01 · 13 upvotes · similarity 0.49
- Simple plugin to get Claude Code to listen to you · hn · 2026-03-13 · 18 upvotes · similarity 0.49
- Total Recall · hn · 2026-02-05 · 67 upvotes · similarity 0.49
- Abralo · hn · 2026-07-08 · 37 upvotes · similarity 0.48
- I built simple and efficient local memory system for Claude Code · hn · 2026-02-02 · 6 upvotes · similarity 0.47
- Claude Code for Visual Studio (native diff with accept/reject) · hn · 2026-06-15 · 20 upvotes · similarity 0.46
- Nori CLI, a better interface for Claude Code (no flicker) · hn · 2026-01-14 · 37 upvotes · similarity 0.46
Other launches for this product
Same idea, different domain
Nobody's really built a agent / copilot tool for Agriculture yet.