AutoBot
live voice control for long-running AI work
Details
- External ID
- 49743478
- Source
- HN
- Company
- —
- Product
- AutoBot
- Website domain
- github.com
- Launched
- Sept. 17, 2026
- Cohort
- —
- Upvotes
- 20
- Upvotes percentile
- 0.7440191387559809
- Tags
- —
- Fetched at
- Sept. 21, 2026, 5:02 p.m.
- Updated at
- Sept. 21, 2026, 5:02 p.m.
Description
I wanted to manage long horizon agentic workstreams via voice, then put my phone down, and have a harness manage completion - extending into full computer use.I was trying to build a personal Jarvis, so I benchmarked AutoBot to see how close I could get:- OSWorld: 32.41% (moved Sol Max from 4th to 1st, beating Opus 5)- AssistantBench: 50.70%Hermes and OpenClaw, but without needing a weekend and VMs. And with the ability to manage deep personalization AND keep strict privacy rules, storing data on encrypted disk.AutoBot is a passion project that grew out of trying to make this all work for myself.I’m sharing it here because I suspect other here have the same frustration. And I’m curious what else everyone is doing for this.It’s an MIT-licensed harness that lives within a project.Native voice lets me discuss tasks, check progress, and steer work; a local ledger tracks unfinished outputs and the evidence needed to call them done. Memory drives more autonomy over time, and defrags and locks in learning nightly while a heartbeat system persists execution.I’m not selling anything. If you’re building something similar for yourself, I’d love to compare notes.GitHub: https://github.com/demeyer1/Autobot
Enrichment
- Theme
- AI agent frameworks and developer tools
- Vertical
- Horizontal
- Function
- Communication
- Audience
- Prosumer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- voice interface for long-running ai tasks
- Manually corrected
- False
Could you build this?
Partial A basic voice-to-agent interface is straightforward to vibe-code, but achieving state-of-the-art OSWorld benchmark results for long-horizon autonomous computer use requires advanced agent scaffolding.
What it would actually take: The architecture needs a low-latency WebSockets voice pipeline (STT/TTS) connected to a multi-modal agent loop operating in an isolated virtual desktop sandbox (VNC/Docker). To achieve top OSWorld scores (30%+), the harness requires sophisticated screen parsing (grounding models, DOM/accessibility trees), hierarchical planning, state verification, and robust backtrack/retry recovery heuristics.
Discussion
3 comments analyzed.
Competitors mentioned: Claude Code
Concerns raised: Prompt injection risks when giving full permissions, Endless approval and safety prompts interrupt workflows
Competitors
Other products that read as similar to this one — 147 launches clear the similarity bar, closest 8 shown.
Attention rank: #36 of 148 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 314 days after the earliest competitor.
- We open sourced Vapi · hn · 2026-03-12 · 8 upvotes · similarity 0.45
- eBook to audiobook narration with realistic AI voices · hn · 2026-06-24 · 8 upvotes · similarity 0.44
- Voice-first todo list that updates live as you talk · hn · 2025-12-27 · 5 upvotes · similarity 0.43
- I built a sub-500ms latency voice agent from scratch · hn · 2026-03-02 · 570 upvotes · similarity 0.41
- Computer Agents · hn · 2026-03-01 · 7 upvotes · similarity 0.41
- ClawdTalk: Voice Calls for ClawdBots · hn · 2026-02-09 · 20 upvotes · similarity 0.40
- VoxConvo · hn · 2025-11-07 · 10 upvotes · similarity 0.40
- Ava · hn · 2026-03-12 · 7 upvotes · similarity 0.39
Other launches for this product
- No other launches for this product.