Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

I built a sub-500ms latency voice agent from scratch

Details

External ID
47224295
Source
HN
Company
—
Product
I built a sub-500ms latency voice agent from scratch
Website domain
ntik.me
Launched
March 2, 2026
Cohort
—
Upvotes
570
Upvotes percentile
0.997539975399754
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

I built a voice agent from scratch that averages ~400ms end-to-end latency (phone stop → first syllable). That’s with full STT → LLM → TTS in the loop, clean barge-ins, and no precomputed responses.What moved the needle:Voice is a turn-taking problem, not a transcription problem. VAD alone fails; you need semantic end-of-turn detection.The system reduces to one loop: speaking vs listening. The two transitions - cancel instantly on barge-in, respond instantly on end-of-turn - define the experience.STT → LLM → TTS must stream. Sequential pipelines are dead on arrival for natural conversation.TTFT dominates everything. In voice, the first token is the critical path. Groq’s ~80ms TTFT was the single biggest win.Geography matters more than prompts. Colocate everything or you lose before you start.GitHub Repo: https://github.com/NickTikhonov/shuoFollow whatever I next tinker with: https://x.com/nick_tikhonov

Enrichment

Theme
voice AI agents and infrastructure
Vertical
Horizontal
Function
Agent / copilot
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
low-latency conversational ai agent
Manually corrected
False

Could you build this?

Partial Wiring together STT, LLM, and TTS APIs is straightforward, but achieving sub-500ms voice turn-taking requires custom low-latency streaming pipeline orchestration, optimized VAD, and server co-location.

What it would actually take: The system requires a persistent WebRTC or bidirectional WebSocket audio pipeline, often in Go, Rust, or optimized Node.js, combined with a fast client-side/server-side Voice Activity Detection (VAD) model (like Silero VAD) to detect intent and handle instant barge-in cancellations. The critical bottleneck is end-to-end latency: chunking streaming STT (e.g., Deepgram), speculative LLM streaming with token-level phrase chunking, and streaming TTS (e.g., Cartesia/ElevenLabs), with minimal network jitter by co-locating servers in the same cloud region. Building this reliably at scale requires deep real-time systems programming and network audio streaming expertise.

Discussion

20 comments analyzed.

Competitors mentioned: Pipecat, Dograh, Sesame, Moshi, Twilio + ElevenLabs stack

Concerns raised: Barge-in cancellation timing issues across providers, Uncanny valley risk with mistimed filler words, Cultural differences in conversation turn-taking expectations, Language-specific challenges (verb position in Japanese/German)

Feature requests: Long-term memory/context between calls, Echo handling, Tool calls for external services, Semantic VAD improvements, Hardware wake-word models for low-power devices

Competitors

Other products that read as similar to this one — 137 launches clear the similarity bar, closest 8 shown.

Attention rank: #1 of 138 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 115 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a agent / copilot tool for Agriculture yet.