Sparrow-2
Solving the cocktail party problem
Details
- External ID
- 49468613
- Source
- HN
- Company
- —
- Product
- Sparrow-2
- Website domain
- tavus.io
- Launched
- Aug. 27, 2026
- Cohort
- —
- Upvotes
- 8
- Upvotes percentile
- 0.4731182795698925
- Tags
- —
- Fetched at
- Sept. 10, 2026, 5:31 a.m.
- Updated at
- Sept. 10, 2026, 5:31 a.m.
Description
Hey there, I’m Brian. I've been shipping conversational models over here at Tavus for the past two years. I want to tell you about our new audio-understanding model: Sparrow-2! It’s a new category of model and a unique new approach to conversational audio.Earlier this year we launched Sparrow-1, (at the time) our SoTA turn taking model. Since our Sparrow-1 launch, I’ve spent a lot of time listening to humans talking and trying to really understand how people know when to talk, when to listen, and when to wait. I’ve also been hunting down failure modes of the current SoTA models. And while Sparrow-1 is great, there are some failure patterns I see. We tried solving the problems with existing approaches, but solving one problem only created another.Sparrow-1 and a lot of the existing turn taking models require noise cancellation to isolate the speaker’s audio from background noise. After removing the “noise” these models rely on simple prosodic and phonetic cues from spoken words to decide when a turn has ended. They throw out information and then pay attention to a small set of verbal and prosodic cues. This approach ignores a ton of important information.Non-verbal cues, sounds, and environmental noise impact turn taking. Noise cancellation models assume everything non-transcribable is noise, but that’s wrong. Humans use breath, sighs, and other sounds to hold the conversational floor or bid for a turn. There are constant micro-interruptions, affirmations, quiet human and environmental sounds that add to the conversational scenario. That sound is part of what we use to understand the nuanced space of the conversational floor! Cancelling out the “noise” has the ruinous side-effect of cancelling out the flow.Now, with Sparrow-2, instead of modelling just the primary speaker’s transcribable audio, we’ve been training a model on all the sounds and letting it decide what matters and what does not when it comes to conversational flow. Our new model continuously streams in audio and is able to understand semantics, prosody, timing, speaker identity, sighs, breaths, backchannels, interruptions, background speech, and unintelligible audio with the goal of understanding what the agent should do given the state of the audio.Sparrow-2 fits into our larger model family and unlocks new conversational behavior, not seen before in production-ready conversational pipelines. It can pass signals to different parts of the conversation. For example if the user is in a room that’s too noisy, the bot will ask the user to move to a quieter space. Also, Sparrow-2 is semi-duplex: it considers the timing of sound in relation to the AI speech as well as the user’s.Our main objective with Sparrow-2 is to finally crack the Cocktail Party Problem: how can we have a human-like conversation 1:1 with a user in noisy environments.I wrote up some technical details about the model architecture here: https://www.tavus.io/blog/sparrow-2I’d love for you guys to give it a try and let me know what you think and feel.
Enrichment
- Theme
- audio and signal processing tools
- Vertical
- Horizontal
- Function
- Dev tools
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Commercial product
- Normalized one-liner
- audio processing for speech separation
- Manually corrected
- False
Could you build this?
No This is a custom foundational audio-understanding AI model designed to solve whole-scene conversational audio and cocktail-party speaker separation, requiring advanced machine learning research and GPU cluster training.
What it would actually take: Building this requires large-scale multi-channel and multi-speaker audio training datasets with precise alignment, trained on high-performance compute clusters using PyTorch. The architecture involves deep neural speech separation, continuous audio encoders, and low-latency streaming transformer architectures. It demands a specialized team of speech-processing researchers and audio ML engineers.
Discussion
2 comments analyzed.
Competitors mentioned: gpt-oss
Competitors
Other products that read as similar to this one — 107 launches clear the similarity bar, closest 8 shown.
Attention rank: #66 of 108 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 301 days after the earliest competitor.
- Sparrow-2 · hn · 2026-09-08 · 11 upvotes · similarity 0.93
- Production duplex speech model for revenue calls · hn · 2026-07-23 · 14 upvotes · similarity 0.45
- Nari Qwen3-TTS and Qwen3-ASR · hn · 2026-09-14 · 90 upvotes · similarity 0.44
- Multimodal perception system for real-time conversation · hn · 2026-02-10 · 54 upvotes · similarity 0.44
- binaural-voice · github · 2026-09-24 · 85 upvotes · similarity 0.42
- I built a sub-500ms latency voice agent from scratch · hn · 2026-03-02 · 570 upvotes · similarity 0.41
- Real-time avatars that change emotions as you talk · hn · 2026-07-14 · 12 upvotes · similarity 0.41
- Yeittu · ph · 2026-09-24 · 1 upvotes · similarity 0.40
Other launches for this product
Same idea, different domain
Nobody's really built a dev tools tool for Sales yet.