Multimodal perception system for real-time conversation
Details
- External ID
- 46965012
- Source
- HN
- Company
- —
- Product
- Multimodal perception system for real-time conversation
- Website domain
- tavuslabs.org
- Launched
- Feb. 10, 2026
- Cohort
- —
- Upvotes
- 54
- Upvotes percentile
- 0.8268194070080862
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
I work on real-time voice/video AI at Tavus and for the past few years, I’ve mostly focused on how machines respond in a conversation.One thing that’s always bothered me is that almost all conversational systems still reduce everything to transcripts, and throw away a ton of signals that need to be used downstream. Some existing emotion understanding models try to analyze and classify into small sets of arbitrary boxes, but they either aren’t fast / rich enough to do this with conviction in real-time.So I built a multimodal perception system which gives us a way to encode visual and audio conversational signals and have them translated into natural language by aligning a small LLM on these signals, such that the agent can "see" and "hear" you, and that you can interface with it via an OpenAI compatible tool schema in a live conversation.It outputs short natural language descriptions of what’s going on in the interaction - things like uncertainty building, sarcasm, disengagement, or even shift in attention of a single turn in a convo.Some quick specs:- Runs in real-time per conversation- Processing at ~15fps video + overlapping audio alongside the conversation- Handles nuanced emotions, whispers vs shouts- Trained on synthetic + internal convo dataHappy to answer questions or go deeper on architecture/tradeoffsMore details here: https://www.tavus.io/post/raven-1-bringing-emotional-intelli...
Enrichment
- Theme
- voice AI agents and infrastructure
- Vertical
- Horizontal
- Function
- Agent / copilot
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- multimodal perception for real-time conversation
- Manually corrected
- False
Could you build this?
No Building sub-second multimodal conversational perception that processes audio-visual streams directly without transcript reduction requires cutting-edge research and heavy GPU inference infrastructure.
What it would actually take: The system requires training custom multimodal neural networks that fuse continuous audio waveforms and video frames to detect sentiment, gaze, and interruptions in real time. The runtime relies on low-latency WebRTC media servers (e.g., LiveKit or custom C++ pipelines) streaming directly to GPU clusters running TensorRT-optimized models. It demands a specialized research team in computer vision, audio processing, and low-latency infrastructure.
Discussion
14 comments analyzed.
Concerns raised: Bias in emotion detection (e.g., loud speech labeled as 'shrill', poor camera angle misinterpreted as lack of eye contact), Computational cost at scale (30-minute interviews, 7-hour depositions), Risk of over-reliance on AI for subjective judgments in hiring/justice systems, Potential misuse in low-quality, insecure surveillance-adjacent products, Poor integration into workflows could create inequitable outcomes
Feature requests: Ability to adapt to individual speech patterns and camera positioning within single conversation, Sub-80ms latency for real-time results
Competitors
Other products that read as similar to this one — 311 launches clear the similarity bar, closest 8 shown.
Attention rank: #57 of 312 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 103 days after the earliest competitor.
- I built an AI conversation partner to practice speaking languages · hn · 2026-01-30 · 65 upvotes · similarity 0.57
- Real-time avatars that change emotions as you talk · hn · 2026-07-14 · 12 upvotes · similarity 0.57
- Realtime TTS-2 · ph · 2026-05-06 · 152 upvotes · similarity 0.55
- I built a voice AI that responds like a real woman · hn · 2026-03-25 · 6 upvotes · similarity 0.51
- KugelAudio: A multilingual voice AI model you can run in your own cluster. · yc · 2026-05-27 · 8 upvotes · similarity 0.50
- FaceTime-style calls with an AI Companion (Live2D and long-term memory) · hn · 2026-01-25 · 34 upvotes · similarity 0.49
- Realtime_Conversation_Video · github · 2026-09-27 · 51 upvotes · similarity 0.49
- LemonSlice · hn · 2026-01-27 · 133 upvotes · similarity 0.48
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a agent / copilot tool for Agriculture yet.