We built a tool to dub any video in the original voice in 40 languages
Details
- External ID
- 48433947
- Source
- HN
- Company
- —
- Product
- We built a tool to dub any video in the original voice in 40 languages
- Website domain
- vaani.media
- Launched
- June 7, 2026
- Cohort
- —
- Upvotes
- 8
- Upvotes percentile
- 0.4952185792349727
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
We kept seeing the same problem: creators make great content but the moment they dub it into another language, everything falls apart. Robotic voice. Music gone. Meaning lost. Lips not matching. So we built Vaani. Whatever language you create in, wherever you are in the world, your content can reach a global audience. We clone your voice, match your gestures, preserve your music, and optionally sync your lips to the new language. 10+ Indian languages, 20+ global languages, all in minutes. We built it for 2 reasons:So creators never have to sound like a robot in another language again So the same video you already made can reach a global audience without filming twiceWould love your feedback. app.vaani.media
Enrichment
- Theme
- ai translation and localization tools
- Vertical
- Media & entertainment
- Function
- Content generation
- Audience
- B2B
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- video dubbing in original voice for multiple languages
- Manually corrected
- False
Could you build this?
Partial While an orchestration app can be vibe-coded, multi-speaker voice cloning, scene-aware audio track separation/re-mixing, and realistic video lip-syncing require heavy ML pipelines and specialized audio-visual models.
What it would actually take: The architecture combines an audio processing pipeline (demucs/spleeter for background music separation, Whisper for ASR, LLMs for scene-aware translation) paired with expressive TTS cloning models (like XTTS/Bark or ElevenLabs API) and video lip-sync generative models (Wav2Lip or proprietary diffusion sync). The hard part is ensuring precise sync/pacing without drift across 40+ languages and preserving emotion, requiring heavy GPU compute and media engineering.
Discussion
5 comments analyzed.
Feature requests: Add more languages (75+ planned)
Competitors
Other products that read as similar to this one — 207 launches clear the similarity bar, closest 8 shown.
Attention rank: #99 of 208 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 221 days after the earliest competitor.
- Vaani · ph · 2026-06-08 · 318 upvotes · similarity 0.72
- FreeVoiceClone · ph · 2026-09-11 · 1 upvotes · similarity 0.52
- SublyAI · ph · 2026-09-25 · 1 upvotes · similarity 0.51
- Unmixr · ph · 2026-09-09 · 2 upvotes · similarity 0.50
- Visual Translate by Vozo · ph · 2026-03-10 · 718 upvotes · similarity 0.49
- Voxlate · ph · 2026-09-25 · 2 upvotes · similarity 0.49
- SpeakVid · ph · 2026-09-08 · 1 upvotes · similarity 0.47
- Familiar · yc · 2026-08-14 · 12 upvotes · similarity 0.47
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a content generation tool for Government yet.