Live Qwen3-Omni API (open-source speech-to-speech)
Details
- External ID
- 46124900
- Source
- HN
- Company
- —
- Product
- Live Qwen3-Omni API (open-source speech-to-speech)
- Website domain
- hathora.dev
- Launched
- Dec. 2, 2025
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.10400763358778627
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
We just deployed Qwen3-Omni to production. As far as we know, this is the only place you can hit an open-source speech-to-speech model via playground with zero setup.The S2S landscape right now: OpenAI (GPT-Realtime), Hume (EVI), and now this. The first two are closed-source. Qwen3-Omni is open.What we built: real-time inference stack optimized for voice, deployed across multiple regions. You can test latency directly at the link.Honest take: we've seen faster results chaining ASR/LLM/TTS compared to native S2S. But the progress on end-to-end models in the last few months has been impressive, and we wanted to make it easy for people to experiment.Would love feedback from anyone who tries it, particularly on latency, voice quality, and where it breaks.
Enrichment
- Theme
- voice dictation and control tools
- Vertical
- —
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- open-source speech-to-speech api
- Manually corrected
- False
Could you build this?
No Deploying a low-latency, real-time speech-to-speech omni model requires massive GPU compute clusters, custom real-time audio streaming infrastructure, and deep ML systems engineering.
What it would actually take: To replicate this service, you need a high-end GPU cluster (e.g., NVIDIA H100s/A100s) with optimized inference runtimes (such as vLLM or custom TensorRT-LLM pipelines modified for multimodal audio streaming). The architecture requires sub-hundred-millisecond WebRTC/WebSocket full-duplex audio pipelining, voice activity detection (VAD), and turn-taking orchestration. Building this requires deep ML systems engineering, streaming audio protocol expertise, and significant capital for hardware infrastructure.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 161 launches clear the similarity bar, closest 8 shown.
Attention rank: #154 of 162 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 26 days after the earliest competitor.
- Nari Qwen3-TTS and Qwen3-ASR · hn · 2026-09-14 · 90 upvotes · similarity 0.66
- Audio AI had a wild day · hn · 2026-01-23 · 5 upvotes · similarity 0.48
- GPT-Live · ph · 2026-07-09 · 211 upvotes · similarity 0.47
- Moonshine Open-Weights STT models · hn · 2026-02-24 · 316 upvotes · similarity 0.47
- Lightning V3 · ph · 2026-04-02 · 321 upvotes · similarity 0.45
- Parrot Speech-to-text API · ph · 2026-05-26 · 194 upvotes · similarity 0.43
- We open sourced Vapi · hn · 2026-03-12 · 8 upvotes · similarity 0.43
- Fastest Qwen 3.8 27 on single RTX5090 · ph · 2026-09-07 · 1 upvotes · similarity 0.43
Other launches for this product
- No other launches for this product.