Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Live Qwen3-Omni API (open-source speech-to-speech)

Details

External ID
46124900
Source
HN
Company
—
Product
Live Qwen3-Omni API (open-source speech-to-speech)
Website domain
hathora.dev
Launched
Dec. 2, 2025
Cohort
—
Upvotes
5
Upvotes percentile
0.10400763358778627
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

We just deployed Qwen3-Omni to production. As far as we know, this is the only place you can hit an open-source speech-to-speech model via playground with zero setup.The S2S landscape right now: OpenAI (GPT-Realtime), Hume (EVI), and now this. The first two are closed-source. Qwen3-Omni is open.What we built: real-time inference stack optimized for voice, deployed across multiple regions. You can test latency directly at the link.Honest take: we've seen faster results chaining ASR/LLM/TTS compared to native S2S. But the progress on end-to-end models in the last few months has been impressive, and we wanted to make it easy for people to experiment.Would love feedback from anyone who tries it, particularly on latency, voice quality, and where it breaks.

Enrichment

Theme
voice dictation and control tools
Vertical
—
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
open-source speech-to-speech api
Manually corrected
False

Could you build this?

No Deploying a low-latency, real-time speech-to-speech omni model requires massive GPU compute clusters, custom real-time audio streaming infrastructure, and deep ML systems engineering.

What it would actually take: To replicate this service, you need a high-end GPU cluster (e.g., NVIDIA H100s/A100s) with optimized inference runtimes (such as vLLM or custom TensorRT-LLM pipelines modified for multimodal audio streaming). The architecture requires sub-hundred-millisecond WebRTC/WebSocket full-duplex audio pipelining, voice activity detection (VAD), and turn-taking orchestration. Building this requires deep ML systems engineering, streaming audio protocol expertise, and significant capital for hardware infrastructure.

Discussion

No comments on this launch.

Competitors

Other products that read as similar to this one — 161 launches clear the similarity bar, closest 8 shown.

Attention rank: #154 of 162 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 26 days after the earliest competitor.

Other launches for this product