Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Nari Qwen3-TTS and Qwen3-ASR

High accuracy, low latency and cost

Details

External ID
49699267
Source
HN
Company
—
Product
Nari Qwen3-TTS and Qwen3-ASR
Website domain
narilabs.com
Launched
Sept. 14, 2026
Cohort
—
Upvotes
90
Upvotes percentile
0.9130781499202552
Tags
—
Fetched at
Sept. 18, 2026, 5:02 p.m.
Updated at
Sept. 18, 2026, 5:02 p.m.

Description

Hey HN, Toby from Nari Labs here.We've been working on making OSS speech models super-fast. Last year, we built Dia, the first OSS text-to-speech model capable of doing natural dialogue. Since then, so many more great speech models have been released to the public.But the market is still dominated by closed source models. We think that's an inference problem. Existing systems such as vLLM / SGLang are not well suited for multimodal inference. To prove this, we built an inference engine specialized for Qwen3-TTS and open-sourced it (https://github.com/nari-labs/nari-qwen3-tts). Running at sub-50 ms latency at 10 RPS, this showed open models can be run much faster and cheaper.Since then, we've been working hard to bring cheap, fast, and high quality serving to all. And we've even beat closed models at their game!Measured on the highly cited Coval (YC S24) voice AI benchmarks, our Qwen3-TTS endpoint not just is #2 in latency, but #1 in accuracy (WER) compared to 11Labs, Cartesia etc. while being the cheapest endpoint. Our Qwen3-ASR endpoint has the lowest latency and #2 accuracy, just 0.1% away from #1. It is the second cheapest model on the list.It took a lot of clever inference engineering to make these models quick, perform well while keeping costs low. Interestingly, Alibaba's official endpoints seem to perform worse in terms of accuracy and latency compared to ours. But nonetheless, much love to the Qwen team for OSS-ing these amazing speech models.We want to continue to push prices down to make speech technology a commodity - so that every app can have great TTS and STT without worrying about unit costs. We're also working on other parts of audio such as diarization - as well as video and world model inference. More to come!

Enrichment

Theme
voice AI agents and infrastructure
Vertical
Media & entertainment
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
speech-to-text and text-to-speech models for developers
Manually corrected
False

Could you build this?

No Achieving state-of-the-art TTS/ASR latency and WER requires deep ML research, specialized model architecture tuning, and custom GPU inference kernels.

What it would actually take: Building this requires a dedicated team of audio/speech ML researchers and systems engineers. The stack involves fine-tuning custom 1.7B parameter speech models, compiling custom CUDA/Triton kernels for ultra-low time-to-first-audio (TTFA), and managing high-throughput distributed GPU clusters with streaming audio websockets.

Discussion

19 comments analyzed.

Competitors mentioned: Whisper, AssemblyAI, loudkit

Concerns raised: TTS generations play too fast, voice switches mid-clip, lack of independent quality evals, lack of demos, ASR inference repo not open source

Feature requests: open source ASR inference engine, provide audio demos, independent evaluations

Competitors

Other products that read as similar to this one — 277 launches clear the similarity bar, closest 8 shown.

Attention rank: #38 of 278 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 312 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.