Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Three new Kitten TTS models

smallest less than 25MB

Details

External ID
47441546
Source
HN
Company
—
Product
Three new Kitten TTS models
Website domain
github.com
Launched
March 19, 2026
Cohort
—
Upvotes
561
Upvotes percentile
0.996309963099631
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Kitten TTS (https://github.com/KittenML/KittenTTS) is an open-source series of tiny and expressive text-to-speech models for on-device applications. We had a thread last year here: https://news.ycombinator.com/item?id=44807868.Today we're releasing three new models with 80M, 40M and 14M parameters.The largest model (80M) has the highest quality. The 14M variant reaches new SOTA in expressivity among similar sized models, despite being <25MB in size. This release is a major upgrade from the previous one and supports English text-to-speech applications in eight voices: four male and four female.Here's a short demo: https://www.youtube.com/watch?v=ge3u5qblqZA.Most models are quantized to int8 + fp16, and they use ONNX for runtime. Our models are designed to run anywhere eg. raspberry pi, low-end smartphones, wearables, browsers etc. No GPU required! This release aims to bridge the gap between on-device and cloud models for tts applications. Multi-lingual model release is coming soon.On-device AI is bottlenecked by one thing: a lack of tiny models that actually perform. Our goal is to open-source more models to run production-ready voice agents and apps entirely on-device.We would love your feedback!

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
text-to-speech models
Manually corrected
False

Could you build this?

No Developing tiny, high-quality, expressive TTS neural network architectures from scratch requires specialized deep learning research and high-performance audio synthesis expertise.

What it would actually take: Building novel sub-25MB TTS models requires deep expertise in audio DSP, neural vocoders (like HiFi-GAN or StyleTTS architectures), acoustic modeling, and extensive compute clusters. The team needs curated multi-speaker speech datasets, phonetic aligners, and rigorous loss functions balancing latency and perceptual quality. Quantization and pruning techniques are necessary to achieve fast on-device inference on mobile and edge chips.

Discussion

20 comments analyzed.

Competitors mentioned: Qwen3-TTS (1.7B model), Gemini TTS, Google Translate TTS, CopySpeak

Concerns raised: Inference latency on consumer GPUs for <25MB models unclear, Japanese TTS quality issues with pronunciation consistency and dataset mismatches, Language support limited to English only currently, Latency at 150ms per word becomes problematic for multi-sentence content

Feature requests: Support for tiny model size (beyond nano), Mini model tier, Multilingual support beyond English, Determinism information and stochastic generation options

Competitors

Other products that read as similar to this one — 109 launches clear the similarity bar, closest 8 shown.

Attention rank: #3 of 110 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 132 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.