Moonshine Open-Weights STT models
higher accuracy than WhisperLargev3
Details
- External ID
- 47143755
- Source
- HN
- Company
- —
- Product
- Moonshine Open-Weights STT models
- Website domain
- github.com
- Launched
- Feb. 24, 2026
- Cohort
- —
- Upvotes
- 316
- Upvotes percentile
- 0.977088948787062
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
I wanted to share our new speech to text model, and the library to use them effectively. We're a small startup (six people, sub-$100k monthly GPU budget) so I'm proud of the work the team has done to create streaming STT models with lower word-error rates than OpenAI's largest Whisper model. Admittedly Large v3 is a couple of years old, but we're near the top the HF OpenASR leaderboard, even up against Nvidia's Parakeet family. Anyway, I'd love to get feedback on the models and software, and hear about what people might build with it.
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- open-weight speech-to-text models
- Manually corrected
- False
Could you build this?
No Training frontier-grade speech-to-text models that outperform Whisper Large v3 requires tens of thousands of hours of curated audio datasets, extensive GPU cluster compute, and deep ML research expertise.
What it would actually take: Requires an end-to-end ASR deep learning pipeline implemented in PyTorch/JAX, using conformer or transformer architectures trained on tens of thousands of hours of diverse, multi-speaker audio data across hundreds of GPUs. The difficult challenges are custom acoustic modeling, latency-optimized streaming attention mechanisms, and dataset curation/alignment, requiring dedicated speech research scientists and significant capital for compute.
Discussion
20 comments analyzed.
Competitors mentioned: Whisper, Moonshine, Web Speech API, MacWhisper, VoiceInk
Concerns raised: Lacks streaming support (unlike Moonshine), Slower and less accurate than Parakeet v3 on older Intel CPUs, Model size/memory constraints for edge devices, Battery drain and heat on mobile devices
Feature requests: Add streaming transcription capability, Multilingual performance improvements, M-series and Jetson device optimization, Transformers.js / WebGPU port compatibility
Competitors
Other products that read as similar to this one — 144 launches clear the similarity bar, closest 8 shown.
Attention rank: #15 of 145 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 112 days after the earliest competitor.
- Serve 100 Large AI models on a single GPU with low impact to TTFT · hn · 2025-11-08 · 7 upvotes · similarity 0.51
- Nari Qwen3-TTS and Qwen3-ASR · hn · 2026-09-14 · 90 upvotes · similarity 0.48
- Three new Kitten TTS models · hn · 2026-03-19 · 561 upvotes · similarity 0.47
- Live Qwen3-Omni API (open-source speech-to-speech) · hn · 2025-12-02 · 5 upvotes · similarity 0.47
- Run open-weight OCR, VLM and vision models behind one API · hn · 2026-09-04 · 5 upvotes · similarity 0.47
- OpenGraviton · hn · 2026-03-07 · 13 upvotes · similarity 0.46
- Audio AI had a wild day · hn · 2026-01-23 · 5 upvotes · similarity 0.44
- Run 500B+ Parameter LLMs Locally on a Mac Mini · hn · 2026-03-09 · 17 upvotes · similarity 0.43
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.