Voice gender classifier for European voice AI (1MB, ONNX, 4ms)
Details
- External ID
- 48107558
- Source
- HN
- Company
- —
- Product
- Voice gender classifier for European voice AI (1MB, ONNX, 4ms)
- Website domain
- huggingface.co
- Launched
- May 12, 2026
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.1147011308562197
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Hi, I'm Kamil and I'm a founder of Applied AI agency in Warsaw, Poland.We've trained a small <1MB voice classifier model that runs on CPU in 4ms. Can be run next to silero VAD in voice AI deployments.What we noticed in production deployments of voice assistants in Contact Centers in EU is that human consultants pick up immediately how to inflect verbs and ajdectives after one utterance from the caller. But voice AI agents don't know it until 1-2 minutes into the call when they are either corrected or the caller uses explicitly words with male/female form a couple of times.Our model solves this just from the first utterance of caller speech. The voice AI pipeline can inject the classification as context to the system prompt. We observed a significant impact of this on the adoption of voice AI in practice.Model + paper: https://huggingface.co/syntropicsignal-ai/gender-voice-class...
Enrichment
- Theme
- ai agents for calls and meetings
- Vertical
- Media & entertainment
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- voice gender classification model
- Manually corrected
- False
Could you build this?
No Training an ultra-compact (<1MB), low-latency (4ms) Bi-LSTM audio classifier across multiple European languages requires specialized DSP, speech dataset curation, and deep acoustic ML training pipelines.
What it would actually take: Building this requires acoustic DSP preprocessing (spectrograms, MFCCs/FBanks), curated and balanced multilingual speech corpora (e.g., Common Voice, FLEURS, LibriSpeech with gender labels), and PyTorch model architecture exploration (small Bi-LSTM or 1D-CNN) constrained under 166k parameters. It demands quantization and export to ONNX Runtime with zero-copy C++/Python bindings to achieve <5ms single-threaded CPU execution, requiring audio ML and systems engineering expertise.
Discussion
2 comments analyzed.
Competitors
Other products that read as similar to this one — 273 launches clear the similarity bar, closest 8 shown.
Attention rank: #250 of 274 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 190 days after the earliest competitor.
- KugelAudio: A multilingual voice AI model you can run in your own cluster. · yc · 2026-05-27 · 8 upvotes · similarity 0.53
- ai-voice-agent · github · 2026-09-19 · 8 upvotes · similarity 0.51
- Speak Chinese: AI Voice Chat · ph · 2026-09-17 · 1 upvotes · similarity 0.48
- Sancharya AI · ph · 2026-09-30 · 1 upvotes · similarity 0.46
- Vozon · ph · 2026-09-20 · 1 upvotes · similarity 0.46
- TransVoice Speech API · ph · 2026-09-16 · 1 upvotes · similarity 0.46
- Realtime TTS-2 · ph · 2026-05-06 · 152 upvotes · similarity 0.45
- Three new Kitten TTS models · hn · 2026-03-19 · 561 upvotes · similarity 0.45
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.