Real-time AI (audio/video in, voice out) on an M3 Pro with Gemma E2B
Details
- External ID
- 47652007
- Source
- HN
- Company
- —
- Product
- Real-time AI (audio/video in, voice out) on an M3 Pro with Gemma E2B
- Website domain
- github.com
- Launched
- April 5, 2026
- Cohort
- —
- Upvotes
- 298
- Upvotes percentile
- 0.9781491002570694
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Related: https://news.ycombinator.com/item?id=47653752
Enrichment
- Theme
- ai video generation and editing tools
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- real-time audio and video ai inference on m3 pro
- Manually corrected
- False
Could you build this?
No Real-time multimodal processing (audio and video in, streaming voice out) on local Apple Silicon requires low-latency Metal/MPS kernel optimizations, real-time audio pipeline synchronization, and local multimodal model inference.
What it would actually take: The architecture requires an optimized C++/Metal/MLX pipeline to run quantized multimodal models (like Gemma) alongside low-latency audio capture/streaming (AVFoundation, WebRTC), speech-to-text, and zero-latency neural TTS. The hardest part is streaming concurrent audio/video frame ingestion into a single autoregressive context window while keeping round-trip conversational latency under 500ms on consumer hardware. This requires deep systems programming, real-time DSP, and GPU-level inference optimization expertise.
Discussion
20 comments analyzed.
Competitors mentioned: Google Home voice assistant, Apple Siri, Gemma E2B, Google Gemma4
Concerns raised: Response time too slow for real-time use (2.5s+ latency), Voice recognition performance inadequate on consumer hardware (M1 Max, RTX 5060 Ti), Requires internet connection to function, Image processing adds ~0.5s overhead, blocking video real-time capability, Video understanding limited to frame snapshots, not temporal context
Feature requests: Reduce response time to sub-100ms for true real-time, Disable image input option to improve speed, Add filler speech sounds ("uh", "umm") during processing, Real-time video understanding with temporal context awareness
Competitors
Other products that read as similar to this one — 90 launches clear the similarity bar, closest 8 shown.
Attention rank: #2 of 91 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 153 days after the earliest competitor.
- VIDEO AI ME · ph · 2026-04-27 · 206 upvotes · similarity 0.42
- SpeechifyAI Simba Voice Agents · ph · 2026-07-13 · 161 upvotes · similarity 0.41
- Lucy Labs · ph · 2026-09-22 · 2 upvotes · similarity 0.41
- FaceTime-style calls with an AI Companion (Live2D and long-term memory) · hn · 2026-01-25 · 34 upvotes · similarity 0.40
- Multimodal perception system for real-time conversation · hn · 2026-02-10 · 54 upvotes · similarity 0.40
- MotorCast AI · ph · 2026-09-29 · 1 upvotes · similarity 0.39
- Realtime TTS-2 · ph · 2026-05-06 · 152 upvotes · similarity 0.39
- I built a voice AI that responds like a real woman · hn · 2026-03-25 · 6 upvotes · similarity 0.39
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.