Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Real-time AI (audio/video in, voice out) on an M3 Pro with Gemma E2B

Details

External ID
47652007
Source
HN
Company
—
Product
Real-time AI (audio/video in, voice out) on an M3 Pro with Gemma E2B
Website domain
github.com
Launched
April 5, 2026
Cohort
—
Upvotes
298
Upvotes percentile
0.9781491002570694
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Related: https://news.ycombinator.com/item?id=47653752

Enrichment

Theme
ai video generation and editing tools
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
real-time audio and video ai inference on m3 pro
Manually corrected
False

Could you build this?

No Real-time multimodal processing (audio and video in, streaming voice out) on local Apple Silicon requires low-latency Metal/MPS kernel optimizations, real-time audio pipeline synchronization, and local multimodal model inference.

What it would actually take: The architecture requires an optimized C++/Metal/MLX pipeline to run quantized multimodal models (like Gemma) alongside low-latency audio capture/streaming (AVFoundation, WebRTC), speech-to-text, and zero-latency neural TTS. The hardest part is streaming concurrent audio/video frame ingestion into a single autoregressive context window while keeping round-trip conversational latency under 500ms on consumer hardware. This requires deep systems programming, real-time DSP, and GPU-level inference optimization expertise.

Discussion

20 comments analyzed.

Competitors mentioned: Google Home voice assistant, Apple Siri, Gemma E2B, Google Gemma4

Concerns raised: Response time too slow for real-time use (2.5s+ latency), Voice recognition performance inadequate on consumer hardware (M1 Max, RTX 5060 Ti), Requires internet connection to function, Image processing adds ~0.5s overhead, blocking video real-time capability, Video understanding limited to frame snapshots, not temporal context

Feature requests: Reduce response time to sub-100ms for true real-time, Disable image input option to improve speed, Add filler speech sounds ("uh", "umm") during processing, Real-time video understanding with temporal context awareness

Competitors

Other products that read as similar to this one — 90 launches clear the similarity bar, closest 8 shown.

Attention rank: #2 of 91 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 153 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.