Clippy
screen-aware voice AI in the browser
Details
- External ID
- 47429822
- Source
- HN
- Company
- —
- Product
- Clippy
- Website domain
- rememberclippy.com
- Launched
- March 18, 2026
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.1070110701107011
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
A friend and I built a browser prototype that answers questions about whatever’s on your screen using getDisplayMedia, client-side wake-word detection, and server-side multimodal inference.Hard parts:– Getting the model to point to specific UI elements– Keeping it coherent across multi-step workflows (“Help me create a sword in Tinkercad”)– Preventing the infinite mirror effect and confusion between window vs full-screen sharing– Keeping voice → screenshot → inference → voice latency low enough to feel conversationalWe packaged it as “Clippy” for fun, but the real experiment is letting a model tool-call fresh screenshots to help it gather more context.One practical use case is remote tech support — I'm sending this to my mom next time she calls instead of screen sharing.Curious what breaks.
Enrichment
- Theme
- voice AI agents and infrastructure
- Vertical
- Horizontal
- Function
- Agent / copilot
- Audience
- B2C
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- voice ai assistant aware of screen context
- Manually corrected
- False
Could you build this?
Partial While standard Web APIs handle media capture, grounding real-time multimodal model responses to precise dynamic UI screen coordinates with low latency across complex user interactions exceeds trivial vibe-coding.
What it would actually take: The stack uses WebRTC/LiveKit for real-time video/audio streaming, Porcupine or WebAssembly for client-side wake-word detection, and vision models (like Gemini Live or GPT-4o Realtime) for inference. The hardest problems are spatial grounding (mapping screen bounding boxes to DOM elements or mouse coordinate overlays in real-time) and maintaining reliable conversational state across multi-step browser interactions.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 115 launches clear the similarity bar, closest 8 shown.
Attention rank: #110 of 116 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 132 days after the earliest competitor.
- Assist · ph · 2026-09-07 · 128 upvotes · similarity 0.43
- FaceTime-style calls with an AI Companion (Live2D and long-term memory) · hn · 2026-01-25 · 34 upvotes · similarity 0.43
- Voice-tracked teleprompter using on-device ASR in the browser · hn · 2026-03-15 · 6 upvotes · similarity 0.43
- Multimodal perception system for real-time conversation · hn · 2026-02-10 · 54 upvotes · similarity 0.39
- Hitoku Draft · hn · 2026-06-04 · 21 upvotes · similarity 0.39
- ClipCast · ph · 2026-09-22 · 1 upvotes · similarity 0.38
- Caddy: The Voice Interface for Your Computer · yc · 2025-11-06 · 149 upvotes · similarity 0.38
- Voice-first todo list that updates live as you talk · hn · 2025-12-27 · 5 upvotes · similarity 0.38
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a agent / copilot tool for Agriculture yet.