Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Omi

watches your screen, hears conversations, tells you what to do

Details

External ID
47784914
Source
HN
Company
—
Product
Omi
Website domain
github.com
Launched
April 15, 2026
Cohort
—
Upvotes
19
Upvotes percentile
0.7410025706940874
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Spent 4 months and built Omi for Desktop, your life architect: It sees your screen, hears your conversations and will advise you on what to do nextBasically Cluely + Rewind + Granola + Wisprflow + ChatGPT + Claude in one appI talk to claude/chatgpt 24/7 but I find it frustrating that i have to capture/send screenshots of my screen and that it doesn't help proactively during my workWhenever omi sees something wrong about my workflow, it will send me a proactive notification with advice. It will also point to something I'm missing.The hardest part was to nail proactivity - after trying 20+ similar tools I didn't find a single one with smart proactive notifications based on content on your screen. I made it look at your screen every second with 4 main prompts:1. Is the user productive or distracted?2. Is there anything useful to say right now?3. is there any task to add to do later?4. is there anything important to remember about the user?Full stack: - Swift - Rust backend - Deepgram transcription - Claude code for messaging - GPT 5.4 summaries - Gemini for embeddings and translationOpen source, stores screenshots locally, uses Claude Code for chat. Has cloud to sync with hardware or mobile app but can be disabled in settings

Enrichment

Theme
niche social and community platforms
Vertical
Horizontal
Function
Agent / copilot
Audience
B2C
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
ai assistant that monitors screen and audio
Manually corrected
False

Could you build this?

Partial While wrapping APIs like Claude and whisper is simple, continuous real-time multi-modal desktop capture (native screen frame capture, audio device loopback, OCR, and low-latency contextual processing) requires non-trivial native OS systems programming.

What it would actually take: A functional implementation requires a native desktop client (e.g., Swift/macOS ScreenCaptureKit and CoreAudio, or Rust/C++ cross-platform bindings) that captures audio and screen state efficiently without killing battery life. It needs a local pipeline for VAD (voice activity detection), on-device OCR or visual frame differencing, and a local or remote context-aggregation buffer feeding an LLM. The challenging parts are native OS permission sandboxing, zero-latency streaming audio loopback, and minimizing CPU/GPU load during continuous capture.

Discussion

13 comments analyzed.

Competitors mentioned: Pomodoro and todo list apps, ChatGPT, Claude

Concerns raised: Privacy/personal data capture concerns, Screen recording and text conversion performance/slowness, AI lacks understanding of user goals and context, Produces irrelevant tangential suggestions, Micromanagement/interruption disrupts flow state

Feature requests: Passive listening/transcription mode (previous version feature), Applications for educational use (courses, DIY, homework, surgery)

Competitors

Other products that read as similar to this one — 83 launches clear the similarity bar, closest 8 shown.

Attention rank: #25 of 84 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 147 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a agent / copilot tool for Agriculture yet.