Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Ourguide

OS wide task guidance system that shows you where to click

Details

External ID
46769422
Source
HN
Company
—
Product
Ourguide
Website domain
ourguide.ai
Launched
Jan. 26, 2026
Cohort
—
Upvotes
52
Upvotes percentile
0.8023715415019763
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Hey! I'm eshaan and I'm building Ourguide -an on-screen task guidance system that can show you where to click step-by-step when you need help.I started building this because whenever I didn’t know how to do something on my computer, I found myself constantly tabbing between chatbots and the app, pasting screenshots, and asking “what do I do next?” Ourguide solves this with two modes. In Guide mode, the app overlays your screen and highlights the specific element to click next, eliminating the need to leave your current window. There is also Ask mode, which is a vision-integrated chat that captures your screen context—which you can toggle on and off anytime -so you can ask, "How do I fix this error?" without having to explain what "this" is.It’s an Electron app that works OS-wide, is vision-based, and isn't restricted to the browser.Figuring out how to show the user where to click was the hardest part of the process. I originally trained a computer vision model with 2300 screenshots to identify and segment all UI elements on a screen and used a VLM to find the correct icon to highlight. While this worked extremely well—better than SOTA grounding models like UI Tars—the latency was just too high. I'll be making that CV+VLM pipeline OSS soon, but for now, I’ve resorted to a simpler implementation that achieves <1s latency.You may ask: if I can show you where to click, why can't I just click too? While trying to build computer-use agents during my job in Palo Alto, I hit the core limitation of today’s computer-use models where benchmarks hover in the mid-50% range (OSWorld). VLMs often know what to do but not what it looks like; without reliable visual grounding, agents misclick and stall. So, I built computer use—without the "use." It provides the visual grounding of an agent but keeps the human in the loop for the actual execution to prevent misclicks.I personally use it for the AWS Console's "treasure hunt" UI, like creating a public S3 bucket with specific CORS rules. It’s also been surprisingly helpful for non-technical tasks, like navigating obscure settings in Gradescope or Spotify. Ourguide really works for any task when you’re stuck or don't know what to do.You can download and test Ourguide here: https://ourguide.ai/downloadsThe project is still very early, and I’d love your feedback on where it fails, where you think it worked well, and which specific niches you think Ourguide would be most helpful for.

Enrichment

Theme
browser automation and scraping for AI
Vertical
Horizontal
Function
Agent / copilot
Audience
B2C
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
ai guidance system showing where to click
Manually corrected
False

Could you build this?

Partial The basic Electron overlay and screen-capture pipeline can be vibe coded, but robust desktop-wide GUI element detection and spatial coordinate mapping across arbitrary OS apps require specialized computer vision models.

What it would actually take: The architecture involves a native desktop client (e.g., Swift on macOS or Electron with native C++ node addons) utilizing OS accessibility APIs (AXUIElement) and low-latency screen capture via ScreenCaptureKit. The hard problem is reliably grounding LLM action intents onto dynamic screen pixels when accessibility trees are incomplete or missing, requiring fine-tuned multimodal grounding models (e.g., UI-VLM / SeeClick / OmniParser). Building this requires expertise in operating system accessibility hooks, low-latency video streaming, and vision-language model alignment.

Discussion

20 comments analyzed.

Competitors mentioned: A Cloud Guru / Pluralsight (edtech labs), Google Synergyse (interactive tutorials), ChatGPT (software help queries), Static docs and videos (software learning)

Concerns raised: Privacy risks from sending full screenshots to third party, Cost to run for typical use cases, Reliability concerns before replacing documentation, Outdated knowledge (e.g., Blender 5 not known by AI), Unclear data retention and encryption claims

Feature requests: Windows version / broader OS support, Auto-learning from newer software documentation versions, Automatic screenshot capture instead of manual uploads, Educational copilot for interactive cloud learning

Competitors

Other products that read as similar to this one — 21 launches clear the similarity bar, closest 8 shown.

Attention rank: #8 of 22 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 25 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a agent / copilot tool for Agriculture yet.