Ourguide
OS wide task guidance system that shows you where to click
Details
- External ID
- 46769422
- Source
- HN
- Company
- —
- Product
- Ourguide
- Website domain
- ourguide.ai
- Launched
- Jan. 26, 2026
- Cohort
- —
- Upvotes
- 52
- Upvotes percentile
- 0.8023715415019763
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Hey! I'm eshaan and I'm building Ourguide -an on-screen task guidance system that can show you where to click step-by-step when you need help.I started building this because whenever I didn’t know how to do something on my computer, I found myself constantly tabbing between chatbots and the app, pasting screenshots, and asking “what do I do next?” Ourguide solves this with two modes. In Guide mode, the app overlays your screen and highlights the specific element to click next, eliminating the need to leave your current window. There is also Ask mode, which is a vision-integrated chat that captures your screen context—which you can toggle on and off anytime -so you can ask, "How do I fix this error?" without having to explain what "this" is.It’s an Electron app that works OS-wide, is vision-based, and isn't restricted to the browser.Figuring out how to show the user where to click was the hardest part of the process. I originally trained a computer vision model with 2300 screenshots to identify and segment all UI elements on a screen and used a VLM to find the correct icon to highlight. While this worked extremely well—better than SOTA grounding models like UI Tars—the latency was just too high. I'll be making that CV+VLM pipeline OSS soon, but for now, I’ve resorted to a simpler implementation that achieves <1s latency.You may ask: if I can show you where to click, why can't I just click too? While trying to build computer-use agents during my job in Palo Alto, I hit the core limitation of today’s computer-use models where benchmarks hover in the mid-50% range (OSWorld). VLMs often know what to do but not what it looks like; without reliable visual grounding, agents misclick and stall. So, I built computer use—without the "use." It provides the visual grounding of an agent but keeps the human in the loop for the actual execution to prevent misclicks.I personally use it for the AWS Console's "treasure hunt" UI, like creating a public S3 bucket with specific CORS rules. It’s also been surprisingly helpful for non-technical tasks, like navigating obscure settings in Gradescope or Spotify. Ourguide really works for any task when you’re stuck or don't know what to do.You can download and test Ourguide here: https://ourguide.ai/downloadsThe project is still very early, and I’d love your feedback on where it fails, where you think it worked well, and which specific niches you think Ourguide would be most helpful for.
Enrichment
- Theme
- browser automation and scraping for AI
- Vertical
- Horizontal
- Function
- Agent / copilot
- Audience
- B2C
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- ai guidance system showing where to click
- Manually corrected
- False
Could you build this?
Partial The basic Electron overlay and screen-capture pipeline can be vibe coded, but robust desktop-wide GUI element detection and spatial coordinate mapping across arbitrary OS apps require specialized computer vision models.
What it would actually take: The architecture involves a native desktop client (e.g., Swift on macOS or Electron with native C++ node addons) utilizing OS accessibility APIs (AXUIElement) and low-latency screen capture via ScreenCaptureKit. The hard problem is reliably grounding LLM action intents onto dynamic screen pixels when accessibility trees are incomplete or missing, requiring fine-tuned multimodal grounding models (e.g., UI-VLM / SeeClick / OmniParser). Building this requires expertise in operating system accessibility hooks, low-latency video streaming, and vision-language model alignment.
Discussion
20 comments analyzed.
Competitors mentioned: A Cloud Guru / Pluralsight (edtech labs), Google Synergyse (interactive tutorials), ChatGPT (software help queries), Static docs and videos (software learning)
Concerns raised: Privacy risks from sending full screenshots to third party, Cost to run for typical use cases, Reliability concerns before replacing documentation, Outdated knowledge (e.g., Blender 5 not known by AI), Unclear data retention and encryption claims
Feature requests: Windows version / broader OS support, Auto-learning from newer software documentation versions, Automatic screenshot capture instead of manual uploads, Educational copilot for interactive cloud learning
Competitors
Other products that read as similar to this one — 21 launches clear the similarity bar, closest 8 shown.
Attention rank: #8 of 22 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 25 days after the earliest competitor.
- Give your AI agent on-screen guides that show users where to click · hn · 2026-09-09 · 19 upvotes · similarity 0.40
- Handrail · ph · 2026-09-24 · 2 upvotes · similarity 0.35
- Giggles · hn · 2026-03-03 · 23 upvotes · similarity 0.35
- What should the GUI for AI agents look like? · hn · 2026-07-31 · 139 upvotes · similarity 0.35
- Glance Switch · ph · 2026-09-27 · 3 upvotes · similarity 0.34
- Drive any macOS app in the background without stealing the cursor · hn · 2026-04-28 · 192 upvotes · similarity 0.34
- Do Not Distracted — Focus OS · ph · 2026-09-30 · 2 upvotes · similarity 0.34
- Lightning Image Viewer · hn · 2026-01-05 · 5 upvotes · similarity 0.34
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a agent / copilot tool for Agriculture yet.