Chat with Orion
a visual agent that sees, reasons and acts
Details
- External ID
- 45919778
- Source
- HN
- Company
- —
- Product
- Chat with Orion
- Website domain
- vlm.run
- Launched
- Nov. 13, 2025
- Cohort
- —
- Upvotes
- 22
- Upvotes percentile
- 0.712882096069869
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Hey HN! We’re excited to share Orion [1] — our new visual agent that sees, reasons, and acts across images, videos, and documents.Frontier VLMs (GPT, Claude, Gemini) can describe what they see, but they can’t reliably act on visual inputs. Ask them to detect objects, segment images, or chain visual steps — they’ll fail in surprisingly inconsistent ways. High-res images collapse to ~1024px. And the visual AI ecosystem is fragmented across separate APIs for image understanding, OCR, image-gen, video-gen, etc.We built Orion to fix this.Orion combines VLM reasoning with reliable computer-vision tools inside a unified chat-completions interface. You can chain visual steps, inspect results, and treat visual tasks the same way you treat text workflows. Here’s a quick demo [2].What Orion can do today: - Detect objects, faces, people (with precise, visualized boxes) - Segment objects or salient regions interactively - Edit, remix, and re-imagine images/videos from prompts - Summarize visual content (images or videos) - Transform images: crop, rotate, upscale - Transform videos: trim, sample, highlight scenes - Parse and structure documents: pagination, layout, OCR, extractionOne unified “chat-completions”-like interface — no juggling multiple vision APIs. Check out the tours in the chat [3] or read the announcement [4].API access opens next week. Happy to answer any questions — otherwise, feel free to try the tours and break things![1] Learn more about Orion: https://vlm.run/orion[2] Promo video: https://youtu.be/cPJN4iZz6QQ[3] Chat: https://chat.vlm.run[4] LinkedIn announcement: https://www.linkedin.com/posts/sudeeppillai_ai-computervisio...
Enrichment
- Theme
- interactive simulations and creative experiments
- Vertical
- Horizontal
- Function
- Agent / copilot
- Audience
- B2B
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- visual agent that sees reasons and acts
- Manually corrected
- False
Could you build this?
No Developing an agent that performs precise visual perception, spatial reasoning, segmentation, and grounded action chains requires proprietary multimodal models and computer vision research beyond standard API wrappers.
What it would actually take: The architecture requires vision-language models (VLMs) fine-tuned with grounded coordinates and segmentation masks, coupled with tools like SAM (Segment Anything) and object detectors. The hardest part is training or fine-tuning models on complex spatial/document reasoning benchmarks to reliably output actionable bounding coordinates rather than loose text descriptions. Building this demands deep computer vision ML research, high-end GPU clusters, and curated visual-grounding datasets.
Discussion
10 comments analyzed.
Competitors mentioned: SAM2 (segmentation models), Other chat interfaces
Concerns raised: Astro-turfed responses in post, Still early days, limited capabilities
Feature requests: Video tracking, In-image editing, 3D capabilities
Competitors
Other products that read as similar to this one — 66 launches clear the similarity bar, closest 8 shown.
Attention rank: #25 of 67 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 15 days after the earliest competitor.
- Visual Agents with Code Mode · hn · 2026-06-17 · 8 upvotes · similarity 0.48
- alpha-s · github · 2026-09-17 · 16 upvotes · similarity 0.45
- Detect any object in satellite imagery using a text prompt · hn · 2026-03-08 · 22 upvotes · similarity 0.44
- Run open-weight OCR, VLM and vision models behind one API · hn · 2026-09-04 · 5 upvotes · similarity 0.44
- Satellite imagery object detection using text prompts · hn · 2026-03-09 · 53 upvotes · similarity 0.43
- LemonSlice · hn · 2026-01-27 · 133 upvotes · similarity 0.40
- Vision-Based, Vectorless RAG for Long Douments · hn · 2025-10-31 · 6 upvotes · similarity 0.40
- Awen · yc · 2026-04-01 · 18 upvotes · similarity 0.39
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a agent / copilot tool for Agriculture yet.