Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Chat with Orion

a visual agent that sees, reasons and acts

Details

External ID
45919778
Source
HN
Company
—
Product
Chat with Orion
Website domain
vlm.run
Launched
Nov. 13, 2025
Cohort
—
Upvotes
22
Upvotes percentile
0.712882096069869
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Hey HN! We’re excited to share Orion [1] — our new visual agent that sees, reasons, and acts across images, videos, and documents.Frontier VLMs (GPT, Claude, Gemini) can describe what they see, but they can’t reliably act on visual inputs. Ask them to detect objects, segment images, or chain visual steps — they’ll fail in surprisingly inconsistent ways. High-res images collapse to ~1024px. And the visual AI ecosystem is fragmented across separate APIs for image understanding, OCR, image-gen, video-gen, etc.We built Orion to fix this.Orion combines VLM reasoning with reliable computer-vision tools inside a unified chat-completions interface. You can chain visual steps, inspect results, and treat visual tasks the same way you treat text workflows. Here’s a quick demo [2].What Orion can do today: - Detect objects, faces, people (with precise, visualized boxes) - Segment objects or salient regions interactively - Edit, remix, and re-imagine images/videos from prompts - Summarize visual content (images or videos) - Transform images: crop, rotate, upscale - Transform videos: trim, sample, highlight scenes - Parse and structure documents: pagination, layout, OCR, extractionOne unified “chat-completions”-like interface — no juggling multiple vision APIs. Check out the tours in the chat [3] or read the announcement [4].API access opens next week. Happy to answer any questions — otherwise, feel free to try the tours and break things![1] Learn more about Orion: https://vlm.run/orion[2] Promo video: https://youtu.be/cPJN4iZz6QQ[3] Chat: https://chat.vlm.run[4] LinkedIn announcement: https://www.linkedin.com/posts/sudeeppillai_ai-computervisio...

Enrichment

Theme
interactive simulations and creative experiments
Vertical
Horizontal
Function
Agent / copilot
Audience
B2B
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
visual agent that sees reasons and acts
Manually corrected
False

Could you build this?

No Developing an agent that performs precise visual perception, spatial reasoning, segmentation, and grounded action chains requires proprietary multimodal models and computer vision research beyond standard API wrappers.

What it would actually take: The architecture requires vision-language models (VLMs) fine-tuned with grounded coordinates and segmentation masks, coupled with tools like SAM (Segment Anything) and object detectors. The hardest part is training or fine-tuning models on complex spatial/document reasoning benchmarks to reliably output actionable bounding coordinates rather than loose text descriptions. Building this demands deep computer vision ML research, high-end GPU clusters, and curated visual-grounding datasets.

Discussion

10 comments analyzed.

Competitors mentioned: SAM2 (segmentation models), Other chat interfaces

Concerns raised: Astro-turfed responses in post, Still early days, limited capabilities

Feature requests: Video tracking, In-image editing, 3D capabilities

Competitors

Other products that read as similar to this one — 66 launches clear the similarity bar, closest 8 shown.

Attention rank: #25 of 67 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 15 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a agent / copilot tool for Agriculture yet.