Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Run open-weight OCR, VLM and vision models behind one API

Details

External ID
49568379
Source
HN
Company
—
Product
Run open-weight OCR, VLM and vision models behind one API
Website domain
vlmrun.com
Launched
Sept. 4, 2026
Cohort
—
Upvotes
5
Upvotes percentile
0.12998405103668262
Tags
—
Fetched at
Sept. 10, 2026, 5:31 a.m.
Updated at
Sept. 10, 2026, 5:31 a.m.

Description

Hey HN. We built an openai-compatible API for running open-weight VLMs, OCR VLMs and ViT-based vision models.The motivation was mostly frustration when running these models in production and discovering all the details around serving VLMs, especially around visual accuracy.A few footguns we kept running into:- quantized models served under the same name (this one still drives me nuts): providers often serve models with different quants, environments, vLLM/SGLang serving params with the same model-id. Vision is especially sensitive to this; some quants that look fine on text benchmarks noticeably hurt OCR/small-text/spatial accuracy.- video performance is varied: when we tested with popular routers on video-native VLMs, more than 80% of providers didn't support video inputs, and even fewer let you control FPS. If you care for time-resolution in videos, none of these providers work even if the models themselves are capable of it.- document inference is all pipelining: rasterizing PDFs, parallelizing page workers, retrying when pages fail inference, dealing with rate-limits, etc can get tricky quickly and takes substantial developer time.- standardization making abstractions leaky: this is less about vision per-se, but generally for serving models with high-quality output assurances. context-limits, max resolution, FPS sampling, quants, GPU SM architecture, can all add variability to (vision) quality even if the model-id claims to be the same.We wanted one place to run OCR models, VLMs and ViTs that we could confidently use for our own internal agents and evals. The gateway was born from this need internally, and now we're opening it up to the public - you can swap the model name to compare GLM-OCR, dots.mocr, PaddleOCR VL, Qwen3.8-27B, Gemma4-26B-A4B etc. We handle the serving/runtime/pipelining underneath, with the goal of giving high-quality visual intelligence.Are there any other vision "footguns" people have run into? especially cases where "same model" across two providers gave materially different outputs.Try different models on Gateway simply by updating the model name:uvx vlmrun gw chat <doc>.pdf -m glm-ocruvx vlmrun gw chat <doc>.pdf -m deepseek-ocr-2uvx vlmrun gw chat <doc>.pdf -m pp-ocrv6uvx vlmrun gw chat <video>.mp4 -m qwen/qwen3.5-0.8b -p "describe the video"

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
api for ocr and vision models
Manually corrected
False

Could you build this?

Partial Building an OpenAI-compatible reverse proxy API is easy, but hosting and scaling high-throughput, low-latency vision-language models with dynamic image tiling and visual token optimizations is hard ML systems work.

What it would actually take: Requires an inference serving cluster built on vLLM, SGLang, or TensorRT-LLM orchestrating multi-GPU instances. The hard part is optimizing multi-resolution image patch encoding, preventing visual token blowups from saturating the KV cache, and implementing efficient continuous batching for heterogeneous vision tasks (OCR vs document analysis vs general VLM). Demands high-performance GPU cluster operations and deep familiarity with multi-modal model architectures.

Discussion

No comments on this launch.

Competitors

Other products that read as similar to this one — 119 launches clear the similarity bar, closest 8 shown.

Attention rank: #104 of 120 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 308 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.