Run open-weight OCR, VLM and vision models behind one API
Details
- External ID
- 49568379
- Source
- HN
- Company
- —
- Product
- Run open-weight OCR, VLM and vision models behind one API
- Website domain
- vlmrun.com
- Launched
- Sept. 4, 2026
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.12998405103668262
- Tags
- —
- Fetched at
- Sept. 10, 2026, 5:31 a.m.
- Updated at
- Sept. 10, 2026, 5:31 a.m.
Description
Hey HN. We built an openai-compatible API for running open-weight VLMs, OCR VLMs and ViT-based vision models.The motivation was mostly frustration when running these models in production and discovering all the details around serving VLMs, especially around visual accuracy.A few footguns we kept running into:- quantized models served under the same name (this one still drives me nuts): providers often serve models with different quants, environments, vLLM/SGLang serving params with the same model-id. Vision is especially sensitive to this; some quants that look fine on text benchmarks noticeably hurt OCR/small-text/spatial accuracy.- video performance is varied: when we tested with popular routers on video-native VLMs, more than 80% of providers didn't support video inputs, and even fewer let you control FPS. If you care for time-resolution in videos, none of these providers work even if the models themselves are capable of it.- document inference is all pipelining: rasterizing PDFs, parallelizing page workers, retrying when pages fail inference, dealing with rate-limits, etc can get tricky quickly and takes substantial developer time.- standardization making abstractions leaky: this is less about vision per-se, but generally for serving models with high-quality output assurances. context-limits, max resolution, FPS sampling, quants, GPU SM architecture, can all add variability to (vision) quality even if the model-id claims to be the same.We wanted one place to run OCR models, VLMs and ViTs that we could confidently use for our own internal agents and evals. The gateway was born from this need internally, and now we're opening it up to the public - you can swap the model name to compare GLM-OCR, dots.mocr, PaddleOCR VL, Qwen3.8-27B, Gemma4-26B-A4B etc. We handle the serving/runtime/pipelining underneath, with the goal of giving high-quality visual intelligence.Are there any other vision "footguns" people have run into? especially cases where "same model" across two providers gave materially different outputs.Try different models on Gateway simply by updating the model name:uvx vlmrun gw chat <doc>.pdf -m glm-ocruvx vlmrun gw chat <doc>.pdf -m deepseek-ocr-2uvx vlmrun gw chat <doc>.pdf -m pp-ocrv6uvx vlmrun gw chat <video>.mp4 -m qwen/qwen3.5-0.8b -p "describe the video"
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- api for ocr and vision models
- Manually corrected
- False
Could you build this?
Partial Building an OpenAI-compatible reverse proxy API is easy, but hosting and scaling high-throughput, low-latency vision-language models with dynamic image tiling and visual token optimizations is hard ML systems work.
What it would actually take: Requires an inference serving cluster built on vLLM, SGLang, or TensorRT-LLM orchestrating multi-GPU instances. The hard part is optimizing multi-resolution image patch encoding, preventing visual token blowups from saturating the KV cache, and implementing efficient continuous batching for heterogeneous vision tasks (OCR vs document analysis vs general VLM). Demands high-performance GPU cluster operations and deep familiarity with multi-modal model architectures.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 119 launches clear the similarity bar, closest 8 shown.
Attention rank: #104 of 120 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 308 days after the earliest competitor.
- OCR Arena · hn · 2025-11-21 · 216 upvotes · similarity 0.52
- Open-weight OCR got so cheap I had to share it · hn · 2026-07-24 · 17 upvotes · similarity 0.49
- Vision-Based, Vectorless RAG for Long Douments · hn · 2025-10-31 · 6 upvotes · similarity 0.48
- Moonshine Open-Weights STT models · hn · 2026-02-24 · 316 upvotes · similarity 0.47
- Satellite imagery object detection using text prompts · hn · 2026-03-09 · 53 upvotes · similarity 0.46
- FASHN VTON v1.5 · hn · 2026-01-28 · 6 upvotes · similarity 0.44
- Docker Model Runner Integrates vLLM for High-Throughput Inference · hn · 2025-11-20 · 7 upvotes · similarity 0.44
- DocsRouter · hn · 2025-12-18 · 14 upvotes · similarity 0.44
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.