Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Open-weight OCR got so cheap I had to share it

Details

External ID
49041317
Source
HN
Company
—
Product
Open-weight OCR got so cheap I had to share it
Website domain
openparser.dev
Launched
July 24, 2026
Cohort
—
Upvotes
17
Upvotes percentile
0.6726403823178017
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

This was not supposed to become a product.When PaddleOCR-VL-1.6 dropped, independent benchmarks put it at the top of document parsing models. I had to try it. I needed a provider, but there simply isn't one ready for production that I would trust.So i set one up myself. I assumed that even after getting it running, serving a vision-language model would be expensive.It turns out the opposite is true. Once I had it running properly, the cost was absurdly low. At proper GPU utilization, the cost is only around $1 per 1,000 pages.The nearest competitors are either much lower quality (Azure Read) or absurdly expensive (Extend or Reducto). Even Mistral OCR 4 which is really good and pretty cheap is still 4x more expensive.So I had to share it. I made the endpoitns public, and vibe-coded a simple dashboard.Let me know what you think!

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Dev tools
Audience
Developer
AI stance
Not AI
Project type
Hobby / open-source project
Normalized one-liner
open-weight ocr model
Manually corrected
False

Could you build this?

Partial Setting up a simple wrapper around PaddleOCR is easy, but running an ultra-cheap, highly performant hosted OCR API with autoscaling GPU infrastructure, multi-tenant queues, and low-cost batch processing requires substantial systems and infrastructure expertise.

What it would actually take: A production deployment needs an orchestration layer (e.g., Ray Serve, Triton Inference Server, or custom vLLM/vLLM-like pipelines) hosted across GPU spot instances (RunPod, Lambda, AWS). It requires robust asynchronous request queueing (Redis/SQS), memory-optimized model inference, PDF page splitting/rendering pipelines, and complex load-balancing to keep inference costs low enough to support $1/1,000 pages.

Discussion

11 comments analyzed.

Competitors mentioned: PaddleOCR, AWS, OpenAI, OpenRouter

Concerns raised: Pricing not justified by openness, just GPU costs, Misleading 'open' branding without self-hosting documentation, Vendor lock-in despite open-weight claims, Monetizing open-source model as SaaS without clear transparency

Feature requests: Self-hosting documentation, Better handwritten text parsing

Competitors

Other products that read as similar to this one — 50 launches clear the similarity bar, closest 8 shown.

Attention rank: #23 of 51 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 266 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a dev tools tool for Sales yet.