Ocrbase
pdf → .md/.json document OCR and structured extraction API
Details
- External ID
- 46691454
- Source
- HN
- Company
- —
- Product
- Ocrbase
- Website domain
- github.com
- Launched
- Jan. 20, 2026
- Cohort
- —
- Upvotes
- 99
- Upvotes percentile
- 0.8965744400527009
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Enrichment
- Theme
- document processing and generation tools
- Vertical
- Horizontal
- Function
- Dev tools
- Audience
- Developer
- AI stance
- AI feature
- Project type
- Commercial product
- Normalized one-liner
- pdf to markdown/json document extraction api
- Manually corrected
- False
Could you build this?
Partial The API wrapper and format conversion (PDF to Markdown/JSON) is straightforward, but delivering high-accuracy OCR and table/document layout extraction requires fine-tuned vision-language models or dedicated document OCR pipelines.
What it would actually take: A production version needs a pipeline using open-source vision OCR models (like Nougat, PaddleOCR, or fine-tuned Surya) deployed on GPU instances with layout-parser/Tesseract fallback. The hard part is accurate table extraction, reading order reconstruction, and low-latency inference on multi-page dense documents. It requires ML engineering for vision models and scalable document processing infrastructure.
Discussion
20 comments analyzed.
Competitors mentioned: Surya, RapidOCR, Tesseract, PaddleOCR, Gemini Flash (for structured extraction)
Concerns raised: 12GB+ VRAM requirement seems high for model size, Security risk storing GitHub secrets in plain-text env files, High computational resource requirements, Cost comparison with Gemini Flash (~$0.50 per 100 pages), Data privacy concerns with sending financial documents to cloud services
Feature requests: Local-only deployment option to avoid sending data to cloud, Lower VRAM requirements for ARM/M-series Mac compatibility, Fallback pipeline with cheaper extraction methods before expensive LLM calls, Secrets manager integration instead of plain-text environment variables
Competitors
Other products that read as similar to this one — 219 launches clear the similarity bar, closest 8 shown.
Attention rank: #10 of 220 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 81 days after the earliest competitor.
- Emboss · hn · 2026-07-30 · 6 upvotes · similarity 0.53
- gopd · github · 2026-09-15 · 8 upvotes · similarity 0.51
- PDF to MD · ph · 2026-09-10 · 1 upvotes · similarity 0.49
- StructOCR · hn · 2026-06-06 · 5 upvotes · similarity 0.49
- PDFCraft · ph · 2026-09-13 · 2 upvotes · similarity 0.49
- OCR Buddy: local browser OCR for code, formulas (LaTeX) and tables · hn · 2026-07-07 · 9 upvotes · similarity 0.47
- PDFidelity · hn · 2026-09-12 · 7 upvotes · similarity 0.46
- Pdf2md · hn · 2026-05-18 · 10 upvotes · similarity 0.46
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a dev tools tool for Sales yet.