Irpapers
Visual embeddings vs. OCR trade-offs in scientific PDFs
Details
- External ID
- 47125210
- Source
- HN
- Company
- —
- Product
- Irpapers
- Website domain
- github.com
- Launched
- Feb. 23, 2026
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.10512129380053908
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Hey HN, we are releasing IRPAPERS to answer a highly pragmatic question: when building a RAG pipeline over PDFs, should you OCR the text or just embed the raw page images?Processing PDFs in production usually involves stringing together brittle OCR heuristics. While recent multimodal embeddings (like ColModernVBERT or ColPali) allow you to skip OCR entirely and retrieve directly from visual layouts, we wanted to measure if the computational overhead is actually worth the utility.The short answer: Transformer-based image pipelines won't be perfect for every use-case, but they fix exactly what OCR breaks.Here is what we found benchmarking 3,230 pages of dense scientific literature:Complementary Bottlenecks: Text representations (BM25 + dense vectors) are highly efficient for exact lexical constraints (e.g., finding a specific acronym like "HyDE"). Conversely, image embeddings shine on spatial architecture diagrams and t-SNE plots where OCR serialization just turns into structural garbage.Multimodal Hybrid Search: Because these failure modes are almost perfectly orthogonal, fusing the two signals gives you the best performance out of the box. By combining them, we pushed top-1 recall to 49% (beating text alone at 46%).The Memory Constraint: Late-interaction image embeddings produce thousands of vectors per page, creating a massive storage bottleneck. To address this need, we evaluate MUVERA encoding. Under the hood, this compresses multi-vector representations into a single fixed-dimensional encoding via SimHash, allowing you to use standard HNSW indexing without the paralyzing memory overhead.In practice, if you are building a RAG workflow today, text-based context still provides higher downstream utility for the actual generation step (0.82 vs 0.71 alignment). Instead of picking one modality and dealing with its blind spots, start with hybrid text search as a sensible default, and inject multi-vector image embeddings to catch the visual edge-cases.We’ve open-sourced the benchmark and the evaluation recipes:Paper https://arxiv.org/abs/2602.17687 IRPAPERS dataset on HuggingFace at huggingface.co/weaviate/IRPAPERS and GitHub at github.com/weaviate/IRPAPERSOur experimental code is also available on GitHub at github.com/weaviate/query-agent-benchmarkingHappy to answer any questions about the evaluation pipeline, the cold start problem of visual benchmarks, or the specific retrieval trade-offs we saw.
Enrichment
- Theme
- document processing and generation tools
- Vertical
- Horizontal
- Function
- Search & retrieval
- Audience
- Developer
- AI stance
- AI feature
- Project type
- Hobby / open-source project
- Normalized one-liner
- visual embeddings for scientific pdfs
- Manually corrected
- False
Could you build this?
Yes This is a benchmarking and evaluation research project comparing standard vision embedding models against OCR pipelines on PDF datasets, which is standard data science script work.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 192 launches clear the similarity bar, closest 8 shown.
Attention rank: #176 of 193 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 115 days after the earliest competitor.
- Vision-Based, Vectorless RAG for Long Douments · hn · 2025-10-31 · 6 upvotes · similarity 0.54
- Unified multimodal memory framework, without embeddings · hn · 2026-01-07 · 7 upvotes · similarity 0.50
- Unsiloed AI · hn · 2026-05-25 · 9 upvotes · similarity 0.48
- A graph paper generator that renders vector PDFs in the browser · hn · 2026-07-02 · 107 upvotes · similarity 0.46
- Extract diagrams from PDF to SVG · hn · 2025-12-22 · 11 upvotes · similarity 0.46
- ml-lensvlm · github · 2026-09-22 · 75 upvotes · similarity 0.45
- CPU-only fast OCR for screenshots, images, PDFs, webpages · hn · 2026-05-31 · 9 upvotes · similarity 0.45
- We built an OCR server that can process 270 dense images/s on a 5090 · hn · 2026-04-23 · 8 upvotes · similarity 0.44
Other launches for this product
- No other launches for this product.