Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Irpapers

Visual embeddings vs. OCR trade-offs in scientific PDFs

Details

External ID
47125210
Source
HN
Company
—
Product
Irpapers
Website domain
github.com
Launched
Feb. 23, 2026
Cohort
—
Upvotes
5
Upvotes percentile
0.10512129380053908
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Hey HN, we are releasing IRPAPERS to answer a highly pragmatic question: when building a RAG pipeline over PDFs, should you OCR the text or just embed the raw page images?Processing PDFs in production usually involves stringing together brittle OCR heuristics. While recent multimodal embeddings (like ColModernVBERT or ColPali) allow you to skip OCR entirely and retrieve directly from visual layouts, we wanted to measure if the computational overhead is actually worth the utility.The short answer: Transformer-based image pipelines won't be perfect for every use-case, but they fix exactly what OCR breaks.Here is what we found benchmarking 3,230 pages of dense scientific literature:Complementary Bottlenecks: Text representations (BM25 + dense vectors) are highly efficient for exact lexical constraints (e.g., finding a specific acronym like "HyDE"). Conversely, image embeddings shine on spatial architecture diagrams and t-SNE plots where OCR serialization just turns into structural garbage.Multimodal Hybrid Search: Because these failure modes are almost perfectly orthogonal, fusing the two signals gives you the best performance out of the box. By combining them, we pushed top-1 recall to 49% (beating text alone at 46%).The Memory Constraint: Late-interaction image embeddings produce thousands of vectors per page, creating a massive storage bottleneck. To address this need, we evaluate MUVERA encoding. Under the hood, this compresses multi-vector representations into a single fixed-dimensional encoding via SimHash, allowing you to use standard HNSW indexing without the paralyzing memory overhead.In practice, if you are building a RAG workflow today, text-based context still provides higher downstream utility for the actual generation step (0.82 vs 0.71 alignment). Instead of picking one modality and dealing with its blind spots, start with hybrid text search as a sensible default, and inject multi-vector image embeddings to catch the visual edge-cases.We’ve open-sourced the benchmark and the evaluation recipes:Paper https://arxiv.org/abs/2602.17687 IRPAPERS dataset on HuggingFace at huggingface.co/weaviate/IRPAPERS and GitHub at github.com/weaviate/IRPAPERSOur experimental code is also available on GitHub at github.com/weaviate/query-agent-benchmarkingHappy to answer any questions about the evaluation pipeline, the cold start problem of visual benchmarks, or the specific retrieval trade-offs we saw.

Enrichment

Theme
document processing and generation tools
Vertical
Horizontal
Function
Search & retrieval
Audience
Developer
AI stance
AI feature
Project type
Hobby / open-source project
Normalized one-liner
visual embeddings for scientific pdfs
Manually corrected
False

Could you build this?

Yes This is a benchmarking and evaluation research project comparing standard vision embedding models against OCR pipelines on PDF datasets, which is standard data science script work.

Discussion

No comments on this launch.

Competitors

Other products that read as similar to this one — 192 launches clear the similarity bar, closest 8 shown.

Attention rank: #176 of 193 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 115 days after the earliest competitor.

Other launches for this product