Llama 3.2 3B and Keiro Research achieves 85% on SimpleQA
Details
- External ID
- 47285569
- Source
- HN
- Company
- —
- Product
- Llama 3.2 3B and Keiro Research achieves 85% on SimpleQA
- Website domain
- keirolabs.cloud
- Launched
- March 7, 2026
- Cohort
- —
- Upvotes
- 6
- Upvotes percentile
- 0.2853628536285363
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
ran this over the weekend. stack was Llama 3.2 3B running locally + Keiro Research API for retrieval.85.0% on 4,326 questions. where that lands:ROMA (357B): 93.9% OpenDeepSearch (671B): 88.3% Sonar Pro: 85.8% Llama 3.2 3B + Keiro: 85.0%the systems ahead of us are running models 100-200x larger. that's why they're ahead. not better retrieval, not better prompting — just way more parameters.the interesting part is how small the gap is despite that. 3 points behind a 671B model. 0.8 behind Sonar Pro. at some point you have to ask what you're actually buying with all that compute for this class of task.Want to know how low the reader model can go before it starts mattering. in this setup it clearly wasn't the limiting factor and also if smaller models with web enabled will perform as good( if not better) as larger models for a lot of non coding tasksFull benchmark script + results --> https://github.com/h-a-r-s-h-s-r-a-h/benchmarkKeiro research -- https://www.keirolabs.cloud/docs/api-reference/research
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- small language model with high reasoning performance
- Manually corrected
- False
Could you build this?
Partial Running a local Llama 3.2 3B model and querying an API is simple, but achieving an 85% score on SimpleQA relies heavily on proprietary high-precision retrieval infrastructure (Keiro Research).
What it would actually take: Replicating this benchmark score requires building an advanced web search, indexing, and fact-checking retrieval engine comparable to Keiro or Perplexity. The backend necessitates low-latency web crawlers, state-of-the-art embedding and reranking models, semantic chunking, and hallucination-minimizing RAG pipelines before feeding context into the small local LLM.
Discussion
1 comment analyzed.
Competitors
Other products that read as similar to this one — 49 launches clear the similarity bar, closest 8 shown.
Attention rank: #37 of 50 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 124 days after the earliest competitor.
- Omni · hn · 2026-06-05 · 6 upvotes · similarity 0.41
- PhAIL · hn · 2026-03-31 · 21 upvotes · similarity 0.38
- open-medical-jev · github · 2026-09-25 · 38 upvotes · similarity 0.37
- What is HN thinking? Real-time sentiment and concept analysis · hn · 2026-02-12 · 37 upvotes · similarity 0.37
- A new benchmark for testing LLMs for deterministic outputs · hn · 2026-04-29 · 60 upvotes · similarity 0.37
- Fast NF4 dequantization Triton kernel (1.41x faster than bitsandbytes) · hn · 2026-07-15 · 5 upvotes · similarity 0.36
- Zero downtime embedding model upgrades · hn · 2026-09-08 · 6 upvotes · similarity 0.36
- Llm.sql · hn · 2026-04-24 · 8 upvotes · similarity 0.35
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.