I taught GPT-OSS-120B to see using Google Lens and OpenCV
Details
- External ID
- 46971287
- Source
- HN
- Company
- —
- Product
- —
- Website domain
- —
- Launched
- Feb. 11, 2026
- Cohort
- —
- Upvotes
- 43
- Upvotes percentile
- 0.793800539083558
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
I built an MCP server that gives any local LLM real Google search and now vision capabilities - no API keys needed. The latest feature: google_lens_detect uses OpenCV to find objects in an image, crops each one, and sends them to Google Lens for identification. GPT-OSS-120B, a text-only model with zero vision support, correctly identified an NVIDIA DGX Spark and a SanDisk USB drive from a desk photo. Also includes Google Search, News, Shopping, Scholar, Maps, Finance, Weather, Flights, Hotels, Translate, Images, Trends, and more. 17 tools total. Two commands: pip install noapi-google-search-mcp && playwright install chromium GitHub: https://github.com/VincentKaufmann/noapi-google-search-mcp PyPI: https://pypi.org/project/noapi-google-search-mcp/ Booyah!
Enrichment
- Theme
- multimodal generative ai and developer tools
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- open source model with vision capabilities
- Manually corrected
- False
Could you build this?
Yes This is a Python MCP server that strings together existing OpenCV contour detection routines with reverse-image requests sent to Google Lens.
Discussion
20 comments analyzed.
Competitors mentioned: Local VLMs (8B-30B models), Llama models, Qwen (vision and reasoning variants), Wolfram Alpha, Gemini API
Concerns raised: Misleading framing - uses external APIs rather than true local vision capability, Performance too slow for practical use, Doesn't actually teach the model to see, just adds a tool wrapper, Requires clarification on which model variant (20B vs 120B), Quantization issues cause tool calling errors
Feature requests: Integrate local VLM instead of relying on external APIs, Train a projector with vision encoder on model directly, Get actual image as input, not just descriptions, Improve tool calling support across inference engines
Competitors
Other products that read as similar to this one — 53 launches clear the similarity bar, closest 8 shown.
Attention rank: #18 of 54 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 90 days after the earliest competitor.
- gpt-image-2-api · github · 2026-09-21 · 14 upvotes · similarity 0.43
- prismctl2api · github · 2026-09-16 · 16 upvotes · similarity 0.40
- gptimage_prompts · github · 2026-09-11 · 8 upvotes · similarity 0.39
- MaxAPI · ph · 2026-09-19 · 2 upvotes · similarity 0.39
- gpt-image-2-api-examples · github · 2026-09-21 · 9 upvotes · similarity 0.38
- GPTImagine 2.5 · ph · 2026-09-10 · 1 upvotes · similarity 0.36
- GPTImageGen25 · ph · 2026-09-10 · 1 upvotes · similarity 0.36
- awesome-gpt-image-2-5-prompts · github · 2026-09-10 · 34 upvotes · similarity 0.36
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.