Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Use Claude Code to Query 600 GB Indexes over Hacker News, ArXiv, etc.

Details

External ID
46442245
Source
HN
Company
—
Product
Use Claude Code to Query 600 GB Indexes over Hacker News, ArXiv, etc.
Website domain
exopriors.com
Launched
Dec. 31, 2025
Cohort
—
Upvotes
397
Upvotes percentile
0.9770992366412213
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Paste in my prompt to Claude Code with an embedded API key for accessing my public readonly SQL+vector database, and you have a state-of-the-art research tool over Hacker News, arXiv, LessWrong, and dozens of other high-quality public commons sites. Claude whips up the monster SQL queries that safely run on my machine, to answer your most nuanced questions.There's also an Alerts functionality, where you can just ask Claude to submit a SQL query as an alert, and you'll be emailed when the ultra nuanced criteria is met (and the output changes). Like I want to know when somebody posts about "estrogen" in a psychoactive context, or enough biology metaphors when talking about building infrastructure.Currently have embedded: posts: 1.4M / 4.6M comments: 15.6M / 38M That's with Voyage-3.5-lite. And you can do amazing compositional vector search, like search @FTX_crisis - (@guilt_tone - @guilt_topic) to find writing that was about the FTX crisis and distinctly without guilty tones, but that can mention "guilt".I can embed everything and all the other sources for cheap, I just literally don't have the money.

Enrichment

Theme
utilities for Claude and Claude Code
Vertical
Horizontal
Function
Search & retrieval
Audience
Developer
AI stance
AI feature
Project type
Commercial product
Normalized one-liner
query large indexes with claude
Manually corrected
False

Could you build this?

No Operating a fast, low-latency 600 GB hybrid vector and relational database indexing massive corpora (Hacker News, ArXiv) requires substantial data infrastructure, distributed search engineering, and significant ongoing server resources.

What it would actually take: Requires a scalable distributed database architecture (such as ClickHouse or DuckDB paired with pgvector or Qdrant/Milvus), massive web crawling and ETL pipelines, embedding computation across millions of documents, and a custom Model Context Protocol (MCP) or SQL execution gateway with strict CPU/memory sandboxing to prevent resource exhaustion.

Discussion

20 comments analyzed.

Competitors mentioned: ZenQuery (LLM-generated SQL from schemas), sql.js-httpvfs (client-side SQLite queries)

Concerns raised: Vector manipulation approach lacks rigor and benchmarks showing it works better, Embedding space manifold structure isn't semantically uniform, Security/IP protection concerns in jurisdictions with hackers and lax oversight

Feature requests: Offline mirror capability, Variable text chunk granularity for embeddings, Better embedding strategies and vector manipulation paradigms

Competitors

Other products that read as similar to this one — 148 launches clear the similarity bar, closest 8 shown.

Attention rank: #6 of 149 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 49 days after the earliest competitor.

Other launches for this product