Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Scry, programmable internet search w/ congestion pricing

Details

External ID
49748041
Source
HN
Company
—
Product
Scry, programmable internet search w/ congestion pricing
Website domain
scry.io
Launched
Sept. 17, 2026
Cohort
—
Upvotes
60
Upvotes percentile
0.8724082934609251
Tags
—
Fetched at
Sept. 21, 2026, 5:02 p.m.
Updated at
Sept. 21, 2026, 5:02 p.m.

Description

Meet Scry, a 500 TB NVMe internet index in ClickHouse that you can run ~arbitrary readonly SQL and some of Datalog over, and I handle the problem of resource-contention with congestion-based micro-auction pricing. When there's capacity, the service is free for non-commercial use.---Hello. It's 2026, we're training simulated fruit fly brains to play Beat Saber, do we still have to be stuck with internet (re)search as fn: natural language -> black box we can't do anything about -> ranked_list/summary?There is a long history of people trying to do very fancy things that end up being done in relational databases and a little SQL. There is a gravity to them, a bitter lesson, just like scaling of generalized ml training methods. I mean many, many information products can be built off essentially giant real-time OLAP databases and frontier LLMs writing brilliant SQL+Datalog+vector+Jev etc. queries.Google Search, Tavily, Exa essentially have the problem of mapping your agents' context you are willing to provide, to a tiny subset of their index. You pay a fixed cost to an extremely hard problem that has a distribution of hardness, which means YOU eat the downsides when they are running out of budgeted compute to help you out.Their algorithms are opaque to the caller, there's really not much user control, and there's not a serious opportunity to communally improve search recipes, like the lexical+Jev recipes you trust to select bleeding edge AI builders.Furthermore, search companies aren't even pursuing text-to-SQL anymore (several have talked to me)... they made up their minds during the traumatic 2024 text-to-sql days. They were just too early.I hope you enjoy. I'm intent on scaling this paradigm on differentiated hardware over much more data, so any compelling use cases or queries I could show off, would be much appreciated!

Enrichment

Theme
database infrastructure and developer tools
Vertical
Horizontal
Function
Search & retrieval
Audience
Developer
AI stance
Not AI
Project type
Commercial product
Normalized one-liner
programmable search engine for developers
Manually corrected
False

Could you build this?

No Scry operates a 500 TB NVMe internet search index in ClickHouse with congestion-based micro-auction pricing and multi-billion document queries, requiring enormous infrastructure and crawler scale.

What it would actually take: Requires a petabyte-scale web scraping and data ingestion pipeline, a distributed multi-node ClickHouse cluster optimized for NVMe drives, and custom rate-limiting/congestion pricing execution engines. The barrier is massive physical hardware capital, petabyte-scale data pipelines, and distributed systems engineering.

Discussion

20 comments analyzed.

Competitors mentioned: Pushshift

Concerns raised: Confusing pricing model and undefined terms, LLM-generated copy is hard to read, Legality of scraping Reddit comments, Handling deleted Reddit comments, Cost effectiveness of data ingestion

Feature requests: Rewrite homepage text cleanly without LLM tropes, Make datasets available via P2P torrents, Plain-English explanation of pricing, Addition to DeepSearchQA leaderboard

Competitors

Other products that read as similar to this one — 68 launches clear the similarity bar, closest 8 shown.

Attention rank: #12 of 69 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 318 days after the earliest competitor.

Other launches for this product