Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Crustdata (YC F24)

Web Search API for Token-Efficient AI Agents

Details

External ID
47146819
Source
HN
Company
—
Product
Crustdata (YC F24)
Website domain
crustdata.com
Launched
Feb. 25, 2026
Cohort
—
Upvotes
10
Upvotes percentile
0.5316711590296496
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Hi HN! We’re Abhilash Chowdhary, Chris Pisarski and Manmohit Grewal. We built Crustdata (YC F24). Today we’re launching our web search API for AI agents, which not only returns the most relevant documents from the web but also maps them to the correct entity (person, company or event). Demo video here https://youtu.be/IouWW97hBN8If you run agents at scale, tokens become a line item. The web data is the worst input: long pages, repeated content, mixed entities, stale claims. The usual web search -> scrape -> summarize + structure forces the agent to spend tokens doing janitorial work before it can take action.We’re trying to move that work upstream. We keep a canonical graph (ontology) of people and companies: stable internal IDs, aliases, and relationships. Then we continuously index the web and attach each document to the right entity ID. Example: raw web search for "Stripe pricing changes 2026" returns ~10 results across ~4,000 tokens, mostly redundant. We return 6 deduplicated results in ~1,200 tokens.This is not just about saving tokens. It also matters because the common failure isn’t “search missed something.” It’s “search found something about the wrong entity.” Names collide. Companies rebrand. Domains move. Press releases get syndicated and look like independent sources. If you treat strings as IDs, you eventually attach evidence to the wrong person/company and the agent takes a confident action based on that mistake.Under the hood, we run a continuous pipeline that updates the entity-linked index: discover -> fetch -> extract -> dedupe -> entity resolution -> attach -> index . And we serve you this index via our search API.We didn’t start with web search. We spent ~2 years building verified people + company data from higher-trust sources. That forced us to build identity as a system, not a string. When we tried to bolt on web search and started building our integrated index of documents + people + companies, we ended up with a pile of local fixes: parser tweaks, domain rules, prompt hacks. Each fix helped one case and broke another because identity isn’t local. That’s when we committed to an entity-first index: pay the entity resolution cost once, then reuse it everywhere.If you’re building AI agents for sales, recruiting, or investing that do a lot of web searches for people and companies, we’d love for you to try our web search APIs. https://crustdata.com/demo

Enrichment

Theme
browser automation and scraping for AI
Vertical
Horizontal
Function
Search & retrieval
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
web search api for ai agents
Manually corrected
False

Could you build this?

No Building a custom web-scale crawler and real-time B2B knowledge graph that accurately resolves entities (people, companies) requires massive crawling infrastructure, anti-bot bypasses, and specialized entity resolution models.

What it would actually take: Requires large-scale distributed web scrapers (Puppeteer/Playwright clusters, proxy networks) continually indexing LinkedIn, corporate registries, and news, feeding into a petabyte-scale graph database and vector search engine (like Neo4j and Vespa/Milvus). The hardest technical challenges are entity deduplication across noisy web sources and high-throughput real-time indexing, requiring an experienced data engineering and NLP research team.

Discussion

No comments on this launch.

Competitors

Other products that read as similar to this one — 123 launches clear the similarity bar, closest 8 shown.

Attention rank: #60 of 124 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 106 days after the earliest competitor.

Other launches for this product