Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Trawl

Scrape any site with natural language fields, not CSS selectors

Details

External ID
47296933
Source
HN
Company
—
Product
Trawlx
Website domain
github.com
Launched
March 8, 2026
Cohort
—
Upvotes
8
Upvotes percentile
0.491389913899139
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Every scraper I've written has the same failure mode: it works for three months, a site redesigns, and my CSS selectors silently return empty strings. The data is still right there on the page — a human can find it instantly — but the scraper is blind.Trawl fixes this by splitting the problem. You describe what you want: trawl "https://books.toscrape.com" --fields "title, price, rating, in_stock" The LLM (Claude) looks at one sample item and derives a full extraction strategy — CSS selectors, attribute mappings, type coercion, fallback selectors. That strategy gets cached. Every subsequent page with the same structure is extracted with pure Go + goquery. No API calls, no token cost, full concurrency.The key insight: LLMs are good at understanding HTML structure, but you don't need them to extract 10,000 rows. Use AI for intelligence, Go for throughput.When a site redesigns, the structural fingerprint changes, the cache misses, and trawl re-derives automatically.You can preview exactly what it figured out: $ trawl "https://example.com/products" --fields "name, price" --plan Strategy for https://example.com/products Item selector: div.product-card Fields: name: h2.product-title -> text (string) price: span.price -> text -> parse_price (float) Confidence: 0.95 Some things that took real engineering effort:- JS-rendered SPAs: headless browser with DOM stability detection — polls until element count stabilizes and skeleton loaders resolve, scrolls to trigger lazy loading, auto-clicks "Show more" buttons - Multi-section pages: detects candidate data regions heuristically, target a specific section with --query "Market Share", scopes extraction via container selectors - Self-healing: monitors extraction health (% of fields populated), re-derives the strategy if it drops below 70% - Iframes: auto-detects and extracts from iframes when they contain richer data than the outer pageOutput is JSON, JSONL, CSV, or Parquet. Pipes cleanly: trawl "https://example.com/products" --fields "name, price" --format jsonl | jq 'select(.price > 50)' Written in Go. MIT licensed.

Enrichment

Theme
browser automation and scraping for AI
Vertical
Horizontal
Function
Dev tools
Audience
Developer
AI stance
AI feature
Project type
Commercial product
Normalized one-liner
web scraping using natural language instead of css selectors
Manually corrected
False

Could you build this?

Yes The core tool wraps standard web scraping with an LLM prompt that extracts schema-driven fields from raw HTML or markdown content without relying on brittle CSS selectors.

Discussion

2 comments analyzed.

Concerns raised: Seems overkill for the use case

Competitors

Other products that read as similar to this one — 72 launches clear the similarity bar, closest 8 shown.

Attention rank: #38 of 73 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 96 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a dev tools tool for Sales yet.