Trawl
Scrape any site with natural language fields, not CSS selectors
Details
- External ID
- 47296933
- Source
- HN
- Company
- —
- Product
- Trawlx
- Website domain
- github.com
- Launched
- March 8, 2026
- Cohort
- —
- Upvotes
- 8
- Upvotes percentile
- 0.491389913899139
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Every scraper I've written has the same failure mode: it works for three months, a site redesigns, and my CSS selectors silently return empty strings. The data is still right there on the page — a human can find it instantly — but the scraper is blind.Trawl fixes this by splitting the problem. You describe what you want: trawl "https://books.toscrape.com" --fields "title, price, rating, in_stock" The LLM (Claude) looks at one sample item and derives a full extraction strategy — CSS selectors, attribute mappings, type coercion, fallback selectors. That strategy gets cached. Every subsequent page with the same structure is extracted with pure Go + goquery. No API calls, no token cost, full concurrency.The key insight: LLMs are good at understanding HTML structure, but you don't need them to extract 10,000 rows. Use AI for intelligence, Go for throughput.When a site redesigns, the structural fingerprint changes, the cache misses, and trawl re-derives automatically.You can preview exactly what it figured out: $ trawl "https://example.com/products" --fields "name, price" --plan Strategy for https://example.com/products Item selector: div.product-card Fields: name: h2.product-title -> text (string) price: span.price -> text -> parse_price (float) Confidence: 0.95 Some things that took real engineering effort:- JS-rendered SPAs: headless browser with DOM stability detection — polls until element count stabilizes and skeleton loaders resolve, scrolls to trigger lazy loading, auto-clicks "Show more" buttons - Multi-section pages: detects candidate data regions heuristically, target a specific section with --query "Market Share", scopes extraction via container selectors - Self-healing: monitors extraction health (% of fields populated), re-derives the strategy if it drops below 70% - Iframes: auto-detects and extracts from iframes when they contain richer data than the outer pageOutput is JSON, JSONL, CSV, or Parquet. Pipes cleanly: trawl "https://example.com/products" --fields "name, price" --format jsonl | jq 'select(.price > 50)' Written in Go. MIT licensed.
Enrichment
- Theme
- browser automation and scraping for AI
- Vertical
- Horizontal
- Function
- Dev tools
- Audience
- Developer
- AI stance
- AI feature
- Project type
- Commercial product
- Normalized one-liner
- web scraping using natural language instead of css selectors
- Manually corrected
- False
Could you build this?
Yes The core tool wraps standard web scraping with an LLM prompt that extracts schema-driven fields from raw HTML or markdown content without relying on brittle CSS selectors.
Discussion
2 comments analyzed.
Concerns raised: Seems overkill for the use case
Competitors
Other products that read as similar to this one — 72 launches clear the similarity bar, closest 8 shown.
Attention rank: #38 of 73 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 96 days after the earliest competitor.
- Robust LLM extractor for websites in TypeScript · hn · 2026-03-26 · 72 upvotes · similarity 0.58
- TafScraper · ph · 2026-09-28 · 3 upvotes · similarity 0.50
- I built an SDK that scrambles HTML so scrapers get garbage · hn · 2026-03-12 · 16 upvotes · similarity 0.45
- ScrapeClaw · github · 2026-09-18 · 36 upvotes · similarity 0.45
- scrape.land · ph · 2026-09-21 · 11 upvotes · similarity 0.44
- ClickScrape · ph · 2026-09-09 · 2 upvotes · similarity 0.44
- Draco · hn · 2026-08-02 · 15 upvotes · similarity 0.44
- Selector Forge · hn · 2026-06-22 · 38 upvotes · similarity 0.44
Other launches for this product
Same idea, different domain
Nobody's really built a dev tools tool for Sales yet.