wxpath
Declarative web crawling in XPath
Details
- External ID
- 46618472
- Source
- HN
- Company
- —
- Product
- wxpath
- Website domain
- github.com
- Launched
- Jan. 14, 2026
- Cohort
- —
- Upvotes
- 64
- Upvotes percentile
- 0.8405797101449275
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
wxpath is a declarative web crawler where web crawling and scraping are expressed directly in XPath.Instead of writing imperative crawl loops, you describe what to follow and what to extract in a single expression: import wxpath # Crawl, extract fields, build a Wikipedia knowledge graph path_expr = """ url('https://en.wikipedia.org/wiki/Expression_language') ///url(//main//a/@href[starts-with(., '/wiki/') and not(contains(., ':'))]) /map{ 'title': (//span[contains(@class, "mw-page-title-main")]/text())[1] ! string(.), 'url': string(base-uri(.)), 'short_description': //div[contains(@class, 'shortdescription')]/text() ! string(.), 'forward_links': //div[@id="mw-content-text"]//a/@href ! string(.) } """ for item in wxpath.wxpath_async_blocking_iter(path_expr, max_depth=1): print(item) The key addition is a `url(...)` operator that fetches and returns HTML for further XPath processing, and `///url(...)` for deep (or paginated) traversal. Everything else is standard XPath 3.1 (maps/arrays/functions).Features:- Async/concurrent crawling with streaming results- Scrapy-inspired auto-throttle and polite crawling- Hook system for custom processing- CLI for quick experimentsAnother example, paginating through HN comments (via "follow=" argument) pages and extracting data: url('https://news.ycombinator.com', follow=//a[text()='comments']/@href | //a[@class='morelink']/@href) //tr[@class='athing'] /map { 'text': .//div[@class='comment']//text(), 'user': .//a[@class='hnuser']/@href, 'parent_post': .//span[@class='onstory']/a/@href } Limitations: HTTP-only (no JS rendering yet), no crawl persistence. Both are on the roadmap if there's interest.GitHub: https://github.com/rodricios/wxpathPyPI: pip install wxpathI'd love feedback on the expression syntax and any use cases this might unlock.Thanks!
Enrichment
- Theme
- browser automation and scraping for AI
- Vertical
- Horizontal
- Function
- Dev tools
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- declarative web scraping with xpath
- Manually corrected
- False
Could you build this?
Yes It is a Python library parsing HTML using standard XPath engines (like lxml) and chaining network requests recursively based on XPath query evaluations.
Discussion
9 comments analyzed.
Competitors mentioned: Scrapy and Crawlee frameworks, Ferret declarative web crawling framework, Xidel, LLMs for web content extraction, lxml/XPath
Concerns raised: LLMs have token limits and cost issues, XPath lacked crawling capabilities
Competitors
Other products that read as similar to this one — 42 launches clear the similarity bar, closest 8 shown.
Attention rank: #6 of 43 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 70 days after the earliest competitor.
- PageSieve, a web scraping browser extension · hn · 2026-08-16 · 17 upvotes · similarity 0.45
- SCRAPR · ph · 2026-03-09 · 259 upvotes · similarity 0.44
- Crawio · ph · 2026-09-30 · 1 upvotes · similarity 0.41
- JSPath · ph · 2026-09-19 · 2 upvotes · similarity 0.39
- Fitter · hn · 2026-07-26 · 5 upvotes · similarity 0.37
- 301 Redirect Manager Silentfrog · ph · 2026-09-22 · 2 upvotes · similarity 0.37
- Draco · hn · 2026-08-02 · 15 upvotes · similarity 0.36
- Robust LLM extractor for websites in TypeScript · hn · 2026-03-26 · 72 upvotes · similarity 0.36
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a dev tools tool for Sales yet.