Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

wxpath

Declarative web crawling in XPath

Details

External ID
46618472
Source
HN
Company
—
Product
wxpath
Website domain
github.com
Launched
Jan. 14, 2026
Cohort
—
Upvotes
64
Upvotes percentile
0.8405797101449275
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

wxpath is a declarative web crawler where web crawling and scraping are expressed directly in XPath.Instead of writing imperative crawl loops, you describe what to follow and what to extract in a single expression: import wxpath # Crawl, extract fields, build a Wikipedia knowledge graph path_expr = """ url('https://en.wikipedia.org/wiki/Expression_language') ///url(//main//a/@href[starts-with(., '/wiki/') and not(contains(., ':'))]) /map{ 'title': (//span[contains(@class, "mw-page-title-main")]/text())[1] ! string(.), 'url': string(base-uri(.)), 'short_description': //div[contains(@class, 'shortdescription')]/text() ! string(.), 'forward_links': //div[@id="mw-content-text"]//a/@href ! string(.) } """ for item in wxpath.wxpath_async_blocking_iter(path_expr, max_depth=1): print(item) The key addition is a `url(...)` operator that fetches and returns HTML for further XPath processing, and `///url(...)` for deep (or paginated) traversal. Everything else is standard XPath 3.1 (maps/arrays/functions).Features:- Async/concurrent crawling with streaming results- Scrapy-inspired auto-throttle and polite crawling- Hook system for custom processing- CLI for quick experimentsAnother example, paginating through HN comments (via "follow=" argument) pages and extracting data: url('https://news.ycombinator.com', follow=//a[text()='comments']/@href | //a[@class='morelink']/@href) //tr[@class='athing'] /map { 'text': .//div[@class='comment']//text(), 'user': .//a[@class='hnuser']/@href, 'parent_post': .//span[@class='onstory']/a/@href } Limitations: HTTP-only (no JS rendering yet), no crawl persistence. Both are on the roadmap if there's interest.GitHub: https://github.com/rodricios/wxpathPyPI: pip install wxpathI'd love feedback on the expression syntax and any use cases this might unlock.Thanks!

Enrichment

Theme
browser automation and scraping for AI
Vertical
Horizontal
Function
Dev tools
Audience
Developer
AI stance
Not AI
Project type
Hobby / open-source project
Normalized one-liner
declarative web scraping with xpath
Manually corrected
False

Could you build this?

Yes It is a Python library parsing HTML using standard XPath engines (like lxml) and chaining network requests recursively based on XPath query evaluations.

Discussion

9 comments analyzed.

Competitors mentioned: Scrapy and Crawlee frameworks, Ferret declarative web crawling framework, Xidel, LLMs for web content extraction, lxml/XPath

Concerns raised: LLMs have token limits and cost issues, XPath lacked crawling capabilities

Competitors

Other products that read as similar to this one — 42 launches clear the similarity bar, closest 8 shown.

Attention rank: #6 of 43 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 70 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a dev tools tool for Sales yet.