Partial content web crawling using HTTP/2 and Go
Details
- External ID
- 46766157
- Source
- HN
- Company
- —
- Product
- Partial content web crawling using HTTP/2 and Go
- Website domain
- substack.com
- Launched
- Jan. 26, 2026
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.09617918313570488
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Hi, I wrote a low-level HTTP/2 web crawler in Go, which can scrape partial content to save traffic.Tl;dr e.g. the HTML of a YouTube video contains the video description, views, likes etc. in its first 600KB, the remaining 900KB are of no use for me, but I have to pay my proxies by the gigabyte.My crawler receives packet per packet, and if I got everything I needed I reset the request, and only pay-for-what-i-crawled.This is also potentially useful for large-scale crawling operations, where duplicates matter. You could compute a simHash on the fly, and reset on-the-fly before crawling the entire document (again).
Enrichment
- Theme
- proxy, dns, and networking tools
- Vertical
- Horizontal
- Function
- Dev tools
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- partial content web crawling with http/2
- Manually corrected
- False
Could you build this?
Partial A standard Go web scraper is trivial to build, but custom HTTP/2 connection handling that resets streams mid-flight to avoid downloading full payloads requires low-level network protocol manipulation.
What it would actually take: A Go-based crawler that truncates downloads must bypass standard `net/http` high-level abstractions to interact directly with HTTP/2 framing (`golang.org/x/net/http2`). The hard technical part is safely terminating the stream using `RST_STREAM` (with `NO_ERROR` or `CANCEL`) immediately once target byte thresholds are reached while keeping the underlying TCP/TLS connection alive and proxy multiplexing intact. Building this requires deep familiarity with RFC 7540 (HTTP/2 framing and stream lifecycle state machines) and custom transport layer implementation in Go.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 85 launches clear the similarity bar, closest 8 shown.
Attention rank: #85 of 86 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 61 days after the earliest competitor.
- scrape.land · ph · 2026-09-21 · 11 upvotes · similarity 0.45
- Crawio · ph · 2026-09-30 · 1 upvotes · similarity 0.44
- Scrappy · hn · 2026-02-16 · 10 upvotes · similarity 0.44
- ClipHutch · ph · 2026-09-12 · 2 upvotes · similarity 0.42
- ZipSee · hn · 2026-04-04 · 5 upvotes · similarity 0.42
- Trying to fix the web scraping industry's benchmark problem · hn · 2026-07-16 · 18 upvotes · similarity 0.41
- I wrote a ~2KB executable file HTTP file downloader without Libc · hn · 2026-03-29 · 9 upvotes · similarity 0.41
- Bulk Image Download · ph · 2026-09-22 · 1 upvotes · similarity 0.41
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a dev tools tool for Sales yet.