Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

WhiskeySour

A 10x faster drop-in replacement for BeautifulSoup

Details

External ID
47901770
Source
HN
Company
—
Product
—
Website domain
—
Launched
April 25, 2026
Cohort
—
Upvotes
8
Upvotes percentile
0.4813624678663239
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

The ProblemI’ve been using BeautifulSoup for sometime. It’s the standard for ease-of-use in Python scraping, but it almost always becomes the performance bottleneck when processing large-scale datasets.Parsing complex or massive HTML trees in Python typically suffers from high memory allocation costs and the overhead of the Python object model during tree traversal. In my production scraping workloads, the parser was consuming more CPU cycles than the network I/O. Lxml is fast but again uses up a lot of memory when processing large documents and has can cause trouble with malformed HTML.The SolutionI wanted to keep the API compatibility that makes BS4 great, but eliminates the overhead that slows down high-volume pipelines. It also uses html5ever which That’s why I built WhiskeySour. And yes… I *vibe coded the whole thing*.WhiskeySour is a drop-in replacement. You should be able to swap from "bs4 import BeautifulSoup" with "from whiskeysour import WhiskeySour" and see immediate speedups. Your workflows that used to take more than 30 mins might take less than 5 mins now.I have shared the detailed architecture of the library here: https://the-pro.github.io/whiskeySour/architecture/Here is the benchmark report against bs4 with html.parser: https://the-pro.github.io/whiskeySour/bench-report/Here is the link to the repo: https://github.com/the-pro/WhiskeySourWhy I’m sharing thisI’m looking for feedback from the community on two fronts:1. Edge cases: If you have particularly messy or malformed HTML that BS4 handles well, I’d love to know if WhiskeySour encounters any regressions.2. Benchmarks: If you are running high-volume parsers, I’d appreciate it if you could run a test on your own datasets and share the results.

Enrichment

Theme
niche developer utilities and toolchains
Vertical
Horizontal
Function
Dev tools
Audience
Developer
AI stance
Not AI
Project type
Commercial product
Normalized one-liner
faster html parsing library replacing beautifulsoup
Manually corrected
False

Could you build this?

No Writing a 10x faster drop-in replacement for BeautifulSoup requires implementing low-level HTML tokenization/parsing engines in C/Rust and exposing Python C-bindings with zero-copy semantics.

What it would actually take: To achieve 10x the speed of BeautifulSoup (which already wraps C parsers like lxml), one must build an HTML5 specification-compliant parser in Rust or C++ using SIMD-accelerated string scanning and custom arena allocators. It requires maintaining strict compatibility with BeautifulSoup's dynamic search and traversal API via PyO3 or CPython C-API while minimizing Python object creation overhead.

Discussion

1 comment analyzed.

Competitors

Other products that read as similar to this one — 97 launches clear the similarity bar, closest 8 shown.

Attention rank: #49 of 98 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 172 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a dev tools tool for Sales yet.