Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

JustHTML

A pure Python HTML5 parser that just works

Details

External ID
46124443
Source
HN
Company
—
Product
JustHTML
Website domain
github.com
Launched
Dec. 2, 2025
Cohort
—
Upvotes
6
Upvotes percentile
0.2652671755725191
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

I got frustrated with HTML parsing in Python.I wanted a Python HTML parser that was both correct and easy to install. The C-based ones (lxml, selectolax) are fast but not HTML5 compliant. The pure Python ones (html.parser, BeautifulSoup's default) are easy to install but choke on real-world HTML. html5lib is 80% correct but painfully slow.So I wrote JustHTML. It's:• 100% HTML5 compliant – passes all 8,500+ html5lib tests. If a browser can parse it, JustHTML can.• Pure Python, zero dependencies – pip install and go. Works on PyPy, Pyodide, anywhere.• Fast enough – ~0.1s to parse Wikipedia's homepage. Not C-fast, but 50% faster than html5lib.• Simple API – doc.query("div.foo > p") with CSS selectors. One method to learn.Example: from justhtml import JustHTML doc = JustHTML("<div><p class='intro'>Hello!</p></div>") print(doc.query(".intro")[0].to_html()) I've fuzz-tested it with 3 million malformed documents.Would love feedback, especially on the API design.

Enrichment

Theme
niche developer utilities and toolchains
Vertical
—
Function
Dev tools
Audience
Developer
AI stance
Not AI
Project type
Hobby / open-source project
Normalized one-liner
pure python html5 parser
Manually corrected
False

Could you build this?

Partial Implementing a full HTML5-compliant specification parser in pure Python involves an enormous, notoriously complex state machine with exhaustive edge-case error recovery rules.

What it would actually take: The WHATWG HTML specification requires implementing hundreds of specific tokenization and tree-construction states, insertion modes, and adoption agency algorithms. Building a compliant parser requires writing and debugging against the extensive html5lib-tests suite, tracking complex token-stream mutations, and handling subtle character encoding fallbacks. A developer would need deep knowledge of the WHATWG parsing algorithm and months of rigorous testing against edge cases.

Discussion

2 comments analyzed.

Concerns raised: RSS feed throwing errors

Competitors

Other products that read as similar to this one — 96 launches clear the similarity bar, closest 8 shown.

Attention rank: #80 of 97 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 31 days after the earliest competitor.

Other launches for this product