JustHTML
A pure Python HTML5 parser that just works
Details
- External ID
- 46124443
- Source
- HN
- Company
- —
- Product
- JustHTML
- Website domain
- github.com
- Launched
- Dec. 2, 2025
- Cohort
- —
- Upvotes
- 6
- Upvotes percentile
- 0.2652671755725191
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
I got frustrated with HTML parsing in Python.I wanted a Python HTML parser that was both correct and easy to install. The C-based ones (lxml, selectolax) are fast but not HTML5 compliant. The pure Python ones (html.parser, BeautifulSoup's default) are easy to install but choke on real-world HTML. html5lib is 80% correct but painfully slow.So I wrote JustHTML. It's:• 100% HTML5 compliant – passes all 8,500+ html5lib tests. If a browser can parse it, JustHTML can.• Pure Python, zero dependencies – pip install and go. Works on PyPy, Pyodide, anywhere.• Fast enough – ~0.1s to parse Wikipedia's homepage. Not C-fast, but 50% faster than html5lib.• Simple API – doc.query("div.foo > p") with CSS selectors. One method to learn.Example: from justhtml import JustHTML doc = JustHTML("<div><p class='intro'>Hello!</p></div>") print(doc.query(".intro")[0].to_html()) I've fuzz-tested it with 3 million malformed documents.Would love feedback, especially on the API design.
Enrichment
- Theme
- niche developer utilities and toolchains
- Vertical
- —
- Function
- Dev tools
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- pure python html5 parser
- Manually corrected
- False
Could you build this?
Partial Implementing a full HTML5-compliant specification parser in pure Python involves an enormous, notoriously complex state machine with exhaustive edge-case error recovery rules.
What it would actually take: The WHATWG HTML specification requires implementing hundreds of specific tokenization and tree-construction states, insertion modes, and adoption agency algorithms. Building a compliant parser requires writing and debugging against the extensive html5lib-tests suite, tracking complex token-stream mutations, and handling subtle character encoding fallbacks. A developer would need deep knowledge of the WHATWG parsing algorithm and months of rigorous testing against edge cases.
Discussion
2 comments analyzed.
Concerns raised: RSS feed throwing errors
Competitors
Other products that read as similar to this one — 96 launches clear the similarity bar, closest 8 shown.
Attention rank: #80 of 97 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 31 days after the earliest competitor.
- WhiskeySour · hn · 2026-04-25 · 8 upvotes · similarity 0.52
- Robust LLM extractor for websites in TypeScript · hn · 2026-03-26 · 72 upvotes · similarity 0.46
- Justif · hn · 2026-07-17 · 5 upvotes · similarity 0.45
- Myjs An accidental pure-Python JavaScript interpreter · hn · 2026-09-05 · 7 upvotes · similarity 0.43
- Lambda 0.2 · hn · 2026-03-24 · 9 upvotes · similarity 0.42
- Justif · hn · 2026-07-17 · 198 upvotes · similarity 0.41
- python-data-parser · github · 2026-09-27 · 8 upvotes · similarity 0.41
- Trawl · hn · 2026-03-08 · 8 upvotes · similarity 0.40
Other launches for this product
- No other launches for this product.