Pulpie
Models for Cleaning the Web
Details
- External ID
- 48806575
- Source
- HN
- Company
- —
- Product
- Pulpie
- Website domain
- usefeyn.com
- Launched
- July 6, 2026
- Cohort
- —
- Upvotes
- 106
- Upvotes percentile
- 0.9241338112305855
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Hey HN, I'm Shreyash, founder of Feyn. We built Pulpie, a family of Pareto optimal models for cleaning the web. Pulpie strips boilerplate (ads, footers, sidebars) from raw HTML and returns just the main content as HTML or Markdown.We match SOTA extraction quality while being 20x cheaper. Cleaning 1 billion webpages costs $7,900 with Pulpie versus $159,000 with Dripper, the current leading extractor.The gains come from architecture. Today's leading extractors are decoders that generate output one token at a time. Each step reads the full model from memory to produce a single token. Conversely, Pulpie models are encoders. They run one forward pass over the full input HTML and label each block as boilerplate or content. As a result, Pulpie is compute-bound while decoders are memory-bound. Cheaper GPUs have relatively more compute than memory bandwidth. This makes Pulpie easy to run optimally.Here's Pulpie and Dripper cleaning the same pages side by side: https://www.youtube.com/watch?v=ibd-tIiQECo. You can try a side-by-side comparison yourself: https://huggingface.co/spaces/feyninc/pulpieOur motivation for Pulpie came from building a deep research harness. Every search API returns noisy content that contains ads, nav elements, and sidebars. In one instance, an ad for "Gemini on Pixel" slipped into our search results, got passed into LLM context, and ended up in the final answer served to the user. Pretty embarrassing moment for us but it helped us realize how bad data kills model intelligence. We built Pulpie to get clean data for cheap.All models are open source on Hugging Face. You can read about our training process and how to use Pulpie here: https://usefeyn.com/blog/pulpie-pareto-optimal-models-for-cl...Happy to answer any questions!
Enrichment
- Theme
- document processing and generation tools
- Vertical
- Horizontal
- Function
- Data infrastructure
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- ai models for web data cleaning
- Manually corrected
- False
Could you build this?
No Developing custom ML models that achieve state-of-the-art HTML content extraction at 20x lower cost requires novel model architecture, distillation, and specialized web datasets.
What it would actually take: Requires curating an extensive, representative multi-language web corpus (e.g., Common Crawl subset) with gold-standard main content labels, training a teacher LLM/vision model, and distilling it into small, Pareto-optimal transformers or CNN/GNN architectures over DOM trees. Model optimization for high-throughput inference (ONNX/TensorRT) and custom loss functions for DOM subtree classification are critical. This needs seasoned ML researchers and large-scale data engineering pipelines.
Discussion
20 comments analyzed.
Competitors mentioned: Defuddle and Readability (heuristic-based approaches), Trafilatura, Traditional deterministic HTML to Markdown conversion
Concerns raised: Performance on JavaScript-rendered pages, Running on modest hardware/homelab setups, Cost comparison vs. free HTML-to-Markdown converters
Feature requests: Support for ecommerce product scraping (Amazon, Shopee), Handle shadow DOMs in addition to images and tables, Extract tabular data from web pages
Competitors
Other products that read as similar to this one — 26 launches clear the similarity bar, closest 8 shown.
Attention rank: #3 of 27 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 202 days after the earliest competitor.
- I built an SDK that scrambles HTML so scrapers get garbage · hn · 2026-03-12 · 16 upvotes · similarity 0.39
- CleanMD · ph · 2026-09-15 · 1 upvotes · similarity 0.37
- Watermarks Remover: Clean LLM watermarks from text and files · hn · 2026-08-28 · 5 upvotes · similarity 0.36
- WhiskeySour · hn · 2026-04-25 · 8 upvotes · similarity 0.36
- Watermark Eraser · ph · 2026-09-11 · 1 upvotes · similarity 0.35
- I built a clipboard tool to strip/keep specific formatting like Italics · hn · 2026-01-02 · 38 upvotes · similarity 0.35
- Clean HTML for Semantic Extraction · hn · 2026-01-09 · 5 upvotes · similarity 0.33
- Irpapers · hn · 2026-02-23 · 5 upvotes · similarity 0.33
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a data infrastructure tool for Media & entertainment yet.