Robust LLM extractor for websites in TypeScript
Details
- External ID
- 47526486
- Source
- HN
- Company
- —
- Product
- Robust LLM extractor for websites in TypeScript
- Website domain
- github.com
- Launched
- March 26, 2026
- Cohort
- —
- Upvotes
- 72
- Upvotes percentile
- 0.8788437884378844
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
We've been building data pipelines that scrape websites and extract structured data for a while now. If you've done this, you know the drill: you write CSS selectors, the site changes its layout, everything breaks at 2am, and you spend your morning rewriting parsers.LLMs seemed like the obvious fix — just throw the HTML at GPT and ask for JSON. Except in practice, it's more painful than that:- Raw HTML is full of nav bars, footers, and tracking junk that eats your token budget. A typical product page is 80% noise. - LLMs return malformed JSON more often than you'd expect, especially with nested arrays and complex schemas. One bad bracket and your pipeline crashes. - Relative URLs, markdown-escaped links, tracking parameters — the "small" URL issues compound fast when you're processing thousands of pages. - You end up writing the same boilerplate: HTML cleanup → markdown conversion → LLM call → JSON parsing → error recovery → schema validation. Over and over.We got tired of rebuilding this stack for every project, so we extracted it into a library.Lightfeed Extractor is a TypeScript library that handles the full pipeline from raw HTML to validated, structured data:- Converts HTML to LLM-ready markdown with main content extraction (strips nav, headers, footers), optional image inclusion, and URL cleaning - Works with any LangChain-compatible LLM (OpenAI, Gemini, Claude, Ollama, etc.) - Uses Zod schemas for type-safe extraction with real validation - Recovers partial data from malformed LLM output instead of failing entirely — if 19 out of 20 products parsed correctly, you get those 19 - Built-in browser automation via Playwright (local, serverless, or remote) with anti-bot patches - Pairs with our browser agent (@lightfeed/browser-agent) for AI-driven page navigation before extractionWe use this ourselves in production at Lightfeed, and it's been solid enough that we decided to open-source it.GitHub: https://github.com/lightfeed/extractor npm: npm install @lightfeed/extractor Apache 2.0 licensed.Happy to answer questions or hear feedback.
Enrichment
- Theme
- niche developer utilities and toolchains
- Vertical
- Horizontal
- Function
- Dev tools
- Audience
- Developer
- AI stance
- AI feature
- Project type
- Commercial product
- Normalized one-liner
- llm extractor for websites in typescript
- Manually corrected
- False
Could you build this?
Yes This is a TypeScript library that fetches web pages, cleans HTML/DOM, and passes relevant sections to structured LLM outputs via JSON schema or instructor-style extraction.
Discussion
20 comments analyzed.
Competitors mentioned: Playwright (browser automation library), Cloudflare (anti-bot protection), Datadome (anti-bot protection), Claude Code (LLM tool calling), Gemini 2.5 flash (LLM pricing alternative)
Concerns raised: Anti-bot detection (Cloudflare, recaptcha v2, proof of work), JSON output reliability and false positives, Token limits for large-scale scraping (700 tokens per page unsustainable), Cost at scale (1M pages = ~$210 with Gemini), Ethical concerns about scraping without publisher consent
Feature requests: Extract interactive/rendered data after DOM interactions (collapse bars, etc.), Support for recaptcha and proof-of-work challenges, HTML input mode to avoid browser overhead, Improved JSON output validation and repair, Better handling of nullable/optional fields in nested arrays
Competitors
Other products that read as similar to this one — 214 launches clear the similarity bar, closest 8 shown.
Attention rank: #23 of 215 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 146 days after the earliest competitor.
- Trawl · hn · 2026-03-08 · 8 upvotes · similarity 0.58
- Tabstack Structured Extraction · ph · 2026-06-11 · 199 upvotes · similarity 0.53
- Smelt · hn · 2026-03-07 · 6 upvotes · similarity 0.52
- TSM Tools · ph · 2026-09-09 · 2 upvotes · similarity 0.48
- I built an SDK that scrambles HTML so scrapers get garbage · hn · 2026-03-12 · 16 upvotes · similarity 0.47
- Rawkit · hn · 2026-02-12 · 8 upvotes · similarity 0.47
- Lambda 0.2 · hn · 2026-03-24 · 9 upvotes · similarity 0.46
- JustHTML · hn · 2025-12-02 · 6 upvotes · similarity 0.46
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a dev tools tool for Sales yet.