Smelt
Extract structured data from PDFs and HTML using LLM
Details
- External ID
- 47287378
- Source
- HN
- Company
- —
- Product
- Smelt
- Website domain
- github.com
- Launched
- March 7, 2026
- Cohort
- —
- Upvotes
- 6
- Upvotes percentile
- 0.2853628536285363
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
I built a CLI tool in Go that extracts structured data (JSON, CSV, Parquet) from messy PDFs and HTML pages.The core idea: LLMs are great at understanding structure but wasteful for bulk data extraction. So smelt uses a two-pass architecture:1. A fast Go capture layer parses the document and detects table-like regions 2. Those regions (not the whole document) get sent to Claude for schema inference — column names, types, nesting 3. The Go layer then does deterministic extraction using the inferred schemaThis means the LLM is never in the hot path of actual data processing. It figures out "what is this data?" once, and then Go handles the "extract 10,000 rows" part efficiently.Usage is simple: smelt invoice.pdf --format json smelt https://example.com/pricing --format csv smelt report.pdf --schema # just show the inferred structure You can also pass --query "extract the revenue table" to focus extraction when a document has multiple tables.Still early (no OCR yet, HTML is limited to <table> elements), but it handles the common cases well. Would love feedback on the architecture — especially from anyone who's dealt with PDF table extraction at scale.
Enrichment
- Theme
- browser automation and scraping for AI
- Vertical
- Horizontal
- Function
- Agent / copilot
- Audience
- B2B
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- extract structured data from documents using llm
- Manually corrected
- False
Could you build this?
Yes Smelt is a Go CLI that parses HTML/PDF text and prompts an LLM to extract structured schema output, easily scaffolded with standard PDF parsers and LLM APIs.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 97 launches clear the similarity bar, closest 8 shown.
Attention rank: #66 of 98 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 124 days after the earliest competitor.
- Tabstack Structured Extraction · ph · 2026-06-11 · 199 upvotes · similarity 0.56
- A new benchmark for testing LLMs for deterministic outputs · hn · 2026-04-29 · 60 upvotes · similarity 0.55
- Robust LLM extractor for websites in TypeScript · hn · 2026-03-26 · 72 upvotes · similarity 0.52
- PDFCraft · ph · 2026-09-13 · 2 upvotes · similarity 0.48
- DocumentsAI · ph · 2026-09-09 · 1 upvotes · similarity 0.47
- Misata · hn · 2025-12-16 · 24 upvotes · similarity 0.44
- Trawl · hn · 2026-03-08 · 8 upvotes · similarity 0.43
- OpenFable · hn · 2026-04-08 · 5 upvotes · similarity 0.43
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a agent / copilot tool for Agriculture yet.