Distilled 0.6B text-to-SQL model
Details
- External ID
- 46703884
- Source
- HN
- Company
- —
- Product
- Distilled 0.6B text-to-SQL model
- Website domain
- github.com
- Launched
- Jan. 21, 2026
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.09617918313570488
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
We used our platform to fine-tune a tiny text-to-SQL model using distillation from DeepSeek V3. Repo has instructions for how to replicate this.This is definitely not the best-performing model like this out there! But I found it surprising we were able to get to this much out of it: stone's throw away from a teacher 1000x the size!We also ran the same thing using the 4B Qwen and matched the teacher accuracy, though here the difference is merely 100x :)I find this pretty cool - obviously our distilled models can only do this one task and don't generalize, but that's often exactly what you want when you're building agentic systems.Happy to answer any questions!
Enrichment
- Theme
- database infrastructure and developer tools
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- small text-to-sql language model
- Manually corrected
- False
Could you build this?
Partial While running a fine-tuning script is accessible, generating high-quality synthetic text-to-SQL distillation datasets from a teacher model and tuning a sub-billion parameter model to retain generalization requires ML engineering expertise.
What it would actually take: The core pipeline consists of PyTorch, Hugging Face Transformers, and LoRA/QLoRA or full fine-tuning on a small base model (like Qwen2.5-0.5B). The difficult piece is creating and filtering the distillation dataset: schema serialization, prompt synthesis using DeepSeek-V3, and strict execution-based SQL validation to prune hallucinatory queries. Tuning hyperparameters to prevent catastrophic forgetting on tiny models requires empirical ML evaluation and GPU compute resources.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 115 launches clear the similarity bar, closest 8 shown.
Attention rank: #114 of 116 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 84 days after the earliest competitor.
- Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it · hn · 2026-07-30 · 170 upvotes · similarity 0.53
- Llm.sql · hn · 2026-04-24 · 8 upvotes · similarity 0.45
- Prela · hn · 2026-06-15 · 6 upvotes · similarity 0.45
- JevBench, a reproducible benchmark for typed decision models · hn · 2026-09-22 · 149 upvotes · similarity 0.42
- I implemented a neural network in SQL · hn · 2026-07-13 · 121 upvotes · similarity 0.41
- SQLazy · ph · 2026-09-21 · 2 upvotes · similarity 0.41
- SLayer, a semantic layer maintained by your agent · hn · 2026-05-11 · 12 upvotes · similarity 0.40
- Z80-μLM, a 'Conversational AI' That Fits in 40KB · hn · 2025-12-29 · 514 upvotes · similarity 0.39
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.