ChonkLM
Tiny language models running offline in the browser
Details
- External ID
- 48077627
- Source
- HN
- Company
- —
- Product
- ChonkLM
- Website domain
- chonklm.com
- Launched
- May 9, 2026
- Cohort
- —
- Upvotes
- 6
- Upvotes percentile
- 0.3053311793214863
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
I had been looking to try <500M parameter language models but you wouldn't find an API to try them anywhere, so I built this cloudflare hosted static website that hosts weights and built an inference runtime for these models that uses WebGPU and runs inference from your browser.These are only so useful in a multi-turn conversation but it's still interesting to see what you can pack in a <250mb model.I tried using ONNX versions earlier, but there were too many quirks of using them with language models and the TPS wasn't too impressive. Inspired by svenflow/webgpu-gemma, I put my codex and claude to the task of writing WGSL to run inference for GGUF versions of these models.Once you load this website and a model, it should load offline too, until your browser evicts the model from the cache.
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- small language models running offline in browser
- Manually corrected
- False
Could you build this?
No Writing a custom WebGPU transformer inference runtime from scratch requires specialized expertise in low-level GPU programming, WGSL compute shaders, memory hierarchy management, and quantized tensor operations.
What it would actually take: A real implementation requires a custom WebGPU pipeline using WGSL compute shaders for GEMM (General Matrix Multiply), scaled dot-product attention, RoPE positional embeddings, and KV-cache management. Model weights must be parsed and mapped directly into WebGPU GPUBuffer allocations with appropriate alignment and quantization (e.g., int8/int4 or float16). The hard part is authoring numerically stable, performant compute shaders that avoid GPU timeouts and browser-specific WebGPU driver crashes, requiring a graphics and systems ML engineer.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 81 launches clear the similarity bar, closest 8 shown.
Attention rank: #61 of 82 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 182 days after the earliest competitor.
- RunNburn · hn · 2026-07-30 · 11 upvotes · similarity 0.48
- General Compute · ph · 2026-05-22 · 314 upvotes · similarity 0.45
- Z80-μLM, a 'Conversational AI' That Fits in 40KB · hn · 2025-12-29 · 514 upvotes · similarity 0.44
- WebGCM · hn · 2026-09-21 · 6 upvotes · similarity 0.44
- Wally by RunAnywhere: The fastest inference for open frontier models · yc · 2026-09-29 · 5 upvotes · similarity 0.44
- Neurogrid Community Cloud · ph · 2026-09-15 · 1 upvotes · similarity 0.43
- Serve 100 Large AI models on a single GPU with low impact to TTFT · hn · 2025-11-08 · 7 upvotes · similarity 0.42
- Train a language model in the browser with WebGPU · hn · 2025-11-21 · 6 upvotes · similarity 0.41
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.