I wrote a 1-bit WebGPU runtime to run a 1.7B LLM in the browser
Details
- External ID
- 48820583
- Source
- HN
- Company
- —
- Product
- I wrote a 1-bit WebGPU runtime to run a 1.7B LLM in the browser
- Website domain
- aidekin.com
- Launched
- July 7, 2026
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.1081242532855436
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Enrichment
- Theme
- gpu compute and acceleration tools
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI feature
- Project type
- Hobby / open-source project
- Normalized one-liner
- quantized llm runtime for browser webgpu
- Manually corrected
- False
Could you build this?
No Writing a custom 1-bit quantized inference runtime in WebGPU shader language (WGSL) to execute a 1.7B parameter language model in the browser requires cutting-edge GPU compute and ML systems expertise.
What it would actually take: This requires implementing custom WGSL compute shaders for 1-bit / sub-byte matrix multiplication (such as BitNet b1.58 or ternary/binary GEMM kernels), managing WebGPU memory buffers, and implementing a complete transformer execution pipeline (KV cache, RMSNorm, RoPE, attention). The core difficulty lies in optimizing memory bandwidth and SIMD/subgroup operations within browser GPU constraints to achieve acceptable tokens-per-second inference speeds.
Discussion
2 comments analyzed.
Concerns raised: 300MB initial download may be too large for quick support Q&A use case, Download time delay before model is available for questions
Feature requests: Hosted fallback Q&A while model downloads, Cache model locally after initial download to avoid re-downloading
Competitors
Other products that read as similar to this one — 448 launches clear the similarity bar, closest 8 shown.
Attention rank: #404 of 449 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 250 days after the earliest competitor.
- Running Win32/DirectX games in the browser via x86 emulation and WebGPU · hn · 2026-07-14 · 5 upvotes · similarity 0.57
- Open-source calculator for "will my GPU run this LLM?" · hn · 2026-08-24 · 5 upvotes · similarity 0.56
- INXM // local` OSS for using LLM as compiler and not as runtime · hn · 2026-08-19 · 5 upvotes · similarity 0.54
- Nanbeige 4.1-3B running in the browser via WebGPU · hn · 2026-02-19 · 6 upvotes · similarity 0.53
- I just released v7 Javalin, a JVM web framework · hn · 2026-02-24 · 6 upvotes · similarity 0.48
- The Simpsons Hit and Run Running in the Browser (WASM/WebGL) · hn · 2026-04-15 · 5 upvotes · similarity 0.48
- Train a language model in the browser with WebGPU · hn · 2025-11-21 · 6 upvotes · similarity 0.48
- A tiny LLM running at 21,000 tok/s on a $250 FPGA (Live Demo) · hn · 2026-08-10 · 79 upvotes · similarity 0.46
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.