Z80-μLM, a 'Conversational AI' That Fits in 40KB
Details
- External ID
- 46417815
- Source
- HN
- Company
- —
- Product
- Z80-μLM, a 'Conversational AI' That Fits in 40KB
- Website domain
- github.com
- Launched
- Dec. 29, 2025
- Cohort
- —
- Upvotes
- 514
- Upvotes percentile
- 0.9904580152671756
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
How small can a language model be while still doing something useful? I wanted to find out, and had some spare time over the holidays.Z80-μLM is a character-level language model with 2-bit quantized weights ({-2,-1,0,+1}) that runs on a Z80 with 64KB RAM. The entire thing: inference, weights, chat UI, it all fits in a 40KB .COM file that you can run in a CP/M emulator and hopefully even real hardware!It won't write your emails, but it can be trained to play a stripped down version of 20 Questions, and is sometimes able to maintain the illusion of having simple but terse conversations with a distinct personality.--The extreme constraints nerd-sniped me and forced interesting trade-offs: trigram hashing (typo-tolerant, loses word order), 16-bit integer math, and some careful massaging of the training data meant I could keep the examples 'interesting'.The key was quantization-aware training that accurately models the inference code limitations. The training loop runs both float and integer-quantized forward passes in parallel, scoring the model on how well its knowledge survives quantization. The weights are progressively pushed toward the 2-bit grid using straight-through estimators, with overflow penalties matching the Z80's 16-bit accumulator limits. By the end of training, the model has already adapted to its constraints, so no post-hoc quantization collapse.Eventually I ended up spending a few dollars on Claude API to generate 20 questions data (see examples/guess/GUESS.COM), I hope Anthropic won't send me a C&D for distilling their model against the ToS ;PBut anyway, happy code-golf season everybody :)
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- lightweight conversational ai for embedded systems
- Manually corrected
- False
Could you build this?
No Fitting a 2-bit quantized neural network inference engine and interactive UI inside the strict 40KB-64KB memory boundary of an 8-bit Z80 microprocessor demands extreme low-level embedded engineering and assembly hacking.
What it would actually take: This requires implementing custom quantization schemes (ternary/2-bit arithmetic), packing weights into nibbles, and writing hand-tuned Z80 assembly to execute matrix-vector multiplications without hardware multiplier units. The stack relies on cross-compilers like SDCC and Z80 emulators (or real hardware testing). Deep expertise in embedded architectures, instruction cycle budgeting, and low-bit quantized neural network architectures is mandatory.
Discussion
20 comments analyzed.
Competitors mentioned: Slack, Discord, Teams, Windows 2000/XP, IE 7
Concerns raised: Slow execution speed on retro hardware (Model I, Z-80 machines), Very slow on actual hardware (1 min 9 sec for 2-char response), Modern chat apps bloated/resource-intensive due to Chrome/HTML overhead, Business logic doesn't justify resource consumption, Model size/memory requirements for unlimited context window
Feature requests: Unlimited context window capability, English-only mode to reduce model size, Support for mobile devices under $300 with basic specs, Smaller model footprint for smartphone deployment, Layer-based organization to minimize bank switching overhead
Competitors
Other products that read as similar to this one — 253 launches clear the similarity bar, closest 8 shown.
Attention rank: #9 of 254 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 60 days after the earliest competitor.
- GLM-5.3 · ph · 2026-08-15 · 253 upvotes · similarity 0.49
- Nari Qwen3-TTS and Qwen3-ASR · hn · 2026-09-14 · 90 upvotes · similarity 0.49
- Three new Kitten TTS models · hn · 2026-03-19 · 561 upvotes · similarity 0.47
- I ran a language model on a PS2 · hn · 2026-03-21 · 46 upvotes · similarity 0.46
- ChonkLM · hn · 2026-05-09 · 6 upvotes · similarity 0.44
- giga-embeddings-10B-A1.8B_hybrid · github · 2026-09-25 · 41 upvotes · similarity 0.44
- I built a tiny LLM to demystify how language models work · hn · 2026-04-06 · 915 upvotes · similarity 0.44
- Moonshine Open-Weights STT models · hn · 2026-02-24 · 316 upvotes · similarity 0.43
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.