Getting GLM 5.2 running on my slow computer
Details
- External ID
- 48842459
- Source
- HN
- Company
- —
- Product
- Getting GLM 5.2 running on my slow computer
- Website domain
- github.com
- Launched
- July 9, 2026
- Cohort
- —
- Upvotes
- 937
- Upvotes percentile
- 0.995221027479092
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
A few days ago I found myself trying out GLM 5.2 and was really positively impressed. The capabilities and security I was getting from this LLM are similar to those I've gotten from models like Claude or GPT, and this really surprised me.But then I thought, "I wonder how it would work on a normal computer like mine," and above all, "I wonder if it would work without going into OOM on a computer like mine." So I started working with the help of agents to test this possibility.I started converting the model to int4, understanding MTP usage, and if possible implementing DSA for long context. How it responds in int4 and whether the quality is maintained or not. Until I got to the point, on my computer with 32GB of RAM, I was able to communicate with GLM 5.2 with times that, of course, aren't high in cold start, but even then, we're talking about 0.1 tok/s, but that wasn't important to me. The important thing was the journey to reach this goal. I just wanted it to work at all costs, even slowly.So I created Colibrì, which was born from a very simple idea, to be honest, but tested in every way, where a 744B Mixture-of-Experts model activates only ~40B parameters per token—and only ~11 GB of those change from token to token (the routed experts). So:The dense part (attention, shared experts, embeddings—~17B params) stays resident in RAM at int4 (~9.9 GB); The 21,504 routed experts (75 MoE layers × 256 experts + the MTP head, ~19 MB each at int4) live on disk (~370 GB) and are streamed on demand, with a per-layer LRU cache, an optional pinned hot-store, and the OS page cache as a free L2.The engine is a single C file (c/glm.c, ~1,300 lines) plus small headers. No BLAS, no Python at runtime, no GPU.No GPU or serious hardware because I don't have that hardware so I can't test it on hardware that is more powerful than my computer.Colibrì is a one-person project, written and tested entirely on a 12-core laptop with 25 GB of RAM — the numbers above are the ceiling of what I can measure at home.Any feedback is welcome! (and if anyone wanted to participate in the project I would be delighted)Repo: https://github.com/JustVugg/colibri
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- —
- Audience
- Developer
- AI stance
- —
- Project type
- Hobby / open-source project
- Normalized one-liner
- —
- Manually corrected
- False
Could you build this?
No Developing or adapting high-efficiency LLM runtime engines to execute massive models on resource-constrained hardware requires low-level systems programming and hardware-specific kernel optimization.
What it would actually take: A working implementation requires a custom runtime in C++ or Rust (analogous to llama.cpp), implementing specialized quantization schemes (such as AWQ, GGUF k-quants, or EXL2) and optimized SIMD/GPU compute kernels (AVX-512, Metal, Vulkan). The hard technical problem is designing memory-mapped weight paging, KV-cache quantization, and aggressive layer-offloading to bypass hardware RAM and memory bandwidth constraints. This necessitates deep expertise in machine learning systems, hardware architectures, and low-level numerical computing.
Discussion
20 comments analyzed.
Competitors mentioned: Bedrock, Codex with GPT-5.4mini, Ollama, llama.cpp, LM Studio
Concerns raised: High hardware costs (8 Nvidia B200s needed), Poor performance on standard laptops (0.07 tokens/s reported), Default context window too small for complex prompts, Slow token generation rates make practical use difficult, Ollama's 4-bit quants and short context break on complex tasks
Feature requests: Support for Ubuntu as coding agent, EC2/cloud deployment optimization, Lower RAM requirements (10GB or less), Better memory management for KV cache and weights, Multi-token decoding improvements
Competitors
Other products that read as similar to this one — 89 launches clear the similarity bar, closest 8 shown.
Attention rank: #3 of 90 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 239 days after the earliest competitor.
- Lumabri · hn · 2026-08-09 · 9 upvotes · similarity 0.55
- RunNburn · hn · 2026-07-30 · 11 upvotes · similarity 0.41
- High speed graphics rendering research with tinygrad/tinyJIT · hn · 2026-01-22 · 31 upvotes · similarity 0.40
- Shoehorn, a library to quantize an LLM to fit your Mac's VRAM · hn · 2026-08-14 · 6 upvotes · similarity 0.40
- glm53-tensorfold-spark · github · 2026-09-28 · 76 upvotes · similarity 0.38
- Claude Code skills that build complete Godot games · hn · 2026-03-16 · 337 upvotes · similarity 0.37
- Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac · hn · 2026-07-29 · 919 upvotes · similarity 0.37
- ChonkLM · hn · 2026-05-09 · 6 upvotes · similarity 0.37
Other launches for this product
- No other launches for this product.