Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Getting GLM 5.2 running on my slow computer

Details

External ID
48842459
Source
HN
Company
—
Product
Getting GLM 5.2 running on my slow computer
Website domain
github.com
Launched
July 9, 2026
Cohort
—
Upvotes
937
Upvotes percentile
0.995221027479092
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

A few days ago I found myself trying out GLM 5.2 and was really positively impressed. The capabilities and security I was getting from this LLM are similar to those I've gotten from models like Claude or GPT, and this really surprised me.But then I thought, "I wonder how it would work on a normal computer like mine," and above all, "I wonder if it would work without going into OOM on a computer like mine." So I started working with the help of agents to test this possibility.I started converting the model to int4, understanding MTP usage, and if possible implementing DSA for long context. How it responds in int4 and whether the quality is maintained or not. Until I got to the point, on my computer with 32GB of RAM, I was able to communicate with GLM 5.2 with times that, of course, aren't high in cold start, but even then, we're talking about 0.1 tok/s, but that wasn't important to me. The important thing was the journey to reach this goal. I just wanted it to work at all costs, even slowly.So I created Colibrì, which was born from a very simple idea, to be honest, but tested in every way, where a 744B Mixture-of-Experts model activates only ~40B parameters per token—and only ~11 GB of those change from token to token (the routed experts). So:The dense part (attention, shared experts, embeddings—~17B params) stays resident in RAM at int4 (~9.9 GB); The 21,504 routed experts (75 MoE layers × 256 experts + the MTP head, ~19 MB each at int4) live on disk (~370 GB) and are streamed on demand, with a per-layer LRU cache, an optional pinned hot-store, and the OS page cache as a free L2.The engine is a single C file (c/glm.c, ~1,300 lines) plus small headers. No BLAS, no Python at runtime, no GPU.No GPU or serious hardware because I don't have that hardware so I can't test it on hardware that is more powerful than my computer.Colibrì is a one-person project, written and tested entirely on a 12-core laptop with 25 GB of RAM — the numbers above are the ceiling of what I can measure at home.Any feedback is welcome! (and if anyone wanted to participate in the project I would be delighted)Repo: https://github.com/JustVugg/colibri

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
—
Audience
Developer
AI stance
—
Project type
Hobby / open-source project
Normalized one-liner
—
Manually corrected
False

Could you build this?

No Developing or adapting high-efficiency LLM runtime engines to execute massive models on resource-constrained hardware requires low-level systems programming and hardware-specific kernel optimization.

What it would actually take: A working implementation requires a custom runtime in C++ or Rust (analogous to llama.cpp), implementing specialized quantization schemes (such as AWQ, GGUF k-quants, or EXL2) and optimized SIMD/GPU compute kernels (AVX-512, Metal, Vulkan). The hard technical problem is designing memory-mapped weight paging, KV-cache quantization, and aggressive layer-offloading to bypass hardware RAM and memory bandwidth constraints. This necessitates deep expertise in machine learning systems, hardware architectures, and low-level numerical computing.

Discussion

20 comments analyzed.

Competitors mentioned: Bedrock, Codex with GPT-5.4mini, Ollama, llama.cpp, LM Studio

Concerns raised: High hardware costs (8 Nvidia B200s needed), Poor performance on standard laptops (0.07 tokens/s reported), Default context window too small for complex prompts, Slow token generation rates make practical use difficult, Ollama's 4-bit quants and short context break on complex tasks

Feature requests: Support for Ubuntu as coding agent, EC2/cloud deployment optimization, Lower RAM requirements (10GB or less), Better memory management for KV cache and weights, Multi-token decoding improvements

Competitors

Other products that read as similar to this one — 89 launches clear the similarity bar, closest 8 shown.

Attention rank: #3 of 90 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 239 days after the earliest competitor.

Other launches for this product