A new engine to run Kimi K3 on a laptop
Details
- External ID
- 49098966
- Source
- HN
- Company
- —
- Product
- —
- Website domain
- —
- Launched
- July 29, 2026
- Cohort
- —
- Upvotes
- 7
- Upvotes percentile
- 0.3972520908004779
- Tags
- —
- Fetched at
- Sept. 8, 2026, 8:28 p.m.
- Updated at
- Sept. 8, 2026, 8:28 p.m.
Description
Kimi K3 has 2.78 trillion parameters and ships as 1.42 TB of weights. It clearly does not fit in the memory of a laptop. But K3 is a Mixture-of-Experts model. For each token, only a small fraction of its 896 experts per layer is activated. That changes the problem: the entire model does not need to be resident in RAM, as long as the weights required by each token can be reached quickly enough.We built WASTE — the Weight-Aware Streaming Tensor Engine — to explore that idea.WASTE keeps the dense, repeatedly used part of the model resident in memory, stores the routed experts in an NVMe-optimized container, and streams only the experts selected during inference. The remaining RAM is used as a bounded expert cache.The current Kimi K3 container is 982 GiB. On a 64 GB MacBook Pro, WASTE runs the complete model at around 0.32–0.34 tokens per second, with a measured minimum memory requirement of approximately 29 GB at a 4K context.That is obviously not interactive performance yet. But the result we found interesting is that it works at all: this is the full open-weights model, not a distillation, a pruned version, or a smaller model using the Kimi name.The engine is written in C and has no BLAS, CUDA, ONNX, or Python dependency in the inference path. The same code can be used through the CLI, embedded as a library, or exposed through the included OpenAI-compatible server.Correctness was the first constraint. Every layer was validated against a PyTorch reference, with final logits matching within 3.6e-06. The vision tower is supported as well and matches its reference within 2.3e-06.The current bottleneck is understood: K3 needs roughly 17 GB of expert data per token, and more than half of the decode time is spent reading experts from disk. The engine is already operating close to the measured throughput limit of the laptop’s internal SSD. The next improvements therefore need to reduce the number of bytes read per token and increase useful expert reuse without pushing the operating system into paging.K3 is deliberately the extreme case. The same engine runs Kimi-Linear 48B from a 19 GB container at 8.92 tokens per second with an 8 GB memory budget. The broader goal is to make models that are much larger than available RAM usable locally, without sending private data to an API and without requiring specialized accelerator hardware.We have published the engine, container format, conversion tools, benchmarks, validation suite, and also the experiments that failed rather than quietly removing them.Everything is fully open source. Feedback on the storage layout, quantization, caching strategy, direct I/O, portability, and potential optimizations would be very welcome. Contributions of any kind — code, benchmarks, testing on different hardware, documentation, bug reports, or new ideas — are more than appreciated.Repo: https://github.com/sqliteai/waste
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- engine to run kimi k3 on a laptop
- Manually corrected
- False
Could you build this?
No Running a 2.78T parameter MoE model with 1.42 TB of weights on consumer laptop hardware requires novel low-level systems engineering, custom disk-to-memory streaming, and advanced quantization.
What it would actually take: Building this requires custom C++/Rust and Metal/CUDA kernels optimized for SSD-to-RAM prefetching and dynamic sparse MoE routing. The developer must implement low-latency async I/O pipelines (io_uring or Apple direct storage APIs) and custom quantized matrix multiplication routines to swap expert weights on the fly per token. Deep expertise in systems programming, GPU memory architectures, and deep learning inference optimization is required.
Discussion
3 comments analyzed.
Competitors mentioned: M5 Pro
Competitors
Other products that read as similar to this one — 169 launches clear the similarity bar, closest 8 shown.
Attention rank: #112 of 170 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 267 days after the earliest competitor.
- Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac · hn · 2026-07-29 · 919 upvotes · similarity 0.52
- Kimi K3 · ph · 2026-07-17 · 447 upvotes · similarity 0.49
- Cuts Long Horizon Inference Costs by 50% via external KV Cache Offload · hn · 2026-07-26 · 22 upvotes · similarity 0.48
- Running Gemma-4 26B at 124 tokens/SEC on a CPU, no GPU · hn · 2026-06-30 · 10 upvotes · similarity 0.47
- Run Full Kimi K3 with 29 GB of RAM · hn · 2026-07-30 · 9 upvotes · similarity 0.46
- DenseK3 · github · 2026-09-15 · 8 upvotes · similarity 0.46
- RunNburn · hn · 2026-07-30 · 11 upvotes · similarity 0.46
- I shrank DeepSeek V4 Flash to 57GB and it wrote a compiler on my Mac · hn · 2026-08-16 · 21 upvotes · similarity 0.45
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.