I shrank DeepSeek V4 Flash to 57GB and it wrote a compiler on my Mac
Details
- External ID
- 49321813
- Source
- HN
- Company
- —
- Product
- I shrank DeepSeek V4 Flash to 57GB and it wrote a compiler on my Mac
- Website domain
- huggingface.co
- Launched
- Aug. 16, 2026
- Cohort
- —
- Upvotes
- 21
- Upvotes percentile
- 0.7547043010752689
- Tags
- —
- Fetched at
- Sept. 10, 2026, 5:32 a.m.
- Updated at
- Sept. 10, 2026, 5:32 a.m.
Description
I built a specialized package of DeepSeek V4 Flash 0731 (originally 284B total parameters, 13B active), preserving reasoning, tool calling and coding capabilities:https://huggingface.co/steadfastgaze/DeepSeek-V4-Flash-0731-...I let it write a minimal C compiler targeting ARM64, then test the result with Fibonacci and FizzBuzz programs, and it succeeded in less than 1 hour, with the full recording at:https://youtu.be/XiwSilmV8B0You can run it on Silicon Macs with my engine https://github.com/steadfastgaze/MoEspresso, while one of the core libraries developed to obtain this result is available at https://github.com/steadfastgaze/mlx-iqk.The above recording was on a 128GB memory MacBook M3 Max, but you can also run it on 32GB MacBooks with a very usable context (128K tokens) and projected 5 tok/s. I did try it on a fanless 16GB memory MacBook Air M1 (1.39 tok/s), but unfortunately the available context was very small.How: - First, efficient quantisation: mlx-iqk takes advantage of IQ_K tensor encoding, more efficient than the ones available via llama.cpp or barebones MLX, originally designed by Iwan Kawrakow - I also changed the layout to a k-contiguous one, to make it faster, at least in this Metal setup.- Second, expert pruning: each of the 40 learned-router layers had 256 experts, and not all of them are equally important for the coding use cases. I removed 80B parameters - this is a known technique called REAP, shared at https://www.cerebras.ai/blog/reap.- Third: balancing the cheapest IQ1_S_R4 tensor encoding (~1.5 bits per weight), selectively promoting projections to IQ2_KS or IQ2_K where the measured error reduction justified the bytes.One of the main ideas was not only to save relevant knowledge, but also to not make it forget how to... stop thinking, how to use reasoning. In the first experiments, it would sometimes reason for thousands of tokens without closing its thinking section, or it would go in loops.Then I solved this by heavily weighting tool-calling traces and structured reasoning in the calibration mix.
Enrichment
- Theme
- DeepSeek model deployment and inference
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- quantized deepseek model for local deployment
- Manually corrected
- False
Could you build this?
No Pruning 80 billion parameters from an open-weights MoE model while preserving reasoning/coding performance and writing a custom C++/Metal inference runtime requires elite machine learning systems research.
What it would actually take: This project requires advanced model pruning/slicing techniques on a Mixture-of-Experts (MoE) architecture, selective expert pruning based on activation sparsity, and custom low-bit quantization (e.g., FP8/2.37 bpw). It also entails developing a bespoke inference engine ('MoEspresso') in C++/Metal/CUDA with custom GPU kernels optimized for unified memory on Apple Silicon, requiring deep ML systems and GPU systems engineering expertise.
Discussion
3 comments analyzed.
Feature requests: GGUF format release, llama.cpp integration/PR, support for non-Mac users
Competitors
Other products that read as similar to this one — 124 launches clear the similarity bar, closest 8 shown.
Attention rank: #50 of 125 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 291 days after the earliest competitor.
- deepseek-v41-flash-mac-mini · github · 2026-09-11 · 88 upvotes · similarity 0.61
- deepseek-v4.1-mac · github · 2026-09-14 · 11 upvotes · similarity 0.59
- On the edge of Apple Silicon memory speeds · hn · 2026-01-17 · 5 upvotes · similarity 0.52
- Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s · hn · 2026-09-01 · 240 upvotes · similarity 0.52
- Warp · hn · 2026-09-15 · 13 upvotes · similarity 0.51
- DeepSeek-V4.1-Flash-vLLM-DGX-Spark · github · 2026-09-10 · 65 upvotes · similarity 0.48
- deepseek-v4.1-flash-next-dgx-spark-512k · github · 2026-09-13 · 6 upvotes · similarity 0.46
- DeepSeek V4.1 Flash · ph · 2026-09-10 · 4 upvotes · similarity 0.45
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.