Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

I shrank DeepSeek V4 Flash to 57GB and it wrote a compiler on my Mac

Details

External ID
49321813
Source
HN
Company
—
Product
I shrank DeepSeek V4 Flash to 57GB and it wrote a compiler on my Mac
Website domain
huggingface.co
Launched
Aug. 16, 2026
Cohort
—
Upvotes
21
Upvotes percentile
0.7547043010752689
Tags
—
Fetched at
Sept. 10, 2026, 5:32 a.m.
Updated at
Sept. 10, 2026, 5:32 a.m.

Description

I built a specialized package of DeepSeek V4 Flash 0731 (originally 284B total parameters, 13B active), preserving reasoning, tool calling and coding capabilities:https://huggingface.co/steadfastgaze/DeepSeek-V4-Flash-0731-...I let it write a minimal C compiler targeting ARM64, then test the result with Fibonacci and FizzBuzz programs, and it succeeded in less than 1 hour, with the full recording at:https://youtu.be/XiwSilmV8B0You can run it on Silicon Macs with my engine https://github.com/steadfastgaze/MoEspresso, while one of the core libraries developed to obtain this result is available at https://github.com/steadfastgaze/mlx-iqk.The above recording was on a 128GB memory MacBook M3 Max, but you can also run it on 32GB MacBooks with a very usable context (128K tokens) and projected 5 tok/s. I did try it on a fanless 16GB memory MacBook Air M1 (1.39 tok/s), but unfortunately the available context was very small.How: - First, efficient quantisation: mlx-iqk takes advantage of IQ_K tensor encoding, more efficient than the ones available via llama.cpp or barebones MLX, originally designed by Iwan Kawrakow - I also changed the layout to a k-contiguous one, to make it faster, at least in this Metal setup.- Second, expert pruning: each of the 40 learned-router layers had 256 experts, and not all of them are equally important for the coding use cases. I removed 80B parameters - this is a known technique called REAP, shared at https://www.cerebras.ai/blog/reap.- Third: balancing the cheapest IQ1_S_R4 tensor encoding (~1.5 bits per weight), selectively promoting projections to IQ2_KS or IQ2_K where the measured error reduction justified the bytes.One of the main ideas was not only to save relevant knowledge, but also to not make it forget how to... stop thinking, how to use reasoning. In the first experiments, it would sometimes reason for thousands of tokens without closing its thinking section, or it would go in loops.Then I solved this by heavily weighting tool-calling traces and structured reasoning in the calibration mix.

Enrichment

Theme
DeepSeek model deployment and inference
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
quantized deepseek model for local deployment
Manually corrected
False

Could you build this?

No Pruning 80 billion parameters from an open-weights MoE model while preserving reasoning/coding performance and writing a custom C++/Metal inference runtime requires elite machine learning systems research.

What it would actually take: This project requires advanced model pruning/slicing techniques on a Mixture-of-Experts (MoE) architecture, selective expert pruning based on activation sparsity, and custom low-bit quantization (e.g., FP8/2.37 bpw). It also entails developing a bespoke inference engine ('MoEspresso') in C++/Metal/CUDA with custom GPU kernels optimized for unified memory on Apple Silicon, requiring deep ML systems and GPU systems engineering expertise.

Discussion

3 comments analyzed.

Feature requests: GGUF format release, llama.cpp integration/PR, support for non-Mac users

Competitors

Other products that read as similar to this one — 124 launches clear the similarity bar, closest 8 shown.

Attention rank: #50 of 125 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 291 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.