Qwen3.6-35B-A3B on a 16 GB M1 Pro with SSD-streamed MoE
Details
- External ID
- 48953349
- Source
- HN
- Company
- —
- Product
- Qwen3.6-35B-A3B on a 16 GB M1 Pro with SSD-streamed MoE
- Website domain
- github.com
- Launched
- July 17, 2026
- Cohort
- —
- Upvotes
- 24
- Upvotes percentile
- 0.7508960573476703
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Enrichment
- Theme
- systems tools and desktop utilities
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- run qwen 35b llm on 16gb m1 mac
- Manually corrected
- False
Could you build this?
No Streaming active mixture-of-experts (MoE) weights from an NVMe SSD into Apple Silicon unified memory at token generation speed requires low-level kernel bypass, Metal/MPS programming, and custom memory-mapping paging kernels.
What it would actually take: Requires low-level C++/Metal implementations using POSIX mmap or direct asynchronous I/O (aio) to stream selectively gated expert tensors from disk into VRAM just-in-time during forward passes. The critical difficulty is hiding high SSD read latency behind attention computation and designing custom weight quantizations optimized for Apple Silicon memory bus bandwidth.
Discussion
4 comments analyzed.
Concerns raised: Performance slower than expected, M1 chip compatibility requirements unclear
Feature requests: Support for cached pre-filled system prompt KV-cache with prefix hash in requests, Remote INFER protocol support
Competitors
Other products that read as similar to this one — 239 launches clear the similarity bar, closest 8 shown.
Attention rank: #54 of 240 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 258 days after the earliest competitor.
- qwen3.6-35b-a3b-144T-S · github · 2026-09-20 · 15 upvotes · similarity 0.65
- Samosa Chat · hn · 2026-07-15 · 6 upvotes · similarity 0.63
- Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone · hn · 2026-08-03 · 312 upvotes · similarity 0.56
- Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s · hn · 2026-09-01 · 240 upvotes · similarity 0.53
- Fine-tune an 8B model on a 4 GB laptop GPU · hn · 2026-08-04 · 139 upvotes · similarity 0.53
- S3 compatible store with 1M IOPS(4K-R,p99~5ms), BYOC in 5min with rust · hn · 2025-12-07 · 23 upvotes · similarity 0.49
- Goxe 19k Logs/S on an I5 · hn · 2026-02-08 · 9 upvotes · similarity 0.49
- qwen38-27B-dual-rtx5060 · github · 2026-09-22 · 22 upvotes · similarity 0.49
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.