Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
Details
- External ID
- 49524447
- Source
- HN
- Company
- —
- Product
- Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
- Website domain
- github.com
- Launched
- Sept. 1, 2026
- Cohort
- —
- Upvotes
- 240
- Upvotes percentile
- 0.9760765550239234
- Tags
- —
- Fetched at
- Sept. 10, 2026, 5:31 a.m.
- Updated at
- Sept. 10, 2026, 5:31 a.m.
Description
I built slotstream, a way to run Qwen3.8-Flash-Next 4-bit on a low-memory mac starting from 16GB, a 125B parameter model that would need 100GB+ memory/RAM, thanks to expert-offloading/ssd-streaming. Easy to install/update, and mac-native using MLX and Swift.It ships with auto-mode, which makes a good tradeoff between memory usage and speed. I'll be implementing and porting the MTP module for speculative decoding next
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- run large language models on constrained hardware
- Manually corrected
- False
Could you build this?
No Slotstream implements custom low-level SSD streaming, expert offloading, and memory paging for 125B MoE models directly against Apple Silicon hardware via Swift and MLX.
What it would actually take: Requires low-level Apple Silicon kernel and memory optimization using Swift, Metal Performance Shaders (MPS), and the MLX C++ API. Developing dynamic weight streaming requires direct I/O (O_DIRECT, posix_fadvise) asynchronously pulling expert matrices from NVMe storage straight into unified memory ahead of execution layers. This requires deep systems programming, systems architecture, and GPU compute pipeline optimization skills.
Discussion
20 comments analyzed.
Competitors mentioned: ollama, oMLX, MTP (Medusa Token Prediction), vLLM
Concerns raised: Memory bandwidth constraints limit performance gains, Context length degradation (model behavior degrades above 70-80k tokens), High per-process memory usage (8.1GB per slotserve process), Disk I/O not fully optimized for non-sequential expert weight access, Installation time and complexity (1 hour setup)
Feature requests: Native desktop app wrapper, Improved local hardware optimization tooling (like MTPLX interface), MacOS Swift app integration, Reduce per-process memory footprint, Performance metrics and hardware comparison benchmarks
Competitors
Other products that read as similar to this one — 170 launches clear the similarity bar, closest 8 shown.
Attention rank: #7 of 171 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 296 days after the earliest competitor.
- qwen38-flash-next-w4a16-cmp170hx · github · 2026-09-16 · 10 upvotes · similarity 0.74
- collabosm · github · 2026-09-25 · 90 upvotes · similarity 0.68
- Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone · hn · 2026-08-03 · 312 upvotes · similarity 0.58
- Shoehorn, a library to quantize an LLM to fit your Mac's VRAM · hn · 2026-08-14 · 6 upvotes · similarity 0.56
- qwen3.6-35b-a3b-144T-S · github · 2026-09-20 · 15 upvotes · similarity 0.55
- Samosa Chat · hn · 2026-07-15 · 6 upvotes · similarity 0.55
- Qwen3.6-35B-A3B on a 16 GB M1 Pro with SSD-streamed MoE · hn · 2026-07-17 · 24 upvotes · similarity 0.53
- I shrank DeepSeek V4 Flash to 57GB and it wrote a compiler on my Mac · hn · 2026-08-16 · 21 upvotes · similarity 0.52
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.