Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s

Details

External ID
49524447
Source
HN
Company
—
Product
Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
Website domain
github.com
Launched
Sept. 1, 2026
Cohort
—
Upvotes
240
Upvotes percentile
0.9760765550239234
Tags
—
Fetched at
Sept. 10, 2026, 5:31 a.m.
Updated at
Sept. 10, 2026, 5:31 a.m.

Description

I built slotstream, a way to run Qwen3.8-Flash-Next 4-bit on a low-memory mac starting from 16GB, a 125B parameter model that would need 100GB+ memory/RAM, thanks to expert-offloading/ssd-streaming. Easy to install/update, and mac-native using MLX and Swift.It ships with auto-mode, which makes a good tradeoff between memory usage and speed. I'll be implementing and porting the MTP module for speculative decoding next

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
run large language models on constrained hardware
Manually corrected
False

Could you build this?

No Slotstream implements custom low-level SSD streaming, expert offloading, and memory paging for 125B MoE models directly against Apple Silicon hardware via Swift and MLX.

What it would actually take: Requires low-level Apple Silicon kernel and memory optimization using Swift, Metal Performance Shaders (MPS), and the MLX C++ API. Developing dynamic weight streaming requires direct I/O (O_DIRECT, posix_fadvise) asynchronously pulling expert matrices from NVMe storage straight into unified memory ahead of execution layers. This requires deep systems programming, systems architecture, and GPU compute pipeline optimization skills.

Discussion

20 comments analyzed.

Competitors mentioned: ollama, oMLX, MTP (Medusa Token Prediction), vLLM

Concerns raised: Memory bandwidth constraints limit performance gains, Context length degradation (model behavior degrades above 70-80k tokens), High per-process memory usage (8.1GB per slotserve process), Disk I/O not fully optimized for non-sequential expert weight access, Installation time and complexity (1 hour setup)

Feature requests: Native desktop app wrapper, Improved local hardware optimization tooling (like MTPLX interface), MacOS Swift app integration, Reduce per-process memory footprint, Performance metrics and hardware comparison benchmarks

Competitors

Other products that read as similar to this one — 170 launches clear the similarity bar, closest 8 shown.

Attention rank: #7 of 171 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 296 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.