Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Qwen3.6-35B-A3B on a 16 GB M1 Pro with SSD-streamed MoE

Details

External ID
48953349
Source
HN
Company
—
Product
Qwen3.6-35B-A3B on a 16 GB M1 Pro with SSD-streamed MoE
Website domain
github.com
Launched
July 17, 2026
Cohort
—
Upvotes
24
Upvotes percentile
0.7508960573476703
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Enrichment

Theme
systems tools and desktop utilities
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
run qwen 35b llm on 16gb m1 mac
Manually corrected
False

Could you build this?

No Streaming active mixture-of-experts (MoE) weights from an NVMe SSD into Apple Silicon unified memory at token generation speed requires low-level kernel bypass, Metal/MPS programming, and custom memory-mapping paging kernels.

What it would actually take: Requires low-level C++/Metal implementations using POSIX mmap or direct asynchronous I/O (aio) to stream selectively gated expert tensors from disk into VRAM just-in-time during forward passes. The critical difficulty is hiding high SSD read latency behind attention computation and designing custom weight quantizations optimized for Apple Silicon memory bus bandwidth.

Discussion

4 comments analyzed.

Concerns raised: Performance slower than expected, M1 chip compatibility requirements unclear

Feature requests: Support for cached pre-filled system prompt KV-cache with prefix hash in requests, Remote INFER protocol support

Competitors

Other products that read as similar to this one — 239 launches clear the similarity bar, closest 8 shown.

Attention rank: #54 of 240 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 258 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.