Samosa Chat
Run Qwen3.6-35B-A3B Locally on a 16 GB Mac
Details
- External ID
- 48915561
- Source
- HN
- Company
- —
- Product
- Samosa Chat
- Website domain
- github.com
- Launched
- July 15, 2026
- Cohort
- —
- Upvotes
- 6
- Upvotes percentile
- 0.2873357228195938
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Commercial product
- Normalized one-liner
- run large language model locally on mac
- Manually corrected
- False
Could you build this?
Yes Samosa Chat is a desktop or terminal chat interface wrapping existing local inference engines (such as llama.cpp or MLX) with quantized weights for Apple Silicon, which requires standard UI and CLI glue code.
Discussion
9 comments analyzed.
Competitors mentioned: llama.cpp, vllm
Concerns raised: Mac overheating with local models, Token throughput speed (7 tokens/sec considered slow), Memory usage scaling (jumped from 22GB to 27GB)
Feature requests: Run multiple local models continuously on Mac, Integration with Chat application
Competitors
Other products that read as similar to this one — 282 launches clear the similarity bar, closest 8 shown.
Attention rank: #205 of 283 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 250 days after the earliest competitor.
- Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone · hn · 2026-08-03 · 312 upvotes · similarity 0.64
- Qwen3.6-35B-A3B on a 16 GB M1 Pro with SSD-streamed MoE · hn · 2026-07-17 · 24 upvotes · similarity 0.63
- qwen3.6-35b-a3b-144T-S · github · 2026-09-20 · 15 upvotes · similarity 0.58
- Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s · hn · 2026-09-01 · 240 upvotes · similarity 0.55
- qwen38-flash-next-w4a16-cmp170hx · github · 2026-09-16 · 10 upvotes · similarity 0.50
- I fine-tuned Qwen 3.5 (0.8B–4B) on a Mac for text-to-SQL · hn · 2026-03-05 · 7 upvotes · similarity 0.49
- Tesla_OS_On_QEMU · github · 2026-09-22 · 69 upvotes · similarity 0.48
- qwen36-q4-mtp-cline · github · 2026-09-22 · 8 upvotes · similarity 0.47
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.