Qwen 3.5 running on a $300 Android phone
on-device, open source
Details
- External ID
- 47238519
- Source
- HN
- Company
- —
- Product
- Qwen 3.5 running on a $300 Android phone
- Website domain
- github.com
- Launched
- March 3, 2026
- Cohort
- —
- Upvotes
- 6
- Upvotes percentile
- 0.2853628536285363
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Qwen 3.5 Small dropped two days ago. I had it running on a mid-tier Android phone within hours.It's great seeing the on-device AI community light up around this release. Off Grid brings it to Android: phones with 6GB RAM in the $200-300 range, ~8 tok/sec on the 2B model. Fully offline.Text generation, vision AI, image gen, voice transcription, tool calling, document analysis — all on-device, nothing uploaded, ever. Works in airplane mode.780+ GitHub stars. ~2,000 downloads across Android and iOS. Early days.GitHub: https://github.com/alichherawalla/off-grid-mobile-aiPlay Store: https://play.google.com/store/apps/details?id=ai.offgridmobi...App Store: https://apps.apple.com/us/app/off-grid-local-ai/id6759299882
Enrichment
- Theme
- local AI inference and ComfyUI tools
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- llm on mobile device
- Manually corrected
- False
Could you build this?
Partial A developer can vibe-code a basic Android chat UI that binds to an existing mobile runtime, but achieving stable, fast inference on budget hardware requires low-level optimization to prevent out-of-memory crashes.
What it would actually take: The architecture relies on an Android app (Kotlin) integrating with low-level native inference runtimes (such as llama.cpp or MLC-LLM) compiled via the Android NDK. The hard engineering involves tuning memory footprints to stay strictly under Android's aggressive low-memory killer (LMK), configuring Vulkan/OpenCL compute shaders across fragmented budget mobile SoCs (like MediaTek or Snapdragon), and optimizing weight quantization (e.g., 2-bit to 4-bit GGUF). This requires deep expertise in embedded systems, mobile C++ development, and mobile GPU compute.
Discussion
10 comments analyzed.
Concerns raised: Mediatek NPU support needed, Copy function copies entire message instead of allowing text selection
Feature requests: Allow selective text copying instead of copying entire messages, Add Mediatek NPU support for hardware acceleration
Competitors
Other products that read as similar to this one — 89 launches clear the similarity bar, closest 8 shown.
Attention rank: #79 of 90 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 91 days after the earliest competitor.
- Off Grid: On-device AI-web browsing, tools vision,image,voice–3x faster · hn · 2026-02-24 · 12 upvotes · similarity 0.62
- Off Grid · hn · 2026-02-14 · 124 upvotes · similarity 0.53
- qwen3.6-35b-a3b-144T-S · github · 2026-09-20 · 15 upvotes · similarity 0.48
- Fastest Qwen 3.8 27 on single RTX5090 · ph · 2026-09-07 · 1 upvotes · similarity 0.46
- qwenimage3.0-qwen-image-3.0-api-cn · github · 2026-09-24 · 64 upvotes · similarity 0.46
- qwen38-flash-next-w4a16-cmp170hx · github · 2026-09-16 · 10 upvotes · similarity 0.46
- qwenimage3.0-qwen-image-3.0-api-ja · github · 2026-09-24 · 65 upvotes · similarity 0.46
- qwen21-fast-comfyui · github · 2026-09-23 · 17 upvotes · similarity 0.45
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.