Building a full agentic harness around a 4B model is hard
Details
- External ID
- 49359341
- Source
- HN
- Company
- —
- Product
- Building a full agentic harness around a 4B model is hard
- Website domain
- orvena.app
- Launched
- Aug. 19, 2026
- Cohort
- —
- Upvotes
- 5
- Upvotes percentile
- 0.12634408602150538
- Tags
- —
- Fetched at
- Sept. 10, 2026, 5:32 a.m.
- Updated at
- Sept. 10, 2026, 5:32 a.m.
Description
Around 3 months ago, we were thinking why none of the iPhone apps running an LLM are built as a full harness (as in inference + agentic loop + context management + tools + MCP servers and etc.). It became more interesting when we noticed even the new Siri is not fully on device (and not available in EU for that matter).Having built a few agentic products around a custom harness in the past, we thought this shouldn't be that hard. well, we underestimated how "dumb" a 4B model can be, especially when it comes to tool calling. :DWe tried 8 different models and we settled on Qwen 3.5 4B and we used every trick we knew to make this model behave. well, it works!It's not gonna win in any intelligence or speed benchmark, but it can do actual useful work and it really is an on device, private, full agentic harness. That said, you can still connect your OpenRouter or OpenAI API keys if that's what you prefer.Please go check it out :) You need an iPhone 15 pro or above.
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- agentic framework for small language models
- Manually corrected
- False
Could you build this?
No Running a local quantized 4B model on an iPhone while handling low-memory constraints, CoreML/Metal optimization, MCP tool-calling protocols, and background agent loops requires specialized iOS machine learning systems expertise.
What it would actually take: A real build requires an iOS app written in Swift/C++ leveraging Apple's MLX Swift or llama.cpp/executorch via Metal, fine-tuning or prompt-engineering a 4B model (like SmolLM2 or Phi/Qwen) to reliably emit structured MCP-compliant tool calls within tight 2-3GB memory ceilings, and deep integrations with EventKit, Contacts, and Shortcuts. The hardest parts are preventing iOS OS memory kills (jetsam) during multi-turn generation and context compression while running local inference.
Discussion
No comments on this launch.
Competitors
Other products that read as similar to this one — 92 launches clear the similarity bar, closest 8 shown.
Attention rank: #86 of 93 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 288 days after the earliest competitor.
- I built an open source multi-agent harness in Go · hn · 2026-04-08 · 6 upvotes · similarity 0.48
- 100% native Swift harness (NOT Electron) · hn · 2026-08-10 · 15 upvotes · similarity 0.46
- HarnessRouter: Unified interface for agent harnesses · hn · 2026-08-17 · 10 upvotes · similarity 0.44
- I made a Raspberry with Qwen my local car AI · hn · 2026-08-25 · 146 upvotes · similarity 0.42
- iClaw is part OpenClaw, part Siri, powered by Apple Intelligence · hn · 2026-04-28 · 7 upvotes · similarity 0.40
- Millwright · hn · 2026-07-22 · 10 upvotes · similarity 0.38
- I built a version of Omarchy that runs on Apple Silicon · hn · 2026-09-02 · 27 upvotes · similarity 0.38
- I'm a CEO Coding with AI · hn · 2025-11-14 · 14 upvotes · similarity 0.38
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.