Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Building a full agentic harness around a 4B model is hard

Details

External ID
49359341
Source
HN
Company
—
Product
Building a full agentic harness around a 4B model is hard
Website domain
orvena.app
Launched
Aug. 19, 2026
Cohort
—
Upvotes
5
Upvotes percentile
0.12634408602150538
Tags
—
Fetched at
Sept. 10, 2026, 5:32 a.m.
Updated at
Sept. 10, 2026, 5:32 a.m.

Description

Around 3 months ago, we were thinking why none of the iPhone apps running an LLM are built as a full harness (as in inference + agentic loop + context management + tools + MCP servers and etc.). It became more interesting when we noticed even the new Siri is not fully on device (and not available in EU for that matter).Having built a few agentic products around a custom harness in the past, we thought this shouldn't be that hard. well, we underestimated how "dumb" a 4B model can be, especially when it comes to tool calling. :DWe tried 8 different models and we settled on Qwen 3.5 4B and we used every trick we knew to make this model behave. well, it works!It's not gonna win in any intelligence or speed benchmark, but it can do actual useful work and it really is an on device, private, full agentic harness. That said, you can still connect your OpenRouter or OpenAI API keys if that's what you prefer.Please go check it out :) You need an iPhone 15 pro or above.

Enrichment

Theme
lightweight and on-device AI runtimes
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Hobby / open-source project
Normalized one-liner
agentic framework for small language models
Manually corrected
False

Could you build this?

No Running a local quantized 4B model on an iPhone while handling low-memory constraints, CoreML/Metal optimization, MCP tool-calling protocols, and background agent loops requires specialized iOS machine learning systems expertise.

What it would actually take: A real build requires an iOS app written in Swift/C++ leveraging Apple's MLX Swift or llama.cpp/executorch via Metal, fine-tuning or prompt-engineering a 4B model (like SmolLM2 or Phi/Qwen) to reliably emit structured MCP-compliant tool calls within tight 2-3GB memory ceilings, and deep integrations with EventKit, Contacts, and Shortcuts. The hardest parts are preventing iOS OS memory kills (jetsam) during multi-turn generation and context compression while running local inference.

Discussion

No comments on this launch.

Competitors

Other products that read as similar to this one — 92 launches clear the similarity bar, closest 8 shown.

Attention rank: #86 of 93 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 288 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.