Needle: We Distilled Gemini Tool Calling into a 26M Model
Details
- External ID
- 48111896
- Source
- HN
- Company
- —
- Product
- Needle: We Distilled Gemini Tool Calling into a 26M Model
- Website domain
- github.com
- Launched
- May 12, 2026
- Cohort
- —
- Upvotes
- 776
- Upvotes percentile
- 1.0
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Hey HN, Henry here from Cactus. We open-sourced Needle, a 26M parameter function-calling (tool use) model. It runs at 6000 tok/s prefill and 1200 tok/s decode on consumer devices.We were always frustrated by the little effort made towards building agentic models that run on budget phones, so we conducted investigations that led to an observation: agentic experiences are built upon tool calling, and massive models are overkill for it. Tool calling is fundamentally retrieval-and-assembly (match query to tool name, extract argument values, emit JSON), not reasoning. Cross-attention is the right primitive for this, and FFN parameters are wasted at this scale.Simple Attention Networks: the entire model is just attention and gating, no MLPs anywhere. Needle is an experimental run for single-shot function calling for consumer devices (phones, watches, glasses...).Training: - Pretrained on 200B tokens across 16 TPU v6e (27 hours) - Post-trained on 2B tokens of synthesized function-calling data (45 minutes) - Dataset synthesized via Gemini with 15 tool categories (timers, messaging, navigation, smart home, etc.)You can test it right now and finetune on your Mac/PC: https://github.com/cactus-compute/needleThe full writeup on the architecture is here: https://github.com/cactus-compute/needle/blob/main/docs/simp...We found that the "no FFN" finding generalizes beyond function calling to any task where the model has access to external structured knowledge (RAG, tool use, retrieval-augmented generation). The model doesn't need to memorize facts in FFN weights if the facts are provided in the input. Experimental results to published.While it beats FunctionGemma-270M, Qwen-0.6B, Granite-350M, LFM2.5-350M on single-shot function calling, those models have more scope/capacity and excel in conversational settings. We encourage you to test on your own tools via the playground and finetune accordingly.This is part of our broader work on Cactus (https://github.com/cactus-compute/cactus), an inference engine built from scratch for mobile, wearables and custom hardware. We wrote about Cactus here previously: https://news.ycombinator.com/item?id=44524544Everything is MIT licensed. Weights: https://huggingface.co/Cactus-Compute/needle GitHub: https://github.com/cactus-compute/needle
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- —
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- 26m parameter model with gemini tool calling distilled
- Manually corrected
- False
Could you build this?
No Training and distilling a custom 26M parameter model specifically for tool calling requires deep ML research expertise, model architecture design, synthetic dataset generation, and specialized GPU training pipelines.
What it would actually take: Architecture requires defining a compact transformer architecture (PyTorch/JAX) tailored for ultra-fast inference on edge devices (e.g. using quantization and ONNX/GGML runtimes). The hard part is generating tens of thousands of high-quality execution trajectories and tool-calling data from frontier models (Gemini), designing effective loss functions for distillation, and rigorous benchmark evaluation to maintain reliable JSON/tool adherence at sub-50M parameters.
Discussion
20 comments analyzed.
Competitors mentioned: llamafile, teale.com
Concerns raised: annotation costs are extremely high ($50-$1000+ per annotation), distillation from big LLMs vs human annotations - why not use human outputs instead
Feature requests: commit and push tool, model that runs on 8GB MacBook Air and 6GB Android smartphones
Competitors
Other products that read as similar to this one — 49 launches clear the similarity bar, closest 8 shown.
Attention rank: #1 of 50 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 187 days after the earliest competitor.
- Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash · hn · 2026-09-18 · 236 upvotes · similarity 0.58
- Needle2: 14MB agentic LLM for phones, wearables, smart home and robots · hn · 2026-08-10 · 537 upvotes · similarity 0.56
- Cactus Hybrid: We taught Gemma 4 to know when it's wrong · hn · 2026-07-22 · 191 upvotes · similarity 0.49
- Morph Reflexes · hn · 2026-06-30 · 20 upvotes · similarity 0.42
- InstinctFlash · hn · 2026-09-22 · 27 upvotes · similarity 0.39
- Gemini 3.1 Flash-Lite · ph · 2026-05-16 · 168 upvotes · similarity 0.37
- OpenDecision · hn · 2026-09-21 · 6 upvotes · similarity 0.35
- Nari Qwen3-TTS and Qwen3-ASR · hn · 2026-09-14 · 90 upvotes · similarity 0.35
Other launches for this product
- No other launches for this product.