Cactus Hybrid: We taught Gemma 4 to know when it's wrong
Details
- External ID
- 49010782
- Source
- HN
- Company
- —
- Product
- Cactus Hybrid: We taught Gemma 4 to know when it's wrong
- Website domain
- github.com
- Launched
- July 22, 2026
- Cohort
- —
- Upvotes
- 191
- Upvotes percentile
- 0.9611708482676224
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
Hey HN, Henry & Roman here from Cactus.A small, on-device model is fast and private, but sometimes wrong, but frontier models are getting expensive pretty fast. So, we post-trained Gemma 4 E2B post-trained to know when it's wrong. Every response comes with a confidence score between 0 and 1. Developers can accept the on-device when it's high, hand off to a bigger cloud model when it's low. By routing only 15-35% of queries to Gemini 3.1 Flash-Lite, Gemma-4-E2B matches Gemini 3.1 Flash-Lite on most benchmarks.- ChartQA: 15-20%- LibriSpeech: 25-30%- MMBench, GigaSpeech, MMAU: 30-35%- MMLU-Pro: 45-55%We were always frustrated by the routing signals hybrid apps rely on: asking the model to rate itself in text (unreliable, and you're parsing prose), or token entropy heuristics (barely better than a coin flip in our tests). So we did mechanistic studies on small models, Gemma 4 particularly, and found the hidden state for different layers carry meaningful self-awareness signal for various situations.SO we extended the model with a 68k params probe layer (LayerNorm, low-rank projection, attention pooling, small MLP head) reads one intermediate layer during decoding and predicts p(wrong); confidence = 1 - p(wrong), returned as structured data, never parsed out of the answer text.Across 12 hold-out benchmarks spanning text, vision and audio, the probe averages 0.814 AUROC vs 0.549 for token entropy. The result that convinced us this is real: the probe was trained on zero audio data, yet scores 0.79-0.88 AUROC on four audio benchmarks where entropy is near-random or worse (0.32-0.52). It's reading a modality-independent correctness signal from the hidden state, not memorizing patterns from its training data.We published all weights on HuggingFace and provide copy-pase codes to run it on Transformers, MLX, Llama.cpp or Cactus. With Ollama, vLLM, SGLang etc in the works. For llama.cpp we ship a patch series you compile in once (upstreaming is planned). The code is MIT licensed; Gemma model use remains subject to the Gemma terms.GitHub: https://github.com/cactus-compute/cactus-hybridWeights: https://huggingface.co/collections/Cactus-Compute/cactus-hyb...Some caveats:- The probe scores single-sequence decoding only, up to the first 1024 generated tokens.- Handoff works best when routing per task in a multi-step process, not per step.- Hierarchical routing is still in the works: try on-device, then DeepSeek v4 Flash, before Fable/GPT5.5/Gemini/Muse/Grok.- The technique is boutique for each model, we will share each weights as they roll out.These issues are currently being tackled at Cactus and updated weights will be shipped directly into the HuggingFace collection and GitHub repository straight up. Please let us know your thoughts, it helps us find ways to improve the design progressively.Thanks a million!
Enrichment
- Theme
- lightweight and on-device AI runtimes
- Vertical
- Horizontal
- Function
- Model & infra
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- llm model with uncertainty detection
- Manually corrected
- False
Could you build this?
No Fine-tuning base open-source LLMs to accurately calibrate confidence scores and self-detect hallucination requires advanced machine learning research, specialized post-training datasets, and significant GPU training compute.
What it would actually take: This project requires an ML research stack (PyTorch, DeepSpeed/FSDP, Axolotl, or Hugging Face TRL) and substantial compute clusters (e.g., multi-H100 rigs). The core technical challenge is synthesizing or curating verification datasets with calibrated uncertainty, designing custom loss functions (such as expected calibration error optimization or reinforcement learning with verifiers), and preventing model drift while retaining base model capabilities.
Discussion
20 comments analyzed.
Concerns raised: Battery drain on edge devices in production, Signal quality and SNR unclear in classification loops, Confidence scores may be unreliable garbage, Difficulty assessing correctness outside programming/math domains
Feature requests: Combine conformal prediction for distribution-free calibration guarantees, Compare multiple independent confidence methods (entropy + verbal reporting + consistency checking), Assess output consistency by rerunning and comparing results
Competitors
Other products that read as similar to this one — 49 launches clear the similarity bar, closest 8 shown.
Attention rank: #11 of 50 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 242 days after the earliest competitor.
- Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac · hn · 2026-07-29 · 919 upvotes · similarity 0.50
- Needle: We Distilled Gemini Tool Calling into a 26M Model · hn · 2026-05-12 · 776 upvotes · similarity 0.49
- Gemma 4 Multimodal Fine-Tuner for Apple Silicon · hn · 2026-04-07 · 235 upvotes · similarity 0.48
- Cactus Needle 3: 8-29MB automation models can match DeepSeek V4 Flash · hn · 2026-09-18 · 236 upvotes · similarity 0.47
- Running Gemma-4 26B at 124 tokens/SEC on a CPU, no GPU · hn · 2026-06-30 · 10 upvotes · similarity 0.47
- I benchmarked Gemma 4 E2B · hn · 2026-04-13 · 8 upvotes · similarity 0.46
- Gemma Gem · hn · 2026-04-06 · 156 upvotes · similarity 0.42
- Cuey · ph · 2026-09-27 · 290 upvotes · similarity 0.41
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a model & infra tool for Fintech yet.