Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Morph Reflexes

Multi-head classifiers for agent traces

Details

External ID
48739038
Source
HN
Company
—
Product
—
Website domain
—
Launched
June 30, 2026
Cohort
—
Upvotes
20
Upvotes percentile
0.7472677595628415
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

The most common failures for production agents are behavioral: looping, reasoning leakage, user frustration, and more. Using a frontier model like GPT or Sonnet to judge every turn is too expensive and slow to run at scale.To solve this, we built Reflexes: semantic signals from agent traces, served fast and cheap over API. Built on custom kernels and a custom inference engine forked from vLLM.Under the hood, it is a small LLM architected around multi-head inference. Small models need to be trained for specific tasks, but running 50 separate small models on the same input for 50 tasks makes no sense.How it works: We use a modern LLM with hybrid attention and remove the decode step. We built an inference engine that lets prefill compute be 99% reused from reflex to reflex, similar in spirit to older 2019-era BERT/HYDRA and older multiple-head techniques. we built the inference engine to reuse the KV/cache across inputs and compute across all reflexes. One shared backbone reads the trace once, then many heads classify different signals. Our inference engine reuses the same KV/cache and compute across all reflexes, giving us sub-30ms inference with less than 0.1% overhead for each additional reflex.We took the same high-level idea and did the hard work to make it work with a modern architecture and attention. On it, we can run inference in under 30ms and serve the full request in under 90ms. If you run 4 reflexes or 100, the extra overhead is less than 2ms.Why does optimizing this matter?If you’re even a medium-sized startup, you’re dealing with tens of thousands of agent runs and millions of turns. If you want to track things like user frustration rates over time, frontier LLM-as-judge does not scale.I built a similar stack at Tesla. When ML engineers needed to sample data across petabytes for signals like `is_camera_obfuscated=true`, along with 200 other things, you need to 1) spin them up quickly 2) run at scale efficientlyWhat it is not: A dashboard. 99% of dashboards go unused. 100% API first and made for devs who want to use this to trigger their own stuff.vibetrain a custom reflex in our dashboard, and/or then let it self improve in production: https://www.morphllm.com/dashboard/reflexDocs: https://docs.morphllm.com/sdk/components/reflexes/indexI’d love feedback from people running agents in prod: what sorts of things do you wish you could track over time across 100% of turns but cant right now?TLDR: semantic signals from agent traces, super fast, cheap via API

Enrichment

Theme
Vertical
Horizontal
Function
Observability & eval
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
multi-head classifiers for agent traces
Manually corrected
False

Could you build this?

No Building accurate, low-latency multi-head semantic classifiers for complex behavioral traces requires extensive curated agent trajectory datasets, custom model distillation/fine-tuning, and specialized ML evaluation infrastructure.

What it would actually take: Requires collecting and annotating tens of thousands of real agent execution traces covering failure modes like loops, prompt injection, and reasoning degradation. The architecture typically uses a distilled lightweight encoder backbone (e.g., modern BERT or small Llama variant) with multi-task classification heads optimized for sub-millisecond inference using TensorRT/ONNX. The hard parts are avoiding false positive churn in production and maintaining inference latencies under 5-10ms.

Discussion

2 comments analyzed.

Competitors mentioned: Morph Apply

Competitors

Other products that read as similar to this one — 450 launches clear the similarity bar, closest 8 shown.

Attention rank: #108 of 451 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 243 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a observability & eval tool for Media & entertainment yet.