Morph Reflexes
Multi-head classifiers for agent traces
Details
- External ID
- 48739038
- Source
- HN
- Company
- —
- Product
- —
- Website domain
- —
- Launched
- June 30, 2026
- Cohort
- —
- Upvotes
- 20
- Upvotes percentile
- 0.7472677595628415
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:26 p.m.
- Updated at
- Sept. 7, 2026, 9:26 p.m.
Description
The most common failures for production agents are behavioral: looping, reasoning leakage, user frustration, and more. Using a frontier model like GPT or Sonnet to judge every turn is too expensive and slow to run at scale.To solve this, we built Reflexes: semantic signals from agent traces, served fast and cheap over API. Built on custom kernels and a custom inference engine forked from vLLM.Under the hood, it is a small LLM architected around multi-head inference. Small models need to be trained for specific tasks, but running 50 separate small models on the same input for 50 tasks makes no sense.How it works: We use a modern LLM with hybrid attention and remove the decode step. We built an inference engine that lets prefill compute be 99% reused from reflex to reflex, similar in spirit to older 2019-era BERT/HYDRA and older multiple-head techniques. we built the inference engine to reuse the KV/cache across inputs and compute across all reflexes. One shared backbone reads the trace once, then many heads classify different signals. Our inference engine reuses the same KV/cache and compute across all reflexes, giving us sub-30ms inference with less than 0.1% overhead for each additional reflex.We took the same high-level idea and did the hard work to make it work with a modern architecture and attention. On it, we can run inference in under 30ms and serve the full request in under 90ms. If you run 4 reflexes or 100, the extra overhead is less than 2ms.Why does optimizing this matter?If you’re even a medium-sized startup, you’re dealing with tens of thousands of agent runs and millions of turns. If you want to track things like user frustration rates over time, frontier LLM-as-judge does not scale.I built a similar stack at Tesla. When ML engineers needed to sample data across petabytes for signals like `is_camera_obfuscated=true`, along with 200 other things, you need to 1) spin them up quickly 2) run at scale efficientlyWhat it is not: A dashboard. 99% of dashboards go unused. 100% API first and made for devs who want to use this to trigger their own stuff.vibetrain a custom reflex in our dashboard, and/or then let it self improve in production: https://www.morphllm.com/dashboard/reflexDocs: https://docs.morphllm.com/sdk/components/reflexes/indexI’d love feedback from people running agents in prod: what sorts of things do you wish you could track over time across 100% of turns but cant right now?TLDR: semantic signals from agent traces, super fast, cheap via API
Enrichment
- Theme
- Vertical
- Horizontal
- Function
- Observability & eval
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Commercial product
- Normalized one-liner
- multi-head classifiers for agent traces
- Manually corrected
- False
Could you build this?
No Building accurate, low-latency multi-head semantic classifiers for complex behavioral traces requires extensive curated agent trajectory datasets, custom model distillation/fine-tuning, and specialized ML evaluation infrastructure.
What it would actually take: Requires collecting and annotating tens of thousands of real agent execution traces covering failure modes like loops, prompt injection, and reasoning degradation. The architecture typically uses a distilled lightweight encoder backbone (e.g., modern BERT or small Llama variant) with multi-task classification heads optimized for sub-millisecond inference using TensorRT/ONNX. The hard parts are avoiding false positive churn in production and maintaining inference latencies under 5-10ms.
Discussion
2 comments analyzed.
Competitors mentioned: Morph Apply
Competitors
Other products that read as similar to this one — 450 launches clear the similarity bar, closest 8 shown.
Attention rank: #108 of 451 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 243 days after the earliest competitor.
- LLM agents that write Python to analyze execution traces at scale · hn · 2026-03-07 · 5 upvotes · similarity 0.51
- reflex · github · 2026-09-23 · 34 upvotes · similarity 0.51
- Autofix Bot · hn · 2025-12-11 · 37 upvotes · similarity 0.50
- Spec27 · hn · 2026-04-30 · 13 upvotes · similarity 0.49
- InstinctFlash · hn · 2026-09-22 · 27 upvotes · similarity 0.48
- Context Gateway · hn · 2026-03-13 · 97 upvotes · similarity 0.46
- BentoLabs AI: Monitoring and Learning layer for long-running agents · yc · 2026-06-01 · 83 upvotes · similarity 0.45
- Oodle.ai · hn · 2026-07-14 · 31 upvotes · similarity 0.45
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a observability & eval tool for Media & entertainment yet.