Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Plano

Edge and service proxy with orchestration for AI agents

Details

External ID
46517177
Source
HN
Company
—
Product
Plain
Website domain
github.com
Launched
Jan. 6, 2026
Cohort
—
Upvotes
8
Upvotes percentile
0.41699604743083
Tags
—
Fetched at
Sept. 7, 2026, 9:25 p.m.
Updated at
Sept. 7, 2026, 9:25 p.m.

Description

Hey HN — I’m Adil from Katanemo (with Salman, Shuguang, and Meiyu)We previously shared an early version of this project as ArchGW. Based on customer feedback, the scope expanded from “LLM routing and model access” into something broader: delivery infrastructure for agentic applications. We renamed it to Plano and reworked the architecture accordingly.The problemOn-the-ground AI practitioners will tell you that calling an LLM is not the hard part. The really hard part is delivering agentic applications to production quickly and reliably, then iterating without rewriting system code every time. In practice, teams keep rebuilding the same concerns that sit outside any single agent’s core logic:This includes model agility — the ability to pull from a large set of LLMs and swap providers without refactoring prompts or streaming handlers. They need to learn from production by collecting signals and traces that tell them what to fix. They need consistent policy enforcement for moderation and jailbreak protection, rather than sprinkling hooks across codebases. And they need multi-agent patterns like handoff and specialization without turning their app into orchestration glue.These concerns get rebuilt and maintained inside fast-changing frameworks and application code, coupling product logic to infrastructure decisions. It’s brittle, and pulls teams away from core product work into plumbing they shouldn’t have to own.What Plano doesPlano moves core delivery concerns out of process into a modular proxy and dataplane designed for agents. It supports inbound listeners (agent orchestration, safety and moderation hooks), outbound listeners (hosted or API-based LLM routing), or both together.Plano provides the following capabilities via a unified, protocol-native, framework-friendly dataplane:- Orchestration: Low-latency routing and handoff between agents. Add or change agents without modifying app code, and evolve strategies centrally instead of duplicating logic across services.- Guardrails & Memory Hooks: Apply jailbreak protection, content policies, and context workflows (rewriting, retrieval, redaction) once via filter chains. This centralizes governance and ensures consistent behavior across your stack.- Model Agility: Route by model name, semantic alias, or preference-based policies. Swap or add models without refactoring prompts, tool calls, or streaming handlers.- Agentic Signals™: Zero-code capture of behavior signals, traces, and metrics across every agent, surfacing traces, token usage, and learning signals in one place.The goal is to keep application code focused on product logic while Plano owns delivery mechanics.More on ArchitecturePlano has two main parts:Envoy-based data plane. Uses Envoy’s HTTP connection management to talk to model APIs, services, and tool backends. We didn’t build a separate model server—Envoy already handles streaming, retries, timeouts, and connection pooling. Some of us are core Envoy contributors at Katanemo.Brightstaff, a lightweight controller written in Rust. It inspects prompts and conversation state, decides which upstreams to call and in what order, and coordinates routing and fallback. It uses small LLMs (1–4B parameters) trained for constrained routing and orchestration. These models do not generate responses and fall back to static policies on failure. The models are open sourced here: https://huggingface.co/katanemoPlano runs alongside your app servers (cloud, on-prem, or local dev), doesn’t require a GPU, and leaves GPUs where your models are hosted.Repo https://github.com/katanemo/plano + docs https://docs.planoai.dev/

Enrichment

Theme
AI agent frameworks and developer tools
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI feature
Project type
Commercial product
Normalized one-liner
edge proxy orchestration for ai agents
Manually corrected
False

Could you build this?

No High-performance edge reverse proxies and orchestration layers require deep systems networking, low-latency concurrent programming (e.g. Rust/Go/Envoy), and complex distributed protocol engineering.

What it would actually take: This requires building or heavily extending an enterprise network proxy (e.g., Envoy, Pingora, or an eBPF/Rust-based proxy) capable of streaming token parsing, intelligent retries, token-bucket rate limiting, semantic caching, and real-time agent workflow routing. The hard part is achieving sub-millisecond overhead under massive concurrency while correctly handling stateful agentic loops, requiring senior systems and distributed infrastructure engineers.

Discussion

2 comments analyzed.

Competitors mentioned: MCP tools (Claude/ChatGPT/Cursor), LeanMCP

Concerns raised: Tool backends drag with auth/setup complexity, Fallback to static policies reliability in production

Competitors

Other products that read as similar to this one — 280 launches clear the similarity bar, closest 8 shown.

Attention rank: #149 of 281 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 66 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.