Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Millwright

Rust-based, self-hosted LLM router

Details

External ID
49011806
Source
HN
Company
—
Product
Millwright
Website domain
github.com
Launched
July 22, 2026
Cohort
—
Upvotes
10
Upvotes percentile
0.5418160095579451
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Hey HN,With the news of OpenRouter possibly being acquired and proliferation of hosted LLM routers (i.e. Ramp Router, Vercel’s AI Gateway), I saw the need for a self hosted solution focused on cost savings, transparency, and performance. So, I built an open sourced router with a simple CLI interface that can easily sit between coding agents and GenAI workloads.For the curious and lazy, at the moment, Millwright has the tools for,- Providers: OpenAI-compatible APIs, Anthropic, Amazon Bedrock- Routing: policy-controlled model roles (cheap, mid, frontier), cheapest healthy route selection- Protocols: OpenAI Chat Completions, Anthropic Messages, text and tool translation- Cache Affinity: role-scoped session lanes without serializing concurrent agent traffic- Spend Tracking: per-team costs, cache usage, model/provider mix, request traces- Cost Analysis: measured usage and modeled candidate economics (HTML, Markdown, JSON)- Reliability: bounded failover, circuit breakers, timeouts, concurrency limits- Setup: interactive provider, model, and pricing configuration without storing provider secrets- Deployment: one Rust binary, Docker, SQLite or PostgreSQLFull disclosure: parts of the codebase were built with AI coding agents. All feedback is welcome, I’d especially value feedback on the routing policy, provider coverage, and anything that would block you from self-hosting it. Feel free to open feature/request and/or contribute as well.

Enrichment

Theme
ML inference and model optimization
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
self-hosted llm routing layer
Manually corrected
False

Could you build this?

Partial The basic proxying and routing logic is vibe-codeable, but building a production-grade, low-latency, resilient LLM gateway in Rust with high-throughput streaming, fallbacks, and connection pooling requires solid systems engineering.

What it would actually take: The software is built with Rust using Tokio, Axum/Hyper, and Tower services to handle streaming HTTP SSE connections, token-bucket rate limiting, semantic caching, and dynamic failovers across heterogeneous provider APIs. The hard parts are maintaining near-zero proxy latency overhead under heavy concurrency, handling divergent provider token-counting conventions, and implementing reliable circuit-breaker patterns.

Discussion

7 comments analyzed.

Competitors mentioned: role-model (on-device router alternative)

Concerns raised: Energy consumption/environmental impact, Vendor lock-in risk with foreign companies, Cognitive dependency - losing ability to use own brain, Pricing will increase after undercharging phase, Legal uncertainty around training on copyrighted software

Competitors

Other products that read as similar to this one — 254 launches clear the similarity bar, closest 8 shown.

Attention rank: #123 of 255 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 266 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.