Nicheloom

Market intelligence for builders — see what's gaining traction before it's crowded.

Xybrid

run LLM and speech locally in your app (no back end, Rust)

Details

External ID
47427332
Source
HN
Company
—
Product
Xybrid
Website domain
github.com
Launched
March 18, 2026
Cohort
—
Upvotes
6
Upvotes percentile
0.2853628536285363
Tags
—
Fetched at
Sept. 7, 2026, 9:26 p.m.
Updated at
Sept. 7, 2026, 9:26 p.m.

Description

Hi HN,We built Xybrid, a Rust library for running LLM + speech pipelines directly inside your app, no server, no daemon, just one binary.We started building it while working on a privacy-focused LLM app with Tauri and realized there wasn’t a straightforward way to embed models directly into shipped applications without relying on a separate server process.Xybrid links into your process like any other library. It supports GGUF / ONNX / CoreML and integrates with Flutter, Swift, Kotlin, Unity, and Tauri, letting you run pipelines like speech → LLM → speech in a single call.On recent phones, we’re seeing ~20 tok/s on Android and ~40 tok/s on iOS for small (~3B) quantized models (varies by device, backend, and thermals).The demo that shows it best: a Unity tavern scene where 6 NPCs generate real-time dialogue fully on-device — no API key, no internet, no per-request cost.Unity demo: https://youtu.be/vSPeTyeow6A Desktop demo (Tauri): https://youtu.be/o83YShqV7O4GitHub: https://github.com/xybrid-ai/xybridIt’s still early — there are rough edges, especially around model support and performance tuning. Happy to answer questions about the architecture, backends, or integrations (Flutter, Swift, Kotlin, Unity, Tauri).

Enrichment

Theme
voice dictation and control tools
Vertical
Horizontal
Function
Model & infra
Audience
Developer
AI stance
AI-native
Project type
Commercial product
Normalized one-liner
on-device llm and speech runtime
Manually corrected
False

Could you build this?

No Embedding local LLM and speech-to-text inference directly into a single cross-platform native Rust binary without external daemons requires specialized low-level systems and ML compilation expertise.

What it would actually take: The architecture relies on Rust bindings to C++ inference runtimes like llama.cpp and whisper.cpp (or pure Rust frameworks like Candle/Burn), managing memory-mapped GGUF models, hardware acceleration (Metal, CUDA, Vulkan), and native audio I/O via cpal. Developers must have deep expertise in systems programming, cross-compilation toolchains, audio buffer streaming, and low-level GPU acceleration.

Discussion

2 comments analyzed.

Competitors

Other products that read as similar to this one — 68 launches clear the similarity bar, closest 8 shown.

Attention rank: #47 of 69 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).

Launched 130 days after the earliest competitor.

Other launches for this product

Same idea, different domain

Nobody's really built a model & infra tool for Fintech yet.