Open-source customizable AI voice dictation built on Pipecat
Details
- External ID
- 46264158
- Source
- HN
- Company
- —
- Product
- Open-source customizable AI voice dictation built on Pipecat
- Website domain
- github.com
- Launched
- Dec. 14, 2025
- Cohort
- —
- Upvotes
- 27
- Upvotes percentile
- 0.708969465648855
- Tags
- —
- Fetched at
- Sept. 7, 2026, 9:25 p.m.
- Updated at
- Sept. 7, 2026, 9:25 p.m.
Description
Tambourine is an open source, fully customizable voice dictation system that lets you control STT/ASR, LLM formatting, and prompts for inserting clean text into any app.I have been building this on the side for a few weeks. What motivated it was wanting a customizable version of Wispr Flow where I could fully control the models, formatting, and behavior of the system, rather than relying on a black box.Tambourine is built directly on top of Pipecat and relies on its modular voice agent framework. The back end is a local Python server that uses Pipecat to stitch together STT and LLM models into a single pipeline. This modularity is what makes it easy to swap providers, experiment with different setups, and maintain fine-grained control over the voice AI.I shared an early version with friends and recently presented it at my local Claude Code meetup. The response was overwhelmingly positive, and I was encouraged to share it more widely.The desktop app is built with Tauri. The front end is written in TypeScript, while the Tauri layer uses Rust to handle low level system integration. This enables the registration of global hotkeys, management of audio devices, and reliable text input at the cursor on both Windows and macOS.At a high level, Tambourine gives you a universal voice interface across your OS. You press a global hotkey, speak, and formatted text is typed directly at your cursor. It works across emails, documents, chat apps, code editors, and terminals.Under the hood, audio is streamed from the TypeScript front end to the Python server via WebRTC. The server runs real-time transcription with a configurable STT provider, then passes the transcript through an LLM that removes filler words, adds punctuation, and applies custom formatting rules and a personal dictionary. STT and LLM providers, as well as prompts, can be switched without restarting the app.The project is still under active development. I am working through edge cases and refining the UX, and there will likely be breaking changes, but most core functionality already works well and has become part of my daily workflow.I would really appreciate feedback, especially from anyone interested in the future of voice as an interface.
Enrichment
- Theme
- voice dictation and control tools
- Vertical
- Horizontal
- Function
- Dev tools
- Audience
- Developer
- AI stance
- AI-native
- Project type
- Hobby / open-source project
- Normalized one-liner
- customizable voice dictation for developers
- Manually corrected
- False
Could you build this?
Yes Tambourine wraps the existing Pipecat framework to stream audio to speech-to-text APIs and format text with an LLM before simulated keystroke insertion, which is standard glue code.
Discussion
16 comments analyzed.
Competitors mentioned: Gemini 2.0 Flash, Electron, local large language models
Concerns raised: only works with proprietary cloud LLMs, not truly open-source without local option, Tauri global shortcuts can break other apps, Tauri not production-ready for larger audiences, unclear documentation about local inference support
Feature requests: fully offline/local LLM support without internet access, online customizable backend as platform-as-a-service, macOS compatibility front-and-center in documentation
Competitors
Other products that read as similar to this one — 153 launches clear the similarity bar, closest 8 shown.
Attention rank: #56 of 154 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 38 days after the earliest competitor.
- We open sourced Vapi · hn · 2026-03-12 · 8 upvotes · similarity 0.50
- Dograh · hn · 2025-12-08 · 16 upvotes · similarity 0.48
- TongueType · hn · 2026-05-15 · 5 upvotes · similarity 0.46
- SpeechOS · hn · 2026-01-21 · 12 upvotes · similarity 0.46
- Jargo · hn · 2026-06-26 · 7 upvotes · similarity 0.44
- Walkie · ph · 2026-04-06 · 223 upvotes · similarity 0.43
- VOOG · hn · 2026-02-15 · 96 upvotes · similarity 0.42
- Dictámelo · ph · 2026-09-09 · 4 upvotes · similarity 0.42
Other launches for this product
- No other launches for this product.
Same idea, different domain
Nobody's really built a dev tools tool for Sales yet.