Open-source simulation testing infra for voice agents
Details
- External ID
- 49646928
- Source
- HN
- Company
- —
- Product
- Open-source simulation testing infra for voice agents
- Website domain
- github.com
- Launched
- Sept. 10, 2026
- Cohort
- —
- Upvotes
- 16
- Upvotes percentile
- 0.7001594896331739
- Tags
- —
- Fetched at
- Sept. 14, 2026, 5:28 p.m.
- Updated at
- Sept. 14, 2026, 5:28 p.m.
Description
Hey HN, we’re Nischal & Naman. We’re brothers, and together we’re building an open-source platform for simulation based testing of voice agents (try it out in 5 mins - https://docs.egma.ai/docs/get-started/quickstart, 2 min demo video - https://youtu.be/wgDWEe5UAUY)Platforms that help you do simulation testing already exist. But they all charge a heavy premium on top of inference costs. We believe if the industry truly wants to scale simulation testing of voice agents, we need to stop charging a premium on inference and start providing infrastructure to scale simulations. The project is built in a way that allows you to bring your own STT, LLM, and TTS provider keys while also providing a way for us to provide inference directly within the platform at a 0% markup so you don’t have to wrestle with multiple keys/ rate limits.A bit on why’re we’re building this -We started working on voice agents in early 2024 and since then, have worked on numerous voice ai systems - like screenless voice-powered hardware for kids[1], AI receptionist deployed in healthcare practices like med spas & therapy clinics and personal accountability coaches. Some of these were full-fledged startups others were just side projects.But we kept on encountering similar issues all throughout. One of the most frustrating parts was calling the agent again and again, reciting the same script to test its behavior. We also kept on encountering new issues all the time in production that we couldn't have simulated pre launch.This frustration led us to a bigger question: how can developers trust the voice agents they’re shipping?We believe building that trust takes two things - - First, you need a way to test the major scenarios your agent will face in production before you release it, without having to make every call yourself. But you can’t test for everything. Real world is too messy to predict in advance. - Which brings us to the second: you need a way to find issues once your agent is in production, whether it’s handling a few dozen conversations or millions. Detect drift in known behaviors AND surface unknown-unknown ways in which your agent is going wrong.We think it’s a really hard problem to solve. And having felt it firsthand, we’re deeply motivated to take a shot at it.Today we’re launching the simulation testing side of the platform. We feel it’s mature enough that real teams can depend on it. For eval design, we took inspiration from anthropic’s evals design[2] and extended it to voice systems. For technical & business model design we borrowed ideas from Langfuse[3]. Our stack is postgres, clickhouse & minio. Its easy to self-host[4] & the code has a permissive MIT license. We also have managed cloud version.We’ve written more about our [testing](https://docs.egma.ai/docs/core-philosophies/testing-philosop...) and [monitoring](https://docs.egma.ai/docs/core-philosophies/monitoring-philo...) philosophy in the docs.We’d love feedback from the HN community and people building voice agents - how are you testing today, what’s working and what’s frustrating? We’ll be in the comments. Thanks![1] https://x.com/theBhulawat/status/1966200231705595932?s=20 [2] https://www.anthropic.com/engineering/demystifying-evals-for... [3] https://github.com/langfuse/langfuse [4] https://docs.egma.ai/self-hosting/get-started
Enrichment
- Theme
- developer tools for AI agents
- Vertical
- Horizontal
- Function
- Dev tools
- Audience
- Developer
- AI stance
- Not AI
- Project type
- Hobby / open-source project
- Normalized one-liner
- testing infrastructure for voice agents
- Manually corrected
- False
Could you build this?
Partial While an agent runner and simulation dashboard can be vibe-coded, building reliable low-latency audio streaming pipelines, SIP/WebRTC mock telephony, and realistic real-time conversational voice turn-taking requires non-trivial media streaming infra.
What it would actually take: The architecture requires an orchestration backend (Node/Go or Python) managing concurrent voice sessions via WebRTC/SIP proxies (e.g., LiveKit or Asterisk), coupled with speech-to-text (STT), LLM agent simulation loops, and text-to-speech (TTS) streaming. The hard part is simulating real-world network packet jitter, packet loss, interruptions (barge-in), and sub-200ms latency metrics accurately. Implementing this requires specialized knowledge of VoIP protocols, audio codecs, and real-time streaming infrastructure.
Discussion
4 comments analyzed.
Competitors
Other products that read as similar to this one — 234 launches clear the similarity bar, closest 8 shown.
Attention rank: #91 of 235 (itself plus its competitors, highest first — normalized so YC and Product Hunt are compared fairly).
Launched 310 days after the earliest competitor.
- NovaSynth by Noveum · ph · 2026-09-17 · 175 upvotes · similarity 0.52
- Spec27 · hn · 2026-04-30 · 13 upvotes · similarity 0.47
- We open sourced Vapi · hn · 2026-03-12 · 8 upvotes · similarity 0.47
- MockVerse · ph · 2026-09-15 · 2 upvotes · similarity 0.47
- Teapot · hn · 2026-02-18 · 7 upvotes · similarity 0.46
- Voice driven murder mystery, Interview AI suspects with your voice · hn · 2026-08-10 · 215 upvotes · similarity 0.46
- I built a voice AI that responds like a real woman · hn · 2026-03-25 · 6 upvotes · similarity 0.46
- Echo · hn · 2026-07-23 · 484 upvotes · similarity 0.43
Other launches for this product
Same idea, different domain
Nobody's really built a dev tools tool for Sales yet.