Egma’s builders claim their released platform supports repository-based simulated voice conversations, mocked tool responses, and production grading for LiveKit and Retell agents, enabling repeatable pre-deployment regression testing alongside production monitoring.
state: watchingheat: lowuncertainty: mediumconvergesscott: mediumvoice-agents agent-evaluation agent-harnessesEgmaNischalNaman
What is this?
Egma is a voice-agent testing and monitoring product; its published PyPI SDK description says it connects LiveKit agents to simulation testing and production monitoring, records the agent’s perspective during simulations, and allows mock-tool injection. The case presents an open-source Show HN release and names Nischal and Naman as builders, but the supplied web snippets do not verify their roles or the open-source status. Repository-based suites, Retell support, CLI execution, and production grading remain case claims rather than capabilities established by these snippets; similar features described for other platforms do not substantiate Egma’s implementation.
Why it matters to Scott
Egma’s documented simulation, mock-tool injection and monitoring converge with Scott’s Evaluation-Driven Development position and offer a concrete harness approach to investigate for Practice Trainer and his voice-booking work, rather than merely another argument for testing. The supplied radar pages track adjacent evaluation tools, not Egma itself; compatibility with Scott’s stack, repository-based regression suites and binding release gates remain unestablished, so this is an implementation lead rather than a validated replacement.
ip:concept.evaluation-driven-developmentdev:concept.deterministic-test-seamsdev:project.sales-trainerdev:concept.outbound-voice-booking-agentradar:concept.voice-agentsradar:concept.agent-evaluationradar:concept.agent-observabilityradar:understudy-agent-scenario-testing
queries asked of Scott's wikis
- agent evaluation harnesses mocked tools regression testing
- version-controlled scenarios CI deployment quality gates
- production traces replayable tests evaluation feedback loops
- voice agent projects LiveKit Retell integrations
- simulation fidelity text versus audio agent testing
Measured heat
now 0 pts/hpeak 2 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 743h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion
How the heat travelled
pace: p57 vs 519 stories at the 720h mark (now 743h old) — ahead of claude-subscriber-token-theft (1.0x), behind hydra-local-agentic-terminal (1.0x)
Evidence (3) — ⭐ canonical anchor
| source | object | author | score | comments |
| 🟧 hn | Show HN: Open-source simulation testing infra for voice agentsRetrieved article excerptOpen article · Retrieved 2026-09-10T17:26:22.692495+00:00 Egma Egma is an open-source platform for testing voice agents and monitoring them in production. Docs · Egma Cloud · Discord 🧩 Core features Test before you ship. Build regression suites with simulated voice and text conversations and mocked tool responses. Keep tests in your agent's repository and run them with the Egma CLI or a coding agent. Monitor production. Review real conversations, tool calls, and metrics. Use graders to check your agent's behavior and investigate failures. Bring your own keys. Use your own API keys for model providers on Egma Cloud or when you self-host. You pay providers directly, with no Egma markup on model usage. Customize your tests. Choose from supported models and customize your caller personas. Set grading instructions and pass thresholds for the behavior you want to test. Egma supports LiveKit (JS/TS) and Retell. Contact us on Discord to request another platform or feature. 🚀 Get started Run your first voice simulation in as little as five minutes. Open your voice agent's repository in your coding agent and paste: Set up Egma in this repo. Install its skills with
npx skills add egma-ai/egma, then use them to connect
my voice agent and run my first voice simulation. For Egma Cloud, create an account and use the prompt above.
To self-host, follow the self-hosting guide and include your instance's URL in the prompt. Allow extra time for installation
and the first build. 🖥️ Self-hosting Run Egma on your own machine git clone https://github.com/egma-ai/egma.git cd egma
npm install --global egma-cli
cp .env.example .env
chmod 600 .env Configure .env using the self-hosting guide , then start Egma: egma self-host up Open localhost:3101 and create your account.
See the self-hosting guide for connecting
your agent, server setup, and troubleshooting. 🧭 Where we're headed Egma supports testing and production monitoring today. We plan to add automatic
fixes: Egma would use a production failure to propose a change to your voice
agent's harness, test it, and open a pull request for you to review. 💬 Help and feedback Read the docs Join Discord Open a GitHub issue ⚖️ License This repository is MIT licensed, except for the ee folders. See LICENSE for more details. | Nischalj10 | 16 | 7 |
| 🟧 echo.github ⭐ | Egma provides voice and text simulation suites, mocked tool responses, CLI execution, production monitoring, and bring-your-own-provider key | Egma | — | — |
| 🟧 hn | Show HN: VoiceGremlin, a SaaS tool for automated tests against AI phone agents | ineptech | 3 | 1 |
Interpretation history
2026-10-07T19:19:46Z
VoiceGremlin's launch turns automated voice-agent testing from Egma's lone claim into a two-entrant niche — independent builders are converging on simulated conversations with automated LLM-judged pass/fail — but neither product shows adoption or independent verification, and VoiceGremlin's real-phone-call modality fits Scott's LiveKit/Retell stack less directly than Egma's repo-based simulation. The niche read strengthens; Egma-specific corroboration (users, third-party verification of fidelity) has not arrived.
2026-10-07T18:32:49Z
evidence attached: hn.story.49995804 — Second independent entrant in voice-agent automated testing (bot-vs-bot calls, LLM-judged pass/fail) supporting the niche the Egma case tracks.
2026-09-10T18:02:41Z
The retrieved repository establishes an available evaluation tool with documented repository-based suites, CLI execution and LiveKit/Retell support, but operational fidelity remains untested; the echo adds no independent corroboration. This look adds no material change to the launch package already routed for attention, and automatic harness fixes remain a roadmap item.
2026-09-10T17:49:04Z
grounded: converges/medium — Egma’s documented simulation, mock-tool injection and monitoring converge with Scott’s Evaluation-Driven Development position and offer a concrete harness appro
2026-09-10T17:44:01Z
case created — The retrieved repository establishes an actionable evaluation artifact, but not the scout's comparative affordability claim.
Decision trace
- 10-11 17:34review_screenjev screen: no material development (noul=0.06)
- 10-08 12:22attention_routeThe editor compared this story and chose to keep watching.
- 10-08 06:22attention_routeThe second independent entrant is the fact that moves this from a single-tool lead to a category signal relevant to active voice projects; with no deadline attached it earns a briefing paragraph at 10
- 10-08 06:19attention_candidatematerial_reprice
- 10-08 06:19repriceVoiceGremlin's launch turns automated voice-agent testing from Egma's lone claim into a two-entrant niche — independent builders are converging on simulated conversations with automated LLM-
- 10-08 06:19review_reactivatedVoiceGremlin's launch turns automated voice-agent testing from Egma's lone claim into a two-entrant niche — independent builders are converging on simulated conversations with automated LLM-
- 10-08 05:42attention_routeThe editor compared this story and chose to keep watching.
- 10-08 05:32attention_candidateattach
- 10-08 05:32attachSecond independent entrant in voice-agent automated testing (bot-vs-bot calls, LLM-judged pass/fail) supporting the niche the Egma case tracks.
- 10-08 05:31propose_attachSecond independent entrant in voice-agent automated testing (bot-vs-bot calls, LLM-judged pass/fail) supporting the niche the Egma case tracks.
- 10-04 08:05review_dormantscheduled targets exhausted or 28 quiet days
- 10-04 08:05drop_targetsquiet through full ladder or over cap 8
- 09-20 05:29review_screenThe only change is an additional generic expression of interest, adding no new fact, implementation result, contradiction, or consequential evidence.
- 09-12 07:33review_screenThe new comments express enthusiasm, general interest, and questions about production metrics, but add no independently verified implementation result, contradiction, release, access change, or conseq
- 09-11 13:21sensor_dirtyengagement_update
- 09-11 05:21sensor_dirtyengagement_update
- 09-11 04:02repriceThe retrieved repository establishes an available evaluation tool with documented repository-based suites, CLI execution and LiveKit/Retell support, but operational fidelity remains untested; the echo
- 09-11 04:02alert_silentThe actionable release package was already routed for attention. This pass contains no new capability, access change or independent implementation result, so another notification would repeat the laun
- 09-11 04:02alert_routeThe actionable release package was already routed for attention. This pass contains no new capability, access change or independent implementation result, so another notification would repeat the laun
- 09-11 03:59alert_shadowThe builders’ launch and supplied repository summary establish a concrete tool to investigate today for Scott’s Practice Trainer and voice-booking work: voice/text simulations, mocked tool responses,
- 09-11 03:59alert_routeThe builders’ launch and supplied repository summary establish a concrete tool to investigate today for Scott’s Practice Trainer and voice-booking work: voice/text simulations, mocked tool responses,
- 09-11 03:49groundEgma’s documented simulation, mock-tool injection and monitoring converge with Scott’s Evaluation-Driven Development position and offer a concrete harness approach to investigate for Practice Trainer
- 09-11 03:44createThe retrieved repository establishes an actionable evaluation artifact, but not the scout's comparative affordability claim.