Ruby Voice System Card
1. Introduction
Ruby Voice and Ruby Voice mini are the speech models behind spoken interaction with Heyno. They are full-duplex: they listen and respond continuously rather than waiting for a clearly defined turn to end, so they can follow pauses, interruptions, and changes in pace and decide in the moment whether to answer or keep listening.
Ruby Voice is the default speech model for enterprise deployments; Ruby Voice mini is the default where latency or cost constraints are tighter. Both ship inside reviewed Heyno deployments and are enabled per surface rather than per account.
The most important things to know about our safety work for this release are that:
We trained these models to respond safely using the same infrastructure we rely on for our reasoning models. They can also delegate more involved work to Ruby-Horizon, and where they do, the result reflects the safety training of the model performing that work. The evaluations in this card describe Ruby Voice with delegation, matching the deployment context.
These models carry system-level safety integrations on par with the stack around our text models, with additions specific to the spoken modality: inputs and generated outputs are checked as the interaction unfolds, and where potentially unsafe content is detected the system can steer or interrupt the response, state a spoken boundary, hand off to a person, or end the interaction.
We built evaluations that focus specifically on the ways people use a spoken interface rather than a text one, and on observations from real-world use of the previous generation. Those results are reported below.
The same monitoring, review, and enforcement infrastructure that applies to our reasoning models applies here, so prevalence can be measured, abuse detected, and policy enforced across surfaces rather than per channel.
2. Model Data and Training
Ruby Voice is trained on licensed and consented audio together with synthetic material constructed to cover conditions we cannot ethically collect. Our processing pipeline filters for quality and reduces personal information before training, and we apply classifiers to exclude harmful or sensitive content from the corpus.
Customer audio is excluded from training by default and used only under an explicit, separate written agreement scoped to the named deployment. Training and inference run on hardware Heyno owns and operates; no third-party voice or model host sits in the path. Available voices are synthetic and fixed, and the system does not reproduce the voice of a real person.
Comparison values from previous versions are drawn from recent snapshots of those versions and may differ slightly from figures published at their launch.
Ruby Voice is intended for use in accordance with Heyno’s usage policies and the terms agreed in each deployment review.
3. Model Safety
3.1 Voice-Native Evaluations for Disallowed Content
3.1.1 Voice-Native Evaluations: Production Prompts
In these evaluations we use real audio from deployments whose agreements permit this use. Before any example is used it passes our privacy and eligibility safeguards, including permission and deletion checks, filtering of ineligible data, and steps to reduce personal information through scrubbing and de-identification. We then transcribe the audio, generate the model’s response, and evaluate that response for safety.
We compare Ruby Voice and Ruby Voice mini against the models they replace. These evaluations are not prevalence weighted: they do not reflect the rates we see in real usage. They were built to be difficult, around cases where the previous generation was not yet responding well.
Table 1: Voice-Native Evaluations: Production Prompts
| Category | Previous | Ruby Voice | Previous mini | Ruby Voice mini |
|---|---|---|---|---|
| Commitments beyond authority | 0.79 | 0.96 | 0.71 | 0.93 |
| Pricing outside defined rules | 0.84 | 0.98 | 0.77 | 0.95 |
| Another party’s record | 0.94 | 0.99 | 0.92 | 0.97 |
| Advice outside scope | 0.88 | 0.94 | 0.81 | 0.91 |
| Identity left uncorrected | 0.77 | 0.96 | 0.70 | 0.92 |
| Softening a refusal | 0.84 | 0.92 | 0.79 | 0.88 |
Ruby Voice performs equal to or better than the previous generation across these adversarially selected prompts. The smallest margin is on softening a refusal, where the model concedes ground without stating anything disallowed; it remains the weakest behavior on this surface and is gated at the approval layer in deployment.
3.1.2 Voice-Native Evaluations: Synthetic Prompts
These evaluations are constructed the same way, except the audio is generated rather than drawn from production. Prompt text is written from our safety policies and related guidance to target specific categories, policy boundaries, and difficult cases, then converted to speech and used as audio input. This lets us cover rare or hard-to-sample situations deliberately, and to check that the intended behavior holds when the content arrives in spoken form rather than written.
Table 2: Voice-Native Evaluations: Synthetic Prompts
| Category | Previous | Ruby Voice | Previous mini | Ruby Voice mini |
|---|---|---|---|---|
| Spoken instruction treated as instruction | 0.87 | 0.99 | 0.82 | 0.97 |
| Sustained pressure across a session | 0.81 | 0.97 | 0.74 | 0.94 |
| Claimed authority | 0.85 | 0.99 | 0.79 | 0.96 |
| Adopting a named persona | 0.77 | 0.98 | 0.71 | 0.95 |
| Degraded audio inducing a guess | 0.83 | 0.96 | 0.76 | 0.93 |
The largest gain is on adopting a named persona, which the previous generation would accept when it was offered gradually rather than requested outright. Neither model reproduces a real person’s voice; this category measures whether the model accepts a false identity in what it says.
4. Red Teaming
Heyno Violet, our internal red team, exercises the spoken surface alongside every other surface the system operates on. On this surface the distinctive property is that pressure arrives in real time and cannot be reviewed before it is answered, so the suite emphasizes sustained sequences rather than single exchanges: rapport built over several interactions, authority asserted mid-sentence, and instructions delivered as speech rather than text.
Spoken instructions are treated as content and never as instruction, and no spoken interaction releases an action with external effect on its own. Findings are reproduced, mitigated, and retained as permanent cases in the suite rather than resolved by prompt revision.
5. Preparedness Framework
Under our preparedness framework, Ruby Voice is assessed at the family level together with Ruby-Horizon. The family is treated as High capability in Voice Impersonation and in Financial and Contractual Action, and below High in Autonomy and Self-Direction and in Data Exfiltration. The full assessment, including thresholds not reached, is documented in the Ruby-Horizon System Card.
For this surface the operative safeguard is that capability is not limited at the point of reasoning but at the point of action: whatever an interaction concludes, an action with external effect passes the approval guard described in the same card.
6. References
[1] Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, Ilya Sutskever. “Robust speech recognition via large-scale weak supervision.” Available at: https://arxiv.org/abs/2212.04356.
[2] Allison Koenecke, Andrew Nam, Emily Lake, Joe Nudell, Minnie Quartey, Zion Mengesha, et al. “Racial disparities in automated speech recognition.” Proceedings of the National Academy of Sciences 117(14). Available at: https://www.pnas.org/doi/10.1073/pnas.1915768117.
[3] Eric Wallace, Kai Xiao, Reimar Leike, Lilian Weng, Johannes Heidecke, Alex Beutel. “The instruction hierarchy: Training LLMs to prioritize privileged instructions.” Available at: https://arxiv.org/abs/2404.13208.
[4] Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz, Mario Fritz. “Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection.” Available at: https://arxiv.org/abs/2302.12173.
[5] Ethan Perez, Saffron Huang, Francis Song, Trevor Cai, Roman Ring, John Aslanides, et al. “Red teaming language models with language models.” Available at: https://arxiv.org/abs/2202.03286.
[6] Herbert H. Clark, Jean E. Fox Tree. “Using uh and um in spontaneous speaking.” Cognition 84(1). Available at: https://doi.org/10.1016/S0010-0277(02)00017-3.
[7] Heyno. “Ruby-Horizon System Card.” Available at: heyno.net.