Today we’re introducing Ruby Voice, the voice model behind every call Heyno answers. It is built for one job: sounding like a person on a real phone line, where the caller is in a truck, the line is noisy, and nobody waits politely for a model to finish thinking.
Built for the phone, not the demo
Most voice stacks chain three models together: one to transcribe, one to think, one to speak. Every hand-off costs time, and the caller hears all of it. Ruby Voice is a single model with speech in and speech out, so a turn comes back in a fraction of a second.
- Noise is normal. Trained on calls from trucks, job sites, and speakerphones, not studio audio.
- Barge-in built in. It stops when the caller starts talking, the way people do.
- Your voice, consistently. Tone and greeting stay the same on the thousandth call.
Latency you can hear
On the phone, the gap before a reply is the whole experience. Under 300 ms feels like a conversation; past a second the caller starts talking over you.
Interruptions and handoff
Ruby Voice yields the moment a caller cuts in, and it knows when the call has stopped being its job: a price dispute, an upset customer, anything outside the rules you set. It hands off to a person with the context already summarized, and texts you what happened.
Voices and languages
| Capability | Detail |
|---|---|
| Voices | A set of natural voices; pick one per number or per team |
| Languages | English and Spanish today, with more in preview |
| Bilingual calls | Switches language mid-call when the caller does |
| Custom greeting | Your own opening line, spoken the same way every time |
Availability
Ruby Voice powers Frontdesk on every paid plan, and it is the voice inside Ruby-Horizon for approved enterprise deployments. Personal accounts get it for their own calls in the app. See Models for how to choose.