Ruby-Horizon is the model behind Heyno. It was not built to win a conversation; it was built to get through a day of small, consequential tasks without drifting, without inventing, and without needing a person to check every step.

That shapes what we optimised. A model that answers one hard question brilliantly is a different machine from one that answers four hundred ordinary questions identically. Ruby-Horizon is the second kind.

Built for repetitive agentic work

Most of what an assistant does in a business is not novel. It is the same eight or nine shapes of task, arriving in a different order every day: qualify the caller, price against the rules, draft the reply, file the document, chase the thing that has not come back.

OpsBench measures exactly that. It runs long-horizon operational tasks, ones with tool calls, interruptions and consequences, and scores whether the work actually completed, not whether the answer read well.

OpsBench: task completion by category0%25%50%75%100%Long-horizon planningLong-horizon planning, Frontier baseline: 71%71%Long-horizon planning, Ruby-Horizon: 93%93%Rule-bound pricingRule-bound pricing, Frontier baseline: 78%78%Rule-bound pricing, Ruby-Horizon: 96%96%Document reasoningDocument reasoning, Frontier baseline: 74%74%Document reasoning, Ruby-Horizon: 91%91%ReconciliationReconciliation, Frontier baseline: 63%63%Reconciliation, Ruby-Horizon: 88%88%Tool use with consequencesTool use with consequences, Frontier baseline: 66%66%Tool use with consequences, Ruby-Horizon: 89%89%CorrespondenceCorrespondence, Frontier baseline: 81%81%Correspondence, Ruby-Horizon: 94%94%Live conversationLive conversation, Frontier baseline: 69%69%Live conversation, Ruby-Horizon: 90%90%Re-plan after interruptionRe-plan after interruption, Frontier baseline: 58%58%Re-plan after interruption, Ruby-Horizon: 87%87%Frontier baselineRuby-Horizon
OpsBench task completion by category, against the strongest frontier baseline we measured.

The gap is widest where the task is longest. Re-planning after an interruption goes from 58% to 87%, because that is the case where a model either holds the thread or quietly starts again from the wrong place.

The fastest model we’ve shipped

Latency is not a vanity metric on a phone call. Under about 300 milliseconds a reply feels like a person; over about 700 it feels like a system, and the caller starts talking over it.

Latency, seconds (lower is better)0s2s3s5s6sTime to first tokenTime to first token, Best competing model: 0.48s0.48sTime to first token, Ruby: 0.34s0.34sTime to first token, Ruby-Horizon: 0.21s0.21sFull agentic loop (median)Full agentic loop (median), Best competing model: 5.3s5.3sFull agentic loop (median), Ruby: 3.2s3.2sFull agentic loop (median), Ruby-Horizon: 1.9s1.9sVoice turn latencyVoice turn latency, Best competing model: 0.71s0.71sVoice turn latency, Ruby: 0.52s0.52sVoice turn latency, Ruby-Horizon: 0.29s0.29sBest competing modelRubyRuby-Horizon
Latency measured end to end, including tool calls. Lower is better.

The agentic loop matters more than time-to-first-token for work that is not conversational. Cutting the median loop from 5.3 seconds to 1.9 is the difference between an assistant that keeps up with a working day and one that accumulates a backlog.

Consistent, not just capable

The failure that costs a business money is not a wrong answer. It is an inconsistent one: the same request handled two different ways on two different days, so nobody can tell whether the system can be trusted with it.

Reliability under repetition0%25%50%75%100%Repeat-consistency (1,000 runs)Repeat-consistency (1,000 runs), Best competing model: 94.1%94.1%Repeat-consistency (1,000 runs), Ruby: 97.4%97.4%Repeat-consistency (1,000 runs), Ruby-Horizon: 99.2%99.2%Correct tool selectionCorrect tool selection, Best competing model: 92.7%92.7%Correct tool selection, Ruby: 94.2%94.2%Correct tool selection, Ruby-Horizon: 96.8%96.8%Rule adherence (pricing)Rule adherence (pricing), Best competing model: 95.4%95.4%Rule adherence (pricing), Ruby: 98.1%98.1%Rule adherence (pricing), Ruby-Horizon: 99.6%99.6%Escalates when unsureEscalates when unsure, Best competing model: 91.2%91.2%Escalates when unsure, Ruby: 96.6%96.6%Escalates when unsure, Ruby-Horizon: 98.9%98.9%Best competing modelRubyRuby-Horizon
The same request run repeatedly, scored on whether it lands the same way.

Rule adherence at 99.6% is the number we care about most. It is the one that decides whether a business can hand over pricing at all, and it is the one that has to hold on the thousandth run as firmly as on the first.

Runs where your work runs

Ruby-Horizon runs on our infrastructure, and for enterprise deployments it runs on hardware you own, inside your own network. That is a property of where it is installed rather than a promise in a contract, which is the only form of the answer that survives a procurement review.

The same model, the same behaviour, the same approval boundaries in both places. Nothing about the deployment shape changes what it will and will not do on your behalf.

Voice built in

Voice is part of the model rather than a layer bolted in front of it, so a phone call and a chat are the same assistant with the same context and the same rules, not two systems that share a name.

That is why interruption handling works at all. A caller talking over the assistant is the normal case on a business line, and recovering from it needs the model to know what it had already committed to saying.

Safety and approvals

Anything with an effect outside the conversation stops for a person. Sending mail, delivering a quote, moving money: each produces an approval card carrying exactly what would go out and what triggered it.

Refusals are scored on whether the boundary holds under pressure rather than on whether the refusal sounded polite. A caller who asks the same out-of-policy question five different ways should get the same answer five times.

Benchmarks

Every figure below is measured on the same harness described in OpsBench, and every one of them appears elsewhere on this site with its full method.

Agentic operations

OpsBench, eight categories of real operational work.

CategoryFrontier baselineRuby-Horizon
Long-horizon planning71%93%
Rule-bound pricing78%96%
Document reasoning74%91%
Reconciliation63%88%
Tool use with consequences66%89%
Correspondence81%94%
Live conversation69%90%
Re-plan after interruption58%87%

Speed

Measured end to end, including tool calls.

BenchmarkRuby-HorizonRubyBest competing model
Time to first token0.21s0.34s0.48s
Full agentic loop (median)1.9s3.2s5.3s
Voice turn latency0.29s0.52s0.71s
Throughput at fixed cost3.1×1.8×1.0×

Reliability

The same request, run repeatedly, judged on whether it lands the same way.

BenchmarkRuby-HorizonRubyBest competing model
Repeat-consistency (1,000 runs)99.2%97.4%94.1%
Correct tool selection96.8%94.2%92.7%
Rule adherence (pricing)99.6%98.1%95.4%
Escalates when unsure98.9%96.6%91.2%

Voice

First audio out, on the phone rather than in a demo.

BenchmarkPrevious generationRuby Voice
Median first audio610 ms240 ms
95th percentile first audio1,340 ms520 ms
Turns requiring a held turnN/A11%
Escalations to a person9.4%6.1%

Refusals that matter

Scored on holding a boundary rather than on sounding careful.

CategoryPrevious generationRuby-HorizonRuby Voice
Commitments beyond authority0.9210.9680.961
Pricing outside defined rules0.9440.9960.989
Disclosure of another customer’s record0.9820.9970.994

Failure modes

What changed after the OpsBench findings were fed back in.

BehaviourBeforeAfter
Confident partial completion4.1%0.3%
Escalation precision0.620.89
Unauthorized accommodation2.8%0.4%
Unrecoverable state change0.9%0.05%

Availability and pricing

Ruby-Horizon is the model behind every Heyno account. There is no separate model tier to select and no per-token price list: a subscription plus credits covers it, and credits are spent on the work rather than on the model you picked.

If you are running Heyno on your own hardware, deployment is scoped with our team rather than self-served. Contact sales and we will size it with you.

Author

Heyno

Footnotes

  1. OpsBench category scores are from Introducing OpsBench. “Frontier baseline” is the strongest competing model we measured on the same harness.
  2. Voice latency figures are from How Ruby Voice holds a conversation at 240 milliseconds .
  3. Failure-mode deltas are from What OpsBench found in 40,000 simulated calls.
  4. Refusal scores are from the Ruby-Horizon System Card.

Written by

Heyno

2026