Apr 30, 2026EnterpriseResearch

Expanding into broader research models

From a speech stack to models that carry long-horizon work, and the evaluation discipline to know whether they do.

The interesting problem is no longer whether a model can produce a good paragraph. It is whether it can hold a task across hours, act through tools with consequences, and know when to stop.

The direction

Our research is organized around one question: can a model be trusted to carry work end to end for an organization that cannot afford a confident mistake?

That question decomposes into long-horizon coherence, tool use under real permissions, calibrated uncertainty, and robustness to adversarial input.

Long-horizon coherence: the constraint set four steps ago still binds.
Calibrated uncertainty: knowing when to stop and ask.

What long-horizon means

A task that spans hours, several systems, and multiple people is not a longer version of a short task. Errors compound, context has to be maintained rather than re-read, and the correct action often depends on something established much earlier.

The hard part is not the step. It is that step nineteen depends on step three.

Measuring it

We built OpsBench because existing benchmarks score the final answer and ignore what the model did on the way there. Operational work has side effects, and side effects have to be scored.

Outcome, side effects, and escalation quality are scored separately.

Safety is part of capability

A model that cannot be trusted with an action cannot be given the action, which caps how useful it can be. Red-teaming through Heyno Violet and the published system cards are not a separate workstream from capability. They are what determines how much autonomy is available.

Long-horizonResearch focus
Side effectsScored
System cardsPublished

How we publish

Model releases ship with a system card. Benchmarks ship with methodology. Where results are internal, we say they are internal: the current models are deployed with controlled enterprises and agencies rather than publicly.

More in Enterprise

Keep reading

View all