Expanding into broader research models
From a speech stack to models that carry long-horizon work, and the evaluation discipline to know whether they do.
The interesting problem is no longer whether a model can produce a good paragraph. It is whether it can hold a task across hours, act through tools with consequences, and know when to stop.
The direction
Our research is organized around one question: can a model be trusted to carry work end to end for an organization that cannot afford a confident mistake?
That question decomposes into long-horizon coherence, tool use under real permissions, calibrated uncertainty, and robustness to adversarial input.
What long-horizon means
A task that spans hours, several systems, and multiple people is not a longer version of a short task. Errors compound, context has to be maintained rather than re-read, and the correct action often depends on something established much earlier.
Measuring it
We built OpsBench because existing benchmarks score the final answer and ignore what the model did on the way there. Operational work has side effects, and side effects have to be scored.
Safety is part of capability
A model that cannot be trusted with an action cannot be given the action, which caps how useful it can be. Red-teaming through Heyno Violet and the published system cards are not a separate workstream from capability. They are what determines how much autonomy is available.
How we publish
Model releases ship with a system card. Benchmarks ship with methodology. Where results are internal, we say they are internal: the current models are deployed with controlled enterprises and agencies rather than publicly.
More in Enterprise