Evals
Pin a conversation as a regression test and re-run it on every change.
Coming soon. Evals are not available yet. This page describes what they will do so you can plan for it.
The problem
Editing a skill changes how every agent behaves. Nothing today tells you whether the edit broke a case that used to work. You find out when someone reports it.
What evals will do
Evals turn a conversation into a regression test.
- Pin any Studio conversation as a test case.
- Re-run the suite when a role is published, on demand, or on a schedule.
- Read a pass or fail signal per case, with the differences between runs grouped together.
Until then
Test a role in Studio before you publish it. Studio runs a real conversation against the role and shows which skills and tools it used. That check is manual and it leaves no record, which is the gap evals close.

