Skilder

Evals

Pin a conversation as a regression test and re-run it on every change.

Coming soon. Evals are not available yet. This page describes what they will do so you can plan for it.

The problem

Editing a skill changes how every agent behaves. Nothing today tells you whether the edit broke a case that used to work. You find out when someone reports it.

What evals will do

Evals turn a conversation into a regression test.

  • Pin any Studio conversation as a test case.
  • Re-run the suite when a role is published, on demand, or on a schedule.
  • Read a pass or fail signal per case, with the differences between runs grouped together.

Until then

Test a role in Studio before you publish it. Studio runs a real conversation against the role and shows which skills and tools it used. That check is manual and it leaves no record, which is the gap evals close.