Checks reference

A golden test case is a request plus the checks a correct run must pass. Each check is one explicit assertion about what the agent did (its tool calls) or said (its final reply), graded the same way in the rubric editor's preview and in CI.

Check types

CheckPasses whenNeeds
Tool must be calledthe tool appears among the run's tool callstool calls
Tool must NOT be calledthe tool never appearstool calls
Tools called in this orderthe tools appear in this order; other calls may happen in between (lookup → refund passes for lookup → policy → refund)tool calls
Output must containthe text appears in the reply: case-insensitive, or a regular expression if you tick Regular expressionthe reply
Forbidden in outputthe text (or regex) does not appear in the replythe reply
Output is valid JSONthe reply parses as JSON (surrounding whitespace is ignored)the reply
AI judge: meets a criterionClaude decides the reply meets a plain-English criterion. See AI judge checksthe reply + an Anthropic key

Tool checks match on the tool name. Arguments aren't compared, because argument shapes vary from run to run in ways that don't indicate a regression.

Severity

Every check is critical or minor:

  • Critical: failing it fails the test case.
  • Minor: failing it lowers the score and is reported, but the test case can still pass.

How a test case is graded

  • Score = the share of checks that passed (out of those that could be evaluated).
  • Pass = no critical check failed, and no critical check went unverified.
  • No checks → not graded. A test case without checks verifies nothing, so it's reported as not graded and left out of the pass rate. It never counts as a pass.
  • Content checks need the reply. If a run didn't report the agent's reply, content checks fail and say why, instead of passing unverified.
  • An invalid regex fails the check with the error, instead of crashing the run.
  • AI judge checks without a verdict (no API key, an outage) are not graded. A critical one keeps the test case from passing, and the case is flagged partially graded.
  • Free-text rubric items (plain sentences on older test cases, like "states the refund amount") aren't graded and are listed as such. Convert them into AI judge checks to grade them.

Each check result comes with a human-readable detail, e.g. "stripe.refund was never called. The agent called: orders.lookup." or "Forbidden “guarantee” found: …should arrive in 3-5 business days. I guarantee it."

In the data

Checks are stored on the test case as a versioned list. The CI API returns them with each test case, so your pipeline can see what will be graded:

json
{  "checks_version": 1,  "checks": [    { "id": "c1", "type": "tool_called", "tool": "stripe.refund", "severity": "critical" },    {      "id": "c2",      "type": "tool_order",      "tools": ["orders.lookup", "stripe.refund"],      "severity": "minor"    },    { "id": "c3", "type": "output_contains", "value": "3-5 business days", "severity": "critical" },    {      "id": "c4",      "type": "output_not_contains",      "value": "\\b(?:\\d[ -]?){13,16}\\b",      "regex": true,      "severity": "critical"    },    { "id": "c5", "type": "output_valid_json", "severity": "minor" },    {      "id": "c6",      "type": "llm_judge",      "criterion": "States the refund amount",      "severity": "critical"    }  ]}

Types: tool_called, tool_not_called, tool_order, output_contains, output_not_contains, output_valid_json, llm_judge.

Older test cases that only have expected_tool_calls or a guardrails list are read as the equivalent checks, so they keep working.

Next

Write checks without code in the rubric editor.