Checks reference
A golden test case is a request plus the checks a correct run must pass. Each check is one explicit assertion about what the agent did (its tool calls) or said (its final reply), graded the same way in the rubric editor's preview and in CI.
Check types
Tool checks match on the tool name. Arguments aren't compared, because argument shapes vary from run to run in ways that don't indicate a regression.
Severity
Every check is critical or minor:
- Critical: failing it fails the test case.
- Minor: failing it lowers the score and is reported, but the test case can still pass.
How a test case is graded
- Score = the share of checks that passed (out of those that could be evaluated).
- Pass = no critical check failed, and no critical check went unverified.
- No checks → not graded. A test case without checks verifies nothing, so it's reported as not graded and left out of the pass rate. It never counts as a pass.
- Content checks need the reply. If a run didn't report the agent's reply, content checks fail and say why, instead of passing unverified.
- An invalid regex fails the check with the error, instead of crashing the run.
- AI judge checks without a verdict (no API key, an outage) are not graded. A critical one keeps the test case from passing, and the case is flagged partially graded.
- Free-text rubric items (plain sentences on older test cases, like "states the refund amount") aren't graded and are listed as such. Convert them into AI judge checks to grade them.
Each check result comes with a human-readable detail, e.g. "stripe.refund was never called. The agent called: orders.lookup." or "Forbidden “guarantee” found: …should arrive in 3-5 business days. I guarantee it."
In the data
Checks are stored on the test case as a versioned list. The CI API returns them with each test case, so your pipeline can see what will be graded:
Types: tool_called, tool_not_called, tool_order, output_contains, output_not_contains, output_valid_json, llm_judge.
Older test cases that only have expected_tool_calls or a guardrails list are read as the equivalent checks, so they keep working.
Next
Write checks without code in the rubric editor.