Running your golden dataset as a CI regression gate
Once you have golden test cases in a dataset (Dataset Studio), you can run them against your agent on every PR and have a broken capability fail CI instead of showing up as a support ticket. AgentLasso doesn't invoke your agent itself — your own CI job does, since it's already the thing with credentials and network access to reach it. AgentLasso's job is narrower: hand back the dataset, grade what you send back, and tell you pass/fail.
The loop
POST /api/v1/eval-runswith your dataset's ID. Get back arun_idand the full list of test cases (input,evaluation_criteria_json, andchecks— exactly what will be graded) in that dataset.- Your script calls your agent once per test case, however it normally would (HTTP call, CLI invocation, SDK call — whatever "running your agent" means for you), and collects its final output and the tools it called.
POST /api/v1/eval-run-resultswith therun_idand one entry per test case:{ test_case_id, actual_output?, actual_tool_calls }. AgentLasso grades all of it synchronously and returns the result in the same response — no polling.- Read
ci_passfrom the response and exit non-zero if it'sfalse.
That's the whole contract. One eval-runs call, one eval-run-results call,
one exit code.
How grading works
Every test case carries a list of checks (set up in Dataset Studio). Each one is evaluated deterministically against what your agent actually did — no LLM involved:
Tools match by name, not arguments — real agents vary argument shape in ways that don't indicate a regression. Text matching is case-insensitive.
- Severity. A failing
criticalcheck fails the test case; a failingminorcheck lowers itsscore(the percentage of checks passed) but doesn't fail it. - Send
actual_output. An output check can't pass without it — if you don't report it, those checks fail and say why, rather than passing unverified. - No checks = not graded. A test case with no checks verified nothing:
it comes back
graded: falseand is left out ofpass_rate. A run where nothing was graded returnspass_rate: null,ci_pass: falseand aci_fail_reason. - Every result says exactly what happened.
check_resultslists each check withpassedand adetail(e.g.Forbidden “guarantee” found: …we guarantee it…), and the Test Runs page shows the agent's actual output and tool calls next to them.
Free-text rubric items on a test case (semantic_guardrails /
negative_constraints, e.g. "states the refund amount") are not
evaluated — grading them correctly needs an LLM judge, which is a planned
follow-up, not something we'd fake with keyword matching. judge_reasoning
says when a test case has them, so you always know what a pass did and
didn't check.
pass_threshold (0–1, default 1.0) controls how strict the gate is — the
default means any single failing graded test case fails the run. Pass a
lower value on eval-runs if you want tolerance for a known-flaky subset
while you're building out coverage.
Example: GitHub Actions
Use the same adapter in AgentLasso
If callAgent above is (or wraps) an HTTPS endpoint that takes { input } and
returns { output, tool_calls }, add it under Settings → Your agent.
AgentLasso can then call your agent directly, e.g. to try a test case's rubric
against a live answer before you save it. It also sends source
(agentlasso-test, agentlasso-preview or agentlasso-sandbox) and, when
running a test case, test_case_id, so you can tag those requests. Tool-call
arguments may be an object or a JSON string. Calls time out after 30 seconds
and redirects aren't followed.
Auth
Same project API key as trace ingestion — x-api-key header (or
Authorization: Bearer <key>), from Settings.