Running your golden dataset as a CI regression gate

Once you have golden test cases in a dataset (Dataset Studio), you can run them against your agent on every PR and have a broken capability fail CI instead of showing up as a support ticket. AgentLasso doesn't invoke your agent itself — your own CI job does, since it's already the thing with credentials and network access to reach it. AgentLasso's job is narrower: hand back the dataset, grade what you send back, and tell you pass/fail.

The loop

  1. POST /api/v1/eval-runs with your dataset's ID. Get back a run_id and the full list of test cases (input, evaluation_criteria_json, and checks — exactly what will be graded) in that dataset.
  2. Your script calls your agent once per test case, however it normally would (HTTP call, CLI invocation, SDK call — whatever "running your agent" means for you), and collects its final output and the tools it called.
  3. POST /api/v1/eval-run-results with the run_id and one entry per test case: { test_case_id, actual_output?, actual_tool_calls }. AgentLasso grades all of it synchronously and returns the result in the same response — no polling.
  4. Read ci_pass from the response and exit non-zero if it's false.

That's the whole contract. One eval-runs call, one eval-run-results call, one exit code.

How grading works

Every test case carries a list of checks (set up in Dataset Studio). Each one is evaluated deterministically against what your agent actually did — no LLM involved:

CheckPasses whenNeeds
Tool must be calledthe tool appears in actual_tool_callstool calls
Tool must NOT be calledthe tool never appearstool calls
Tools called in orderthe tools appear in that order (other calls may happen in between)tool calls
Output must containthe text (or regex) appears in actual_outputoutput
Forbidden in outputthe text (or regex) does not appear in actual_outputoutput
Output is valid JSONactual_output parses as JSONoutput

Tools match by name, not arguments — real agents vary argument shape in ways that don't indicate a regression. Text matching is case-insensitive.

  • Severity. A failing critical check fails the test case; a failing minor check lowers its score (the percentage of checks passed) but doesn't fail it.
  • Send actual_output. An output check can't pass without it — if you don't report it, those checks fail and say why, rather than passing unverified.
  • No checks = not graded. A test case with no checks verified nothing: it comes back graded: false and is left out of pass_rate. A run where nothing was graded returns pass_rate: null, ci_pass: false and a ci_fail_reason.
  • Every result says exactly what happened. check_results lists each check with passed and a detail (e.g. Forbidden “guarantee” found: …we guarantee it…), and the Test Runs page shows the agent's actual output and tool calls next to them.

Free-text rubric items on a test case (semantic_guardrails / negative_constraints, e.g. "states the refund amount") are not evaluated — grading them correctly needs an LLM judge, which is a planned follow-up, not something we'd fake with keyword matching. judge_reasoning says when a test case has them, so you always know what a pass did and didn't check.

pass_threshold (0–1, default 1.0) controls how strict the gate is — the default means any single failing graded test case fails the run. Pass a lower value on eval-runs if you want tolerance for a known-flaky subset while you're building out coverage.

Example: GitHub Actions

yaml
name: agent-eval
on: pull_request
jobs:  eval:    runs-on: ubuntu-latest    steps:      - uses: actions/checkout@v4      - uses: actions/setup-node@v4        with:          node-version: 22      - run: npm ci      - name: Run golden dataset against this PR's agent        env:          AGENTLASSO_API_KEY: ${{ secrets.AGENTLASSO_API_KEY }}          AGENTLASSO_DATASET_ID: ${{ vars.AGENTLASSO_DATASET_ID }}        run: node scripts/run-eval.mjs
javascript
// scripts/run-eval.mjsconst API_KEY = process.env.AGENTLASSO_API_KEY;const DATASET_ID = process.env.AGENTLASSO_DATASET_ID;const BASE_URL = "https://agentlasso.dev/api/v1";
async function callAgent(input) {  // Replace with however you actually invoke your agent -- an HTTP call to a  // staging deployment, a local process, an SDK call. Must return its final  // output and the tools it called ({ name, arguments? }), so AgentLasso can  // grade both content and trajectory checks.  const response = await fetch(process.env.AGENT_ENDPOINT, {    method: "POST",    headers: { "content-type": "application/json" },    body: JSON.stringify({ input }),  });  const body = await response.json();  return { output: body.output, toolCalls: body.tool_calls ?? [] };}
async function main() {  const createRes = await fetch(`${BASE_URL}/eval-runs`, {    method: "POST",    headers: { "content-type": "application/json", "x-api-key": API_KEY },    body: JSON.stringify({ dataset_id: DATASET_ID }),  });  const { run_id, test_cases } = await createRes.json();
  const results = [];  for (const testCase of test_cases) {    const { output, toolCalls } = await callAgent(testCase.input);    results.push({      test_case_id: testCase.test_case_id,      actual_output: output,      actual_tool_calls: toolCalls,    });  }
  const reportRes = await fetch(`${BASE_URL}/eval-run-results`, {    method: "POST",    headers: { "content-type": "application/json", "x-api-key": API_KEY },    body: JSON.stringify({ run_id, results }),  });  const graded = await reportRes.json();
  const rate = graded.pass_rate === null ? "n/a" : `${(graded.pass_rate * 100).toFixed(1)}%`;  console.log(`Pass rate: ${rate} (threshold: ${graded.pass_threshold * 100}%)`);  for (const r of graded.results) {    const label = !r.graded ? "SKIP" : r.pass_fail ? "PASS" : "FAIL";    console.log(`${label} ${r.test_case_id}: ${r.judge_reasoning}`);  }
  if (!graded.ci_pass) {    console.error(graded.ci_fail_reason ?? "Eval run failed the configured pass threshold.");    process.exit(1);  }}
main().catch((err) => {  console.error(err);  process.exit(1);});

Use the same adapter in AgentLasso

If callAgent above is (or wraps) an HTTPS endpoint that takes { input } and returns { output, tool_calls }, add it under Settings → Your agent. AgentLasso can then call your agent directly, e.g. to try a test case's rubric against a live answer before you save it. It also sends source (agentlasso-test, agentlasso-preview or agentlasso-sandbox) and, when running a test case, test_case_id, so you can tag those requests. Tool-call arguments may be an object or a JSON string. Calls time out after 30 seconds and redirects aren't followed.

Auth

Same project API key as trace ingestion — x-api-key header (or Authorization: Bearer <key>), from Settings.