Continuous behavioral coverage for AI agents

Discover what your agent actually does in production, identify what you're not testing, turn real behavior into golden tests, and prevent regressions before they ship.

AI agents change constantly. A prompt, model, tool or routing change can fix one problem and silently break another, and you usually find out from users. AgentLasso watches your agent's real production behavior and keeps your regression suite in step with it, so you can ship changes with confidence that what already worked still works.

AgentLasso in 20 seconds: observe production traces, discover capabilities and edge cases, convert them into tests, catch the regression in CI

The loop

1. Discover what your agent actually does. Send your production traces (a flat JSON POST or OpenTelemetry). AgentLasso groups them into capabilities, the jobs your agent performs, like process_refund or track_shipment, and scores each one on usage, success, risk and business impact.

2. Identify what you're not testing. The Agent Quality Coverage number tells you how much of your agent's real behavior your golden tests check, broken down into risk, traffic, failure-mode and capability coverage. Every gap is listed with a real example request.

3. Turn real behavior into golden tests. Approve real conversations from the Review Queue and define what "correct" means with checks: which tools must or must never be called, what the reply must or must never say, and plain-English criteria graded by an AI judge.

4. Prevent regressions before they ship. Your CI runs the golden suite against your agent on every pull request. AgentLasso grades every check and fails the build when a critical one fails.

Core concepts

TermMeaning
TraceOne request your agent handled: the user's message, the agent's reply, and the tools it called (with arguments).
CapabilityA job your agent performs, e.g. process_refund. Traces are classified into capabilities, first by matching their tool calls (free), then by Claude for the rest.
Tool pathThe tools a trace called, in order, e.g. orders.lookup → stripe.refund. Coverage is measured on tool paths.
Golden test caseA real (or written) request plus the checks a correct answer must pass. Test cases live in golden datasets.
CheckOne assertion about a run: a tool is (or isn't) called, tools run in order, the reply contains (or never contains) some text, the reply is valid JSON, or an AI judge criterion. A check is critical (failing it fails the test) or minor (lowers the score).
CoverageHow much of your agent's real behavior your golden tests check, not whether the agent passes them.
RegressionBehavior that used to work and broke after a change, caught when a golden test fails in CI.

How it fits your stack

  • In: production traces from your agent, either a simple JSON POST or standard OpenTelemetry (GenAI conventions). If you already send traces to LangSmith, Langfuse or Datadog, keep doing that and dual-export.
  • Out: a pass/fail signal in your CI. Your pipeline fetches the golden dataset, runs it through your agent, and sends back what the agent said and did. See Run golden tests in CI.
  • Your agent, optionally: connect an HTTPS endpoint to grade test cases against your agent's live answers from the rubric editor, and to chat with it in the Live Sandbox.

PII (emails, phone numbers, SSNs, card numbers) is redacted from traces before storage by default.

Open source

AgentLasso is open-core: everything except the ee/ folder is Apache-2.0, so you can self-host it. Intent classification and AI judge checks run on your own Anthropic API key. Paid plans add hosted infrastructure and team governance.

Next

Follow the Quickstart to go from your first trace to a golden test running in CI.