Quickstart

From your first production trace to a golden test running in CI. You'll need an agent that calls tools (or any LLM app you can describe as request → reply), and about 15 minutes.

1. Create a project

Sign up at agentlasso.dev. Your workspace starts with a sample agent and data so every screen has something to show; your own traces appear alongside it.

In Settings:

  • Copy your project's API key. You'll send traces with it.
  • Write a sentence or two in Agent description about what your agent does. Classification uses it to name capabilities and judge edge cases in your agent's real domain.

2. Send your first trace

POST one request your agent handled. No SDK, no OpenTelemetry envelope needed:

sh
curl -X POST "https://agentlasso.dev/api/v1/telemetry" \  -H "x-api-key: <your-project-api-key>" \  -H "Content-Type: application/json" \  -d '{    "user_input": "Can you cancel my subscription?",    "agent_output": "Done — cancelled effective end of your billing period.",    "tool_calls": [{ "name": "stripe.subscription.cancel", "arguments": { "subscription_id": "sub_123" } }]  }'
  • user_input: what the user asked.
  • agent_output: the agent's final reply.
  • tool_calls: the tools it called, in order, with their arguments. This is what trajectory checks and coverage are measured on, so include it whenever your agent uses tools.

The trace shows up under Production Sync right away. Emails, phone numbers, SSNs and card numbers are redacted before storage (you can turn this off per project in Settings).

Already using OpenTelemetry? Point your exporter at the same endpoint instead:

sh
export OTEL_EXPORTER_OTLP_TRACES_ENDPOINT="https://agentlasso.dev/api/v1/telemetry"export OTEL_EXPORTER_OTLP_HEADERS="x-api-key=<your-project-api-key>"

To send real traffic, add the POST to wherever your agent finishes a request, without blocking the response on it. All fields and response codes: Send traces. If you already send traces somewhere else, see Send traces you already have.

3. See what your agent does

Open the Semantic Agent Map. Each trace is classified into a capability:

  • Free, by tool matching: a trace whose tool calls match a registered capability's tools is placed there immediately.
  • By Claude: traces the rules can't place are classified by Claude. This needs an Anthropic API key on the server; until then they wait, and the map says how many.

New traffic the rules can't place shows up under Unmapped log volume, where you can register it as a new capability.

Click a capability to see its scorecard: usage, success, risk (with the reasons), business impact, eval coverage, and the recommended next step. More in Capabilities and classification and The capability scorecard.

4. Turn real behavior into a golden test

  1. Open Dataset Studio → Review Queue. High-value production edge cases wait here. Approve one to add it to a golden dataset.
  2. In Dataset Studio, find the new test case and click Add checks. The rubric editor opens.
  3. Add checks, or start from a template:
    • Trajectory: the tools the original trace called, as one-click suggestions ("Calls stripe.refund", "In order: orders.lookup → stripe.refund").
    • Q&A answer: facts the reply must mention.
    • Forbidden content: phrases and personal data the reply must never contain.
    • Refusal / red team: tools that must never be called.
    • AI judge: plain-English criteria graded by Claude (needs your Anthropic key in Settings → AI judge).
  4. Under Try it before saving, grade the draft against the original production trace (or a pasted reply, or your agent's live answer if you've connected it). You'll see Would pass or Would fail and why, check by check.
  5. Save rubric.

A test case without checks isn't graded: it verifies nothing. The test case card says so. Every check type and grading rule: Checks reference.

5. Check your coverage

Open Coverage. Agent Quality Coverage shows how much of your agent's real behavior your golden tests check, and the drill-downs list exactly what's missing:

  • the capabilities and money flowing through untested tool paths,
  • the busiest untested tool paths, with a real example request each,
  • known failure modes from production without a test,
  • capabilities with no tested path at all.

Close the biggest gap first. The Top priorities panel on the Semantic Agent Map ranks them for you. How each part is measured: Agent Quality Coverage.

6. Run it in CI

Every dataset card in Dataset Studio shows its dataset ID and a Run in CI snippet. Your pipeline:

  1. creates an eval run for the dataset and receives its test cases,
  2. sends each test case's input to your agent,
  3. reports back what the agent said and which tools it called,
  4. gets a graded result, and fails the build if a critical check failed.

Run golden tests in CI has a ready-to-adapt GitHub Actions example; the API reference has every request and response, and Read test results explains the Test Runs page.

Next steps

  • Keep traffic flowing: coverage and the scorecard are only as current as your traces.
  • Add a capability's risk and money field in its Semantic Agent Map deep dive, so coverage weighs what matters most.
  • Prefer to run it yourself? See Self-hosting.