API reference

Base URL: https://agentlasso.dev/api/v1 (self-hosted: your own origin + /api/v1).

Authentication

Every request carries your project's API key (from Settings), as either header:

x-api-key: <your-project-api-key>Authorization: Bearer <your-project-api-key>

Missing or invalid key: 401 with { "error": "…" }. All errors are JSON with an error message.

POST /telemetry

Send a production trace, as flat JSON or OpenTelemetry (OTLP/HTTP JSON), auto-detected. Fields, response and status codes: Send traces.

POST /eval-runs

Start a run of a golden dataset. You get back the test cases to send through your agent.

Request

json
{ "dataset_id": "41111111-0000-0000-0000-000000000001", "pass_threshold": 0.9 }
FieldRequiredMeaning
dataset_idyesThe golden dataset. Its ID is on its card in Dataset Studio.
pass_thresholdno0–1. The run passes when this share of graded test cases passes. Default 1.0.

Response 201

json
{  "run_id": "…",  "pass_threshold": 0.9,  "test_cases": [    {      "test_case_id": "…",      "input": "I want my money back for order 44182",      "evaluation_criteria_json": { "checks_version": 1, "checks": [ … ] },      "checks": [        { "id": "c1", "type": "tool_called", "tool": "stripe.refund", "severity": "critical" }      ]    }  ]}

checks is the normalized list of what will be graded (see Checks reference). Use it to tell which test cases need the agent's reply and which only need its tool calls.

StatusMeaning
201Run created.
400Missing dataset_id, or an invalid pass_threshold.
404No dataset with that ID in this project.

POST /eval-run-results

Report what your agent did for each test case. AgentLasso grades every check and returns the result. Results can be reported once per run.

Request

json
{  "run_id": "…",  "results": [    {      "test_case_id": "…",      "actual_output": "Your refund of $89 is on its way.",      "actual_tool_calls": [        { "name": "orders.lookup", "arguments": { "order_id": "44182" } },        { "name": "stripe.refund" }      ]    }  ]}
FieldRequiredMeaning
run_idyesFrom POST /eval-runs.
results[].test_case_idyesMust belong to the run's dataset.
results[].actual_outputfor content checksThe agent's final reply (string). Content checks fail without it, saying why.
results[].actual_tool_callsfor tool checks[{ name, arguments? }], in the order called.

Response 200

json
{  "run_id": "…",  "pass_rate": 0.92,  "pass_threshold": 0.9,  "ci_pass": true,  "graded_count": 12,  "ungraded_count": 1,  "results": [    {      "test_case_id": "…",      "graded": true,      "score": 100,      "pass_fail": true,      "judge_reasoning": "2/2 checks passed.",      "check_results": [        {          "check_id": "c1",          "type": "tool_called",          "label": "Calls stripe.refund",          "severity": "critical",          "passed": true,          "detail": "stripe.refund was called."        }      ]    }  ]}
  • ci_pass is what your pipeline should act on: exit non-zero when it's false.
  • pass_rate is over graded test cases only. It's null, and ci_pass is false, when nothing in the run had checks; ci_fail_reason then says so.
  • judge_reasoning is a one-line summary of the checks. Despite the name, it doesn't involve an AI judge: AI judge checks aren't graded in CI yet and come back as not graded (see AI judge checks).
StatusMeaning
200Graded.
400Invalid body, or a test_case_id that isn't part of the run's dataset (they're listed).
404No run with that ID in this project.
409Results were already reported for this run.

A complete example script and GitHub Actions workflow: Run golden tests in CI.

POST /sync-cluster (self-hosting)

The background classification job. A scheduler calls it with a shared secret (Authorization: Bearer <LOVABLE_CRON_SECRET>), not a project key. See Self-hosting.