The capability scorecard
The Semantic Agent Map gives every capability five numbers and one recommended action, in terms a product manager and an engineer can both act on:

Click a capability for its deep dive, which shows how every number was calculated.

Usage
Requests classified into the capability in the last 30 days, with the change against the 30 days before.
Success
The share of the capability's traffic not flagged as a failure or edge case during classification, over the last 30 days, with the change in percentage points against the previous 30 days.
It's a proxy for task success, not a graded outcome: it can only be as good as the failure signals classification sees. Below 90% the capability is marked as drifting.
Risk
What the capability can do if it gets things wrong. It's suggested from the tools it calls and from what production shows:
- High: it calls tools that move money or make irreversible changes (refund, charge, payment, transfer, cancel, delete, …).
- Medium: it calls other tools that change data (create, update, send, …).
- Low: it only reads (lookup, get, search, track, …), or calls no tools.
- Adversarial traffic (e.g. prompt-injection attempts) raises the level by one. Escalations to a human are listed as a signal.
Every level comes with its reasons. Your team can override it in the deep dive (Risk → Change) with a required reason, e.g. "Regulated: refunds above $500 need a compliance review". The suggestion stays visible next to the override, and Use the suggestion removes it.
Business impact
Value. In the deep dive's Business value, say what one successful outcome is worth:
- Deflection: each success is a case no human had to handle. Enter the cost of a human-handled case.
- Revenue: each success is a revenue event, e.g. a recovered cart. Enter the revenue per outcome.
- No $ outcome: a quality or safety capability, such as a human handoff.
Value over 30 days = usage × success × your $ per outcome. The $ per outcome is your assumption, and it's shown as one.
Money at stake. How much money actually flows through the capability, measured from the amounts in its tool calls:
- AgentLasso finds money-looking numeric arguments in the capability's tool calls (e.g.
stripe.refund → amount), with how many traces use each and a typical value. IDs, counts and percentages are skipped. - You confirm the field and its unit. Many payment APIs, Stripe included, send cents (
8900= $89.00), so the wrong unit is a 100× error. The preview shows what the typical value means in each unit. - AgentLasso estimates: the average amount per request in recent traces × 30-day usage. It also reports how much of it flows through tool paths no golden test checks.
Eval coverage
The share of the capability's traced traffic whose tool path a golden test checks. The deep dive lists every path, with its share of traffic, whether it's tested, and a real example request for each untested one. See Agent Quality Coverage for exactly how a path counts as tested.
Priority
One recommended action per capability, from transparent rules. The first one that matches wins, and each comes with a one-line reason:
Top priorities (above the table) ranks the capabilities that need action: roughly traffic × risk weight (high 3, medium 2, low 1) × the untested share, boosted by quality drops. It puts what's busy, risky and untested first.
Next
Measure your whole agent, not one capability at a time: Agent Quality Coverage.