Work with your capability map
Your capability map is the list of things your agent does in production, like "process a refund" or "track a shipment". Everything else in AgentLasso is built on it: the scorecard, coverage, risk and the golden tests you prioritize. This guide shows how the four pages under Semantic Agent Map work together to keep that map complete, stable and correct.
Four ideas to know first
- Capability: one thing your agent does for users, with a name (like
process_refund), the tools it uses, and settings such as risk and business value. You decide which capabilities exist. - Assignment: every production trace is placed in exactly one capability, or left unassigned. The free rules place a trace when its tool calls clearly point to one capability. Otherwise the AI picks from your list. The AI never creates capabilities.
- Unassigned: a trace that fits none of your capabilities. It keeps the name the AI proposed for it and waits in Discovery for you to decide.
- Alias: when you merge capability A into B, A's name becomes another name for B. Anything that still uses A's name, like an old trace or an AI answer, counts toward B.
Why this matters: if the AI could invent names, the same behavior could be called "refund_request" one day and "process_refund" the next, splitting its history and its numbers in two. Keeping assignment to your list, and naming new capabilities by hand, keeps the map stable.
How the pages fit together
Semantic Agent Map
What it shows: one row per capability, with its scorecard: usage over 30 days, success rate, risk (suggested, or set by your team), business impact, money at stake, eval coverage and a priority with the recommended next step. The top of the page lists your top priorities.
Use it to:
- Decide what to test next. Start with the top priorities. They weigh traffic, quality, risk and missing coverage.
- Set what the numbers can't know:
- override a suggested risk level, giving a reason;
- set the business value per outcome;
- pick the tool field that carries money.
- See how much is waiting: the Discovery button shows how many proposed capabilities and unassigned traces are waiting.
More: The capability scorecard.
Discovery
What it shows: unassigned traffic grouped by the set of tools it called. Tools don't change between runs the way an AI's wording can. Traces that called no tools are grouped by the AI's proposed name. Each group shows:
- its size and 30-day trend;
- the tools, and example requests;
- the names the AI proposed, with counts. Mixed names can mean the group is really two behaviors;
- how many traces are flagged for review;
- the closest existing capability. If the proposed name already belongs to a capability, the card says so.
For each group, choose one:
- Accept as new capability: when this is a real, distinct thing users need.
- Give it a clear name, such as
exchange_itemorredeem_loyalty_points. Names are saved in lowercase with underscores. - Optionally add a business goal.
- Keep the tools that define it checked. Future traces calling them are assigned automatically by the free rules, without an AI call.
- Give it a clear name, such as
- Add to existing: when it's a variation of something you already have. The closest capability is preselected. You can also add the group's tools to it, so future traces like these are placed automatically.
- Ignore: test traffic, noise or one-offs. You can Undo right away, or find it later under Show ignored.
What happens when you accept or add:
- All of the group's past traces move into the capability.
- Its scorecard history is filled in on the days they happened, so nothing starts from zero and nothing is counted twice.
- The new capability appears on the map.
Tips:
- Busy and growing groups first: the list is sorted by the last 30 days.
- If a group mixes clearly different requests, accept what most of it is, and check later on Validation whether it should be split.
- A group with flagged traces may be a gap where your agent struggles. Look at them in the Review Queue before deciding.
Assignment
What it shows (last 30 days):
- How traces were assigned: by the rules (tools), by AI, by a person (Discovery or a merge), or not at all (waiting in Discovery). Shown overall and per capability.
- Traces from before this was tracked show as "not recorded".
- Rules vs AI: about 10% of the traces the rules place are also sent to the AI as a check. This shows how often they agree, and the latest disagreements.
- This needs an Anthropic API key on the server; without one, the section says so.
- Capabilities: each with its traffic, how its traces were assigned, its rules-vs-AI agreement, its aliases, and Merge into…
Use it to:
- Keep an eye on "Unassigned". A rising share means new behavior is appearing: go to Discovery.
- Read the disagreements. When the rules and the AI disagree, two capabilities usually share tools. Fix it one of two ways:
- give each capability its own distinctive tools;
- if they're really the same behavior, merge them.
- Merge duplicates: Merge into…, pick the capability to keep, confirm. What happens:
- The merged capability's traces, scorecard history, tools and test-case links move to the one you keep.
- Its name becomes an alias.
- Its own settings (risk, business value, money field, goal) are not kept. Note anything you want to keep first.
- A merge can't be undone.
Validation
What it shows: how well the map matches what people think, measured from your answers to one question about two real traces: "Do these two need the same capability from the agent?"
How to label:
- Click Start labeling. Pairs appear side by side: each one's request, reply and tools.
- Answer Same (S), Different (D) or Can't tell (U). K skips. Add an optional note on why.
- You then see what the system thinks and why the pair was picked. Press Enter for the next pair.
What the system thinks stays hidden until you answer, so it can't sway you. Answer based on what the agent has to do, not how the request is worded: "where's my order?" and "has my package shipped?" are the same; "where's my order?" and "change my address" are not.
Which pairs you get are the places the map is most likely wrong:
- Inside one capability: its two most different traces. Checks that the capability is one behavior.
- Across capabilities: capabilities that share tools, or where the rules and AI disagreed. Checks that they're really different.
- Unassigned vs closest: a Discovery trace and its closest capability. Checks whether "Add to existing" would be right.
Reading the results:
- Agreement: how often the assignment matched your answers, overall and per kind of pair, with the number of pairs.
- Under 20 answered pairs, it's marked as a first look.
- "Can't tell" answers are kept but not counted.
- Suggestions need at least 3 answers:
- May be two behaviors: at least half the pairs inside a capability were judged different. Look at its traces and give each behavior its own capability and tools.
- May be one behavior: at least 75% of the pairs across two capabilities were judged the same. Merge them on Assignment.
Answers are never edited. If you answer a pair again, the newest answer counts. Each answer keeps a snapshot, so later merges don't rewrite past results.
Routines
When you connect a new agent
- Send traces: Send traces.
- Open Discovery. Your first traces show up here as proposed capabilities. Accept the real ones, with their tools.
- Check the Semantic Agent Map: set risk and business value for the capabilities that matter.
- Label about 20 pairs on Validation to get a first agreement number.
Every week (about 10 minutes)
- Discovery: handle new groups, busiest first.
- Assignment: check the unassigned share and any new rules-vs-AI disagreements.
- Validation: label 10–20 pairs. Act on any suggestion.
- Semantic Agent Map: work on the top priorities: turn uncovered behavior into golden tests.
After a prompt, model or tool change
- Watch Assignment for a jump in unassigned traffic or disagreements: the agent may now behave differently.
- Label a few pairs on Validation for the capabilities you changed.
When two capabilities seem to overlap
- Check Assignment for disagreements between them.
- Label pairs across them on Validation.
- If people say they're the same, merge; if different, give each its own tools.
When one capability seems to be two behaviors
- Validation says "may be two behaviors", or its traces look mixed on the map.
- Note which requests belong together, and which tools each behavior uses.
- Splitting a capability isn't available yet: its existing traces can't be moved into a new capability from the app. For now, keep the suggestion in mind when you read that capability's numbers.
Questions
- Why is a trace unassigned when one of my capabilities obviously fits? The capability may have no tools set, so the rules can't place it, and the AI didn't judge it a clear match. Add it to that capability from Discovery, and include its tools.
- Can the AI create a capability by itself? No. It can only propose a name. A person decides in Discovery.
- Where did my merged capability go? It's now an alias of the one you kept, listed under that capability on Assignment.
- The Rules vs AI section is empty. It needs
ANTHROPIC_API_KEYon the server and new traffic. - Can I undo a merge? No. Accept and add can't be undone from the page either. Ignore can be.
Next
See how each capability is scored: The capability scorecard.