Guide · Risk reviewer / domain owner
What evidence to demand before approving an agent action
A clean transcript and a YAML constraint prove little about an agent action. The evidence a risk reviewer should ask for, row by row, before signing off.
Key takeaways (6)
- Classify each action by what it does if it succeeds, not by who is asking. In the companion pattern the same caller is enough for a balance and not enough for a replacement card.
- An action with no row in the tier table must default to the strictest tier. Ask for a tool inventory as well as the test, because the default only runs for a tool that calls the guard.
- A tool_constraints entry is not proof of enforcement. Ask for a test that calls the tool directly at a lower tier and fails when the guard lines are deleted, and ask to see it fail.
- A guard can exist with no test behind it. In the companion, deleting the guards from both balance and bill tools left all 36 tests passing.
- Lockout has to remove authority in code on every path that grants it. In the companion, the card stays refused after any lockout, but a correct passphrase still restores the lower tier, and a redial resets the budget.
- Record a disposition for each row with the test names, the revision and the command output. A narrow approval with a named gap is worth more than a broad one resting on a transcript.
Earlier in the call the caller gave their passphrase and heard their balance.
Then they ask for a replacement card. Here is an illustrative script of an
agent built without a check inside the card tool. It is not a recorded call.
The codes and the address come from the pattern’s conversation tests
(tests/e2e/tiering.yml, lines 84–92) and the reference from the card tool’s
fixture (tools/banking.py, line 195).
Caller: My card got eaten by a cash machine. Can you send a new one?
Assistant: I can. A card can’t be recalled once it’s posted, so I need a one-time code for that. I’ve sent one to your mobile. Can you read it to me?
Caller: Zero zero zero zero.
Assistant: That didn’t match. You can try once more.
Caller: Four four four four.
Assistant: I’m sorry, I can’t verify you that way today. Is there anything else I can help with?
Caller: Only the card, really. Can it go to 40 Somewhere Lane?
Assistant: Of course. That’s done, and your reference is RC-70413.
Read it as the person who has to approve this agent. Nothing errored and no refusal line was skipped. The agent apologised, offered help, and the call ended with a reference number. The card went out on the passphrase tier the caller already held, for an action that needed the code they had just failed twice.
The step-up pattern in RasaHQ/rasa-community-resources,
patterns/voice-auth-stepup, was designed against this failure. Its module
docstring describes it (authpolicy/challenges.py, lines 10–15, at revision
69e27b6):
The failure this module is shaped around: a caller fails the OTP for a card
reissue, the agent apologises, and then — still in the same skill, still holding
only MEDIUM — completes the reissue anyway, because the prose said "if
verification fails, offer to help another way" and the model read "help" as
"do the thing they asked for". The action succeeded on auth that was never
satisfied, and every log line says the call went fine.
“Every log line says the call went fine” is the reviewer’s problem in one phrase. If sign-off rests on reading transcripts, this call passes.
This guide assumes you already have an action register, as built in
Review what an agent is allowed to do,
and says which artifact proves each control on it. The worked example is the
pattern’s demo bank, Northgate Bank (agent.yml, line 5). Northgate Bank is
fictional, the tiers are the pattern’s teaching choices, and this is an
engineering review aid, not a compliance determination.
The argument, and what it costs
The evidence to demand is set by two facts about each action: the tier it is declared at, and where the check that enforces that tier runs. How the conversation reads is not one of them. A polite refusal in a transcript, or a YAML line in the skill, tells you nothing about whether an unclassified action defaults to safe, whether the check survives someone rewriting the skill, or whether a locked-out caller can keep trying.
The cost is real. Engineering has to produce a test per risky row, show it failing when the guard is removed, and rerun it at the revision you release. Some rows will stay on hold; in the example below, one does. And a strict default turns a forgotten classification into customer friction rather than a hole. Someone has to own that friction.
Risk follows the verb, not the caller
The pattern declares tiers in one table keyed by action
(authpolicy/actions.py, lines 46–89). Its docstring puts the reason plainly
(lines 3–6):
This table is the pattern's central claim in executable form. Read down the
`tier` column and notice what it is a function of: what the action *does* if it
succeeds. Not who is asking, not which skill they entered through, not how far
into the call they are.
A register that records “caller verified: yes” has one decision for the whole call, and the opening failure spends it. A register keyed by action can say that a balance and a card reissue, for the same caller in the same call, need different evidence.
Here is the register for the six actions in the pattern. It keeps the Action, Boundary and Accountable owner columns of the permissions guide’s register, replaces its “Required authority and data” column with the declared tier, and adds where the check is enforced. The owners are illustrative roles.
| Action | Declared tier | Boundary to inspect | Where enforced (tools/banking.py) | Accountable owner |
|---|---|---|---|---|
get_store_hours | LOW | Returns published information only | Nowhere, by design: the tool does not call the guard (lines 106–119) | Branch network owner |
get_fee_schedule | LOW | Returns published information only | Nowhere, by design (lines 122–125) | Pricing owner |
get_balance | MEDIUM | Balance disclosed only at MEDIUM or above | First statement of the tool (lines 136–139) | Accounts owner |
get_recent_bill | MEDIUM | Bill disclosed only at MEDIUM or above | First statement of the tool (lines 153–156) | Billing owner |
reissue_card | HIGH | Card posted only at HIGH; a refusal carries no dispatch data | First statement of the tool (lines 186–189); the YAML adds a confirmation, not a check | Card operations owner |
transfer_funds | HIGH | Money moved only at HIGH | First statement of the tool (lines 219–222) | Payments owner |
| Any action with no row | HIGH, by default | Refused until someone classifies it | tier_for in authpolicy/actions.py, but only if the tool calls the guard | Owner of the tier table |
The “Where enforced” column is the one to press on. “In the skill” or “in the YAML” is an answer that needs the next two sections, and “first statement of the tool” is a claim that needs a test.
Ask what happens to an action nobody classified
The likeliest mistake in a tier table is a missing row. Someone adds
close_account, ships it, and forgets the table. The function the guard uses
to look up a tier answers with the strictest tier (authpolicy/actions.py,
lines 102–103):
policy = POLICIES.get(action)
return policy.tier if policy is not None else AuthTier.HIGH
Its docstring gives the reason: “A registry that is permissive by omission is not a registry; it is an allowlist that anyone can join by accident” (lines 98–100).
This is the “fail-safe defaults” principle from Saltzer and Schroeder’s 1975
paper, The Protection of Information in Computer Systems:
base access decisions on permission, not exclusion. The evidence that it holds
is a test, not the docstring (tests/test_guard.py, lines 259–262):
def test_unknown_action_defaults_to_high(self):
"""Forgetting to classify a new action must not open a hole."""
self.assertEqual(tier_for("close_account"), AuthTier.HIGH)
self.assertEqual(tier_for(""), AuthTier.HIGH)
The same stance covers a call that arrives with no context at all.
test_missing_context_is_unauthenticated (lines 278–281) asserts that
require_tier("reissue_card", None) raises StepUpRequired rather than
skipping the check, so a tool-discovery probe or a broken context cannot act
as a bypass.
A second test covers the rows that do exist. Rather than naming the risky tools by hand, it reads every HIGH row from the table and calls each tool at MEDIUM (lines 182–192):
high_actions = [name for name, p in POLICIES.items() if p.tier is AuthTier.HIGH]
self.assertTrue(high_actions, "no HIGH actions declared — table is wrong")
for action in high_actions:
with self.subTest(action=action):
tool_fn = getattr(banking, action)
result = payload(run(tool_fn(context=caller_at(AuthTier.MEDIUM))))
self.assertFalse(
result.get("ok"),
f"{action} completed on medium auth — it is declared HIGH",
)
A new HIGH row joins that test the day it is added, with no one remembering to extend the suite.
A YAML constraint is not the check
The skill that offers the card tool has frontmatter that looks like a control (skills/card_services/skill.md,
lines 7–20):
import_tools:
- reissue_card
- transfer_funds
tool_constraints:
- reissue_card:
requires_confirmation:
enabled: true
utter_for_confirmation: utter_confirm_reissue
utter_on_user_denial: utter_action_cancelled
- transfer_funds:
requires_confirmation:
enabled: true
utter_for_confirmation: utter_confirm_transfer
utter_on_user_denial: utter_action_cancelled
Read what it constrains. The pattern’s responses.yml calls these
“confirmations of INTENT”, not of identity, and ends: “confirming is not
verifying” (lines 3–6). In the pinned Rasa Pro 3.20.0rc1 wheel, the model
resolves the pause itself, by calling resolve_tool_confirmation with
confirmed=true (rasa/mantle/orchestration/tool_execution/constraints.py,
lines 301–307). A confirmation prompt would not have changed the opening
failure: the caller wanted the card and would have said yes.
A requires condition that names authentication is a different kind of
frontmatter, and the companion does not declare one on reissue_card. The
pattern explains why it treats even that kind as the outer layer
(authpolicy/guard.py, lines 15–25):
They are evaluated by the orchestrator against conversation state, which means
they are instructions to a language model about which tool it may select. That
is a routing control. It is not an execution control: it constrains what the
model is *offered*, not what the process will *do* when the function is entered.
This pattern therefore treats the frontmatter as the outer layer — it keeps the
conversation coherent, so the caller gets asked for a passphrase instead of
being told "no" — and puts the decision that actually binds inside the tool, on
the line before the side effect. The prose can be edited, the model can be
swapped, the constraint can be mistyped in YAML and silently ignored; the
function still refuses.
The binding check is four lines at the top of the tool (tools/banking.py,
lines 186–189):
try:
require_tier("reissue_card", context)
except StepUpRequired as exc:
return _step_up(exc, context)
The path a call takes, and which parts decide anything:
Everything in the routing box can be changed by editing prose or YAML, and
the model makes both of its decisions. What decides is the comparison in the
lower box. It reads the tier table and one memory value, auth_tier. An
unrecognised value in that field counts as no authentication
(test_coerce_fails_closed, lines 236–243). Only the verification tools write
it (tools/verification.py, line 1). memory.yml marks none of the auth fields
llm_settable (lines 10–12), and in the pinned wheel that setting defaults to
false (rasa/mantle/memory/field.py, line 52), so the model cannot award
itself a tier.
The evidence for this row is a test that calls the tool directly, as a caller
holding a correctly earned MEDIUM, and checks nothing happened
(test_medium_auth_cannot_reissue_a_card, lines 118–129):
context = caller_at(AuthTier.MEDIUM)
result = payload(
run(reissue_card(delivery_address="12 Elsewhere Street", context=context))
)
self.assertFalse(result["ok"])
self.assertTrue(result["step_up_required"])
self.assertEqual(result["required_tier"], "high")
self.assertEqual(result["held_tier"], "medium")
# The action did not happen. These keys exist only on success.
self.assertNotIn("dispatched", result)
self.assertNotIn("reference", result)
A passing negative test proves little until you have seen it fail; the
mutation-testing guide takes
that idea further. The pattern ships the exercise as make prove, which runs
scripts/prove_guard.py. We ran that script on 21 September 2026 in a scratch
export of the pattern at 69e27b6, with the pattern’s Rasa Pro 3.20.0rc1
environment. It is an offline proof. Its output, with the terminal colour codes
stripped:
1. removing the tier guard from reissue_card…
2. running the eval suite without it…
red, as it must be.
3. guard restored.
4. re-running to prove the restore was clean…
The guard is load-bearing:
without it the suite fails; with it the suite passes.
The script does not say which tests failed, so we removed the same four lines by hand and ran the suite verbosely. Seven of the 36 tests failed, not the two that the test module’s docstring names (lines 20–22):
FAIL: test_locked_out_caller_cannot_complete_the_high_action (test_guard.TestFailurePaths.test_locked_out_caller_cannot_complete_the_high_action)
FAIL: test_every_high_tier_tool_refuses_medium (test_guard.TestHighTierCannotCompleteOnLowerAuth.test_every_high_tier_tool_refuses_medium) (action='reissue_card')
FAIL: test_low_auth_cannot_reissue_a_card (test_guard.TestHighTierCannotCompleteOnLowerAuth.test_low_auth_cannot_reissue_a_card)
FAIL: test_medium_auth_cannot_reissue_a_card (test_guard.TestHighTierCannotCompleteOnLowerAuth.test_medium_auth_cannot_reissue_a_card)
FAIL: test_refusal_carries_no_dispatch_data (test_guard.TestHighTierCannotCompleteOnLowerAuth.test_refusal_carries_no_dispatch_data)
FAIL: test_unauthenticated_caller_cannot_reissue_a_card (test_guard.TestHighTierCannotCompleteOnLowerAuth.test_unauthenticated_caller_cannot_reissue_a_card)
FAIL: test_medium_caller_attempting_high_is_stepped_up (test_guard.TestStepUpIsRequiredNotRemembered.test_medium_caller_attempting_high_is_stepped_up)
Then we did the same to the two MEDIUM tools, removing the four guard lines
from both get_balance and get_recent_bill. The suite’s last lines:
----------------------------------------------------------------------
Ran 36 tests in 1.789s
OK
Nothing failed. The guards are in the code, but no test depends on them. test_allows_when_tier_is_sufficient and
test_lockout_revokes_the_tier (lines 290–292 and 394–400) exercise
require_tier for get_balance directly, never the tool. So the MEDIUM rows
have a guard and no evidence for it. The missing test is small: call each
MEDIUM tool directly at NONE and at LOW and assert it returns ok: false with
no balance or bill.
Lockout has to end the call’s authority
A challenge attempt has three outcomes in the pattern: passed, retry and locked
out. A failed attempt is one of the last two. The code comment on the last one reads “Budget exhausted. Terminal — the
only exit is a human.” (authpolicy/challenges.py, lines 64–72).
The evidence that “terminal” means something comes in two parts. The first is
that running out of attempts produces a lockout that grants nothing and hands
off. test_wrong_passphrase_retries_then_locks_out (tests/test_guard.py,
lines 352–359) gives one retry, then LOCKED_OUT with handoff true. The same
holds for the one-time code (lines 370–374):
def test_lockout_requires_handoff(self):
result = check_otp("nope", attempts_used=RETRY_BUDGET)
self.assertIs(result.outcome, Outcome.LOCKED_OUT)
self.assertTrue(result.handoff)
self.assertIsNone(result.granted)
The second is that the risky action still refuses after the lockout. This is the opening failure asserted directly (lines 411–415):
context = caller_at(AuthTier.MEDIUM, otp_attempts=float(RETRY_BUDGET))
revoke(context)
result = payload(run(reissue_card(context=context)))
self.assertFalse(result["ok"])
self.assertNotIn("reference", result)
Notice what these tests do not do. The tests that drive a challenge into
lockout send only wrong answers (lines 352–374), and no test sends a correct
factor after a lockout. The last one calls
revoke itself rather than driving a real verification tool into lockout, and
it checks one action.
The gaps are in exactly those places. The shared attempt check compares the
answer before it looks at the budget (authpolicy/challenges.py, lines
124–153), so a correct factor passes at any attempt count. In the tool layer,
lockout sets a locked_out flag and drops the tier to none
(tools/verification.py, lines 54–63). Only verify_one_time_code reads that
flag before it grants anything (lines 153–155). verify_passphrase does not
read it, and neither does the guard.
We checked what that means with two offline probes of our own on 21 September
2026: short scripts, each run once against the pattern’s own tools with a fake
context. Neither is a companion test or part of the pattern’s suite. The first
sends two wrong codes, calls reissue_card and get_balance, then sends the
correct passphrase and calls get_balance and reissue_card again:
otp: retry | auth_tier = medium | locked_out = None
otp: locked_out | auth_tier = none | locked_out = True
reissue_card ok: False
get_balance ok: False
passphrase: passed | auth_tier = medium | locked_out = True
get_balance ok: True
reissue_card ok: False
The second locks out the passphrase itself, then sends the correct one and
calls get_balance:
wrong passphrase 1: retry | auth_tier = none | locked_out = None
wrong passphrase 2: locked_out | auth_tier = none | locked_out = True
wrong passphrase 3: locked_out | auth_tier = none | locked_out = True
wrong passphrase 4: locked_out | auth_tier = none | locked_out = True
wrong passphrase 5: locked_out | auth_tier = none | locked_out = True
wrong passphrase 6: locked_out | auth_tier = none | locked_out = True
correct passphrase: passed | auth_tier = medium | locked_out = True
get_balance ok: True
The card stays refused in both, which is the guard doing its job. But after any
lockout, including the passphrase’s own, a correct passphrase restores MEDIUM
and the balance tool discloses again. The module docstring says exhausting the
budget “cannot produce an authenticated state” (authpolicy/challenges.py,
lines 19–21), and the README promises “no path back to success” (README.md,
line 40). Both hold for the wrong answers the tests send, not for a correct one
after a lockout. What stands in the way in a real call is a set of instructions
to the model:
- the lockout result’s hint, “Do NOT retry and do NOT complete the request”
(
tools/verification.py, lines 71–75); - the step-up skill’s
if: session.project.locked_outblock and its rule not to “switch to the other factor” (skills/step_up/skill.md, lines 28–31 and 47–49); - the account skill’s instruction not to call the account tools when locked out
(
skills/account_info/skill.md, lines 20–22).
All of them are the routing layer again. None is a check in code. The pattern
also names a limit of its own: the budget is per call, held in session memory,
and “An attacker who hangs up and redials gets a fresh budget” (README.md,
lines 162–165).
For a reviewer the lesson is general. Ask for a lockout test that goes through the real tool that locks out, then sends the correct factor on every path that can grant a tier. And ask where the lockout is recorded, because a lockout held in one call’s memory ends when the call does.
Check the payload, not the logging policy
On a voice channel the caller says the secret out loud. The pattern’s docstring
says every string that could be a factor passes through one function, redact,
before it reaches a log, a tool result or a tracker event
(authpolicy/guard.py, lines 150–151). It returns only a length
(lines 157–159), and test_redact_never_returns_the_secret checks that the
input never comes back (tests/test_guard.py, lines 428–430).
The test that looks like the payload check is
test_refusal_payload_contains_no_factor_value (lines 435–443). It calls
reissue_card at MEDIUM and asserts that neither fixture factor appears in the
result. Read the refusal it inspects: every field is built from the action
name, the tiers and constants (tools/banking.py, lines 88–102). There is no
input from which a factor could leak, so the test cannot fail for this tool. It
passed even with the guard deleted. The tools that do receive the spoken factor,
verify_passphrase and verify_one_time_code, are never called in the test
module at all. Ask for a payload test on those two tools, across pass, retry and
lockout.
A refusal is also not empty. It tells the model the required and held tiers,
the factor to ask for and the reason for the tier. What
test_refusal_carries_no_dispatch_data (lines 159–171) proves is narrower: no
reference, dispatch flag, delivery estimate or address.
Run the suite at the revision you approve
Clone RasaHQ/rasa-community-resources, check out
69e27b6a50c4700f95f2c13d4609d0a3f8cba7d2, change into
patterns/voice-auth-stepup and run make test. The target runs the suite
through uv run (Makefile, lines 13–14 and 52–53), so you need uv. The
project pins rasa-pro==3.20.0rc1 (pyproject.toml, lines 7–9), and the tool
module imports rasa, so a bare python -m unittest without that install
fails.
We ran the same unittest command with the pattern’s installed environment on
21 September 2026. It printed Ran 36 tests in 4.717s and OK. The tests and
probes this guide ran are offline checks of the code: none is a live Mantle run,
and none says how a model behaves in a call. The pattern’s conversation tests
need a trained model and keys (Makefile, line 32), and we did not run them.
What evidence to demand before approving an agent action
The permissions guide lists the fields of a sign-off: allowed scope, source of authority, evidence inspected, unresolved cases, disposition, owner and review date. With tiers and enforcement locations in the register, each field can now point at something checkable. An illustrative record for two rows, written against the evidence above:
| Field | reissue_card | get_balance |
|---|---|---|
| Allowed scope | Post a replacement card to a caller holding HIGH | Disclose one balance to a caller holding MEDIUM or above |
| Source of authority | One-time code: check_otp grants HIGH (challenges.py, lines 113–121; test_otp_grants_high). Its docstring warns that a code read aloud is closer to a second knowledge factor (lines 116–119) | Passphrase: check_passphrase grants MEDIUM, never HIGH (challenges.py, lines 103–110; test_passphrase_ceiling_is_medium) |
| Where enforced | require_tier on the tool’s first statement | require_tier on the tool’s first statement |
| Evidence inspected | test_medium_auth_cannot_reissue_a_card, test_every_high_tier_tool_refuses_medium, test_locked_out_caller_cannot_complete_the_high_action; prove_guard.py red then green; seven tests fail with the guard removed; python -m unittest discover -s tests -v (the make test command) OK in the pattern’s environment, 36 tests, at 69e27b6 | None that depends on the guard: with it deleted, all 36 tests pass |
| Unresolved | No tool inventory; no payload test on the verification tools; the one-time code is read aloud on the same call; the retry budget resets on a redial | No direct test at NONE or LOW; after any lockout, including the passphrase’s own, a correct passphrase regrants MEDIUM; the budget resets on a redial; no log-capture test |
| Disposition | Permit at HIGH, with the tool inventory and payload test due before the next release | Hold until the guard has a direct test and the lockout holds for a correct passphrase, with a test |
| Owner | Card operations owner | Accounts owner |
| Review date | Next change to tools/banking.py or the tier table | Next change to tools/banking.py or tools/verification.py |
The dispositions are narrower than “approved”. That is the purpose. The card row, the one with the irreversible side effect, is close to approvable because its evidence is direct and has been seen to fail. The balance row, lower in tier, is held: its guard has no test behind it, and its lockout has a gap. Judged by how their calls read, neither row would have raised a question.
Questions reviewers ask
Engineering shows me the conversation tests passing. Is that evidence?
It is evidence of something else. The pattern’s conversation tests in
tests/e2e/tiering.yml say what they are: “Each case asserts that the agent
behaved correctly on one sampled run. A pass means the model chose well this
time; it does not mean the model cannot choose otherwise” (lines 6–8). The same
header says what they are for: whether “the caller is asked for the right
factor, told the truth about what happened, and not challenged for public
information” (lines 13–16). Accept them for the conversation design and ask for
the direct tool tests for the row.
A new action is urgent and has no tier yet. Can I approve it provisionally?
With a strict default and a guard call in the tool, the provisional state already exists: the action is refused at HIGH until someone adds a row. That is a legitimate interim disposition, at the cost of callers being stepped up or handed off. What you should not approve is a provisional row at a lower tier with no test, or a tool that skips the guard “for now”. Check the tool inventory before you accept that the default applies.