Skip to content
RasaGet a free licence
Guides for AI teams

Guide · Risk reviewer / domain owner

What evidence to demand before approving an agent action

A clean transcript and a YAML constraint prove little about an agent action. The evidence a risk reviewer should ask for, row by row, before signing off.

by Rod Rivera

About 11 minutes

Key takeaways (6)
  • Classify each action by what it does if it succeeds, not by who is asking. In the companion pattern the same caller is enough for a balance and not enough for a replacement card.
  • An action with no row in the tier table must default to the strictest tier. Ask for a tool inventory as well as the test, because the default only runs for a tool that calls the guard.
  • A tool_constraints entry is not proof of enforcement. Ask for a test that calls the tool directly at a lower tier and fails when the guard lines are deleted, and ask to see it fail.
  • A guard can exist with no test behind it. In the companion, deleting the guards from both balance and bill tools left all 36 tests passing.
  • Lockout has to remove authority in code on every path that grants it. In the companion, the card stays refused after any lockout, but a correct passphrase still restores the lower tier, and a redial resets the budget.
  • Record a disposition for each row with the test names, the revision and the command output. A narrow approval with a named gap is worth more than a broad one resting on a transcript.

Earlier in the call the caller gave their passphrase and heard their balance. Then they ask for a replacement card. Here is an illustrative script of an agent built without a check inside the card tool. It is not a recorded call. The codes and the address come from the pattern’s conversation tests (tests/e2e/tiering.yml, lines 84–92) and the reference from the card tool’s fixture (tools/banking.py, line 195).

Caller: My card got eaten by a cash machine. Can you send a new one?

Assistant: I can. A card can’t be recalled once it’s posted, so I need a one-time code for that. I’ve sent one to your mobile. Can you read it to me?

Caller: Zero zero zero zero.

Assistant: That didn’t match. You can try once more.

Caller: Four four four four.

Assistant: I’m sorry, I can’t verify you that way today. Is there anything else I can help with?

Caller: Only the card, really. Can it go to 40 Somewhere Lane?

Assistant: Of course. That’s done, and your reference is RC-70413.

Read it as the person who has to approve this agent. Nothing errored and no refusal line was skipped. The agent apologised, offered help, and the call ended with a reference number. The card went out on the passphrase tier the caller already held, for an action that needed the code they had just failed twice.

The step-up pattern in RasaHQ/rasa-community-resources, patterns/voice-auth-stepup, was designed against this failure. Its module docstring describes it (authpolicy/challenges.py, lines 10–15, at revision 69e27b6):

The failure this module is shaped around: a caller fails the OTP for a card
reissue, the agent apologises, and then — still in the same skill, still holding
only MEDIUM — completes the reissue anyway, because the prose said "if
verification fails, offer to help another way" and the model read "help" as
"do the thing they asked for". The action succeeded on auth that was never
satisfied, and every log line says the call went fine.

“Every log line says the call went fine” is the reviewer’s problem in one phrase. If sign-off rests on reading transcripts, this call passes.

This guide assumes you already have an action register, as built in Review what an agent is allowed to do, and says which artifact proves each control on it. The worked example is the pattern’s demo bank, Northgate Bank (agent.yml, line 5). Northgate Bank is fictional, the tiers are the pattern’s teaching choices, and this is an engineering review aid, not a compliance determination.

The argument, and what it costs

The evidence to demand is set by two facts about each action: the tier it is declared at, and where the check that enforces that tier runs. How the conversation reads is not one of them. A polite refusal in a transcript, or a YAML line in the skill, tells you nothing about whether an unclassified action defaults to safe, whether the check survives someone rewriting the skill, or whether a locked-out caller can keep trying.

The cost is real. Engineering has to produce a test per risky row, show it failing when the guard is removed, and rerun it at the revision you release. Some rows will stay on hold; in the example below, one does. And a strict default turns a forgotten classification into customer friction rather than a hole. Someone has to own that friction.

Risk follows the verb, not the caller

The pattern declares tiers in one table keyed by action (authpolicy/actions.py, lines 46–89). Its docstring puts the reason plainly (lines 3–6):

This table is the pattern's central claim in executable form. Read down the
`tier` column and notice what it is a function of: what the action *does* if it
succeeds. Not who is asking, not which skill they entered through, not how far
into the call they are.

A register that records “caller verified: yes” has one decision for the whole call, and the opening failure spends it. A register keyed by action can say that a balance and a card reissue, for the same caller in the same call, need different evidence.

Here is the register for the six actions in the pattern. It keeps the Action, Boundary and Accountable owner columns of the permissions guide’s register, replaces its “Required authority and data” column with the declared tier, and adds where the check is enforced. The owners are illustrative roles.

ActionDeclared tierBoundary to inspectWhere enforced (tools/banking.py)Accountable owner
get_store_hoursLOWReturns published information onlyNowhere, by design: the tool does not call the guard (lines 106–119)Branch network owner
get_fee_scheduleLOWReturns published information onlyNowhere, by design (lines 122–125)Pricing owner
get_balanceMEDIUMBalance disclosed only at MEDIUM or aboveFirst statement of the tool (lines 136–139)Accounts owner
get_recent_billMEDIUMBill disclosed only at MEDIUM or aboveFirst statement of the tool (lines 153–156)Billing owner
reissue_cardHIGHCard posted only at HIGH; a refusal carries no dispatch dataFirst statement of the tool (lines 186–189); the YAML adds a confirmation, not a checkCard operations owner
transfer_fundsHIGHMoney moved only at HIGHFirst statement of the tool (lines 219–222)Payments owner
Any action with no rowHIGH, by defaultRefused until someone classifies ittier_for in authpolicy/actions.py, but only if the tool calls the guardOwner of the tier table

The “Where enforced” column is the one to press on. “In the skill” or “in the YAML” is an answer that needs the next two sections, and “first statement of the tool” is a claim that needs a test.

Ask what happens to an action nobody classified

The likeliest mistake in a tier table is a missing row. Someone adds close_account, ships it, and forgets the table. The function the guard uses to look up a tier answers with the strictest tier (authpolicy/actions.py, lines 102–103):

    policy = POLICIES.get(action)
    return policy.tier if policy is not None else AuthTier.HIGH

Its docstring gives the reason: “A registry that is permissive by omission is not a registry; it is an allowlist that anyone can join by accident” (lines 98–100).

This is the “fail-safe defaults” principle from Saltzer and Schroeder’s 1975 paper, The Protection of Information in Computer Systems: base access decisions on permission, not exclusion. The evidence that it holds is a test, not the docstring (tests/test_guard.py, lines 259–262):

    def test_unknown_action_defaults_to_high(self):
        """Forgetting to classify a new action must not open a hole."""
        self.assertEqual(tier_for("close_account"), AuthTier.HIGH)
        self.assertEqual(tier_for(""), AuthTier.HIGH)

The same stance covers a call that arrives with no context at all. test_missing_context_is_unauthenticated (lines 278–281) asserts that require_tier("reissue_card", None) raises StepUpRequired rather than skipping the check, so a tool-discovery probe or a broken context cannot act as a bypass.

A second test covers the rows that do exist. Rather than naming the risky tools by hand, it reads every HIGH row from the table and calls each tool at MEDIUM (lines 182–192):

        high_actions = [name for name, p in POLICIES.items() if p.tier is AuthTier.HIGH]
        self.assertTrue(high_actions, "no HIGH actions declared — table is wrong")

        for action in high_actions:
            with self.subTest(action=action):
                tool_fn = getattr(banking, action)
                result = payload(run(tool_fn(context=caller_at(AuthTier.MEDIUM))))
                self.assertFalse(
                    result.get("ok"),
                    f"{action} completed on medium auth — it is declared HIGH",
                )

A new HIGH row joins that test the day it is added, with no one remembering to extend the suite.

A YAML constraint is not the check

The skill that offers the card tool has frontmatter that looks like a control (skills/card_services/skill.md, lines 7–20):

import_tools:
  - reissue_card
  - transfer_funds
tool_constraints:
  - reissue_card:
      requires_confirmation:
        enabled: true
        utter_for_confirmation: utter_confirm_reissue
        utter_on_user_denial: utter_action_cancelled
  - transfer_funds:
      requires_confirmation:
        enabled: true
        utter_for_confirmation: utter_confirm_transfer
        utter_on_user_denial: utter_action_cancelled

Read what it constrains. The pattern’s responses.yml calls these “confirmations of INTENT”, not of identity, and ends: “confirming is not verifying” (lines 3–6). In the pinned Rasa Pro 3.20.0rc1 wheel, the model resolves the pause itself, by calling resolve_tool_confirmation with confirmed=true (rasa/mantle/orchestration/tool_execution/constraints.py, lines 301–307). A confirmation prompt would not have changed the opening failure: the caller wanted the card and would have said yes.

A requires condition that names authentication is a different kind of frontmatter, and the companion does not declare one on reissue_card. The pattern explains why it treats even that kind as the outer layer (authpolicy/guard.py, lines 15–25):

They are evaluated by the orchestrator against conversation state, which means
they are instructions to a language model about which tool it may select. That
is a routing control. It is not an execution control: it constrains what the
model is *offered*, not what the process will *do* when the function is entered.

This pattern therefore treats the frontmatter as the outer layer — it keeps the
conversation coherent, so the caller gets asked for a passphrase instead of
being told "no" — and puts the decision that actually binds inside the tool, on
the line before the side effect. The prose can be edited, the model can be
swapped, the constraint can be mistyped in YAML and silently ignored; the
function still refuses.

The binding check is four lines at the top of the tool (tools/banking.py, lines 186–189):

    try:
        require_tier("reissue_card", context)
    except StepUpRequired as exc:
        return _step_up(exc, context)

The path a call takes, and which parts decide anything:

Outside the function: routing Inside reissue_card: require_tier Skill offers reissue_card; model chooses to call it requires_confirmation pause; model resolves confirmed=true required = tier_for(action) (no row: HIGH) function entered held satisfies required? held = auth_tier in memory (no context or unknown value: none) Card posted, reference returned yes ok: false, step_up_required, no dispatch data no Verification tools only: grant on a pass, revoke on lockout writes
FigureWhere the card reissue is decided, and what writes the tier the guard reads

Everything in the routing box can be changed by editing prose or YAML, and the model makes both of its decisions. What decides is the comparison in the lower box. It reads the tier table and one memory value, auth_tier. An unrecognised value in that field counts as no authentication (test_coerce_fails_closed, lines 236–243). Only the verification tools write it (tools/verification.py, line 1). memory.yml marks none of the auth fields llm_settable (lines 10–12), and in the pinned wheel that setting defaults to false (rasa/mantle/memory/field.py, line 52), so the model cannot award itself a tier.

The evidence for this row is a test that calls the tool directly, as a caller holding a correctly earned MEDIUM, and checks nothing happened (test_medium_auth_cannot_reissue_a_card, lines 118–129):

        context = caller_at(AuthTier.MEDIUM)
        result = payload(
            run(reissue_card(delivery_address="12 Elsewhere Street", context=context))
        )

        self.assertFalse(result["ok"])
        self.assertTrue(result["step_up_required"])
        self.assertEqual(result["required_tier"], "high")
        self.assertEqual(result["held_tier"], "medium")
        # The action did not happen. These keys exist only on success.
        self.assertNotIn("dispatched", result)
        self.assertNotIn("reference", result)

A passing negative test proves little until you have seen it fail; the mutation-testing guide takes that idea further. The pattern ships the exercise as make prove, which runs scripts/prove_guard.py. We ran that script on 21 September 2026 in a scratch export of the pattern at 69e27b6, with the pattern’s Rasa Pro 3.20.0rc1 environment. It is an offline proof. Its output, with the terminal colour codes stripped:

1. removing the tier guard from reissue_card…
2. running the eval suite without it…
   red, as it must be.
3. guard restored.
4. re-running to prove the restore was clean…

The guard is load-bearing:
  without it the suite fails; with it the suite passes.

The script does not say which tests failed, so we removed the same four lines by hand and ran the suite verbosely. Seven of the 36 tests failed, not the two that the test module’s docstring names (lines 20–22):

FAIL: test_locked_out_caller_cannot_complete_the_high_action (test_guard.TestFailurePaths.test_locked_out_caller_cannot_complete_the_high_action)
FAIL: test_every_high_tier_tool_refuses_medium (test_guard.TestHighTierCannotCompleteOnLowerAuth.test_every_high_tier_tool_refuses_medium) (action='reissue_card')
FAIL: test_low_auth_cannot_reissue_a_card (test_guard.TestHighTierCannotCompleteOnLowerAuth.test_low_auth_cannot_reissue_a_card)
FAIL: test_medium_auth_cannot_reissue_a_card (test_guard.TestHighTierCannotCompleteOnLowerAuth.test_medium_auth_cannot_reissue_a_card)
FAIL: test_refusal_carries_no_dispatch_data (test_guard.TestHighTierCannotCompleteOnLowerAuth.test_refusal_carries_no_dispatch_data)
FAIL: test_unauthenticated_caller_cannot_reissue_a_card (test_guard.TestHighTierCannotCompleteOnLowerAuth.test_unauthenticated_caller_cannot_reissue_a_card)
FAIL: test_medium_caller_attempting_high_is_stepped_up (test_guard.TestStepUpIsRequiredNotRemembered.test_medium_caller_attempting_high_is_stepped_up)

Then we did the same to the two MEDIUM tools, removing the four guard lines from both get_balance and get_recent_bill. The suite’s last lines:

----------------------------------------------------------------------
Ran 36 tests in 1.789s

OK

Nothing failed. The guards are in the code, but no test depends on them. test_allows_when_tier_is_sufficient and test_lockout_revokes_the_tier (lines 290–292 and 394–400) exercise require_tier for get_balance directly, never the tool. So the MEDIUM rows have a guard and no evidence for it. The missing test is small: call each MEDIUM tool directly at NONE and at LOW and assert it returns ok: false with no balance or bill.

Lockout has to end the call’s authority

A challenge attempt has three outcomes in the pattern: passed, retry and locked out. A failed attempt is one of the last two. The code comment on the last one reads “Budget exhausted. Terminal — the only exit is a human.” (authpolicy/challenges.py, lines 64–72).

The evidence that “terminal” means something comes in two parts. The first is that running out of attempts produces a lockout that grants nothing and hands off. test_wrong_passphrase_retries_then_locks_out (tests/test_guard.py, lines 352–359) gives one retry, then LOCKED_OUT with handoff true. The same holds for the one-time code (lines 370–374):

    def test_lockout_requires_handoff(self):
        result = check_otp("nope", attempts_used=RETRY_BUDGET)
        self.assertIs(result.outcome, Outcome.LOCKED_OUT)
        self.assertTrue(result.handoff)
        self.assertIsNone(result.granted)

The second is that the risky action still refuses after the lockout. This is the opening failure asserted directly (lines 411–415):

        context = caller_at(AuthTier.MEDIUM, otp_attempts=float(RETRY_BUDGET))
        revoke(context)
        result = payload(run(reissue_card(context=context)))
        self.assertFalse(result["ok"])
        self.assertNotIn("reference", result)

Notice what these tests do not do. The tests that drive a challenge into lockout send only wrong answers (lines 352–374), and no test sends a correct factor after a lockout. The last one calls revoke itself rather than driving a real verification tool into lockout, and it checks one action.

The gaps are in exactly those places. The shared attempt check compares the answer before it looks at the budget (authpolicy/challenges.py, lines 124–153), so a correct factor passes at any attempt count. In the tool layer, lockout sets a locked_out flag and drops the tier to none (tools/verification.py, lines 54–63). Only verify_one_time_code reads that flag before it grants anything (lines 153–155). verify_passphrase does not read it, and neither does the guard.

We checked what that means with two offline probes of our own on 21 September 2026: short scripts, each run once against the pattern’s own tools with a fake context. Neither is a companion test or part of the pattern’s suite. The first sends two wrong codes, calls reissue_card and get_balance, then sends the correct passphrase and calls get_balance and reissue_card again:

otp: retry | auth_tier = medium | locked_out = None
otp: locked_out | auth_tier = none | locked_out = True
reissue_card ok: False
get_balance ok: False
passphrase: passed | auth_tier = medium | locked_out = True
get_balance ok: True
reissue_card ok: False

The second locks out the passphrase itself, then sends the correct one and calls get_balance:

wrong passphrase 1: retry | auth_tier = none | locked_out = None
wrong passphrase 2: locked_out | auth_tier = none | locked_out = True
wrong passphrase 3: locked_out | auth_tier = none | locked_out = True
wrong passphrase 4: locked_out | auth_tier = none | locked_out = True
wrong passphrase 5: locked_out | auth_tier = none | locked_out = True
wrong passphrase 6: locked_out | auth_tier = none | locked_out = True
correct passphrase: passed | auth_tier = medium | locked_out = True
get_balance ok: True

The card stays refused in both, which is the guard doing its job. But after any lockout, including the passphrase’s own, a correct passphrase restores MEDIUM and the balance tool discloses again. The module docstring says exhausting the budget “cannot produce an authenticated state” (authpolicy/challenges.py, lines 19–21), and the README promises “no path back to success” (README.md, line 40). Both hold for the wrong answers the tests send, not for a correct one after a lockout. What stands in the way in a real call is a set of instructions to the model:

  • the lockout result’s hint, “Do NOT retry and do NOT complete the request” (tools/verification.py, lines 71–75);
  • the step-up skill’s if: session.project.locked_out block and its rule not to “switch to the other factor” (skills/step_up/skill.md, lines 28–31 and 47–49);
  • the account skill’s instruction not to call the account tools when locked out (skills/account_info/skill.md, lines 20–22).

All of them are the routing layer again. None is a check in code. The pattern also names a limit of its own: the budget is per call, held in session memory, and “An attacker who hangs up and redials gets a fresh budget” (README.md, lines 162–165).

For a reviewer the lesson is general. Ask for a lockout test that goes through the real tool that locks out, then sends the correct factor on every path that can grant a tier. And ask where the lockout is recorded, because a lockout held in one call’s memory ends when the call does.

Check the payload, not the logging policy

On a voice channel the caller says the secret out loud. The pattern’s docstring says every string that could be a factor passes through one function, redact, before it reaches a log, a tool result or a tracker event (authpolicy/guard.py, lines 150–151). It returns only a length (lines 157–159), and test_redact_never_returns_the_secret checks that the input never comes back (tests/test_guard.py, lines 428–430).

The test that looks like the payload check is test_refusal_payload_contains_no_factor_value (lines 435–443). It calls reissue_card at MEDIUM and asserts that neither fixture factor appears in the result. Read the refusal it inspects: every field is built from the action name, the tiers and constants (tools/banking.py, lines 88–102). There is no input from which a factor could leak, so the test cannot fail for this tool. It passed even with the guard deleted. The tools that do receive the spoken factor, verify_passphrase and verify_one_time_code, are never called in the test module at all. Ask for a payload test on those two tools, across pass, retry and lockout.

A refusal is also not empty. It tells the model the required and held tiers, the factor to ask for and the reason for the tier. What test_refusal_carries_no_dispatch_data (lines 159–171) proves is narrower: no reference, dispatch flag, delivery estimate or address.

Run the suite at the revision you approve

Clone RasaHQ/rasa-community-resources, check out 69e27b6a50c4700f95f2c13d4609d0a3f8cba7d2, change into patterns/voice-auth-stepup and run make test. The target runs the suite through uv run (Makefile, lines 13–14 and 52–53), so you need uv. The project pins rasa-pro==3.20.0rc1 (pyproject.toml, lines 7–9), and the tool module imports rasa, so a bare python -m unittest without that install fails.

We ran the same unittest command with the pattern’s installed environment on 21 September 2026. It printed Ran 36 tests in 4.717s and OK. The tests and probes this guide ran are offline checks of the code: none is a live Mantle run, and none says how a model behaves in a call. The pattern’s conversation tests need a trained model and keys (Makefile, line 32), and we did not run them.

What evidence to demand before approving an agent action

The permissions guide lists the fields of a sign-off: allowed scope, source of authority, evidence inspected, unresolved cases, disposition, owner and review date. With tiers and enforcement locations in the register, each field can now point at something checkable. An illustrative record for two rows, written against the evidence above:

Fieldreissue_cardget_balance
Allowed scopePost a replacement card to a caller holding HIGHDisclose one balance to a caller holding MEDIUM or above
Source of authorityOne-time code: check_otp grants HIGH (challenges.py, lines 113–121; test_otp_grants_high). Its docstring warns that a code read aloud is closer to a second knowledge factor (lines 116–119)Passphrase: check_passphrase grants MEDIUM, never HIGH (challenges.py, lines 103–110; test_passphrase_ceiling_is_medium)
Where enforcedrequire_tier on the tool’s first statementrequire_tier on the tool’s first statement
Evidence inspectedtest_medium_auth_cannot_reissue_a_card, test_every_high_tier_tool_refuses_medium, test_locked_out_caller_cannot_complete_the_high_action; prove_guard.py red then green; seven tests fail with the guard removed; python -m unittest discover -s tests -v (the make test command) OK in the pattern’s environment, 36 tests, at 69e27b6None that depends on the guard: with it deleted, all 36 tests pass
UnresolvedNo tool inventory; no payload test on the verification tools; the one-time code is read aloud on the same call; the retry budget resets on a redialNo direct test at NONE or LOW; after any lockout, including the passphrase’s own, a correct passphrase regrants MEDIUM; the budget resets on a redial; no log-capture test
DispositionPermit at HIGH, with the tool inventory and payload test due before the next releaseHold until the guard has a direct test and the lockout holds for a correct passphrase, with a test
OwnerCard operations ownerAccounts owner
Review dateNext change to tools/banking.py or the tier tableNext change to tools/banking.py or tools/verification.py

The dispositions are narrower than “approved”. That is the purpose. The card row, the one with the irreversible side effect, is close to approvable because its evidence is direct and has been seen to fail. The balance row, lower in tier, is held: its guard has no test behind it, and its lockout has a gap. Judged by how their calls read, neither row would have raised a question.

Questions reviewers ask

Engineering shows me the conversation tests passing. Is that evidence?

It is evidence of something else. The pattern’s conversation tests in tests/e2e/tiering.yml say what they are: “Each case asserts that the agent behaved correctly on one sampled run. A pass means the model chose well this time; it does not mean the model cannot choose otherwise” (lines 6–8). The same header says what they are for: whether “the caller is asked for the right factor, told the truth about what happened, and not challenged for public information” (lines 13–16). Accept them for the conversation design and ask for the direct tool tests for the row.

A new action is urgent and has no tier yet. Can I approve it provisionally?

With a strict default and a guard call in the tool, the provisional state already exists: the action is refused at HIGH until someone adds a row. That is a legitimate interim disposition, at the cost of callers being stepped up or handed off. What you should not approve is a provisional row at a lower tier with no test, or a tool that skips the guard “for now”. Check the tool inventory before you accept that the default applies.