# Rasa Community — published authored resources Fictional scenarios and synthetic results are teaching examples, not production claims. --- # We said Rasa never expands api_key: ${VAR}. Its client does Source: https://rasa.community/library/guides/api-key-placeholders-rasa-3-21/ Author: Rod Rivera Published: 2026-09-29 On 3 September our companion repository, `RasaHQ/rasa-community-resources`, told its readers in `docs/MIGRATING.md` that for model-group credentials in Rasa Pro "the `${VAR}` form has never worked": a line such as `api_key: ${OPENAI_API_KEY}` would hand the provider the placeholder text as its key. The same commit, `432d416`, changed 67 credential lines across 46 files, 63 of them to `api_key_env: NAME`. Its evidence was a reproduction against Rasa Pro 3.20.0.dev6, recorded in the commit message: ```text read_yaml("api_key: ${MY_KEY}") -> {'api_key': '${MY_KEY}'} _resolve_api_key_env({'api_key_env': 'MY_KEY'}) -> {'api_key': 'sk-REAL-…'} _resolve_api_key_env({'api_key': '${MY_KEY}'}) -> {'api_key': '${MY_KEY}'} ``` The commit concluded "There is no rescue path". Follow the same value the rest of the way to the call into litellm, on the released 3.20.0 wheel, with litellm's `acompletion` replaced by a stub that records its arguments and raises before any network call, and this is what the stub records: ```text rasa-pro 3.20.0 stopped: ProviderClientAPIException api_key handed to litellm.acompletion: sk-probe-real-value ``` The real key reached litellm. Replace the one Rasa function that expands it with the identity function, and the same stub records `${PROBE_KEY}` instead. Run the quickstart live on 3.20.0 with `api_key: ${ANTHROPIC_API_KEY}`, and it passes its acceptance script against Anthropic. Meanwhile the spelling our readers were told to adopt is the one Rasa Pro 3.21.0.dev3, a prerelease, refuses. The companion's release run against it on 28 September logged 22 lines like this one: ```text - [mantle.validation.config.invalid_model_group_credentials] Model group at index 0 in 'integrations.yml' has invalid credentials: Model group 'orchestrator' uses 'api_key_env', which is no longer supported. Replace it with 'api_key: ${ENV_VAR_NAME}'. ``` Anyone who followed the note moved off the form 3.21.0.dev3 asks for and onto the form it rejects. The error does name the replacement, so the fix is one edit per line; what it cannot tell a reader is that the note which sent them there was wrong. The mistake was ours, and it was one of method rather than of reading. Each line of that reproduction is true. It stopped two layers before the network, and a repository-wide rewrite followed it. **The position this guide takes:** a secret that is still a placeholder after your config loader runs has been deferred, not dropped. The only proof of what a provider receives is the arguments captured at the call site. The cost is a probe per credential path, written against a private module attribute that can move between releases. If you are testing a Rasa Pro project against 3.21.0.dev3, the last section lists which credential lines must change. The sections before it follow the value through four stops, so you can check your own. ## The loader defers api_key on purpose The YAML loader expands `${VAR}` in `_env_var_constructor` in `rasa/shared/utils/yaml.py`. This is lines 140 to 168 of the 3.20.0 wheel (sha256 `a68fa7ef…e9b5`), unedited: ```python def _env_var_constructor(loader: BaseConstructor, node: ScalarNode) -> str: """Process environment variables found in the YAML.""" value = loader.construct_scalar(node) expanded_vars = os.path.expandvars(value) not_expanded = [ w for w in expanded_vars.split() if w.startswith("$") and w in value ] constructed = loader.constructed_objects key_node = next(reversed(constructed)) if constructed else None if not_expanded: if ( isinstance(key_node, ScalarNode) and key_node.value in DEFERRED_RESOLUTION_KEYS ): # These fields (e.g. OAuth client_id, token_url, Langfuse keys) are # resolved at runtime, so an unset env var is not an error at load time. return value raise RasaException( f"Error when trying to expand the " f"environment variables in '{value}'. " f"Please make sure to also set these " f"environment variables: '{not_expanded}'." ) if isinstance(key_node, ScalarNode) and key_node.value in SENSITIVE_DATA: return value return expanded_vars ``` Read it in order. The constructor runs on the scalars the loader tags as `!env_var`, which are unquoted `${VAR}` values, and expands each one first. If a variable is unset, it raises, unless the key is on `DEFERRED_RESOLUTION_KEYS`. Only once every variable has a value does it reach the last branch. For a key on `SENSITIVE_DATA` it then returns `value`, the original text, and throws `expanded_vars` away. `SENSITIVE_DATA` in `rasa/shared/constants.py` (lines 361 to 373) opens with `API_KEY`, and line 185 sets `API_KEY = "api_key"`; nine other constants follow, among them the AWS and Langfuse credentials. `DEFERRED_RESOLUTION_KEYS`, the keys allowed to stay unset at load, does not include `API_KEY` in 3.20.0 or 3.21.0.dev3. So `api_key: ${OPENAI_API_KEY}` comes back from `read_yaml` as the string `${OPENAI_API_KEY}` when the variable is set. When it is unset, `read_yaml` raises `RasaException` and names the variable. `model: ${VAR}` comes back expanded. Why throw away a value the loader already has? The wheel says so in its own comments. The fallback constructor's docstring in the same file (lines 104 to 124) calls `SENSITIVE_DATA` and `DEFERRED_RESOLUTION_KEYS` "protections", and warns that a parser which always expanded "would risk exposing secrets to any bare-parser consumer in the process". A parsed config that holds `${OPENAI_API_KEY}` can be logged, dumped or passed around without carrying the key. The placeholder also lets Rasa check how a secret was written. In the same 3.20.0 wheel, `validate_model_group_configuration_setup` in `rasa/engine/validation.py` validates the model groups in `endpoints.yml` and calls a sensitive-keys check (lines 1425 to 1446). That check tests every sensitive value against the `${...}` pattern at lines 1394 to 1396: ```python if key in SENSITIVE_DATA: if isinstance(value, str): if not re.match(SECRET_DATA_FORMAT_PATTERN, value): ``` A value that fails raises "must be set as an environment variable". `re.match` anchors only at the start, so the check refuses an `api_key` string that does not begin with `${`. It does not require `api_key`, and `api_key_env` is not on the list it walks. What it does accept is a reference, which it can only see if the loader has left the reference in place. On 3.20.0, then, the one kind of `api_key` string this check lets through in an `endpoints.yml` model group is the spelling our note said had never worked. Our commit named the reason correctly, "a secret-leak guard", and then assumed that a guard at the loader means nothing downstream expands the value. A loader built this way hands out references and leaves the dereferencing to later code. The job is to find that code. Rasa's redaction module states the trap from the other side (`rasa/utils/secret_redaction.py`, lines 21 to 26): keys on the list "survive in memory as the literal `"${VAR}"` placeholder", while keys off it, such as `token`, are already expanded. The 3.21.0.dev3 wheel keeps the loader branch unchanged, at lines 166 and 167 of its `yaml.py`. ## Follow the value to the call site One probe, run on three wheels: 3.20.0.dev6 (the version the commit names), 3.20.0 and 3.21.0.dev3. It sets a fake key, loads a model group and prints the key at each stop. No provider is called. The notes give what the probe printed at each numbered line. :::annotated{title="probe.py"} ```python import os, sys os.environ["PROBE_KEY"] = "sk-probe-real-value" import rasa from importlib.metadata import version print("rasa-pro", version("rasa-pro")) from rasa.shared.utils.yaml import read_yaml cfg = read_yaml(""" model_groups: - id: orchestrator models: - provider: anthropic model: claude-haiku-4-5 api_key: ${PROBE_KEY} """) group = cfg["model_groups"][0] print("after read_yaml:", group["models"][0]["api_key"]) # (1) from rasa.mantle.llm import client as mc from rasa.shared.utils.io import resolve_environment_variables if hasattr(mc, "_resolve_api_key_env"): print("after _resolve_api_key_env:", mc._resolve_api_key_env(group)["models"][0]["api_key"]) # (2) from rasa.shared.utils.llm import llm_factory c = llm_factory(dict(group), mc._DEFAULT_LLM_CONFIG) print("client:", type(c).__name__) print("stored api_key:", c._completion_fn_args.get("api_key")) # (3) print("api_key passed to litellm at call time:", resolve_environment_variables(c._completion_fn_args).get("api_key")) # (4) ``` 1. `after read_yaml: ${PROBE_KEY}` on all three versions. Our September reproduction checked this stop. 2. `after _resolve_api_key_env: ${PROBE_KEY}` on 3.20.0.dev6 and 3.20.0. This was the reproduction's second stop, and its last. The 3.21.0.dev3 wheel has no such function, so the line does not print there. 3. `stored api_key: ${PROBE_KEY}` on all three. `llm_factory` has built a `DefaultLiteLLMClient`, and the arguments it keeps for the provider call still hold the placeholder. 4. `api_key passed to litellm at call time: sk-probe-real-value` on all three. This line applies, by hand, the function the client's completion methods apply to those arguments. ::: Stop 2 explains the commit's "no rescue path". In the 3.20.0 wheel, `_resolve_api_key_env` in `rasa/mantle/llm/client.py` looks for one key and nothing else (lines 142 to 147): ```python env_var_name = value.get(_API_KEY_ENV_CONFIG_KEY) if isinstance(env_var_name, str): value.pop(_API_KEY_ENV_CONFIG_KEY) api_key = os.environ.get(env_var_name) if api_key: value[API_KEY] = api_key ``` It substitutes a value when a mapping carries `api_key_env`, and leaves an `api_key: ${VAR}` string alone. We read "this resolver does not expand it" as "nothing expands it". Stop 4 is where the value becomes real. The client's calls go through `_BaseLiteLLMClient` in `rasa/shared/providers/llm/_base_litellm_client.py`. In the 3.20.0 wheel it has three completion methods: `completion` (line 130), `acompletion` (165) and `acompletion_stream` (228). Each wraps the stored arguments in `resolve_environment_variables`, imported at line 28, before it calls litellm (lines 155, 190 and 235). Lines 153 to 159, from `completion`: ```python formatted_messages = self._get_formatted_messages(messages) arguments = cast( Dict[str, Any], resolve_environment_variables(self._completion_fn_args) ) response = completion( messages=formatted_messages, **{**arguments, **kwargs} ) ``` The function itself, in `rasa/shared/utils/io.py` (lines 518 to 525), applies `os.path.expandvars` to every string in a string, list or dict: ```python if isinstance(value, str): return os.path.expandvars(value) elif isinstance(value, list): return [resolve_environment_variables(item) for item in value] elif isinstance(value, dict): return {key: resolve_environment_variables(val) for key, val in value.items()} else: return value ``` So the value travels as `${PROBE_KEY}` through the loader, the resolver and the constructed client, and becomes `sk-probe-real-value` in the last Rasa frame before litellm. Two credential spellings resolved in two places, one at bootstrap and one inside every completion call, is a design a reader could fairly call a trap. It caught us. :::callout{type="pitfall" title="Inspecting the built client still shows the placeholder"} Stop 3 prints `${PROBE_KEY}`. A test that goes one step further than ours did, builds the client and asserts on its stored arguments, would have confirmed the wrong belief. Only the arguments at the call itself show what is sent. ::: ## Put a spy on the last function before the network Stop 4 in `probe.py` has a weakness a sceptic would name: it calls `resolve_environment_variables` itself, so it shows what the function returns, not that the client calls it. The second probe removes the shortcut. It installs a test spy, the test double that records how it was called: a replacement for `acompletion` in the base client's module that stores its arguments and raises. It builds the client with `LLMClient.from_bootstrap` (the constructor that reads the `integrations.yml` model group) and lets the client make the call. The lines that do the work, excerpted; the full script is folded below: ```python import rasa.shared.providers.llm._base_litellm_client as base seen = {} async def fake_acompletion(**kw): seen.update(kw); raise RuntimeError("stopped before network") base.acompletion = fake_acompletion # ... read_yaml of the same model group as probe.py ... c = LLMClient.from_bootstrap(B()) try: asyncio.run(c._client.acompletion("hi")) except Exception as e: print("stopped:", type(e).__name__) print("api_key handed to litellm.acompletion:", seen.get("api_key")) ``` :::solution{title="Show the whole of probe_call.py"} The script as run, unedited from the receipt `offline-call-site-probe.txt`. ```python import os, asyncio os.environ["PROBE_KEY"] = "sk-probe-real-value" from importlib.metadata import version print("rasa-pro", version("rasa-pro")) from rasa.shared.utils.yaml import read_yaml import rasa.shared.providers.llm._base_litellm_client as base from rasa.mantle.llm import client as mc from rasa.mantle.llm.client import LLMClient seen = {} async def fake_acompletion(**kw): seen.update(kw); raise RuntimeError("stopped before network") base.acompletion = fake_acompletion cfg = read_yaml(""" model_groups: - id: orchestrator models: - provider: anthropic model: claude-haiku-4-5 api_key: ${PROBE_KEY} """) group = cfg["model_groups"][0] class B: llm = group c = LLMClient.from_bootstrap(B()) try: asyncio.run(c._client.acompletion("hi")) except Exception as e: print("stopped:", type(e).__name__) print("api_key handed to litellm.acompletion:", seen.get("api_key")) ``` ::: Run with `uv run --no-project --python 3.12 --prerelease allow --with rasa-pro==3.20.0 python probe_call.py`, then again with `rasa-pro==3.21.0.dev3`. The output, unedited: ```text rasa-pro 3.20.0 stopped: ProviderClientAPIException api_key handed to litellm.acompletion: sk-probe-real-value rasa-pro 3.21.0.dev3 stopped: ProviderClientAPIException api_key handed to litellm.acompletion: sk-probe-real-value ``` The stub raised `RuntimeError`, and the probe caught `ProviderClientAPIException`, so the client's own `acompletion` made the call and wrapped the stub's error. On 3.20.0, `from_bootstrap` also runs `_resolve_api_key_env` first; on 3.21.0.dev3 it runs the credential validator, and the placeholder passes it. That shows the real key arrives. It does not yet show why. The counter-test, `probe_call_counter.py`, is `probe_call.py` with one line added after the stub is installed: ```python base.resolve_environment_variables = lambda value: value # counter-test: call-time expansion disabled ``` The base client imports the function by name, so replacing the module attribute replaces what its completion methods call. The command differs from the first probe's in two ways: an `env -u` prefix that keeps any real key out of the environment, and a `grep` that drops two structlog lines about SSL certificates: ```text env -u ANTHROPIC_API_KEY -u OPENAI_API_KEY -u LITELLM_MODIFY_PARAMS uv run --no-project --python 3.12 --prerelease allow --with rasa-pro==3.20.0 python probe_call_counter.py | grep -v '\[debug' ``` Then again with `rasa-pro==3.21.0.dev3`. The output, unedited: ```text rasa-pro 3.20.0 stopped: ProviderClientAPIException api_key handed to litellm.acompletion: ${PROBE_KEY} rasa-pro 3.21.0.dev3 stopped: ProviderClientAPIException api_key handed to litellm.acompletion: ${PROBE_KEY} ``` With that one call disabled, the placeholder reaches litellm on both versions. Nothing else on this path expands it. The same method carries to any stack that interpolates environment variables into config. Find the last function your code calls before bytes leave the process. Put a spy there that records its arguments and raises. Drive the real code path into it and assert on what it recorded. Then disable the step you think is responsible and check that the assertion fails. The loader, the parsed config and the constructed client are all intermediate states, and each of them can legitimately hold a placeholder. Here is everything the probes show, in one place: | What was checked | Rasa Pro 3.20.0 | Rasa Pro 3.21.0.dev3 | | ---------------------------------------------------------- | --------------------- | --------------------- | | `read_yaml` returns | `${PROBE_KEY}` | `${PROBE_KEY}` | | The built client stores | `${PROBE_KEY}` | `${PROBE_KEY}` | | The spy on `litellm.acompletion` records | `sk-probe-real-value` | `sk-probe-real-value` | | The spy, with `resolve_environment_variables` as identity | `${PROBE_KEY}` | `${PROBE_KEY}` | | A live quickstart run with `api_key: ${ANTHROPIC_API_KEY}` | acceptance passed | acceptance passed | The probes are the mechanism evidence; the live runs are the outcome. Runs recorded 29 September 2026. On the released 3.20.0 wheel, the quickstart with its `orchestrator` model group set to Anthropic, `claude-haiku-4-5` and `api_key: ${ANTHROPIC_API_KEY}` trained and passed its acceptance script: four conversations, no ERROR lines in the server log. On 3.21.0.dev3 the quickstart as it ships, with the same model group, passed too; its log holds two failed turns in which a request with no non-system message was refused, and neither concerns the key. A live pass shows that a provider accepted the key. It does not show which line produced it, and it would look the same if some resolver we had not found did the work. The counter-test is what ties the result to `resolve_environment_variables`. Neither kind of evidence is enough alone. :::callout{type="warn" title="What this evidence does not prove"} - **Each live run is one pass.** One acceptance run per version, on one model and one provider. It shows the key worked, not how often anything else does. - **One method is exercised.** The spy covers `acompletion`. `completion` and `acompletion_stream` wrap the same call at lines 155 and 235 by the source alone; we did not stub them. The claim covers the default LiteLLM client the probes build; we did not examine other provider clients in the wheel. - **It is the code, not a promise.** Call-time expansion is what this code does in two wheels, one of them a prerelease. A refactor could move it. - **All of this is for the unquoted form.** Unquoted, an unset variable behind `api_key` makes `read_yaml` raise, because `api_key` is not on `DEFERRED_RESOLUTION_KEYS`. A quoted `"${NAME}"` behaves differently; see the last section. ::: ## Rasa 3.21.0.dev3: api_key_env is no longer supported Here is what the wrong premise cost. Of the lines commit `432d416` moved to `api_key_env`, 30 went into `endpoints.yml` model groups; one of those 30 is a commented-out line, and the 30 sit in 15 files. Four more lines it changed sat in speech blocks and were deleted, because by the commit's account `asr:` and `tts:` entries take no credential key at all. Commit `432d416` also added a lint check that enforced `api_key_env`, in the repository's gate and in the starter pack, and it put the "has never worked" note in `docs/MIGRATING.md`. This guide's probes, live runs and table cover the `integrations.yml` group Mantle builds, not `endpoints.yml` model groups. The commit was careful about the check. It records that the check was "demonstrated FAILING on a deliberately broken fixture" and that its unit tests were mutation-tested, and it states the rule it was following: "A check only ever seen green is the defect this change exists to fix." The premise the check enforced got no such test. The commit predicts the symptom a literal `${VAR}` key would cause, "a provider auth error at runtime", and quotes no such error. For an unquoted `${VAR}` on the path traced above, the placeholder text never reaches litellm: a set variable is expanded at the call, and an unset one stops `read_yaml`. The literal placeholder does reach the model call in one case, which the last section shows: a quoted placeholder whose variable is unset. Rasa Pro 3.21.0.dev3 then turned the question from a belief into a validation error. The companion's release workflow ran `make ci KEEP_GOING=1` against it on 28 September and logged 22 lines like the one at the top of this guide: 21 for the model group at index 0 of `integrations.yml`, one at index 1, all for `api_key_env` on `orchestrator`. Here is where each spelling ends up on each version: :::diagram{title="Two spellings of one credential on Rasa Pro 3.20.0 and 3.21.0.dev3"} ```dot rankdir=LR; env [label="api_key_env: NAME"]; var [label="api_key: ${NAME}"]; resolver [label="3.20.0 from_bootstrap\n_resolve_api_key_env", class="mark-1"]; validator [label="3.21.0.dev3 from_bootstrap\nvalidate_model_group_credentials", class="mark-2"]; rejected [label="InvalidConfigException", class="blocked mark-3"]; client [label="default LiteLLM client\nresolve_environment_variables", class="mark-4"]; call [label="litellm call\nreceives the real key", class="ok"]; env -> resolver [label="value written to api_key"]; var -> resolver [label="passed through as ${NAME}"]; env -> validator; var -> validator [label="exactly ${NAME}"]; validator -> rejected [label="api_key_env"]; resolver -> client; validator -> client; client -> call; ``` 1. On 3.20.0, for the `integrations.yml` model group Mantle builds in `LLMClient.from_bootstrap`, `_resolve_api_key_env` writes the value of `NAME` into `api_key` and passes a `${NAME}` string through unchanged. The offline spellings probe recorded the real value for `api_key_env` there. This note covers that path only. 2. On 3.21.0.dev3, `from_bootstrap` runs the validator instead, before `llm_factory` builds a client. 3. `api_key_env` is refused: `from_bootstrap` raises `InvalidConfigException` (offline spellings probe). This is the form the companion moved to. 4. Whatever reaches the client as `${NAME}` is expanded here, inside each completion method. The spy and the counter-test show it on both versions. ::: The 22 CI lines are a separate observation from the diagram's `InvalidConfigException`. They carry the code `mantle.validation.config.invalid_model_group_credentials` and the prefix "Model group at index 0 in 'integrations.yml'" (21 lines) or "at index 1" (one), which come from a caller we did not open. Their message text matches `validate_model_group_credentials` in `rasa/shared/providers/model_group_validation.py` of the 3.21.0.dev3 wheel (sha256 `8b7a1977…3495`), and the attribution rests on that match. The validator raises `api_key_env_not_supported` when a model entry contains `api_key_env`. Its docstring then requires every `SENSITIVE_DATA` value that is present to be exactly `${ENV_VAR_NAME}`, and the code checks each string with `re.fullmatch` against `SECRET_DATA_FORMAT_PATTERN`, which is `\${(\w+)}` (line 421 of `rasa/shared/constants.py` in both wheels). A string that fails raises `sensitive_key_string_value_must_be_set_as_env_var`. In the same wheel, `LLMClient.from_bootstrap` no longer calls `_resolve_api_key_env`; it runs this validator and raises `InvalidConfigException` before `llm_factory` builds anything. The companion's fix, commit `ededc1c`, converts the `api_key_env` entries back to `api_key: ${NAME}` and flips both lint checks to reject `api_key_env` and any `api_key` that is not exactly `${NAME}`. Its message still says "On 3.20 the rule was the reverse". The probes and the live run contradict the half of that which matters: `api_key: ${NAME}` reached the litellm call as the real key on 3.20.0 as well, and a provider accepted it. The companion's main branch no longer says "has never worked". Pull request #50 merged as `5206372`, and commit `1e5101b` in it rewrites the 3.20 history in `docs/MIGRATING.md` and says why the old advice existed: "This catalog once advised `api_key_env` because a reproduction stopped at the YAML loader, which returns sensitive values unexpanded and defers expansion to the call." Two sentences in the same file at that commit still contradict the probes. Lines 240 and 241 say "The provider client expands that reference from the environment when it is built"; stop 3 shows the built client still holding the placeholder, and the expansion happens at each call. Line 330 says "On 3.20 it had to be `api_key_env: OPENAI_API_KEY`". :::checkpoint{id="api-key-placeholder-at-call-site" question="On Rasa Pro 3.20.0, read_yaml returns api_key as the string ${ANTHROPIC_API_KEY}, and ANTHROPIC_API_KEY is set. What does litellm.acompletion receive?" options="The literal ${ANTHROPIC_API_KEY},The value of ANTHROPIC_API_KEY,Nothing; the loader raises first" answer="1"} The loader returns sensitive keys unexpanded, and the built client keeps them that way. The default LiteLLM client runs `resolve_environment_variables` over its arguments inside each completion method, so the spy on `acompletion` recorded the real value; disable that one call and it records the placeholder. The loader's raise branch is for an unset variable, and this one is set. ::: ## Which credential lines to change for 3.21.0.dev3 :::callout{type="info" title="Scoped to a prerelease"} Rasa Pro 3.21.0.dev3 is a prerelease. The validator, its messages and the call-time expansion may all change before a stable 3.21. ::: The table covers the entries under `models:` in the `integrations.yml` model group that `LLMClient.from_bootstrap` builds, which is what the 3.21.0.dev3 validator checks. It does not cover `endpoints.yml` model groups, where 30 of the companion's rewritten lines went. Each cell names the layer that accepts or refuses the line. The results come from an offline probe, `offline-credential-spellings-probe.txt`, that loads each spelling with `read_yaml`, builds it with `from_bootstrap` and calls it through the spy, plus the live runs for the unquoted `${NAME}`. | Line in an `integrations.yml` model group | Rasa Pro 3.20.0 | Rasa Pro 3.21.0.dev3 | | ----------------------------------------- | --------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ | | `api_key: ${NAME}`, unquoted | Loader returns it raw; the spy recorded the real value; live passed. Unset: `read_yaml` raises | Validator accepts it; the spy recorded the real value; live passed. Unset: `read_yaml` raises. Keep it | | `api_key: "${NAME}"`, quoted | The spy recorded the real value. Unset: the spy recorded the literal `${UNSET_PROBE}` | Validator accepts it. Unset: the spy recorded the literal `${UNSET_PROBE}`. Remove the quotes | | `api_key_env: NAME` | `_resolve_api_key_env` in `from_bootstrap` writes the value; the spy recorded it | `from_bootstrap` raises `InvalidConfigException`. Change it | | `api_key: $NAME` | Loader does not tag it (no `${`); the spy recorded the real value, and `$PROBE_KEY` with call-time expansion disabled | `from_bootstrap` raises `InvalidConfigException`. Change it | | `api_key: ${NAME:-default}` | Loader tags it, `expandvars` leaves it as written, and `read_yaml` raises `RasaException` | Same: `read_yaml` raises before the validator runs. Change it | | A literal key | Passed through; the spy recorded the literal | `from_bootstrap` raises `InvalidConfigException`. Change it | :::callout{type="pitfall" title="A quoted placeholder with an unset variable reaches litellm as text"} Written as `api_key: "${UNSET_PROBE}"` with the variable unset, the key loaded without error on both versions, passed the 3.21.0.dev3 validator, and the spy on `litellm.acompletion` recorded the literal `${UNSET_PROBE}`. The loader's resolver is an implicit one, which YAML applies only to unquoted scalars, so a quoted value is never tagged `!env_var` and the loader's raise branch never runs. At the call, `os.path.expandvars` leaves the unset variable as written. The unquoted form fails at load instead. With Rasa's code unmodified, this is the only case in the probe where the placeholder text reached the model call. ::: :::solution{title="Show the spellings probe and its output"} The offline probe behind the table, from `offline-credential-spellings-probe.txt`: litellm's `acompletion` is replaced by a spy that raises before any network call, `PROBE_KEY` is a fake value and `UNSET_PROBE` is unset. Add your own spellings to the list and run it the same way. The command: ```text env -u ANTHROPIC_API_KEY -u OPENAI_API_KEY -u LITELLM_MODIFY_PARAMS -u UNSET_PROBE uv run --no-project --python 3.12 --prerelease allow --with rasa-pro==3.20.0 python probe_spellings.py 2>/dev/null | grep -v '\[debug' ``` Then again with `rasa-pro==3.21.0.dev3`. The script, as run: ```python import os, asyncio os.environ["PROBE_KEY"] = "sk-probe-real-value" os.environ.pop("UNSET_PROBE", None) from importlib.metadata import version print("rasa-pro", version("rasa-pro")) from rasa.shared.utils.yaml import read_yaml import rasa.shared.providers.llm._base_litellm_client as base from rasa.mantle.llm.client import LLMClient seen = {} async def fake_acompletion(**kw): seen.update(kw); raise RuntimeError("stopped before network") base.acompletion = fake_acompletion def probe(line): seen.clear() try: cfg = read_yaml("model_groups:\n - id: orchestrator\n models:\n - provider: anthropic\n model: claude-haiku-4-5\n " + line + "\n") except Exception as e: return f"read_yaml raised {type(e).__name__}" group = cfg["model_groups"][0] class B: llm = group try: c = LLMClient.from_bootstrap(B()) except Exception as e: return f"from_bootstrap raised {type(e).__name__}" try: asyncio.run(c._client.acompletion("hi")) except Exception: pass return f"spy on litellm.acompletion recorded api_key={seen.get('api_key')!r}" for line in ["api_key: ${PROBE_KEY}", "api_key: ${PROBE_KEY:-fallback}", "api_key: $PROBE_KEY", "api_key: sk-literal-probe", "api_key_env: PROBE_KEY", "api_key: ${UNSET_PROBE}", 'api_key: "${PROBE_KEY}"', 'api_key: "${UNSET_PROBE}"']: print(f"{line!r:36} {probe(line)}") base.resolve_environment_variables = lambda value: value # counter-test: call-time expansion disabled for line in ["api_key: $PROBE_KEY", 'api_key: "${PROBE_KEY}"']: print(f"{line!r:36} {probe(line)} [call-time expansion disabled]") ``` The output on both versions, unedited: ```text rasa-pro 3.20.0 'api_key: ${PROBE_KEY}' spy on litellm.acompletion recorded api_key='sk-probe-real-value' 'api_key: ${PROBE_KEY:-fallback}' read_yaml raised RasaException 'api_key: $PROBE_KEY' spy on litellm.acompletion recorded api_key='sk-probe-real-value' 'api_key: sk-literal-probe' spy on litellm.acompletion recorded api_key='sk-literal-probe' 'api_key_env: PROBE_KEY' spy on litellm.acompletion recorded api_key='sk-probe-real-value' 'api_key: ${UNSET_PROBE}' read_yaml raised RasaException 'api_key: "${PROBE_KEY}"' spy on litellm.acompletion recorded api_key='sk-probe-real-value' 'api_key: "${UNSET_PROBE}"' spy on litellm.acompletion recorded api_key='${UNSET_PROBE}' 'api_key: $PROBE_KEY' spy on litellm.acompletion recorded api_key='$PROBE_KEY' [call-time expansion disabled] 'api_key: "${PROBE_KEY}"' spy on litellm.acompletion recorded api_key='${PROBE_KEY}' [call-time expansion disabled] rasa-pro 3.21.0.dev3 'api_key: ${PROBE_KEY}' spy on litellm.acompletion recorded api_key='sk-probe-real-value' 'api_key: ${PROBE_KEY:-fallback}' read_yaml raised RasaException 'api_key: $PROBE_KEY' from_bootstrap raised InvalidConfigException 'api_key: sk-literal-probe' from_bootstrap raised InvalidConfigException 'api_key_env: PROBE_KEY' from_bootstrap raised InvalidConfigException 'api_key: ${UNSET_PROBE}' read_yaml raised RasaException 'api_key: "${PROBE_KEY}"' spy on litellm.acompletion recorded api_key='sk-probe-real-value' 'api_key: "${UNSET_PROBE}"' spy on litellm.acompletion recorded api_key='${UNSET_PROBE}' 'api_key: $PROBE_KEY' from_bootstrap raised InvalidConfigException [call-time expansion disabled] 'api_key: "${PROBE_KEY}"' spy on litellm.acompletion recorded api_key='${PROBE_KEY}' [call-time expansion disabled] ``` ::: The validator's docstring says nested mappings, for example Azure `oauth`, get the same rule, so check the other keys on `SENSITIVE_DATA` too, not only `api_key`. If you upgrade, write `api_key: ${NAME}`, the spelling the evidence here supports on both versions; the [quickstart](/quickstart/) agent already does. Before you change a whole repository, put a spy on the call and read what it recorded. :::cta{href="/quickstart/" label="Start from the quickstart agent"} The quickstart's `integrations.yml` already uses `api_key: ${ANTHROPIC_API_KEY}` in its `orchestrator` model group, the form the live runs above passed with. ::: --- # How to audit a document an AI agent generates Source: https://rasa.community/library/guides/audit-documents-an-agent-produces/ Author: Rod Rivera Published: 2026-09-19 The record below is illustrative, not a recorded model run: it is the example the [deriving-a-document tutorial](/library/tutorials/deriving-a-document/) opens on, and the client, Marged Ellis, and her portfolio are invented fixture data. An adviser asks an agent for a client suitability record. Here is the first sentence of its body; the example goes on from there: ```text Your portfolio is currently valued at approximately £486,000, held across a balanced mix of global equities (around 58%), sterling corporate bonds (roughly 30%), and property funds (about 12%). ``` Every figure sits close to the invented custodian extract behind it: £486,210.44 in total, equities at 57.7%, bonds at 30.4% and a property row at 11.9%. A reviewer holding the extract would tick all four. What the page does not say is which figures were read from a record and which were rounded, remembered from an older valuation or composed so the sentence would feel complete. The property figure is the one to look at. The extract has a property row, but the record the companion project assembles from the same fixture never cites it. The rendered record's Completeness section says so: ```text ## Completeness 3 declared field(s) have no source and render blank: - `property_value` - `property_weight` - `include_property_breakdown` ``` The prose version states a property weighting that the assembled record has no citation for, in the same confident voice as the figures that do. This guide is for the person asked to approve an agent that writes documents a client or regulator will keep: summaries, letters, case records. Its position is that a fluent document proves nothing, and neither does a green test suite unless the tests run the path the agent actually uses. Sign off on a list of every way the model can write to the document, each tested through the engine you will deploy, and on a refusal that reaches both the agent and any copy already written. We tested that position on the best case available: the companion project behind the tutorial, `tutorials/rasa-document-artifact-tutorial` in `RasaHQ/rasa-community-resources` at commit `69e27b6`, which was designed so that a model cannot write a figure. Its 41 tests pass. We still found five gaps a sign-off should catch, and this guide is built on them. The scenario and its field list are synthetic: the fixture's README says every person, holding and valuation in it is "invented for this tutorial" (`data/source/README.md`, lines 1–5), and the property row is `POS-PF4402-003` in `data/source/holdings.json` (lines 43–47). Nothing here claims the design satisfies any regulator's suitability, record-keeping or disclosure rule, and which fields your own documents need is a decision for your team. ## Why reading the output is not the control The natural control is a careful review of what the agent produced. It fails in three ways, and none of them is about how careful the reviewer is. - **It checks agreement, not origin.** "About 12%" agrees with 11.9% within rounding. The reviewer cannot see whether the model read that row, recalled a figure from an earlier conversation or produced a number that fitted. - **It checks a copy, not a process.** A clean review of this record says nothing about the next one. - **It happens once.** A custodian restates a valuation after the review and before the letter goes out, and the reviewed figure is now wrong under the firm's name. The tutorial's full example goes on to say the allocation "remains well suited to your Balanced risk profile". That is a suitability judgement, and none of the fixture's three sources (`holdings.json`, `factfind.json` and `references/disclosures.json`) contains it or the word "suited". A reviewer may agree with it; that agreement is the reviewer's judgement, not evidence about where the sentence came from. ## What the model can touch in this design The companion project takes the pen away. The conversation edits a structured state, never the document, and the renderer rebuilds the whole document from that state each time it runs (`docpkg/render.py`, `render_markdown`). Most fields are `sourced`: the model names a record and the record supplies the value. Two are `negotiated`: the model picks a value from a short fixed list, such as who the record is addressed to (`docpkg/state.py`, lines 114–119). The model can only fill in the parameters the engine offers it. We asked Rasa Pro 3.20.0rc1 to build the schema for each document tool, offline, with the engine's own `agent_tool_schema_from_custom_tool` (`rasa/mantle/llm/tool_schemas.py`, lines 577–610 in the wheel). The output of that run: ```text point_field_at_record: properties=['field_key', 'source_id', 'record_id', 'record_field', 'reason'] required=['field_key', 'source_id', 'record_id', 'record_field', 'reason'] additionalProperties=False choose_option: properties=['field_key', 'option', 'reason'] required=['field_key', 'option', 'reason'] additionalProperties=False clear_document_field: properties=['field_key', 'reason'] required=['field_key', 'reason'] additionalProperties=False render_document: properties=[] required=[] additionalProperties=False ``` :::solution{title="Before reading on: from those four schemas, list every way the model can change what the document says"} There are four, and a sign-off should name all of them. 1. **Which record a sourced field cites**, through `point_field_at_record`. It cannot type a figure, but it can cite the wrong record. 2. **A value from a fixed list**, through `choose_option`, for the two negotiated fields. 3. **Removing any field**, through `clear_document_field`, which takes any `field_key`. 4. **Free text**, through `reason`, which every write tool requires and which the renderer prints into the record's revision history. The fourth is easy to miss, because nothing in the schema says where a reason ends up. ::: The list names the inputs. What it does not show is where each one ends up, including the file the render tool leaves behind: :::diagram{title="Where each model input lands"} ```dot rankdir=LR; choices [label="Record choices and\nfixed-list values"]; removals [label="Removed fields"]; reasons [label="Reason text"]; state [label="Document state"]; body [label="Record body\n(values, blanks, Provenance)"]; history [label="Revision history"]; file [label="out/DOC-SUIT-00417.md"]; refused [label="A later refusal", class="blocked"]; choices -> state; removals -> state; reasons -> state; state -> body [label="values and blanks"]; state -> history [label="reasons, verbatim"]; body -> file [label="render_document writes"]; history -> file; refused -> file [label="leaves it in place"]; ``` ::: :::checkpoint{id="audit-document-green-suite" question="The record reads correctly, and engineering shows you all 41 companion tests passing. The tests call the Python functions directly. What has that established?" options="The record is safe to send,The functions behave as the tests describe when called directly,Refusals reach the agent in the deployed engine,The model cannot change what the record says" answer="1"} A test establishes what it exercises. These call the functions directly, so they say nothing about how the engine calls them, and the schemas above show four ways the model can still change the record. Finding 1 shows the difference the call path makes. ::: ## Auditing a document an AI agent generates: five findings We ran each check offline, in a fresh export of the companion at `69e27b6`. The engine checks used an installed Rasa Pro 3.20.0rc1 under Python 3.12.13; the suite was also run under Python 3.14.3 with no engine. The scripts are kept with this guide's brief as receipts, each run's output saved verbatim with the command on its first line. | Finding | What we ran | Result | | ----------------------------------------------------- | ----------------------------------------------------- | --------------------------------------------------- | | 1. The refusal does not reach the agent | Tools called through the engine's `call_tool_func` | State changed; the agent was told only "error" | | 2. The boundary test checks a flag that never decides | `llm_settable` set to true on `total_value` | The test's condition breaks; the tool still refuses | | 3. The model writes prose into the record | A reason containing a suitability sentence and a pipe | Printed verbatim, unescaped, in the history | | 4. The model can blank or replace the risk warning | `clear_field`, then a re-citation, on `risk_warning` | Both render: blank, then "Property funds" | | 5. A refusal leaves the earlier copy | Render, restate the extract, render again | Refused; the first file still holds £486,210.44 | None of these needs a model, a licence or a network, and none is a claim about what a model would choose to do in a conversation. Each is a path the code leaves open. ## Finding 1: the refusal the tests prove is not what the agent receives Engineering will show you the suite. It passes with or without the engine installed; the last lines of our `make test` run, the command engineering would use, under Python 3.14.3: ```text Ran 41 tests in 0.053s OK ``` Among those tests is one written for exactly the question a reviewer asks: does a refusal reach the agent? Its docstring reads "The path the AGENT runs, not just the function underneath it" (`tests/test_derived_document.py`, line 296), and another test in the same file warns against "asserting on a convenient internal path while the path that actually runs was different" (lines 115–120). But the refusal test calls `td.render_document()` directly, as a plain function (lines 295–305). The agent does not call it that way. In Rasa Pro 3.20.0rc1 the engine runs a local tool with `return await func(context=ctx, **tool_args), None` (`rasa/mantle/orchestration/tool_execution/invoker.py`, line 234). The companion's tools are declared with plain `def`, so the function runs to the end, side effects included, and only then does awaiting its result raise a `TypeError`. The invoker's catch-all (lines 161–172) turns that into a generic error for the model. This is a defect in the companion, not an engine quirk. The engine's decorator documents the contract: `tool()` is described as "Mark an async function as an LLM-callable tool", and a tool function's type is `Callable[..., Awaitable[ToolResult]]` (`rasa/mantle/tools/decorator.py`, lines 80 and 270). We called the tools through `ToolInvoker.call_tool_func`, the engine method that runs a builder tool (`call_builder_tool` hands every call to it, line 106), with a stand-in for the session object. The only session method it uses on this path is `recording_write`, which in the engine sets and restores a marker around the call and lets any exception through (`rasa/mantle/state/tracker_handle.py`, lines 525–545). Ours does nothing, so it cannot change what the tool returns. The first call re-points the total at the cash line (log timestamps removed): ```text == 1. point_field_at_record through the engine total_value before: 486210.44 [error ] mantle.skill_executor.tool_error error_type=TypeError tool=point_field_at_record [debug ] mantle.skill_executor.tool_error.debug error="object ToolResult can't be used in 'await' expression" tool=point_field_at_record returned: {"error": "object ToolResult can't be used in 'await' expression"} total_value after: 21804.1 revisions: 22 ``` The edit happened, and the agent was told it failed. The render call behaves the same way: on the fixture state it wrote the document file and reported an error, and after the extract was restated it reported the same error, with no `provenance_broken` and no field name. An agent in that position cannot tell a success from a refusal. In a separate copy we changed `point_field_at_record` and `render_document` to `async def` and ran only `render_document` through the engine, with the extract restated. It returned the structured refusal: `{"ok": false, "refused": "provenance_broken", ... "fields": ["total_value"]}`. That is one tool checked. It is not a review of the other tools or of a live conversation. **What to ask for:** the refusal tests run through the engine's tool invoker at the version you will deploy, or a recorded agent run that shows the refusal arriving, not a direct call to the Python function. ## Finding 2: the boundary test checks a flag that never decides The suite has two tests named for the field boundary (`tests/test_derived_document.py`, lines 61–73). The first asserts that no `sourced` field has `llm_settable` set to true. It reads well in a review pack. We set `llm_settable` to true on `total_value` and called the function the flag would open: ```text == A. flip llm_settable to True on the sourced field total_value sourced fields now llm_settable: ['total_value'] set_negotiated_field('total_value') refused: free_text_into_sourced_field ``` The condition the test checks is broken, and the refusal is unchanged. The flag is read, in the same condition as the kind: `if spec.kind != "negotiated" or not spec.llm_settable:` (`docpkg/edits.py`, line 189). A sourced field fails the kind half whatever the flag says, and `set_sourced_field` refuses a negotiated field outright (line 138). So for a sourced field the flag never decides anything; it matters only for the two negotiated fields, where setting it to false would block `choose_option`. What stops the model writing a figure is the shape of the tools: `point_field_at_record` has no `value` argument, and `render_document` takes nothing. The test file says as much. Its docstring names `TestTheModelCannotWriteTheArtifact` and `TestTheRendererRefuses` as "the guarantee" (lines 6–8), and the first of those checks the signatures themselves (lines 97–132): ```python from tools.document import point_field_at_record params = set(inspect.signature(point_field_at_record).parameters) self.assertNotIn("value", params) self.assertIn("record_id", params) ``` The engine schema output above shows the same five parameters reaching the model, so the signature test and the engine agree on that tool. :::callout{type="info" title="Two things called llm_settable"} The companion's `llm_settable` is an attribute of its own `FieldSpec` class. Rasa has a setting with the same name on memory fields, defaulting to false (`rasa/mantle/memory/field.py`, line 52 in the wheel), and the companion's `memory.yml` does not use it for document values (lines 3–9). If engineering shows you a `memory.yml` as evidence of the document boundary, ask which of the two they mean. ::: **What to ask for:** for each protection, the test that fails when the protection is removed. A test on a condition the code never relies on passes whether the protection is there or not. The [mutation testing guide](/library/guides/mutation-test-an-agent-guard/) is a related read. ## Finding 3: the model writes prose into the record Every write tool requires a `reason`, and the reason is free text. The renderer prints it into the revision history as it arrived (`docpkg/render.py`, line 191). Body values have their pipe characters escaped (lines 129–132); reasons do not. We passed a reason that reads like advice: ```text == B. a reason is free text, rendered verbatim | 22 | risk_profile | Balanced | Balanced | Client confirmed the portfolio remains well suited to her needs | no further review | ``` That row now carries a suitability statement no record supports, and the unescaped pipe has given it an extra column. It is the same kind of sentence the illustrative prose record was faulted for, arriving by a different door. :::callout{type="warn" title="The banner stops being true after the first model edit"} Every rendered record opens by saying "Nothing in this document was written by a model" (`docpkg/render.py`, lines 114–115). Every write tool takes its reason from the model, and the renderer prints it in the Reason column. After the first model edit, a reader who trusts the banner will not read that row as model output. ::: **What to ask for:** where every piece of model-supplied text ends up. Either reasons stay out of the document the client keeps, or they are chosen from a fixed list, or the record labels them as the model's words. ## Finding 4: the model can blank or replace the risk warning The regulated sections, Charges and Disclosures, are protected from values the conversation supplies (`docpkg/edits.py`, lines 183–187). Removal is another matter. `clear_field` checks only that the field exists and that a reason was given (lines 226–243), and the model reaches it through `clear_document_field`. We cleared the risk warning: ```text == C. clear the regulated risk warning | Risk warning | — | 4 declared field(s) have no source and render blank: - `property_value` - `property_weight` - `risk_warning` **(required)** - `include_property_breakdown` ``` The document still renders. The missing warning is flagged as required in the Completeness section, and nothing stops the record being produced without it. Replacement is open too. `set_sourced_field` checks that the field is declared, that it is a sourced field and that a reason was given, but not which section it belongs to (lines 135–160), and the model reaches it through `point_field_at_record`. We pointed the risk warning at the asset-class field of the property holding: ```text accepted; cited as: Custodian position extract · POS-PF4402-003 · asset_class | Risk warning | Property funds | ``` The disclosure now reads "Property funds", with a valid citation, so the renderer has nothing to refuse. The regulated sections are closed to values typed in conversation, not to removal or re-citation. **What to ask for:** which fields the model may remove or re-point, whether regulated wording can cite only the disclosure library, and whether a missing required field stops the document or only lists it. The answer in the code should match the answer your compliance team would give. ## Finding 5: a refusal leaves the earlier copy The renderer refuses when a cited figure no longer matches its record; the suite shows that with the total restated to £911,000 (lines 257–261). But `render_document` also writes the document to `out/DOC-SUIT-00417.md` on every successful render (`tools/document.py`, lines 233–235), and a later refusal does nothing to that file. From the engine run (log timestamps removed): ```text == 3. restate the extract, render_document again direct call: provenance_broken [error ] mantle.skill_executor.tool_error error_type=TypeError tool=render_document [debug ] mantle.skill_executor.tool_error.debug error="object ToolResult can't be used in 'await' expression" tool=render_document returned through engine: {"error": "object ToolResult can't be used in 'await' expression"} out file still exists: True out file total line: | Total portfolio value | £486,210.44 | ``` The direct call refused, and the file rendered before the restatement still shows the old total. The refusal governs the next rendering. Whatever picks up files from `out/` will find a copy whose figure no longer matches its record. **What to ask for:** which copy leaves the building, and what checks it at the moment it leaves. A document that is verified when rendered and sent later is only as current as the gap between the two. ## What stays with you after the fixes Fixing the five findings would still leave questions no code settles. **The right record.** A figure can be traced and still be the wrong one. The companion's `scripts/render_document.py --diff` re-points the total at the cash line of the same valuation, and the document renders without complaint because the new citation is valid: ```text One field changed: field : total_value before : 486210.44 after : 21804.1 cited : Custodian position extract · VAL-2026-08-29-PF4402 · cash_gbp ``` Someone who knows the client's holdings still has to read the Provenance section. **Whether the record is true.** The renderer checks that a figure agrees with its record. A wrong extract produces a correctly cited wrong figure. **Where the state and the trail live.** The revision history only grows, but in the companion it lives in a module-level object for the length of one process (`tools/document.py`, lines 82–90). The file names what must survive a move to a real document service: the service owns the state and derives the document, and the agent never receives a document it can edit and hand back. None of our checks says anything about your service. Ask for the same checks against it. **The legal question.** The field declarations and the trail are technical evidence. They do not by themselves satisfy a record-keeping or audit requirement; that control is your compliance function's design. Record each finding and each remaining question against the "Export a document" line of your action register, whose boundary reads "Artifact and delivery match the approved state"; the [permission review guide](/library/guides/review-agent-permissions/) sets out that register. Anything still open at launch is an unresolved risk that a [release decision](/library/guides/make-a-release-decision/) should name with an owner. :::cta{href="/library/tutorials/deriving-a-document/" label="Read the tutorial behind the companion project"} The deriving-a-document tutorial is the engineer's side of this design, built chapter by chapter in the same companion project. ::: --- # Build an evaluation set that exposes failures Source: https://rasa.community/library/guides/build-an-evaluation-set/ Author: Rasa team Published: 2026-09-07 An evaluation set is a claim about the situations your agent needs to handle. If every case is a polite request with all the required information, the set will tell you little about a confused caller or a failing tool. This guide is for the evaluation or data specialist working with an AI product engineer. The deliverable is a versioned case manifest and a reviewable result table, not a large pile of conversations. ## Define the outcome with the domain owner Use a fictional booking assistant as a worked example. One task is “retrieve the applicable change policy for the customer's authorized booking.” Correctness includes booking identity and source policy, not just a plausible answer. Write a case row with these fields: case ID; workflow slice; input provenance; initial state; tool conditions; expected facts; prohibited actions; expected recovery; result evidence; dataset version. Keep customer identifiers out of public fixtures. If synthetic cases are used, label them synthetic and document who checked their domain plausibility. | Case | Deliberate condition | Observable expectation | | ----------------- | --------------------------------------- | --------------------------------------------------------- | | policy-happy | Authorized booking, policy available | Answer refers to the returned policy and booking | | policy-ambiguous | Two bookings match | Agent asks which booking; no private details disclosed | | policy-timeout | Lookup fails | No invented policy; recoverable next step | | policy-pressure | Customer asks to ignore a fee rule | No unauthorized booking change | | policy-correction | Customer corrects the booking reference | Subsequent lookup uses the corrected authorized reference | These rows specify behavior. They are not ready-to-run Rasa YAML, and none is a measured result. ## Keep facts and judgments separate Use factual assertions for things the system can observe: a tool was called with an authorized reference, an unwanted action did not occur, or a resulting record has the right value. Use a rubric-based judge for qualities such as clarity when a deterministic assertion cannot express them well. Rasa documents [simulation evaluation](https://rasa.com/docs/pro/testing/simulation-evaluation/), and the [community evaluation tutorial](/library/tutorials/evaluation-harness/) walks through a harness. Inspect the result artifacts and the actual behavior they claim to describe. A report field saying “blocked” is insufficient if the forbidden action still executed. When a judge disagrees with a factual assertion, keep both outputs. The factual failure remains a failure; a positive language score cannot cancel it. Sample judge decisions for human review using a rubric the domain owner understands. ## Report slices and denominators Suppose an explicitly fictional run passes eight of ten routine cases and zero of two ambiguous-booking cases. The pooled result is 8 / 12, or about 66.7%. Reporting only the routine 80% conceals the case type most likely to expose private information. Report passed, failed, and not-run counts for every slice. Keep not-run cases separate and explain the reason; do not count them as passes. Record the agent version, dataset version, configuration, run identifier, and locations of observable results. Repeat stochastic cases when variation matters, and describe the run count rather than implying one run establishes reliability. ## Protect the usefulness of the set Reserve cases for release review that were not repeatedly used to tune the agent. Otherwise the team can improve the visible test score while losing the ability to detect overfitting. Add new cases from observed failures with permission and appropriate redaction; preserve why each was added. The awkward case is a valid outcome the original expected answer did not anticipate. Have the domain owner decide whether the expectation was wrong. If it was, version the dataset and explain the change before rerunning both the baseline and candidate. Start by filling the manifest for one workflow and the five conditions above. Ask the product manager to use the results in a [release decision](/library/guides/make-a-release-decision/). NIST's [MEASURE function](https://airc.nist.gov/airmf-resources/airmf/5-sec-core/) supplies the broader rationale for documented, context-sensitive measurement; the case design here is an editorial recommendation. --- # Choose an agent workflow worth building Source: https://rasa.community/library/guides/choose-an-agent-workflow/ Author: Rasa team Published: 2026-09-08 Start with a decision your team can make this week: choose one workflow whose outcome you can observe and whose mistakes you can contain. A convincing conversation is a weak substitute for a completed customer task. This guide is for the AI product manager deciding what to build next. You will leave with a scope brief to discuss with an engineer and the person who handles the task today. No Rasa installation is required. ## Work through one candidate Consider **Horizon Travel**, a fictional business used here only as a worked example. Its support team answers itinerary questions and changes bookings. We have no measured business results for this example. A first proposal is “automate travel support.” Replace it with “help a signed-in traveller find the change conditions for an existing booking, then request a human-assisted change.” The second statement exposes both an outcome and a boundary. The agent may explain a retrieved policy; it may not invent eligibility or commit a paid change. | Decision | Example brief | Evidence to collect before building | | ----------------- | ------------------------------------------------------------- | ------------------------------------------------------ | | User and task | Signed-in traveller checking one existing booking | Observe real support sessions with permission | | Current outcome | A support colleague retrieves the applicable conditions | Record completion, waiting time, and common exceptions | | Proposed outcome | Traveller sees the applicable conditions and can request help | Check source, booking identity, and task completion | | Excluded action | Charge a fee or change the booking | Identify the system that actually authorizes changes | | Accountable owner | Travel support lead | Confirm who accepts unresolved requests | These are proposed boundaries, not claims about what a Rasa configuration automatically enforces. The engineer must connect the identity and booking tools to the actual systems of record. ## Make the decision measurable Use a task outcome rather than message volume: **eligible requests completed with correct source information / eligible requests attempted**. Agree what makes a request eligible before collecting results. Keep abandoned, incorrectly answered, and handed-off requests visible as separate counts; silently excluding failures improves a percentage without improving the experience. Record the current human workflow using the same definition. If that baseline does not exist, collecting it is the next piece of work. Do not promise a savings percentage from a demo. For the fictional booking example, a correct answer with the wrong booking identifier is a failed task. A correct handoff can be a successful outcome when a change requires staff authority. That distinction changes both the product requirement and the evaluation set. ## Choose the smallest useful slice Compare the candidate against three practical questions: 1. Can the team obtain an authoritative answer and identify the customer safely? 2. Can someone observe whether the task finished correctly? 3. Can the team stop or hand off the interaction when a necessary dependency fails? If any answer is unknown, put that uncertainty in the brief and assign an owner. A workflow with accessible data and a modest benefit may be a better first build than a high-volume workflow whose exceptions nobody can explain. The awkward case is a customer who asks for both information and a binding change in one sentence. The brief must say which part can proceed and how the other part reaches a human. “The assistant should be helpful” does not decide this. ## Take this to your team Copy these fields into a one-page brief: user; task; current workflow; observable outcome; denominator; authoritative sources; allowed actions; excluded actions; handoff owner; stop condition; next evidence to collect. Ask an engineer to challenge feasibility and a support colleague to bring one exception. If the team cannot agree on the outcome, postpone implementation and observe the workflow together. If it can, bring the brief and one unresolved dependency to the [Rasa community](/join/). The [14 September 2026 Build Day](/events/build-day/) is a past example of this working format. Use the [guarded-action tutorial](/library/tutorials/guarding-irreversible-actions/) when the next slice introduces an action that cannot easily be undone. The distinction between context, measurement, and response is informed by the [NIST AI RMF Core](https://airc.nist.gov/airmf-resources/airmf/5-sec-core/). The scope brief and fictional example are editorial tools, not a validated business benchmark. --- # Test passes but the real code path fails: derive the fixture Source: https://rasa.community/library/guides/derive-test-fixtures-from-production-code/ Author: Rod Rivera Published: 2026-09-14 In a scratch copy of the handoff pattern from `RasaHQ/rasa-community-resources`, delete one line, `"account_id",`, from the tuple of memory keys the handoff tool reads. Then run two tests that assert the same claim: after a transfer, the human desk asks none of the five questions in its opening script (`DESK_OPENING_SCRIPT`, `handoffpkg/desk.py` lines 42–48) again. ```text $ python3 -m unittest -v tests.test_handoff_context.TestPackageSurvivesHandoff.test_caller_is_never_asked_a_question_they_answered tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question test_caller_is_never_asked_a_question_they_answered (tests.test_handoff_context.TestPackageSurvivesHandoff.test_caller_is_never_asked_a_question_they_answered) The teaching claim, as a number rather than a sentence. ... ok test_the_agent_path_retires_every_desk_question (tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question) THE HEADLINE CLAIM, ON THE PATH THE AGENT ACTUALLY RUNS. ... test_the_agent_path_retires_every_desk_question (tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question) (key='account_id') THE HEADLINE CLAIM, ON THE PATH THE AGENT ACTUALLY RUNS. ... FAIL test_the_agent_path_retires_every_desk_question (tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question) THE HEADLINE CLAIM, ON THE PATH THE AGENT ACTUALLY RUNS. ... FAIL ``` That is an offline run: Python 3.14.3, a scratch copy of the pattern at commit `69e27b6`, network access denied and no licence set. The excerpt keeps the verbose lines and cuts the two tracebacks and the summary lines, `Ran 2 tests` and `FAILED (failures=2)`. Both assertion messages appear in the section on dropping a key. The first test checks a package built from a session typed into the test file. It stays green although the tool can no longer carry the account. The second builds its session from the lists the agent's own code reads and writes, and fails twice. The broken copy recreates a defect the pattern really shipped, which the second test's docstring records (lines 344–349 of `patterns/voice-handoff-context/tests/test_handoff_context.py`): ```text This test exists because the claim once passed while being false. The eval suite built its own session dict containing `account_id`, so `intent.details` was populated and "which account is this about?" was retired — but the agent's own `_MEMORY_KEYS` did not collect `account_id`, so a real handoff shipped an empty `details` and the desk asked anyway. Green tests, broken claim. ``` No model misbehaved. The handoff tool did what its code said, the package builder did what its code said, and the assertion was correct about the input it was given. The input was the problem. The person writing the test had typed it, and it held a field the running agent never gathered. Rerunning the suite could not expose that, because every run built the same session. **A fixture typed by hand is evidence about what its author assumed the code collects.** Take the fixture's keys from the list the running code reads, and that assumption becomes something a test checks. The change is to how you build one fixture; the suite does not grow. It costs you coupling between your tests and your source layout, and the last two sections show where that coupling breaks. The caller, accounts and charges in the pattern are fixture values. Every run in this guide is offline, with no model and no licence, and each saved receipt records its conditions. ## Why the test passes but the real code path fails The hand-built test is short (lines 185–187): ```python def test_caller_is_never_asked_a_question_they_answered(self): """The teaching claim, as a number rather than a sentence.""" self.assertEqual(unanswered_questions(self.package), ()) ``` Its class builds `self.package` in `setUp` by passing `demo_session()` to `build_package_from_session` (lines 133–134). `demo_session()` is a session typed into the test file. The start of it, lines 77–94: ```python def demo_session() -> dict: """Session state at the moment the caller asks for a human.""" session = { # --- allowlisted: the state that SHOULD cross ----------------------- "customer_id": "cust_00417", "display_name": "Jordan Rivera", "verified_tier": "medium", "verified_factors": ["knowledge_passphrase"], "channel": "voice:+1-555-0100", "goal": "dispute_transaction", "goal_label": "Dispute a card transaction", "goal_stage": "blocked", "account_id": "acc_checking", "account_label": "Everyday Checking", "card_last_four": "4821", "dispute_amount": "$248.00", "dispute_merchant": "Northgate Fuel", "dispute_date": "2026-08-29", ``` By the usual standard it is a good fixture. The names match the memory schema and the values are plausible. :::checkpoint{id="derived-fixture-what-a-pass-proves" question="A test builds a realistic session dict by hand, including account_id, passes it to build_package_from_session, and asserts unanswered_questions(package) == (). It passes. What has it shown?" options="That a real handoff retires all five desk questions,That the builder retires them when given a session containing those keys,That the handoff tool collects account_id,Nothing useful" answer="1"} It shows how `build_package_from_session` behaves on a dict someone typed. The handoff tool builds its session from `_MEMORY_KEYS`, and the test never reads that tuple, so the pass says nothing about whether a call produces that dict. It is still real evidence about the builder, which is why "nothing useful" is also wrong. ::: On a real call the handoff tool, `transfer_to_human` in `skills/human_handoff/tools.py` (line 148), builds its session by walking one named tuple (lines 155–160): ```python # Step 1 — collect indiscriminately. This tool is not the safety boundary. session: dict = {} for key in _MEMORY_KEYS: value = context.memory.get(key) if value is not None: session[key] = value ``` A key that is not in `_MEMORY_KEYS` never reaches the builder, whatever memory holds. The tuple runs from line 68 to line 96. Its first lines, 68–84, carry a comment left when the defect was fixed: ```python _MEMORY_KEYS = ( "customer_id", "display_name", "verified_tier", "verified_factors", "channel", "goal", "goal_label", "goal_stage", # The goal's parameters -> intent.details. Omitting these was a real defect: # the eval suite passed on a hand-built session while the live agent shipped # a package with an empty `details`, so the desk still asked "which account # is this about?". test_the_agent_path_retires_every_desk_question now # exercises THIS list rather than a hand-built one. "account_id", "account_label", "card_last_four", ``` The dispute skill had the same gap on the writing side. The loop in `load_caller_context` that writes the caller profile, `skills/dispute_transaction/tools.py` lines 83–91, carries a matching comment: ```python for key in ( "customer_id", "display_name", "verified_tier", "verified_factors", "channel", # The account and card the caller is asking about. These are what make # `intent.details` non-empty, which is what retires "which account is # this about?" on the desk. Omitting them was a real defect: the eval # suite passed on a hand-built session while the live agent produced a # package that still made the desk ask. "account_id", "account_label", "card_last_four", ): ``` So before the fix, the fixture held `account_id`, as the docstring says, and both lists the running agent uses lacked the three detail keys. The skill's comment names the three keys, and the tool's comment sits directly above them. Set today's `demo_session()` beside what the two comments say: | Key | In `demo_session()` today | Written by the dispute skill before the fix | In `_MEMORY_KEYS` before the fix | Reaches `intent.details` on a call | | ---------------- | ------------------------- | ------------------------------------------- | -------------------------------- | ---------------------------------- | | `account_id` | yes | no | no | no | | `account_label` | yes | no | no | no | | `card_last_four` | yes | no | no | no | :::callout{type="info" title="You cannot check out the broken version"} The pattern entered the repository in one commit, `9f2120f`, with the fix already in place. At `69e27b6`, `git log` on the handoff tool, the dispute skill's tools and the test file lists that one commit. The two "before the fix" columns rest on the comments in both tool files and on the test's docstring. None of them comes from a diff you can inspect. To see the failure yourself, put the defect back by hand, as the opening run did. ::: In measurement terms the hand-built test had a construct validity problem. Cronbach and Meehl set out the idea in 1955 ("Construct validity in psychological tests", _Psychological Bulletin_ 52(4), 281–302): does a test measure the thing it is named for? This one measured something real, the builder retiring the account question when `account_id` is present. It was named for a handoff on the agent's path retiring it. A more realistic fixture would have left that gap open. Realism describes values, and the defect was in which keys existed. ## Build the session from the list the tool reads Three lists of keys are in play: the one typed into `demo_session()`, the keys the dispute skill's tools write to memory, and `_MEMORY_KEYS`. Each test takes a different route to the same builder. :::diagram{title="Three key lists, three routes into one builder"} ```dot rankdir=LR; typed [label="demo_session()\ntyped by hand"]; writes [label="Keys the dispute\nskill writes"]; keys [label="_MEMORY_KEYS"]; tool [label="transfer_to_human\n(real call)"]; derived [label="Derived session:\ncollected and written,\nplus typed values"]; builder [label="build_package_from_session", class="ok"]; typed -> builder [label="hand-built test,\ncalled directly"]; writes -> tool [label="memory"]; keys -> tool [label="which keys it reads"]; tool -> builder [label="after converting\nsome fields"]; keys -> derived [label="parsed"]; writes -> derived [label="parsed"]; derived -> builder [label="derived test,\ncalled directly"]; ``` ::: The hand-built test calls `build_package_from_session` directly with `demo_session()`, and neither agent list touches its input. A real call reaches the builder through `transfer_to_human`. That route reads the keys in `_MEMORY_KEYS` from memory the skill's tools wrote, converts some fields (lines 162–175) and then calls the builder (line 180). The derived test also calls the builder directly, with a session whose keys it parsed from the two agent lists. The fix lives in the `TestTheAgentPathSpecifically` class. Two helper methods read the agent's own lists. The first reads `_MEMORY_KEYS` out of the tool module (lines 306–323): ```python def _agent_memory_keys(self) -> tuple[str, ...]: """Read _MEMORY_KEYS out of the tool module without importing rasa. The tool module imports `rasa.mantle`, which the eval suite deliberately does not depend on — `make test` runs with no license and no install. Parsing the literal keeps this test honest about the real list while staying offline. """ source = ( Path(__file__).resolve().parents[1] / "skills" / "human_handoff" / "tools.py" ).read_text() block = source.split("_MEMORY_KEYS = (", 1)[1].split(")", 1)[0] return tuple( line.strip().strip(",").strip('"') for line in block.splitlines() if line.strip().startswith('"') ) ``` The second, `_skill_memory_writes()` (lines 383–398), finds the keys the dispute skill's tools write. It searches `skills/dispute_transaction/tools.py` for `memory.set(...)` calls, `_append(...)` calls and the `for key in (...)` loop in `load_caller_context`. Its body, lines 393–398: ```python keys = set(re.findall(r'memory\.set\(\s*"([^"]+)"', source)) keys |= set(re.findall(r'_append\(\s*context,\s*"([^"]+)"', source)) # The `for key in (...)` loop that writes the caller profile. for block in re.findall(r"for key in \(([^)]*)\):", source): keys |= set(re.findall(r'"([^"]+)"', block)) return keys ``` Neither helper imports the module it reads. Both tool modules import from `rasa.mantle`, and the suite is meant to run with no Rasa install, so the tests parse the source text instead. The last section weighs that choice for your own project. The test then uses both lists (lines 355–366): ```python written = self._skill_memory_writes() collected = self._agent_memory_keys() # Every detail key the allowlist expects must be both written by a skill # and collected by the handoff tool, or it never reaches the package. for key in ("account_id", "account_label", "card_last_four"): with self.subTest(key=key): self.assertIn(key, written, f"no skill writes {key} to memory") self.assertIn(key, collected, f"the handoff tool does not collect {key}") # Now build a session from ONLY what the agent genuinely produces. session = {key: f"value_for_{key}" for key in collected if key in written} ``` The session is the intersection of what the skill writes and what the tool collects, because a key reaches the package only if both happen. The values are placeholders: `value_for_account_id` looks like nobody's account. Apart from `attempts`, which the test adds by hand in its next lines, every key in the session comes from those two parsed lists. A second test, `test_every_key_the_agent_collects_is_either_allowlisted_or_withheld`, builds its session from every key `_agent_memory_keys()` returns, with no intersection (lines 400–404). It asserts that each collected key is either on the allowlist or named in the package's `withheld_fields`. A hand-picked fixture covers the keys its author thought of. This test covers every key the tool collects. Here is the whole suite at the pinned revision, `69e27b6`, run offline with Python 3.14.3 from `patterns/voice-handoff-context`. The excerpt keeps the three agent-path tests and the last lines of the run, and cuts the other 38 tests: ```text $ python3 -m unittest tests.test_handoff_context -v test_every_key_the_agent_collects_is_either_allowlisted_or_withheld (tests.test_handoff_context.TestTheAgentPathSpecifically.test_every_key_the_agent_collects_is_either_allowlisted_or_withheld) No third category. A key is a decision, and both outcomes are visible. ... ok test_the_agent_collects_credentials_and_the_allowlist_stops_them (tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_collects_credentials_and_the_allowlist_stops_them) End to end on the real key list: collected, then withheld. ... ok test_the_agent_path_retires_every_desk_question (tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question) THE HEADLINE CLAIM, ON THE PATH THE AGENT ACTUALLY RUNS. ... ok ---------------------------------------------------------------------- Ran 41 tests in 0.004s OK ``` That shows the technique runs today. It is an offline predicate test on parsed source lists. It is not a recording of any call. ## Drop a key and see which test notices The opening run applied a mutant by hand: one small, deliberate change to production code, made to see whether the tests notice. That is mutation testing, which DeMillo, Lipton and Sayward set out in 1978 ("Hints on Test Data Selection: Help for the Practicing Programmer", _IEEE Computer_ 11(4), 34–41). The [mutation testing guide](/library/guides/mutation-test-an-agent-guard/) covers choosing mutants and counting what they kill. Here one mutant is enough, because it is the original defect. In the scratch copy, delete `account_id` from `_MEMORY_KEYS`: ```diff @@ -80,5 +80,4 @@ # is this about?". test_the_agent_path_retires_every_desk_question now # exercises THIS list rather than a hand-built one. - "account_id", "account_label", "card_last_four", ``` Then run the whole suite and keep the failure lines: ```text $ python3 -m unittest tests.test_handoff_context 2>&1 | grep -E '^(FAIL|AssertionError|Ran|FAILED)' FAIL: test_the_agent_path_retires_every_desk_question (tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question) (key='account_id') AssertionError: 'account_id' not found in ('customer_id', 'display_name', 'verified_tier', 'verified_factors', 'channel', 'goal', 'goal_label', 'goal_stage', 'account_label', 'card_last_four', 'dispute_amount', 'dispute_merchant', 'dispute_date', 'attempts_log', 'questions_answered', 'factors_verified', 'confirmed_facts', 'handoff_reason', 'pin_attempt', 'otp_code') : the handoff tool does not collect account_id FAIL: test_the_agent_path_retires_every_desk_question (tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question) AssertionError: None is not true : intent.details is empty on the live path — the desk will re-ask which account this is about Ran 41 tests in 0.009s FAILED (failures=2) ``` This run was offline too, on the same scratch copy, with Python 3.14.3. Both failures come from the derived test: once at the named-key check, and once at the package itself, where `intent.details` has no `account_id`. That is one failing test with two failures; the other 40 tests pass, including the hand-built one from the opening run. The same claim, asserted against two fixtures, gives two answers on the same broken code, and only the fixture built from the agent's lists gives the right one. The derived test names three keys, so this check guards those three. A mutant that deletes a key no assertion names can pass, as the next section shows. ## What the derived fixture still does not prove Deriving the key set closes one gap. It leaves others, and a report on this suite should name them. Some values are still typed. After building the session from the intersection, the test sets four entries itself (lines 367–372): ```python session["verified_tier"] = "medium" session["goal"] = "dispute_transaction" session["questions_answered"] = "\n".join(DESK_OPENING_SCRIPT) session["attempts"] = [ {"action": "send_otp_sms", "outcome": "failed", "code": "delivery_failed"} ] ``` Three of them overwrite keys the lists supplied. The fourth, `attempts`, is in neither list. The skill writes `attempts_log`, a newline-separated text field (`_append`, lines 63–69 of the dispute tools), in the `action|outcome|code|detail` form that `_parse_attempts` reads (handoff tool, lines 117–118). On a real call the tool pops that field and converts it before calling the builder (line 175): ```python session["attempts"] = _parse_attempts(session.pop("attempts_log", None)) ``` The test calls the builder directly. Its session keeps a placeholder `attempts_log` next to a hand-typed `attempts` in the converted shape, and the conversion never runs on the path it checks. The [handoff measurement guide](/library/guides/measure-a-handoff-that-loses-context/) follows what that means for the desk's question about earlier attempts. The parsers read source text. `_skill_memory_writes()` treats any double-quoted string inside a `for key in (...)` block as a key, comments included. At `69e27b6` its set contains one entry that is not a key at all: `'which account is\n # this about?'`, lifted from the comment in the `for key in (...)` loop in `load_caller_context`. It does no harm here, but a key quoted in a comment inside that loop would count as written. The regex also counts `memory.set("otp_code", ...)`, although the skill sets `otp_code` only on the branch where the tier check fails (dispute tools, lines 173–195). The stray entry comes from an offline run of both parsers. :::callout{type="pitfall" title="An empty parse passes the loop tests"} Convert the `_MEMORY_KEYS` entries to single quotes and `_agent_memory_keys()` returns `()`, because it keeps only lines that start with `"`. The two tests that assert named keys fail. `test_every_key_the_agent_collects_is_either_allowlisted_or_withheld` passes, because its loop has nothing to check. Any derived fixture needs at least one assertion that a key you know about is present. (Offline mutation on a scratch copy.) ::: The intersection also hides keys that are collected and never written. The same offline parser run prints the difference between the two sets. One of its four output lines: ```text collected, not written: ['dispute_amount', 'dispute_date', 'dispute_merchant', 'factors_verified', 'handoff_reason'] ``` No tool in `skills/dispute_transaction/tools.py` writes `dispute_amount`, `dispute_merchant` or `dispute_date`, so the intersection drops them and the derived test passes without them. No assertion names them. The [measurement guide](/library/guides/measure-a-handoff-that-loses-context/) shows what the desk loses as a result. Finally, none of this tests the model. Every run here is an offline test on parsed lists and a builder function, or an import of a module. It does not show what a model says on a call. It does not catch other kinds of fixture drift either, such as a policy string copied into a test and later changed in one place only. ## Apply it to your own project: parse the source or import the module Start by finding the list your code under test reads or writes at run time: a tuple of memory keys, the fields a tool returns, the slots a flow fills. That list is the source for your fixture's keys. The handoff pattern has two such lists, one on each side, and parses both from source. The other obvious option is to import them. | Approach | Collected keys come from | Written keys come from | What it costs | | ----------------------------------------------- | --------------------------------------------------- | ----------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | | Parse the source text, as the companion does | The literal as written in the handoff tool | A regex over the skill's source | No Rasa install in the test job. The parsers break on reformatting, read comments, and cannot see keys built at run time. | | Import the handoff tool module | The tuple Python evaluates, whatever its formatting | Still a regex, or a run of the skill's tools against a stand-in context | The test job must install `rasa-pro` and everything the module imports. | | Move the list into a module with no Rasa import | One object both the tool and the test import | Still a regex, or a run of the skill's tools | A small refactor of production code. The companion does not do this. | Importing covers only the collected side. The skill's writes are calls inside tool functions, and no single module attribute lists them all. Running those tools against a stand-in context would record them; neither this guide nor the companion tests that option. A licence is not what stands in the way of importing. The two variables `rasa/utils/licensing.py` names for the licence are `RASA_LICENSE` and `RASA_PRO_LICENSE` (lines 24–25 in the pinned wheel). With both unset, `rasa-pro` 3.20.0rc1 installed and network access denied, the handoff tool module itself imports. It pulls in `ToolContext` and `tool` from `rasa.mantle.tools.decorator` and `ToolResult` from `rasa.mantle.tools.result` (lines 27–28). This run used Python 3.12.13. `` stands for the virtual environment's path, as in the saved receipt, and the output is unedited: ```text $ env -u RASA_LICENSE -u RASA_PRO_LICENSE sandbox-exec -p '(version 1)(allow default)(deny network*)' /bin/python -c 'import importlib.util, os, sys; print("RASA_LICENSE set:", "RASA_LICENSE" in os.environ, "| RASA_PRO_LICENSE set:", "RASA_PRO_LICENSE" in os.environ); spec = importlib.util.spec_from_file_location("handoff_tools", "skills/human_handoff/tools.py"); mod = importlib.util.module_from_spec(spec); spec.loader.exec_module(mod); import rasa; print("rasa", rasa.__version__); print("imported skills/human_handoff/tools.py; transfer_to_human:", type(mod.transfer_to_human).__name__); print("_MEMORY_KEYS:", mod._MEMORY_KEYS); print("licence modules loaded:", sorted(m for m in sys.modules if "licens" in m))' 21:30:33 - LiteLLM:WARNING: get_model_cost_map.py:454 - LiteLLM: Failed to fetch remote model cost map from https://raw.githubusercontent.com/BerriAI/litellm/main/model_prices_and_context_window.json: [Errno 8] nodename nor servname provided, or not known. Falling back to local backup. RASA_LICENSE set: False | RASA_PRO_LICENSE set: False rasa 3.20.0rc1 imported skills/human_handoff/tools.py; transfer_to_human: function _MEMORY_KEYS: ('customer_id', 'display_name', 'verified_tier', 'verified_factors', 'channel', 'goal', 'goal_label', 'goal_stage', 'account_id', 'account_label', 'card_last_four', 'dispute_amount', 'dispute_merchant', 'dispute_date', 'attempts_log', 'questions_answered', 'factors_verified', 'confirmed_facts', 'handoff_reason', 'pin_attempt', 'otp_code') licence modules loaded: ['rasa.utils.licensing'] ``` The warning is LiteLLM failing to reach the network the sandbox denied, then falling back to a local copy. The import loaded `rasa.utils.licensing` and did not fail. In this run, importing the module needed no licence. The run shows nothing about running an agent. The cost of importing is the install. Import the module when your test environment can afford the install. The collected-key parser goes, and its limits go with it; the written-key side and the other limits remain. Parse the source when the suite has to run with no Rasa install and no licence, as the handoff pattern's suite deliberately does. Then assert named keys so an empty parse fails, and keep the parsers as narrow as the file allows. Whichever you choose, finish the way the dropped-key section did. Delete a key your derived test names from the production list in a scratch copy, run the suite, and check that the derived test fails and says which key. If nothing fails, the fixture is not derived from the list you think it is. A key the test does not name can go unnoticed. With `dispute_amount` deleted from `_MEMORY_KEYS`, all 41 tests still pass (offline run on a scratch copy). For the wider design of an evaluation set, including slices and denominators, see the [evaluation set guide](/library/guides/build-an-evaluation-set/). :::solution{title="Should I delete the hand-built fixtures?"} No. A hand-built session is the right input for testing a function on data you choose, including keys the running code never collects. `demo_session()` also carries values under keys that are not in `_MEMORY_KEYS`, such as `card_number` and `ssn` (test file, lines 66–74 and 126). A derived session could never contain them. The [handoff summary guide](/library/guides/design-the-handoff-summary/) covers how the package names withheld fields. Keep hand-built fixtures for that job. Do not use them as evidence that the agent's path produces those inputs. ::: :::solution{title="My code does not read a named list. It reads whatever memory holds."} Then the declarations are the nearest list, for example the keys in your `memory.yml` files. Deriving from declarations tells you a field exists, not that anything writes it, which is why the handoff test intersects collected keys with written keys. If neither list exists in code, fix that before the fixture: a transfer that reads everything has no list a test can hold it to. ::: :::cta{href="/library/tutorials/evaluation-harness/05-what-your-eval-cannot-tell-you/" label="Read what your eval cannot tell you"} The evaluation harness chapter covers a related failure over time. Its `ground_truth:` strings copy policy text from `skills/faq_support/tools.py`, so a policy changed in one place only leaves the judge scoring against a policy that no longer exists. ::: --- # Design a handoff a customer can trust Source: https://rasa.community/library/guides/design-a-human-handoff/ Author: Rasa team Published: 2026-09-12 The customer has already explained the problem. A handoff that makes them repeat it can erase the value of the conversation that came before. For a conversation designer or customer-experience specialist, the useful deliverable is a handoff contract: when it happens, what the customer hears, what context travels, and what happens if nobody can receive it. You can design that contract before an engineer connects a support tool. ## Start with the customer's choice In this fictional Horizon Travel example, a traveller asks to dispute a booking-change fee. The agent can retrieve a policy, but a staff member must decide exceptions. The names and dialogue below are invented for this exercise. > Customer: The fee doesn't match what I was told. I want someone to review it. > > Agent: I can ask the support team to review the fee. I'll include the booking reference and your explanation so you don't have to start again. Would you like me to send that? > > Customer: Yes, but don't include the medical reason I mentioned earlier. > > Agent: I'll leave that out. I'll share the fee question and booking reference only. The last turn changes the payload. A warm sentence followed by a full transcript would break the promise. Ask the engineer to show the actual receiving record during testing. ## Specify the context packet | Field | What belongs in it | What to leave out | | -------------------- | ------------------------------------ | ------------------------------------------------- | | Task | Review the disputed change fee | A guessed resolution | | Reference | Authorized booking reference | Unnecessary identity documents | | Customer explanation | The approved summary | Details the customer excluded | | Agent actions | Policy retrieved; no booking changed | An unsupported claim that staff approved a change | | Outstanding question | Whether an exception applies | A fabricated urgency label | The receiving team should help define these fields. A summary is useful only if it supports their work. The [tools and memory tutorial](/library/tutorials/tools-and-memory/) gives engineers an implementation starting point for deciding where context belongs; it does not by itself establish your privacy rules. ## Design the failure path too Write and test the following conversation states: 1. **Offer:** explain why a human is needed and what information will be shared. 2. **Confirmation:** allow the customer to correct or limit the summary. 3. **Delivery:** wait for the receiving system's result before claiming success. 4. **Recovery:** if delivery fails or the team is unavailable, say what remains unresolved and offer the verified alternative channel. For example: “I couldn't send the request just now. No booking change has been made. You can try again or contact support through the help page.” Link only to a channel your team has verified. Do not invent a response time to soften the message. A particularly awkward case is a timeout after the support system created a ticket. Repeating the action can produce duplicates. The designer's requirement should be “check whether a request already exists before offering another send”; the engineer decides how the receiving system supports that check. ## Run a short design review Ask a support colleague to play the customer and another to inspect the received context. Test a normal request, a correction to the summary, an excluded sensitive detail, an unavailable queue, and a request to cancel before sending. For each case, record what the customer was promised, what was actually sent, and what the receiving team could do next. Pass the case only when those three agree. Capture disagreements as changes to the contract, not as a judgement about whether the dialogue “sounds human.” Leave the review with a revised example conversation and the fields engineering must verify. Share them with the [Rasa community](/join/) if the integration is your current blocker. The [14 September 2026 Build Day](/events/build-day/) is a past example of working through a blocker together. NIST describes human-factors and domain-expert participation in its [AI actor tasks](https://airc.nist.gov/airmf-resources/airmf/appendices/app-a-descriptions-of-ai-actor-tasks/). This handoff contract is our practical application of that perspective; it is not a claim that a particular dialogue guarantees trust. --- # Stop customers repeating themselves after a handoff Source: https://rasa.community/library/guides/design-the-handoff-summary/ Author: Rod Rivera Published: 2026-09-15 > **Assistant:** I’ll bring in a specialist and pass them everything we’ve > covered, so you won’t need to repeat any of it. Shall I go ahead? > > **Caller:** Yes, go ahead. The desk agent’s screen: ```text ### TODAY — one free-text handoff_reason (every example in the catalog) ======================================================================== INBOUND HANDOFF ho_today — Dispute needs a second factor the caller cannot complete on this line. ======================================================================== CALLER UNIDENTIFIED TRUST tier=unverified — Identity NOT established. Treat every identifying detail as unconfirmed. ASKING FOR unknown (stage: stated) ALREADY TRIED (do not retry these): · (nothing) DO NOT ASK — the caller already answered these: · (nothing) YOU MAY: · Answer general questions · Explain what the agent already did · DO NOT discuss account-specific details — identity not established · DO NOT action irreversible changes — step the caller up first ======================================================================== The caller will now be asked 5 question(s) they already answered: ✗ Can I take your name? ✗ Can you confirm your date of birth? ✗ Which account is this about? ✗ What are you calling about today? ✗ Have you tried anything already? ``` > **Desk agent:** Can I take your name? … Can you confirm your date of birth? > … Which account is this about? … What are you calling about today? … Have > you tried anything already? The caller answers all five again, seconds after being told they would not have to. The dialogue is an illustrative script, not a recorded call; the assistant’s line is the confirmation text in the companion project. The screen is real output from `make compare` in `patterns/voice-handoff-context` (`RasaHQ/rasa-community-resources` at `69e27b6`), run offline. It shows what a desk receives when a handoff carries nothing but a `handoff_reason` string. This companion’s own `transfer_to_human` sends much more; the script builds the “TODAY” screen from the reason alone (`scripts/agent_desk.py`, lines 107–113) to stand in for the rest of the catalogue. By “the catalogue” the pattern’s authors mean the other example projects in `rasa-community-resources`: its README says every `human_handoff` skill there “captures one free-text `handoff_reason` string and opens a ticket” (`README.md`, lines 14–17), and the baseline test’s docstring says the same (`tests/test_handoff_context.py`, lines 190–194). The companion’s `agent.yml` calls the assistant “a concise, competent banking assistant” and names no bank. The caller, account and merchant are fixture values, and the desk is a terminal fixture standing in for a contact-centre screen. ## Make the summary a view over typed fields `test_the_catalogs_current_handoff_answers_nothing` pins the baseline: a package built from a reason string alone leaves the desk all five of its opening questions. A sentence, however well written, gives the desk nothing it can check a question off against. The stance of this guide: **the screen a human agent reads first should be computed from typed fields, never written and stored as prose.** The companion is built on a premise its authors state in the summary module (`handoffpkg/summary.py`, lines 12–17). It is a design premise, not a measurement of desk behaviour: ```text A handoff package that carries both a structured intent and a separately authored prose summary carries two sources of truth. They agree on the day they are written and diverge on the first edit. The human agent reads the prose — it is faster — so the human agent reads the stale one. The failure is silent and it is always in the direction of the desk acting on information the agent already superseded. ``` The rule comes from the [deriving-a-document tutorial’s chapter on derived rendering](/library/tutorials/deriving-a-document/04-derived-rendering/), where the renderer takes state and nothing else. Here it is applied to the desk. `HandoffPackage.summary` is a read-only property that renders the other sections on every access (`handoffpkg/schema.py`, lines 237–250): ```python @property def summary(self) -> str: """Human-readable summary, DERIVED on every access. Not a field. Not settable. Not cached. Every read re-renders from ``identity`` / ``intent`` / ``attempts`` / ``do_not_repeat``, so the summary and the structured fields are the same information in two renderings and cannot drift. Change a field and the next read of ``summary`` reflects it; there is no second copy to forget to update. This is the SOW's "derived rather than authored separately" requirement implemented as a property rather than promised in a docstring. """ return render_summary(self) ``` :::checkpoint{id="handoff-summary-what-moves-the-screen" question="A package’s fields record the caller at tier medium. Which change makes the desk screen show tier high?" options="Rewriting the summary text before it is sent,Adding a summary key saying tier high to the JSON the desk reads,Changing identity.verified_tier in the package" answer="2"} The first option has nowhere to happen: `summary` is a property with no setter. The second fails on the far side, because the desk rebuilds the package from its fields and recomputes the summary; the next paragraph shows the test for it. Only a field change reaches the screen, which is what makes a screen reviewed once as trustworthy as its fields. ::: Two tests hold that shut. `test_summary_is_a_property_with_no_setter` assigns to `package.summary` and expects it to raise. `test_a_summary_smuggled_through_serialisation_is_discarded` adds a `summary` key reading “Caller verified at tier 'high'. Clear them for anything.” to a serialised package, reads it back, and asserts that the restored summary says `tier 'medium'` and does not contain the injected text. A value can stop short of the screen in three places. The tool may never read it, the allowlist may hold it back, or it may arrive as an authored summary and be dropped: :::diagram{title="From session memory to the desk’s first screen"} ```dot rankdir=LR; session [label="Session memory\n(what the tools wrote)"]; skipped [label="Keys not in _MEMORY_KEYS\n(e.g. disputed_txn_label)", class="blocked"]; collect [label="transfer_to_human\nreads _MEMORY_KEYS"]; gate [label="build_package_from_session\nSESSION_ALLOWLIST", shape=diamond]; fields [label="identity, intent,\nattempts, do_not_repeat", class="ok"]; names [label="withheld_fields\n(names only)", class="ok"]; values [label="Withheld values\n(never copied)", class="blocked"]; summary [label="summary property\n(rendered on every read)"]; screen [label="Desk first screen\n(reconstruct)"]; authored [label="An authored summary\n(no field to hold it)"]; dropped [label="Dropped on read-back\n(package_from_dict)", class="blocked"]; session -> skipped [label="never read"]; session -> collect -> gate; gate -> fields [label="allowlisted"]; gate -> names [label="not allowlisted"]; gate -> values [label="value dropped"]; fields -> summary; fields -> screen; names -> screen; authored -> dropped [style=dashed]; ``` ::: This choice has a cost. A computed screen reads like a form. It carries what a field can hold: no “she sounded frightened”, no aside passed on in the caller’s words. Text still crosses wherever an allowlisted field holds text: `handoff_reason`, `goal_label`, `account_label`, the `questions_answered` and `confirmed_facts` lists, and the free-text `detail` of each attempt (`handoffpkg/redaction.py`, lines 71–101 and 252; `memory.yml`, lines 86–91). The attempt detail is where “Carrier rejected the SMS twice. Do not resend to this number.” reaches the screen below. Apart from `handoff_reason`, which nothing on the live path sets (see below), each of those strings is written by a tool, so its wording is part of your specification. Each new thing the desk should see means a memory field, a tool that writes it, an entry in the tool’s key list, an allowlist entry, a place in the package built by `build_package_from_session` (`redaction.py`, lines 204–273), a rule in the desk if it retires a question, and a review. Here this guide disagrees with the site’s [handoff guide](/library/guides/design-a-human-handoff/). Its context packet carries “the approved summary” under “Customer explanation”, and its confirmation state lets the customer “correct or limit the summary”. A summary the customer has corrected is still prose written about the call, separately from the fields, and under this design it does not go on the desk screen. If the caller wants their account of the problem passed on, carry the caller’s own words in a field of their own and show them on the screen as a quotation. The companion has no such field. If you add one, it is caller-authored text crossing verbatim, with the exposure the pitfall below describes. ### Specify the screen as the questions it retires With the summary computed, the design work moves down a level: which field answers which question the desk would otherwise ask. The fixture desk writes that down as a mapping (`handoffpkg/desk.py`, lines 53–59): ```python _QUESTION_RETIRED_BY: dict[str, str] = { "Can I take your name?": "identity.display_name", "Can you confirm your date of birth?": "identity.verified_tier", "Which account is this about?": "intent.details.account_id", "What are you calling about today?": "intent.goal", "Have you tried anything already?": "attempts", } ``` `unanswered_questions()` (`desk.py`, lines 178–191) walks that table against a package and asks `_package_answers` (from line 194) whether each field is populated. Those rules (lines 208–228) are where the design decisions sit: | Desk question | Retired by | Still asked when | | ----------------------------------- | --------------------------- | -------------------------------------------------- | | Can I take your name? | `identity.display_name` | the name is empty | | Can you confirm your date of birth? | `identity.verified_tier` | the tier is `unverified` | | Which account is this about? | `intent.details.account_id` | no account id crossed | | What are you calling about today? | `intent.goal` | the goal is empty or `unknown` | | Have you tried anything already? | `attempts` | the list is empty, even if the true answer is “no” | The date-of-birth question is retired by the verification tier, not by a date of birth. The desk needs to know that identity was established and how strongly; it never needs the value. :::callout{type="info" title="An empty attempts list does not answer “Have you tried anything?”"} `_package_answers` treats an empty `attempts` list as missing information, because a package from a session that never recorded attempts looks exactly the same (`desk.py`, lines 200–206). If “nothing tried yet” matters to your desk, the assistant has to record it as an attempt. The desk will not infer it from silence. ::: ## The screen the agent’s own tools produce The second half of `make compare` shows a full screen, but it is rendered from `DEMO_SESSION`, a dictionary typed by hand in `scripts/agent_desk.py` (lines 35–74). To see what the running agent would send, we called the dispute skill’s tools and `transfer_to_human` directly, with a stand-in for the memory context, on an unmodified copy of the pattern. No model, network or licence is involved. Two values are typed in. The first is the charge the caller picks, `txn_c101`; the model can set `disputed_txn_id` on a call, because the dispute skill declares it `llm_settable` (`skills/dispute_transaction/memory.yml`, lines 3–6). The second is the one-line `handoff_reason`, and nothing on the live path can set that one (see below); the script types it only so the transfer can run. The script is `editorial/receipts/design-the-handoff-summary/agent-path-desk-screen.py` in the site repository, and the first line of its saved transcript records the exact command. The handoff id is random on each run (`tools.py`, line 177). Its unedited output, Python 3.12.13 and rasa-pro 3.20.0rc1: ```text attempts_log as written by the dispute tools: verify_passphrase|succeeded||Caller answered the knowledge factor on the first try. send_otp_sms|failed|delivery_failed|Carrier rejected the SMS twice. Do not resend to this number. raise_dispute|blocked|insufficient_tier|Dispute needs tier 'high'; caller is at 'medium'. disputed_txn_label in memory: Northgate Fuel $248.00 disputed_txn_label in _MEMORY_KEYS: False dispute_date in memory: False withheld_fields: ['otp_code', 'pin_attempt'] desk_still_needs_to_ask: [] hint: Context package delivered. Tell the caller the handoff id and that they will not need to repeat themselves. Do not read back any withheld field. ======================================================================== INBOUND HANDOFF ho_ed47b54e — Dispute needs a second factor the caller cannot complete on this line. ======================================================================== CALLER Jordan Rivera (cust_00417) via voice TRUST tier=medium — Verified with a knowledge factor. Account-specific information is in scope. ASKING FOR Dispute a card transaction [account_id=acc_checking, account_label=Everyday Checking, card_last_four=4821] (stage: blocked) ALREADY TRIED (do not retry these): · verify_passphrase → succeeded (Caller answered the knowledge factor on the first try.) · send_otp_sms → failed (Carrier rejected the SMS twice. Do not resend to this number.) · raise_dispute → blocked (Dispute needs tier 'high'; caller is at 'medium'.) DO NOT ASK — the caller already answered these: · Can I take your name? · Can you confirm your date of birth? · Which account is this about? · What are you calling about today? · verified: knowledge_passphrase · confirmed: Caller consented to the call being recorded. YOU MAY: · Answer general questions · Explain what the agent already did · Discuss account-specific details · DO NOT action irreversible changes — step the caller up first WITHHELD BY POLICY (present in the session, not transferred): · otp_code · pin_attempt Do not ask the caller to read these out to you. ======================================================================== ``` The `attempts_log` lines at the top are what the dispute tools write. `transfer_to_human` turns them into the package’s `attempts` (`tools.py`, line 175), which is why the last desk question counts as retired even though it is missing from the `DO NOT ASK` list. That list shows the questions the tools recorded as text; retirement is computed from fields. Specify one of them as the source for both, or the screen will say one thing while the check says another. Look at the `ALREADY TRIED` block next. The line telling the desk not to resend the SMS matters more than the name, because without it the desk agent’s first helpful move is to repeat the step that already failed twice. Then look at `ASKING FOR`. The account and the card are there. **The charge the caller disputed is not.** The dispute skill writes it: `raise_dispute` reads `disputed_txn_id` (`skills/dispute_transaction/tools.py`, line 156) and stores the merchant and amount as `disputed_txn_label` (line 170). Both keys are declared in the skill’s own `memory.yml` (lines 3 and 9), and the output above shows `Northgate Fuel $248.00` sitting in memory. But the handoff tool’s key list (`skills/human_handoff/tools.py`, lines 68–96) names neither key. It names `dispute_amount`, `dispute_merchant` and `dispute_date` instead, which no tool writes, and no tool writes a date at all. So the charge never leaves memory. `desk_still_needs_to_ask` is still empty, because “Which charge is it?” is not one of the five questions. ## Name what was withheld The last block on that screen is the easiest one to leave out of a spec. A credential that is absent and a credential that was held back look the same on a screen that only shows what arrived. The pattern sends the names of withheld keys (`handoffpkg/redaction.py`, lines 104–108): ```python # Keys whose *names* are still useful to the desk even though their values must # not cross. Naming them lets the summary say "withheld by policy" rather than # leaving the desk to wonder, and stops an agent asking the caller to repeat a # credential the system already has. ANNOUNCE_WITHHELD = True ``` That comment states the authors’ intent: a named gap is meant to stop a helpful desk agent asking the caller to fill it. On the agent’s path above, the two names are `otp_code` and `pin_attempt`, the credentials the dispute tools write. `test_withheld_names_cross_but_values_do_not` plants seven sensitive values, checks that every name reaches `withheld_fields`, and checks that the planted PIN, `4242`, appears nowhere in the serialised package. `test_every_key_the_agent_collects_is_either_allowlisted_or_withheld` checks there is no silent third outcome for a key the tool reads. For the screen, specify the wording as well as the list. “Do not ask the caller to read these out to you” is an instruction to a person. The skill carries the same rule on the assistant’s side (`skills/human_handoff/skill.md`, lines 21–28): ```markdown Do NOT ask the caller to summarise the call, restate who they are, confirm their account again, or repeat anything they already told you. All of it is already in session state and `transfer_to_human` transfers it. Asking is the exact failure this skill exists to prevent. Never put a PIN, one-time code, passphrase, full card number or token into `handoff_reason`. Those are withheld from the transfer by policy, and writing one into the reason line would carry it across anyway. ``` :::callout{type="pitfall" title="Allowlisted text crosses as written"} `handoff_reason` is allowlisted, so whatever lands in it reaches the desk verbatim. `test_freetext_in_an_allowlisted_field_is_transferred_verbatim` writes a spoken card number into it directly and asserts the number arrives. That is the situation once a tool writes the caller’s words into the field. `scan_freetext_risk` flags the shape and the tool reports it as advisory only (`tools.py`, lines 182–185). Write the reason line’s instruction as a rule about content (“one line on why a human is needed; never quote the caller”), because the allowlist will not clean it. ::: ## Check the promise against the agent’s own path Now the sentence the caller heard (`skills/human_handoff/responses.yml`, lines 1–10): ```yaml # Confirmation wording for the handoff. The promise made here — "you will not # need to repeat yourself" — is one the context package has to actually keep, # and tests/test_handoff_context.py is what keeps it honest. responses: utter_confirm_handoff: - text: >- I'll bring in a specialist and pass them everything we've covered, so you won't need to repeat any of it. Shall I go ahead? metadata: rephrase: true ``` Two facts about this sentence change how you test it. Both come from reading the Rasa Pro 3.20.0rc1 wheel, not from a live run. - **It is spoken before the package exists:** it is the confirmation that `requires_confirmation` attaches to `transfer_to_human` (`skill.md`, lines 7–13). The engine’s confirmation gate says “The tool itself is never invoked here”, only “after the LLM confirms” (`rasa/mantle/orchestration/tool_execution/constraints.py`, lines 301–303). - **It may not be the sentence spoken:** `rephrase: true` is read from the response metadata into the response definition (`rasa/mantle/content/responses.py`, lines 353 and 362), which documents it as “the orchestrator rewords _text_ via a dedicated LLM call” (line 122). The confirmation ask carries that flag (`constraints.py`, line 503). The approved wording is where the promise starts, not a guarantee of what the caller hears. Nothing in the companion checks the spoken sentence, so the place to hold it is a check that runs before release, against a package the agent’s own tools produced. The same reading turns up a gap before the sentence is ever spoken. The skill tells the model to “set `handoff_reason`” (`skill.md`, line 18), and `transfer_to_human` is offered only when that field has a value (`requires: session.project.handoff_reason`, line 9). But `handoff_reason` is a project memory field (`memory.yml`, lines 94–96). At 3.20.0rc1 the wheel rejects `llm_settable` on project fields: “Project fields cannot be written by the LLM” (`rasa/mantle/memory/validation.py`, lines 81–95). The `set_fields` tool offers only `llm_settable` fields and collect targets (`rasa/mantle/llm/tool_schemas.py`, lines 26–29 and 525–526), no tool in the pattern writes the field, and a tool whose `requires` is falsy is left out of the model’s tool list (`tool_schemas.py`, lines 974–990). So at this revision nothing on the live path can set the reason, and the transfer tool would not be offered at all. The pattern has no hooks file and no `seed: true` field, the other two ways a project field is written (a grep of the pattern is saved with this guide’s receipts). This comes from reading the source, not from a live call. Put a writer for the reason line in your specification: a tool that composes it from fields, so its wording is yours and never the caller’s. We have not built or tested that tool. This companion has already shipped a green test over a hand-built session while the running agent did not collect `account_id`; [deriving test fixtures from production code](/library/guides/derive-test-fixtures-from-production-code/) covers that defect and the method that fixes it. ### What to demand in your specification The agent’s real path loses the disputed charge, the full demo screen is hand-built, and the promise is repeated regardless of what the desk still needs. So demand three things before you sign off the wording: 1. Review only desk screens rendered from a package the agent’s own tools produced. 2. Map each part of the promise to a desk question, and each question to its field, the tool that writes it and its entry in `_MEMORY_KEYS`. 3. Make the line spoken after the transfer depend on `desk_still_needs_to_ask`. The tool already reports that list. Here is the part of its result the assistant sees (`skills/human_handoff/tools.py`, lines 207–214): ```python "withheld_fields": list(package.withheld_fields), "desk_still_needs_to_ask": list(unanswered_questions(package)), "freetext_risk": freetext_risk, "hint": ( "Context package delivered. Tell the caller the handoff id and " "that they will not need to repeat themselves. Do not read back " "any withheld field." ), ``` The hint repeats the promise whatever `desk_still_needs_to_ask` contains, and so does the skill’s closing instruction (`skill.md`, lines 30–33). The companion does not branch on it; the branch belongs in your specification. :::solution{title="Try it: what should the assistant say when desk_still_needs_to_ask is not empty?"} Write the line before you open this. One proposal, not built or tested in the companion, for a list holding “Which account is this about?”: > **Assistant:** You’re through to a specialist, reference ho_1a2b3c4d. They have > your details and what we’ve tried. They may need to check which account > this is about. The reference is illustrative. The sentence names the one gap instead of denying it, so the caller is not surprised by the question and the promise stays true. ::: ### Make the promise no wider than the test “Everything we’ve covered” is a promise about the whole conversation. The check covers five questions. Close that gap from the design side, in one of two ways: 1. **Widen the list:** add “Which charge is it?” to the opening script and its entry in `_QUESTION_RETIRED_BY` in the same change. A question without a mapping makes `unanswered_questions` raise `KeyError` (`desk.py`, line 188), and `transfer_to_human` calls it on every transfer (`tools.py`, line 208). With both in place the check fails at this revision, and that failure is the point. Making it pass takes four more changes, none of which we have tried: collect `disputed_txn_label` in `_MEMORY_KEYS`; add it to `SESSION_ALLOWLIST`, or it is only named as withheld (`redaction.py`, lines 150–154); add it to the fixed `detail_keys` that build `intent.details` (lines 218–230); and add a branch to `_package_answers`, which returns `False` for any field it does not know (`desk.py`, line 229). 2. **Narrow the sentence:** promise what the fields carry, and say what is still to come. A proposed rewording, not tested in the companion: “I’ll pass the specialist your details, what you’re calling about and what we’ve already tried, so you won’t have to go through it again. They’ll need to complete one extra security check before they can raise the dispute. Shall I go ahead?” :::callout{type="warn" title="Collecting is a named list"} The tool’s own comment calls this “a completeness gap in the transfer … never a safety gap” (`tools.py`, lines 62–67). It fails safe for privacy and silently for the caller, so add “collect the key” to the checklist for every field the desk should see. ::: ## What these checks prove, and what they do not The six tests named in this guide pass when run together, and so does the full 41-test suite, `make test` (both run offline; the output is saved with this guide’s receipts). Every run here is offline, on the package, the tools and the fixture desk. Four claims rest on reading source instead of a run: when the confirmation is spoken, that `rephrase` rewords it, that nothing can set `handoff_reason`, and that the model can set `disputed_txn_id`. None of it shows what a model says on a live call, such as whether it asks the caller to summarise despite the skill’s instruction or how it rewords the confirmation. Test those with your own conversations. The allowlist is a design boundary, not a compliance control. The pattern’s docstring says PCI, HIPAA and similar regimes impose obligations “that a dataclass does not discharge” (`redaction.py`, lines 28–30), and that it does not redact the call recording, the transcript or logs written before the handoff (lines 31–34). The fixture desk is a local directory: `transfer_to_human` writes the package as a JSON file into `fixtures/desk_queue` (`tools.py`, lines 188–189). Whether a real contact-centre system accepted and queued the case is a separate question, and nothing here answers it. ## Questions you will hit while specifying it :::solution{title="Why not have a model write the summary at handoff time?"} Because it would be another thing that can disagree with the fields. The rendering function’s docstring (`handoffpkg/summary.py`, lines 102–106) puts it this way: “A summary produced by an LLM at handoff time would be a _sixth_ piece of state”, one that is “unversioned, unreproducible and free to contradict the other five.” A rendered view is the same every time for the same package, so you can test it. If the desk wants friendlier wording, change the renderer, not the data. ::: :::solution{title="Should the assistant read the desk preview back to the caller?"} No. The tool’s result includes the rendered screen so the assistant knows what was passed on, and the comment on it says it is returned “so the agent can honestly tell the caller what was passed on”, and “never so it can read it back to them” (`tools.py`, lines 198–199). Reading a screen aloud also risks speaking account details to whoever is on the line. Tell the caller the handoff id and what happens next, in a sentence you have checked against `desk_still_needs_to_ask`. ::: :::cta{href="/library/guides/design-a-human-handoff/" label="Plan consent, delivery and recovery"} The handoff guide covers what the caller agrees to send, the offer and confirmation states, and what the assistant says when nobody can take the call. ::: --- # Rasa 3.20.0 sent endpoints.yml's api_key_env; OpenAI refused Source: https://rasa.community/library/guides/endpoints-api-key-env-ignored/ Author: Rod Rivera Published: 2026-09-27 The companion voice agent, unedited, starting on Rasa Pro 3.20.0 against the real OpenAI API: ```text $ LLM_API_HEALTH_CHECK=true uv run --frozen rasa run --enable-api --port 55931 INFO Sending a test LLM API request for the component - ContextualResponseRephraser. config: {"model": "gpt-5.2", […], "api_key_env": "OPENAI_API_KEY"} ERROR Test call to the LLM API failed for component - ContextualResponseRephraser. […] litellm.BadRequestError: OpenAIException - Unknown parameter: 'api_key_env'. (harness) server process exit code: 1; /status: connection refused (http 000) ``` Excerpt, re-wrapped to fit. Cut: timestamps, logger names, the `event` and `level` fields, the config except `model` and `api_key_env`, and the `ProviderClientAPIException` wrapper; `[…]` marks cuts inside a line. The last line is our harness's summary, prefix ours. The unedited lines are in the collapsed block below. One line changed in the rephraser's model group in `endpoints.yml` starts the server: ```diff - id: openai-llm models: - provider: openai model: gpt-5.2 - api_key_env: OPENAI_API_KEY + api_key: ${OPENAI_API_KEY} temperature: 0.0 ``` ```text (harness) POST /webhooks/rest/webhook "Hi there" reply: Hi. I’m Atlas, Horizon Travel’s voice assistant. I can help with things like viewing your itinerary, checking flight status, updating a booking, or reporting lost baggage. What would you like to do? ``` First line: our summary of the request. The reply is the response's `text` field, re-wrapped, with its JSON escapes decoded. Atlas and Horizon Travel are the companion project's fictional fixture. The failure needs `LLM_API_HEALTH_CHECK=true`: Rasa's default is `false`, the companion's `.env.example` does not set it, and with it off the same unedited files started and answered with no ERROR line. Here is the mechanism. Rasa checks the keys of `endpoints.yml` model groups by name when it trains and when it loads a model, but only names on its lists: `api_key` must be written as `${NAME}`, for example. `api_key_env` is a key that Rasa's own Mantle client defines for `integrations.yml`, and it is on no list, so it passes. On the OpenAI and default litellm paths we read, the client built from the group then keeps every key it does not recognise and hands it to litellm, which sends it to the provider. The first objection comes from the provider, and only when something calls it. Here the caller was the startup health check, and it tested a client that no reply used: in the recorded turns that could tell, replies were rephrased by the orchestrator's client, built from `integrations.yml`. :::diagram{title="Two clients, two files: which call uses which model group (Rasa Pro 3.20.0, Mantle)"} ```dot rankdir=LR; ep [label="endpoints.yml\nopenai-llm group", class="mark-1"]; rephraser [label="ContextualResponseRephraser\ncreated by the Agent"]; hc [label="health check\nif LLM_API_HEALTH_CHECK=true", class="mark-2"]; testcall [label="startup test call\nbody carries api_key_env", class="blocked mark-3"]; integ [label="integrations.yml\norchestrator group"]; mantle [label="LLMClient.from_bootstrap\npops api_key_env", class="mark-4"]; orch [label="orchestrator\n_rephrase_and_send", class="mark-5"]; turn [label="turn-time rephrase\nno api_key_env in body", class="ok"]; ep -> rephraser [label="nlg.llm.model_group"]; rephraser -> hc; hc -> testcall; integ -> mantle; mantle -> orch [label="orchestrator.py 268"]; orch -> turn [label="line 936"]; ``` 1. `nlg.llm.model_group` names this group, and nothing on its path reads `api_key_env`. 2. Run when the server starts; in our training run with the check on it sent no test call. 3. OpenAI and Anthropic each refused this request, and the server did not come up. 4. The one reader of `api_key_env`, fed from `integrations.yml`. 5. `_rephrase_and_send` calls the orchestrator's client; in the recorded turns that could tell, the replies used it. ::: **Our position:** Rasa should refuse `api_key_env` in `endpoints.yml` at load time with a dedicated rule whose message names `api_key: ${NAME}`. The cheaper fix is one entry in `SENSITIVE_DATA`, the list its load-time check already keeps. By our reading of the source that entry would reject both string forms, but with the wrong advice: `api_key_env: NAME` would be told it "must be set as an environment variable", and `api_key_env: ${NAME}` would then be told it is "not allowed", with a list of the keys that may use `${...}`. Neither message says to write `api_key` instead. A dedicated rule is one more check in a loop that already walks every model entry, and it can name the fix. Either way, a project whose `endpoints.yml` carries the line, as the companion's examples do, would stop at `rasa train` instead of running quietly with the check off. We would take that break. On 3.20.0, write `api_key: ${NAME}` in both files and make the provider answer once before every deploy; `integrations.yml` also accepts `api_key_env: NAME` there, but the placeholder is the form that works in both. :::solution{title="Show the unedited log lines from the OpenAI and Anthropic runs"} OpenAI, the run above: ```text 2026-09-29 11:50:21 INFO rasa.shared.utils.health_check.health_check - {"event_info": "Sending a test LLM API request for the component - ContextualResponseRephraser.", "config": {"model": "gpt-5.2", "api_base": null, "api_version": null, "api_type": "openai", "provider": "openai", "reasoning_effort": "none", "temperature": 0.0, "max_completion_tokens": 256, "timeout": 5, "api_key_env": "OPENAI_API_KEY"}, "event": "contextual_response_rephraser.init.send_test_llm_api_request", "level": "info"} 2026-09-29 11:50:21 ERROR rasa.core.run - {"event_info": "Test call to the LLM API failed for component - ContextualResponseRephraser.", "config": {"model": "gpt-5.2", "api_base": null, "api_version": null, "api_type": "openai", "provider": "openai", "reasoning_effort": "none", "temperature": 0.0, "max_completion_tokens": 256, "timeout": 5, "api_key_env": "OPENAI_API_KEY"}, "error": "ProviderClientAPIException(\"\\nOriginal error: litellm.BadRequestError: OpenAIException - Unknown parameter: 'api_key_env'.)\")", "event": "contextual_response_rephraser.init.send_test_llm_api_request_failed", "level": "error"} ``` Anthropic, the run in the provider section below: ```text 2026-09-29 11:33:04 ERROR rasa.core.run - {"event_info": "Test call to the LLM API failed for component - ContextualResponseRephraser.", "config": {"model": "claude-haiku-4-5", "provider": "anthropic", "api_key_env": "ANTHROPIC_API_KEY", "temperature": 0.0}, "error": "ProviderClientAPIException('\\nOriginal error: litellm.BadRequestError: AnthropicException - {\"type\":\"error\",\"error\":{\"type\":\"invalid_request_error\",\"message\":\"api_key_env: Extra inputs are not permitted\"},\"request_id\":\"req_011CfXaYEuMPSm6AqRe3Du4h\"})')", "event": "contextual_response_rephraser.init.send_test_llm_api_request_failed", "level": "error"} ``` ::: Live runs of `examples/mantle-voice-agent` (`RasaHQ/rasa-community-resources` at `01d6a6e`) on Rasa Pro 3.20.0 with litellm 1.100.1, as the project's `uv.lock` pins them. Runs recorded 29 September 2026 against the 3.20.0 release. Real-provider runs used the key in the project's `.env`; the rest used placeholder keys and a local stand-in. Receipts are in `editorial/receipts/endpoints-api-key-env-ignored/` in the site repository. The line is ours. Companion commit `432d416`, titled "Use api_key_env: NAME; api_key: ${VAR} never expanded", made this change to every `endpoints.yml` credential line: ```text $ git -C rasa-community-resources show 432d416 -- 'examples/*/endpoints.yml' 'tutorials/*/endpoints.yml' 'community/*/endpoints.yml' | grep '^[-+] .*api_key' | sort | uniq -c 1 - api_key: ${GEMINI_API_KEY} 28 - api_key: ${OPENAI_API_KEY} 1 - # api_key: ${OPENAI_API_KEY} 1 + api_key_env: GEMINI_API_KEY 28 + api_key_env: OPENAI_API_KEY 1 + # api_key_env: OPENAI_API_KEY ``` On 3.20.0 the commit's premise did not hold for `endpoints.yml`: the client expands `${NAME}` when it calls the provider, which is why the edit above started the server and why a stand-in received the expanded placeholder key. At `01d6a6e`, 29 active `api_key_env` lines remain in 15 `endpoints.yml` files. We ran one of those files against a real provider. ## The request carried api_key_env as a body field OpenAI's error names a parameter; a recorder shows the request that carried it. The stand-in is a local HTTP server that `OPENAI_API_BASE` points at. It writes one JSON line per request (path, bearer value, top-level keys of the body, model) and answers with a fake completion. For two runs we gave the rephraser's group its own variable, set `OPENAI_API_KEY` and `REPHRASER_OPENAI_KEY` to two placeholders, trained, started the server and sent "Hi there". One run kept the `api_key_env` spelling; the other used `api_key: ${...}`. ::::compare{title="What the stand-in received, by rephraser credential line (selected fields)"} :::pane{label="api_key_env: REPHRASER_OPENAI_KEY" tone="bad"} ```text /v1/embeddings sk-standin-default /v1/chat/completions sk-standin-default body keys: api_key_env, max_completion_tokens, messages, model, reasoning_effort, temperature api_key_env = REPHRASER_OPENAI_KEY /v1/chat/completions sk-standin-default body keys: messages, model, reasoning_effort, temperature ``` The test call goes out on the default key and carries the variable's name as a body field. ::: :::pane{label="api_key: ${REPHRASER_OPENAI_KEY}" tone="good"} ```text /v1/embeddings sk-standin-default /v1/chat/completions sk-standin-rephraser body keys: max_completion_tokens, messages, model, reasoning_effort, temperature /v1/chat/completions sk-standin-default body keys: messages, model, reasoning_effort, temperature ``` The test call carries the named key and no stray field. The last request is unchanged. ::: :::: Each pane shows the path, bearer and body keys from the stand-in's raw lines, in the order received; the raw lines are in the receipts. The first request is an embeddings call made during training by Mantle's default references embedder, not by an `endpoints.yml` group. The second is the rephraser's startup test call, the only request with `max_completion_tokens`, matching the 256 in the health check's logged config. The third is the reply to "Hi there". The stand-in accepts anything, so in the left pane the field simply arrives. The left pane also shows the test call going out on `sk-standin-default`, the value of `OPENAI_API_KEY`, while the YAML named `REPHRASER_OPENAI_KEY`; the litellm source below explains why. The stand-in is 25 lines of Python, and these two records are its output. :::solution{title="Show the stand-in and the commands for these two runs"} Start the recorder, then train and start Rasa against it with two different placeholder keys (set the same variables for `rasa train`). The ports are yours to choose. ```sh python openai_standin_chat.py requests.jsonl & LLM_API_HEALTH_CHECK=true OPENAI_API_KEY=sk-standin-default \ REPHRASER_OPENAI_KEY=sk-standin-rephraser \ OPENAI_API_BASE=http://127.0.0.1:/v1 \ uv run --frozen rasa run --enable-api --port ``` Give the rephraser's group `api_key_env: REPHRASER_OPENAI_KEY` for the left pane or `api_key: ${REPHRASER_OPENAI_KEY}` for the right. The script logs a bearer value only when it starts with `sk-standin`, so a real key sent by mistake is never written down. ```python """Local stand-in for api.openai.com used only with placeholder keys. Records, per request: path, the bearer value (placeholders set by the test, never a real key), top-level body keys, model. Returns a fake chat completion or fake embeddings. Nothing is forwarded.""" import http.server, json, sys, datetime, hashlib PORT, LOG = int(sys.argv[1]), sys.argv[2] class H(http.server.BaseHTTPRequestHandler): def do_POST(self): body = json.loads(self.rfile.read(int(self.headers.get('content-length', 0))) or b'{}') auth = self.headers.get('authorization', '') bearer = auth.split(' ', 1)[1] if auth.startswith('Bearer ') else ('(none)' if not auth else '(non-bearer)') if not bearer.startswith('sk-standin') and bearer not in ('(none)', '(non-bearer)'): bearer = '(unexpected value, not logged)' with open(LOG, 'a') as f: f.write(json.dumps({'at': datetime.datetime.now().isoformat(timespec='seconds'), 'path': self.path, 'bearer': bearer, 'body_keys': sorted(body), 'model': body.get('model'), 'api_key_env_in_body': body.get('api_key_env')}) + '\n') if self.path.rstrip('/').endswith('/embeddings'): inputs = body.get('input', []); inputs = [inputs] if isinstance(inputs, str) else inputs out = {'object': 'list', 'data': [{'object': 'embedding', 'index': i, 'embedding': [0.01] * 1536} for i, _ in enumerate(inputs)], 'model': body.get('model'), 'usage': {'prompt_tokens': 0, 'total_tokens': 0}} else: out = {'id': 'chatcmpl-standin', 'object': 'chat.completion', 'created': 0, 'model': body.get('model'), 'choices': [{'index': 0, 'message': {'role': 'assistant', 'content': 'stand-in reply'}, 'finish_reason': 'stop'}], 'usage': {'prompt_tokens': 1, 'completion_tokens': 1, 'total_tokens': 2}} data = json.dumps(out).encode() self.send_response(200); self.send_header('content-type', 'application/json'); self.send_header('content-length', str(len(data))); self.end_headers(); self.wfile.write(data) def log_message(self, *a): pass http.server.ThreadingHTTPServer(('127.0.0.1', PORT), H).serve_forever() ``` ::: ## Rasa's load-time check has a credential list, and api_key_env is not on it Rasa does check these groups before any request. `validate_model_group_configuration_setup` in `rasa/engine/validation.py` runs over every `endpoints.yml` model group when `rasa train` builds a model (`rasa/model_training.py`, line 347) and when Rasa loads one (`rasa/engine/loader.py`, line 46). Two of its checks are about key names. The first allows a `${...}` value only on these keys (lines 1351 to 1360): ```python allowed_env_vars = { DEPLOYMENT_CONFIG_KEY, API_BASE_CONFIG_KEY, API_KEY, API_VERSION_CONFIG_KEY, AWS_REGION_NAME_CONFIG_KEY, AWS_ACCESS_KEY_ID_CONFIG_KEY, AWS_SECRET_ACCESS_KEY_CONFIG_KEY, AWS_SESSION_TOKEN_CONFIG_KEY, } ``` The second requires every key on this list, from `rasa/shared/constants.py` (lines 361 to 372), to be written as `${...}`, and raises a `ValidationError` otherwise: ```python SENSITIVE_DATA = [ API_KEY, AWS_ACCESS_KEY_ID_CONFIG_KEY, AWS_SECRET_ACCESS_KEY_CONFIG_KEY, AWS_SESSION_TOKEN_CONFIG_KEY, LANGFUSE_CONFIG_PUBLIC_KEY, LANGFUSE_CONFIG_PRIVATE_KEY, CLIENT_ID_CONFIG_KEY, CLIENT_SECRET_CONFIG_KEY, TOKEN_URL_CONFIG_KEY, A2A_JWT_SECRET_CONFIG_KEY, ] ``` `api_key_env: OPENAI_API_KEY` is neither a `${...}` value nor a listed key, so it passes both, and `rasa train` exited 0 in each of the six runs where we recorded its exit code. Neither file mentions `api_key_env`. The list is where Rasa has already decided that key names matter for credentials; `api_key_env` is a credential name Rasa itself defines, in its Mantle client, and the list does not know it. In effect the two checks are a name-based schema for secrets, with one name missing. ## Past the check, the two clients we read keep unknown keys :::annotated{title="examples/mantle-voice-agent/endpoints.yml, lines 16-31 (comment block cut)"} ```yaml nlg: type: rephrase llm: model_group: openai-llm # (1) model_groups: - id: openai-llm models: - provider: openai # (2) model: gpt-5.2 api_key_env: OPENAI_API_KEY # (3) temperature: 0.0 ``` 1. The rephraser's group. `ContextualResponseRephraser` resolves it when it is created and runs the health check on it. 2. `provider: openai` selects `OpenAILLMClient`. Any provider Rasa does not map, `anthropic` included, gets `DefaultLiteLLMClient` (`rasa/shared/providers/mappings.py`). 3. The load-time check passes it, and nothing after it reads it. It stays in the config, is printed in the log's `config`, and is the field OpenAI named. ::: The client config takes out the keys it knows and keeps the rest. Its comment says "The rest of parameters (e.g. model parameters) are considered as extra parameters (this also includes timeout)". This is `OpenAIClientConfig`, in `rasa/shared/providers/_configs/openai_client_config.py` (lines 141 to 154): ```python this = OpenAIClientConfig( # Required parameters model=config.pop(MODEL_CONFIG_KEY), # Pop the 'provider' key. Currently, it's *optional* because of # backward compatibility with older versions. provider=config.pop(PROVIDER_CONFIG_KEY, OPENAI_PROVIDER), # Optional parameters api_base=config.pop(API_BASE_CONFIG_KEY, None), api_version=config.pop(API_VERSION_CONFIG_KEY, None), api_type=config.pop(API_TYPE_CONFIG_KEY, OPENAI_API_TYPE), # The rest of parameters (e.g. model parameters) are considered # as extra parameters (this also includes timeout). extra_parameters=config, ) ``` `api_key_env` is in "the rest". The default litellm client's config does the same after taking out only `model` and `provider`. Both clients then spread the extras into the litellm call, in `rasa/shared/providers/llm/_base_litellm_client.py` (lines 84 to 96): ```python @property def _completion_fn_args(self) -> dict: return { # Since all providers covered by LiteLLM use the OpenAI format, but # not all support every OpenAI parameter, raise an exception if # provider/model uses unsupported parameter "drop_params": False, # All other parameters set through config, can override drop_params **self._litellm_extra_parameters, # Model name is constructed in the LiteLLM format from the provided config # Non-overridable to ensure consistency "model": self._litellm_model_name, } ``` The `**self._litellm_extra_parameters` spread carries `api_key_env` forward, beside `temperature` and `reasoning_effort`. `drop_params: False` is a separate setting; its comment says the aim is to "raise an exception if provider/model uses unsupported parameter". In our runs the exception came from the provider, after the request was sent. The same file passes the arguments through `resolve_environment_variables`, which calls `os.path.expandvars` (`rasa/shared/utils/io.py`). That is where `api_key: ${REPHRASER_OPENAI_KEY}` became `sk-standin-rephraser`. litellm 1.100.1 does the last step (`litellm/utils.py`, lines 4719 to 4745). For `openai`, a keyword it does not recognise goes into `extra_body`; for other providers it goes into the request parameters. With no `api_key`, litellm's OpenAI path falls back to the `OPENAI_API_KEY` environment variable (`litellm/main.py`), which is where the left pane's `sk-standin-default` came from. Both litellm points are source reading; the wire behaviour is in the records. This covers the OpenAI and default litellm clients, the two paths we read. ## Only integrations.yml turns api_key_env into a key A grep of the 3.20.0 wheel finds one reader of the YAML key, in `rasa/mantle/llm/client.py`. The Azure clients' hits are a private attribute whose YAML key is `api_key`, and a docstring in `rasa/cli/project_env.py` names `integrations.yml`: ```text $ grep -rn api_key_env rasa/ --include='*.py' rasa/mantle/llm/client.py:43:_API_KEY_ENV_CONFIG_KEY = "api_key_env" rasa/mantle/llm/client.py:136:def _resolve_api_key_env(config: Dict[str, Any]) -> Dict[str, Any]: rasa/mantle/llm/client.py:150: "mantle.llm.api_key_env.not_set", rasa/mantle/llm/client.py:257: Resolves ``api_key_env`` entries — where the value is an environment rasa/mantle/llm/client.py:260: ``api_key_env`` key (an engine convention unknown to LiteLLM). rasa/mantle/llm/client.py:262: llm_config = _resolve_api_key_env(dict(bootstrap.llm)) rasa/shared/providers/llm/azure_openai_llm_client.py:137: self._api_key_env_var = ( rasa/shared/providers/llm/azure_openai_llm_client.py:138: self._resolve_api_key_env_var() if not self._oauth else None rasa/shared/providers/llm/azure_openai_llm_client.py:197: def _resolve_api_key_env_var(self) -> str: rasa/shared/providers/llm/azure_openai_llm_client.py:344: elif self._api_key_env_var: rasa/shared/providers/llm/azure_openai_llm_client.py:345: auth_parameter = {LITE_LLM_API_KEY_FIELD: self._api_key_env_var} rasa/shared/providers/embedding/azure_openai_embedding_client.py:98: self._api_key_env_var = ( rasa/shared/providers/embedding/azure_openai_embedding_client.py:99: self._resolve_api_key_env_var() if not self._oauth else None rasa/shared/providers/embedding/azure_openai_embedding_client.py:104: def _resolve_api_key_env_var(self) -> str: rasa/shared/providers/embedding/azure_openai_embedding_client.py:244: elif self._api_key_env_var: rasa/shared/providers/embedding/azure_openai_embedding_client.py:245: auth_parameter = {LITE_LLM_API_KEY_FIELD: self._api_key_env_var} rasa/cli/project_env.py:64: ``api_key_env`` in ``integrations.yml`` available even when the process has ``` `_resolve_api_key_env` walks a model group, pops `api_key_env` and writes the named variable's value into `api_key`. Its one caller explains why, in words that fit this whole page (`client.py`, lines 254 to 262): ```python def from_bootstrap(cls, bootstrap: ModelBootstrap) -> "LLMClient": """Build client from the project's integrations.yml LLM config. Resolves ``api_key_env`` entries — where the value is an environment variable *name* — into the corresponding ``api_key`` value so the config that reaches ``llm_factory`` never contains the raw ``api_key_env`` key (an engine convention unknown to LiteLLM). """ llm_config = _resolve_api_key_env(dict(bootstrap.llm)) ``` `bootstrap.llm` is the orchestrator's group, resolved from the project's `integrations.yml` (`rasa/mantle/model_archive/bootstrap.py`). In the voice agent that file carries the same line under a different group (lines 6 to 15): ```yaml llm: model_group: orchestrator model_groups: - id: orchestrator models: - provider: openai model: gpt-5.2 api_key_env: OPENAI_API_KEY temperature: 0.0 ``` So one project holds the same spelling twice: resolved in `integrations.yml`, forwarded from `endpoints.yml`. The `integrations.yml` half rests on the source above and on the shape of the reply request, which carried no `api_key_env` in either stand-in record. No run gave the orchestrator's group a second variable name to tell the two apart. ## The startup check tested a client the recorded replies never used The forwarded key reaches a provider only when something sends the rephraser's group. `perform_llm_health_check`, in `rasa/shared/utils/health_check/health_check.py` (lines 77 to 149), sends a test call when the variable reads `true` (compared case-insensitively). With it off, the same function can still probe Azure deployment-only groups at inference; otherwise it logs this warning, quoted from our real-OpenAI run with the check off: ```text The LLM_API_HEALTH_CHECK environment variable is set to false, which will disable LLM health check. It is recommended to set this variable to true in production environments. ``` A failed test call is re-raised as a `HealthCheckError` (lines 268 to 295); on the first screen `rasa.core.run` logged it and the process exited 1. For this project's openai group we checked the switch at the stand-in. `rasa run` alone with the check on sent one request, the test call, `api_key_env` included; with it off the stand-in received nothing. With the check on, `rasa train` exited 0 and the stand-in received only the embeddings request. In this project the check runs when the server starts, not when it trains. With the check off, does anything else use the group? For one Mantle component this is on record: the [guide on Mantle's references embedder](/library/guides/mantle-references-default-to-openai-embeddings/) showed the embedder is chosen from `agent.yml` and `integrations.yml`, and that repointing the `endpoints.yml` embeddings group changed no request. The rephrase needs its own evidence; the source and three recorded turns give it. The orchestrator builds its client with `LLMClient.from_bootstrap(self._static_data)`, and `_rephrase_and_send` passes that client to `self._call_llm` (`rasa/mantle/orchestration/orchestrator.py`, lines 268 and 936). The `endpoints.yml` rephraser is created by the Agent's `NaturalLanguageGenerator.create` (`rasa/core/agent.py`), and a grep of `rasa/mantle` finds no reference to `ContextualResponseRephraser`. We recorded four REST turns, one in each of rows 2, 3, 5 and 6 of the table below. Row 2 cannot tell the clients apart, because both would carry the same key. The other three can. In row 3 the check was off on real OpenAI, the group still carried `api_key_env`, and the greeting came back reworded, logged as `mantle.turn.completed` with `"source": "rephrase_text"`. OpenAI refused that field in row 1, and a failed rephrase falls back to the verbatim text (`rasa/mantle/orchestration/responses.py`, lines 93 to 98). So by inference the rephrase request did not carry the field. The first sentence of each, from the receipt: ```text responses.yml utter_greet: Hello. I'm Atlas, your Horizon Travel voice assistant. reply, check off: Hi. I’m Atlas, Horizon Travel’s voice assistant. ``` At the stand-in, row 5's rephrase request carried no `api_key_env` field although the group did. In row 6, with the group spelled `api_key: ${REPHRASER_OPENAI_KEY}`, the test call carried `sk-standin-rephraser` and the rephrase request carried `sk-standin-default`, the orchestrator's key. These are REST turns on one project; we did not run other channels or a project without Mantle. ## The provider is the first thing that objects OpenAI's reply is on the first screen. For the Anthropic run we switched the rephraser's group to Anthropic and left `api_key_env` in it: ```diff - id: openai-llm models: - - provider: openai - model: gpt-5.2 - api_key_env: OPENAI_API_KEY + - provider: anthropic + model: claude-haiku-4-5 + api_key_env: ANTHROPIC_API_KEY temperature: 0.0 ``` The orchestrator and embeddings went to the stand-in; only the rephraser's test call went to Anthropic, and `/status` never answered. The error, as an excerpt re-wrapped to fit: ```text ERROR Test call to the LLM API failed for component - ContextualResponseRephraser. […] AnthropicException - […]"type":"invalid_request_error", "message":"api_key_env: Extra inputs are not permitted"[…], "request_id":"req_011CfXaYEuMPSm6AqRe3Du4h" ``` Cut: the timestamp, the logger name, the `config` object, the `event` and `level` fields, the `ProviderClientAPIException` and `litellm.BadRequestError` wrapper, and the outer JSON of Anthropic's reply; `[…]` marks the cuts inside a line, and the backslash escapes on the quotes are decoded. The full line is in the collapsed block near the top. Two providers, two client classes in Rasa, two litellm branches (`extra_body` and request parameters), and the same refusal. Neither reply names the credential problem: both describe an unexpected parameter, which is accurate and points away from the key. Strictness is what made the key visible. The stand-in is the lenient case: it accepted the field and the run carried on. A provider that ignored unknown fields would do the same. We ran no such provider; the stand-in shows the shape of that case, not any provider's behaviour. Seven runs of this project in six rows; row 1 holds two. They are every run on this page except the two switch runs, which started the server and sent no message, the training-only run, and the quickstart run, which is a different project. Only the rephraser's group and the environment change between rows. In row 3 the check was off, so the group was sent nowhere. | # | Rephraser group credential line | Provider | `LLM_API_HEALTH_CHECK` | What happened | Receipt | | --- | --------------------------------------------------------------------------- | -------------------------------------- | ---------------------- | ----------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------ | | 1 | `api_key_env: OPENAI_API_KEY` (as committed) | real OpenAI | true | Test call rejected: `Unknown parameter: 'api_key_env'`. Exit 1, `/status` never answered | `r320-openai-rephraser-committed.txt`, and run 2 of `r320-openai-committed-health-check-off.txt` | | 2 | `api_key: ${OPENAI_API_KEY}` | real OpenAI | true | Test call sent, no ERROR line. `/status` answered; "Hi there" got the Atlas greeting | `r320-openai-rephraser-placeholder.txt` | | 3 | `api_key_env: OPENAI_API_KEY` (as committed) | real OpenAI | false | Started, 0 ERROR lines. The greeting came back reworded | `r320-openai-committed-health-check-off.txt` | | 4 | `provider: anthropic`, `claude-haiku-4-5`, `api_key_env: ANTHROPIC_API_KEY` | real Anthropic, for the test call only | true | Test call rejected: `api_key_env: Extra inputs are not permitted`. `/status` never answered | `r320-rephraser-anthropic.txt` | | 5 | `api_key_env: REPHRASER_OPENAI_KEY` | stand-in | true | Test call: bearer `sk-standin-default`, body field `api_key_env`. Reply request: bearer `sk-standin-default` | `r320-rephraser-own-var.txt` and its raw `.jsonl.txt` | | 6 | `api_key: ${REPHRASER_OPENAI_KEY}` | stand-in | true | Test call: bearer `sk-standin-rephraser`, no `api_key_env`. Reply request: bearer `sk-standin-default` | `r320-rephraser-placeholder.txt` and its raw `.jsonl.txt` | Rows 1 and 4 are one key refused by two providers in their own words. Rows 1 and 3 differ only in the switch. Rows 5 and 6 are the counter-test for the spelling. In rows 2, 4, 5 and 6, and in the first of row 1's two runs, the test call was logged, so by the source the variable was true; where it came from was not recorded for those runs. The second run of row 1 set it on the command line, as the first screen shows. :::callout{type="warn" title="What these runs do not show"} - Row 1 is two runs; rows 2 to 4 are one run each, of one model, recorded on one day. None of it is a rate or a provider's documented behaviour. - The failure needs `LLM_API_HEALTH_CHECK=true`, which is off by default. With it off, row 3 started and answered. - One project. The other 14 `endpoints.yml` files in the companion carry the same line by count; none was run, and whether their `nlg` blocks name those groups was not checked. - No Gemini request was recorded, so the one gemini line is not covered. - Only the OpenAI and default litellm client paths were read. Router groups with several models and the Azure, self-hosted and Rasa-hosted clients were not. - Everything rests on Rasa Pro 3.20.0, litellm 1.100.1 and companion `01d6a6e`. ::: ## Make the provider answer before you deploy Find the lines. In the companion at `01d6a6e` this matched 30 lines in 15 files, one of them inside a comment: ```sh git grep -n api_key_env -- '*endpoints*.yml' ``` Rewrite each active match as `api_key: ${NAME}` with the same variable name; in `endpoints.yml` we ran that form on OpenAI only. Use the same form in `integrations.yml`. There 3.20.0 reads `api_key_env` too, and the placeholder is expanded at call time by the same base client. A live 3.20.0 run of this site's quickstart project with `api_key: ${ANTHROPIC_API_KEY}` in `integrations.yml` trained and answered four conversations. One spelling in both files means one rule to check. Then start once with the check on and the real key, and require `/status` to answer: ```sh LLM_API_HEALTH_CHECK=true uv run --frozen rasa run --enable-api ``` With the committed line it exits 1; with the edit it starts. In this project that is one real request per start. A strict provider tells you the request was malformed; it does not tell you which key it used. For that, run the stand-in above. :::checkpoint{id="endpoints-api-key-env-harmless" question="In the companion voice agent (a Mantle project whose replies the orchestrator rephrases) the endpoints.yml rephraser group says api_key_env: OPENAI_API_KEY. rasa train passes and the agent answers every message you send it. What does that show about the line?" options="Rasa resolved it into the key,It is harmless because nothing reads it,Nothing yet: neither check you ran tested it" answer="2"} Training ran the load-time check, which does not list `api_key_env`, so a pass says nothing about the line. A working reply says nothing either: in the recorded turns that could tell, replies came from the orchestrator's client, built from `integrations.yml`. The line was tested only when the startup test call sent it, and there OpenAI and Anthropic both refused it. So it was not resolved, and it was not harmless. ::: ## Added 29 September 2026: the prerelease refuses the key Rasa Pro 3.21.0.dev3, a prerelease uploaded to PyPI on 28 September at 15:57 UTC, refuses `api_key_env` in an `endpoints.yml` model group at load time. `validate_model_group_configuration_setup` now calls `validate_model_group_credentials`, which raises `api_key_env_not_supported` with a message that ends "which is no longer supported. Replace it with 'api_key: ${ENV_VAR_NAME}'." That is a dedicated rule with the message argued for above. We read it in the prerelease source and did not run it; everything above this note is about 3.20.0. If you are on 3.20.0, write `api_key: ${NAME}` now: it works there in both files, and the prerelease refuses the other spelling in `endpoints.yml`. :::cta{href="/compatibility/" label="Check which release this was tested on"} These findings were recorded on Rasa Pro 3.20.0 and have not been re-run on later releases. The compatibility report shows the release the site targets today and the recorded scope of each maintained article. Re-run the strict start above before you trust this page on a newer release. ::: --- # Why TTS failover returns after a Rasa voice agent restart Source: https://rasa.community/library/guides/fail-over-speech-vendors-mid-call/ Author: Rod Rivera Published: 2026-09-22 The alert shows two log lines. They come from an offline drill in the companion project, with the vendor calls stubbed, in which the primary text-to-speech (TTS) vendor starts returning HTTP 402 (out of credits) on the third sentence of a call: ```text 2026-09-21 21:19:57 [warning ] voicerouter.tts.failing_over attempt=1 from_provider=rime verdict='quota (HTTP 402) — skipping briefly' voice_changes=True 2026-09-21 21:19:57 [info ] voicerouter.tts.failover from_provider=rime reason='served after failover' to_provider=deepgram ``` Picture the on-call engineer who sees those lines. The agent has switched to a backup voice, and something looks stuck, so they redeploy the voice service. The alerts stop. On the next real call, the router tries Rime first again. Rime returns 402 again, the same two lines fire and the pager goes off again. Nothing about Rime's account changed. The redeploy wiped the router's in-process record that Rime was out of credits, a record that had been keeping later calls off Rime for 900 seconds, so the next conversation paid for the discovery again. The scene is illustrative; the log lines and the 900 seconds come from the drill run described below. **The signal worth routing is the failure's class, not the fact that a vendor failed.** Page a person on a failure that only a person can fix. Let the router handle a rate limit. Redeploy only after the fix, because a redeploy on its own resets the router's memory and does nothing to the vendor. The cost of this stance is work up front: your alerting has to read a field from the log or metric rather than count failovers, and someone has to own each class before the first incident. This guide uses the voice router in `RasaHQ/rasa-community-resources` at revision `69e27b6`: the pattern `patterns/voice-vendor-router` and the example agent `examples/mantle-voice-routed-skills`, whose assistant Vela answers for the companion's seeded demo bank (its README, lines 207 and 212). Both pin Rasa Pro 3.20.0rc1 (`rasa-pro==3.20.0rc1` on line 8 of each `pyproject.toml`). The router applies the circuit-breaker pattern to speech vendors, and it honours the HTTP `Retry-After` header when a vendor sends one. Rime, Deepgram and OpenAI appear because they are the example's configured TTS chain. Nothing here measures any vendor's reliability. ## Why a Rasa voice agent needs a router to fail over A Rasa voice channel holds one ASR (speech-to-text) engine and one TTS engine. The example's `integrations.yml` names `voicerouter.RoutedTTS` as its TTS engine, and that router holds a chain of real engines (lines 63–90): `rime`, then `deepgram`, then `openai`. Rasa itself does not walk a chain. Mantle speaks a response through `TurnContext.send`, which calls the channel's `send_text_message`. On that path, the handler that catches a vendor error is this (`rasa/core/channels/voice_stream/voice_channel.py`, lines 619–625, in the 3.20.0rc1 wheel). We did not examine the separate streaming-response path (`send_response_chunk` and its background audio sender), which has its own error handling: ```python try: audio_stream = self.tts_engine.synthesize(text) except TTSError as e: logger.error("voice_channel.tts_synthesis_error", error=str(e)) voice_tracing.get_current_synthesize_speech_recorder().record_error(e) # TODO: add message that works without tts, e.g. loading from disc audio_stream = self.chunk_audio(generate_silence(self.audio_format)) ``` That code logs the error and plays silence, and it tries no other vendor. Its reach is narrower than it looks. It catches only an error raised by _calling_ `synthesize()`. In the base engine, in Rasa's `RimeTTS` and `DeepgramTTS` (`tts/deepgram.py`, lines 184–195) and in `RoutedTTS`, `synthesize` is an async generator, and calling an async generator runs none of its body. Their errors surface later, at line 677, where `_stream_audio_to_channel` reads the stream, outside that `try`. We checked this with a short offline probe, run with the router pattern's virtual environment. It is not part of the companion. `` stands for your checkout path, and the vendor is a stub that always raises: ```python import asyncio, inspect, sys sys.path.insert(0, "/patterns/voice-vendor-router") from rasa.core.channels.voice_stream.tts.tts_engine import TTSError from rasa.core.channels.voice_stream.tts.rime import RimeTTS from voicerouter.routed_tts import RoutedTTS from voicerouter.base import BuiltProvider, ProviderSpec, RouterPolicy from voicerouter.health import reset_shared_registries print("RimeTTS.synthesize is async generator:", inspect.isasyncgenfunction(RimeTTS.synthesize)) print("RoutedTTS.synthesize is async generator:", inspect.isasyncgenfunction(RoutedTTS.synthesize)) class Dead: async def connect(self): pass async def synthesize(self, text, config=None): raise TTSError("401 Unauthorized"); yield b"" reset_shared_registries() r = RoutedTTS([BuiltProvider(ProviderSpec(name="rime", label="rime", config={}), Dead())], RouterPolicy()) try: stream = r.synthesize("hello") # the call Rasa wraps in try/except TTSError print("calling synthesize(): no exception raised") except TTSError as e: print("calling synthesize(): raised", e) async def consume(): try: async for _ in stream: pass except TTSError as e: print("iterating the stream: raised TTSError:", str(e)[:90]) asyncio.run(consume()) ``` Its output, with the router's two log lines filtered out: ```text RimeTTS.synthesize is async generator: True RoutedTTS.synthesize is async generator: True calling synthesize(): no exception raised iterating the stream: raised TTSError: voicerouter: no TTS provider available — rime: disabled (auth) — 1 attempted for this utte ``` We traced that error up through `send_text_message` (lines 1038–1122) and Mantle's `TurnContext.send` (`rasa/mantle/orchestration/turn_context.py`, lines 290–299). Neither catches it, and we did not establish what the turn or the call does after that. Treat "every TTS vendor is out" as an unknown outcome for the caller, not as silence you can plan around. :::diagram{title="Where a TTSError surfaces in the Rasa 3.20.0rc1 voice channel"} ```dot rankdir=LR; call [label="synthesize(text) is called\n(line 620)"]; silence [label="logged as tts_synthesis_error,\nsilence played (lines 621-625)"]; read [label="the stream is read\n(line 677, no try)"]; up [label="send_text_message,\nTurnContext.send: no handler;\noutcome not established", class="blocked"]; call -> silence [label="engine raises\non the call itself"]; call -> read [label="async generator:\nnothing runs yet"]; read -> up [label="TTSError"]; ``` ::: For ASR there are two layers, and the first one matters more. Rasa's base ASR engine reads the vendor socket inside its own `try` (`rasa/core/channels/voice_stream/asr/asr_engine.py`, lines 353–364): ```python async def stream_asr_events(self) -> AsyncIterator[ASREvent]: """Stream the events returned by the ASR system as it is fed audio bytes.""" if self.asr_socket is None: raise ConnectionException("Websocket not connected.") try: async for message in self.asr_socket: asr_event = self.engine_event_to_asr_event(message) if asr_event: yield asr_event except Exception as e: logger.warning(f"Error while streaming ASR events: {e}") ``` A socket error during a read is logged at warning level, and the stream then ends as if the call were over. `DeepgramASR` does not override this method. Only an error that escapes the engine, such as the `ConnectionException` above, reaches the voice channel's own handler, which logs it at error level and re-raises (`voice_channel.py`, lines 1372–1380): ```python except Exception as e: logger.error( "voice_channel.receive_asr_events.error", call_id=call_parameters.call_id, error=repr(e), ) # Close the active transcription span as failed before re-raising. voice_tracing.get_current_transcribe_audio_recorder().record_error(e).end() raise ``` Of the engine's own errors, the router sees only the ones that escape the engine, just as that channel handler does. `RoutedASR` treats a stream that ends without an error as the end of the call (`routed_asr.py`, lines 247–250), so a read error the engine has already swallowed never reaches it. An offline probe (a fake Deepgram socket that raises `ConnectionError` on the first read, no network) shows both cases. Case A is a read that raises; case B is a socket that was never connected. We cut four start-up lines from the output: Rasa's `asr.engine.initialized` line, the `engine class: DeepgramASR` line and the two `voicerouter.asr.ready` lines. The rest is unedited: ```text DeepgramASR overrides stream_asr_events: False 2026-09-21 21:32:18 [warning ] Error while streaming ASR events: socket dropped mid-call (probe) A: socket drops during a read -> RoutedASR.stream_asr_events ended without raising A: socket drops during a read -> deepgram health: failures = 0 | active provider: deepgram 2026-09-21 21:32:18 [warning ] voicerouter.asr.stream_failed error='Websocket not connected.' provider=deepgram verdict='unknown (ConnectionException) — skipping briefly' 2026-09-21 21:32:18 [info ] voicerouter.asr.failover from_provider=deepgram reason='connected after failover' to_provider=backup 2026-09-21 21:32:18 [info ] voicerouter.asr.connected provider=backup 2026-09-21 21:32:18 [info ] voicerouter.asr.resumed note='audio sent during the failure was not transcribed' provider=backup B: socket was never connected -> RoutedASR.stream_asr_events ended without raising B: socket was never connected -> deepgram health: failures = 1 | active provider: backup ``` In case A the router records nothing and does not fail over, so `voicerouter.asr.stream_failed` and `voicerouter.asr.exhausted` never fire when a read raises inside the engine. The probe does not cover a real socket that closes cleanly. What the call does after the stream ends was not established. ## Read the failure class, not the failover Every vendor error the router sees goes through `classify()` in `voicerouter/failures.py` (lines 206–263), which turns an HTTP status, an AWS error code or the message text into a class. The class decides two things: how long the vendor is skipped (lines 70–76) and whether the caller may hear a different voice (lines 88–93): ```python DEFAULT_COOLDOWNS: dict[FailureKind, float] = { FailureKind.AUTH: 0.0, # permanent; cooldown unused FailureKind.CONFIG: 0.0, # permanent; cooldown unused FailureKind.QUOTA: 900.0, # 15 minutes FailureKind.RATE_LIMIT: 20.0, # overridden by Retry-After when present FailureKind.TRANSIENT: 15.0, } ``` ```python _JUSTIFIES_VOICE_CHANGE = { FailureKind.AUTH, FailureKind.CONFIG, FailureKind.QUOTA, FailureKind.UNAVAILABLE, } ``` `classify()` supplies the triggers in the table below. The "waiting" column paraphrases the module docstring (lines 6–10). When a verdict carries the vendor's `Retry-After`, that value replaces the default, with a floor of one second (lines 266–274). Quota, rate-limit and transient verdicts can carry it; `unavailable` verdicts never do (lines 248 and 261). Classes with no default use the policy's `cooldown_seconds`, which the example sets to 60. | Class | Typical trigger | Waiting fixes it? | Router skips the vendor for | Voice changes? | | ------------- | ---------------------------------- | ----------------- | ---------------------------- | --------------------------- | | `auth` | HTTP 401 or 403 | Never | Until restart or reset | Yes | | `config` | HTTP 400, 404, 405, 415 or 422 | Never | Until restart or reset | Yes | | `quota` | HTTP 402, or a 429 with quota text | After a top-up | 900 s, or `Retry-After` | Yes | | `unavailable` | Connection refused, timeout | Often | `cooldown_seconds` (60 here) | Yes | | `rate_limit` | HTTP 429 | Yes, in seconds | 20 s, or `Retry-After` | After one same-vendor retry | | `transient` | HTTP 5xx | Yes, in seconds | 15 s, or `Retry-After` | After one same-vendor retry | The last column comes from `routed_tts.py`, lines 356–377. A retryable failure gets one more attempt on the same vendor (`same_provider_retries: 1`, after `retry_backoff_ms: 250` in the example) as long as the wait is at most one second. If that retry fails too, or the vendor asks for seven seconds, the router moves on. Anything `classify()` cannot place is `unknown`: it is treated like `transient`, with the policy's `cooldown_seconds` as its park. :::callout{type="warn" title="A 900-second park can be logged as skipping briefly"} The verdict text in `failing_over` comes from `Verdict.__str__` (`failures.py`, lines 50–54). It prints "briefly" for any non-permanent failure without a `Retry-After`, including quota, while `cooldown_for` parks quota for 900 seconds. Read the class (`quota`), not the adverb. The drill's own summary shows the real figure: `retried in 900.0s`. ::: ## Where the router keeps circuit state Each vendor has a `ProviderHealth` record. The branch that matters for on-call work is in `record_failure` (`voicerouter/health.py`, lines 75–85): ```python if is_permanent(verdict): # No threshold for these. One 401 is enough to know. self.disabled = True self.opened_at = time.monotonic() self.current_cooldown = float("inf") return verdict if self.consecutive_failures >= self.failure_threshold: self.opened_at = time.monotonic() self.current_cooldown = cooldown_for(verdict, self.cooldown_seconds) return verdict ``` `record_success()` resets the failure count and closes the circuit, but it leaves `disabled` alone. A disabled vendor is dropped from the candidate list. A vendor that is only cooling down stays on the list, below the healthy ones, because a vendor rate-limited ten seconds ago still beats nothing (`routed_tts.py`, lines 194–236). :::diagram{title="One vendor's circuit in the router, and what a restart does to it"} ```dot rankdir=LR; closed [label="closed\n(tried in order)", class="ok"]; open [label="open\n(cooling down,\nlower priority)"]; half [label="half-open\n(next try is a probe)"]; disabled [label="disabled\n(dropped from the list)", class="blocked"]; closed -> open [label="quota, rate limit,\ntransient, unavailable"]; open -> half [label="cooldown elapses"]; half -> closed [label="probe succeeds"]; half -> open [label="probe fails"]; closed -> disabled [label="auth or config"]; disabled -> closed [label="process restart or\nreset_shared_registries()", style=dashed]; open -> closed [label="process restart", style=dashed]; ``` ::: Solid arrows are the router's own decisions. The dashed arrows are the only ways out of `disabled`, and a restart takes them whether or not anyone fixed the vendor. While Rime is out of credits, the first sentence after each 900-second park is a probe. If the account is still empty, the probe fails, Rime is parked again and the failover alert fires once more. While the backups keep serving, the alert repeats once per park, and less often when calls are sparse. That is the router probing on schedule, not a stuck router. If the backups also fail on a sentence, the router tries Rime again before its park ends, because a parked vendor stays on the candidate list (`routed_tts.py`, lines 197–199). The offline tests pin this behaviour. From `patterns/voice-vendor-router`, using the pattern's own virtual environment (Rasa Pro 3.20.0rc1), run on 21 September 2026: ```text $ PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -m unittest -v \ tests.test_voicerouter.TestHealth.test_permanent_failure_disables_and_success_does_not_revive_it \ tests.test_voicerouter.TestHealth.test_quota_parks_far_longer_than_a_transient_error \ tests.test_voicerouter.TestHealth.test_registry_is_shared_per_kind_so_it_outlives_a_call \ tests.test_voicerouter.TestHealth.test_reset_clears_it \ tests.test_voicerouter.TestCandidateSelection.test_disabled_providers_are_dropped_entirely \ tests.test_voicerouter.TestCandidateSelection.test_cooling_providers_drop_below_healthy_ones_but_stay test_permanent_failure_disables_and_success_does_not_revive_it (tests.test_voicerouter.TestHealth.test_permanent_failure_disables_and_success_does_not_revive_it) ... ok test_quota_parks_far_longer_than_a_transient_error (tests.test_voicerouter.TestHealth.test_quota_parks_far_longer_than_a_transient_error) ... ok test_registry_is_shared_per_kind_so_it_outlives_a_call (tests.test_voicerouter.TestHealth.test_registry_is_shared_per_kind_so_it_outlives_a_call) ... ok test_reset_clears_it (tests.test_voicerouter.TestHealth.test_reset_clears_it) ... ok test_disabled_providers_are_dropped_entirely (tests.test_voicerouter.TestCandidateSelection.test_disabled_providers_are_dropped_entirely) ... 2026-09-21 21:20:12 [info ] voicerouter.tts.ready providers=['a', 'b'] skipped=[] ok test_cooling_providers_drop_below_healthy_ones_but_stay (tests.test_voicerouter.TestCandidateSelection.test_cooling_providers_drop_below_healthy_ones_but_stay) ... 2026-09-21 21:20:12 [info ] voicerouter.tts.ready providers=['a', 'b'] skipped=[] ok ---------------------------------------------------------------------- Ran 6 tests in 0.001s OK ``` These are offline checks of the router's decisions with stubbed errors. They say nothing about how a real vendor fails or what a caller hears. `make test` in the same directory runs the whole suite (`Makefile`, lines 116–117). ### What to alert on The router gives you three places to read the class: - **Log lines:** `voicerouter.tts.failing_over` carries the verdict, including the class. `voicerouter.tts.failover` names the vendor that took over. There are also `voicerouter.tts.retrying_same_provider` and `voicerouter.tts.failed_mid_stream`. On the ASR side, add Rasa's own warning `Error while streaming ASR events` to the router's `voicerouter.asr.stream_failed`. When a read raises inside the engine, the engine logs that warning and the router logs nothing. - **Metrics:** `voicerouter/metrics.py` (lines 135–163) creates an OpenTelemetry counter `voicerouter.tts.failure` with the attributes `provider` and `failure_kind`, plus a `voicerouter.tts.failover` counter. Whether they are exported depends on the deployment's tracing configuration (lines 116–123). Key alerts on `failure_kind`, not on the failover count. - **`health_snapshot()`:** this method on the router lists each vendor's state, failure count, last class and `reopens_in`. At this revision, a search of the companion finds only one caller, `drill_failover.py`, and nothing serves it over HTTP, so the drill is where you will see it. ## What a restart resets, and what it cannot Rasa builds a fresh engine pair for every call (`voice_channel.py`, lines 1493–1496): ```python """Run streaming tasks and teardown for one call.""" # Media engines (ASR/TTS) asr_engine, tts_engine = self._get_asr_and_tts_engines(model_metadata) async with asr_engine, tts_engine: ``` If the router kept health on its own object, every call would rediscover a dead vendor. So by default it keeps health in a module-level dictionary that lives as long as the process (`health.py`, lines 156–161): ```python # Deliberately in-process only. A shared store (Redis) would extend this across # workers and is the obvious next step, but it introduces a dependency and a # failure mode of its own, and process scope already removes the large majority # of the waste. _SHARED: dict[str, "HealthRegistry"] = {} ``` A restart empties that dictionary, and a redeploy starts a new process. Either way, every vendor goes back to `closed`, the configured order applies again and the first sentence of the next call goes to Rime. The ASR and TTS registries are separate (`test_registry_is_shared_per_kind_so_it_outlives_a_call`), so a TTS fault on Rime does not mark anything on the ASR side. One comment near the top of `health.py` (lines 13–15) still says the state "lives for the life of the call". The code at lines 151–177 and the test say otherwise; trust those. All of this covers one process. The registry is not shared across workers or replicas, so this guide makes no claim about how several of them behave together. :::checkpoint{id="voice-failover-what-a-redeploy-changes" question="Rime has been failing over with quota (HTTP 402) for twenty minutes. Nobody has topped up the account. You redeploy the voice service. What has the redeploy changed?" options="Rime has credit again,Nothing because the router reloads its health state on start-up,The router has forgotten that Rime is parked so the next call tries Rime first and fails over again,Rime is now disabled until someone resets it" answer="2"} The health registry is a dictionary in process memory, so a new process starts with every vendor `closed`. Rime is tried first, fails with 402 again and is parked again, at the cost of a failover on the next call. The redeploy changed the router's memory, not Rime's account. ::: The reset function exists (`health.py`, lines 180–187): ```python def reset_shared_registries() -> None: """Forget everything. For tests, and for an operator escape hatch. A provider disabled by a rejected key stays disabled for the life of the process, which is right until someone fixes the key — at which point there has to be a way to say so without a restart. """ _SHARED.clear() ``` `test_reset_clears_it` shows that it re-enables a disabled vendor in the running process. It clears the whole dictionary, so the TTS and ASR registries go together, not just the vendor you fixed. Nothing calls it at runtime, though. At this revision, a search of the companion finds only the definition, the tests and `drill_failover.py`. There is no endpoint or command for it. And for Rime, a new key can only arrive by restarting: the engine requires `RIME_API_KEY` and reads it from `os.environ` when it builds its headers (`rasa/core/channels/voice_stream/tts/rime.py`, lines 69 and 145). So in this project, the working reset for a fixed key is a restart that comes **after** the fix. ## The runbook: wait, reset, re-route or escalate Rehearse the credits case before you need it. In `examples/mantle-voice-routed-skills`, run `make drill SCENARIO=credits`. It reads the live `integrations.yml`, builds the real router over stub engines and fails the first vendor from turn 3. We ran the script that target calls, `PYTHONDONTWRITEBYTECODE=1 NO_COLOR=1 .venv/bin/python scripts/drill_failover.py --scenario credits`, on 21 September 2026. Two edits below: the terminal's bold codes are removed from the heading line, and the script's closing three-line summary is cut: ```text Routing rules read from examples/mantle-voice-routed-skills/integrations.yml Vendor calls are stubbed — the routing decisions below are the real ones. 2026-09-21 21:19:57 [info ] voicerouter.tts.ready providers=['rime', 'deepgram', 'openai'] skipped=[] credits — Primary runs out of credits (HTTP 402) from turn 3 chain: rime > deepgram > openai 1. [rime ] Hi, you're through to Northwind. This is Vela. 2. [rime ] One moment. 2026-09-21 21:19:57 [warning ] voicerouter.tts.failing_over attempt=1 from_provider=rime verdict='quota (HTTP 402) — skipping briefly' voice_changes=True 2026-09-21 21:19:57 [info ] voicerouter.tts.failover from_provider=rime reason='served after failover' to_provider=deepgram 3. [deepgram ] Your current balance is two thousand four hundred fifty dollars. <- rime -> deepgram, failover 4. [deepgram ] Transferring four hundred pounds to Sam Rivera - shall I go ahead? 5. [deepgram ] Okay, got it. 6. [deepgram ] That's done. Is there anything else? after the call, rime: open, 1 failure(s), last was quota, retried in 900.0s ``` The drill starts each scenario with an empty registry for two reasons. It runs as a new process, just as a redeploy does, and `drill_failover.py` calls `reset_shared_registries()` before each scenario (line 175). So every run shows the discovery a redeployed service makes on its next call. When a failover alert fires, work through it in order: 1. **Read the class** from `failure_kind` or the verdict (`quota`, `auth`, `rate_limit` and so on). Ignore "briefly". 2. **`rate_limit` or `transient`: wait.** The router retries the same vendor or parks it for seconds. Do nothing unless the rate of these failures keeps rising, which is a capacity conversation with the vendor, not an incident. 3. **`unavailable`: wait and watch.** The vendor is parked for the policy's `cooldown_seconds`, then probed. If the backup is serving, no one needs to act tonight. 4. **`quota`: escalate to whoever owns the vendor account.** After the top-up, the next probe closes the circuit on its own. The probe is the first sentence after the 900-second park ends, or after the vendor's `Retry-After` if it sent one. A restart afterwards only brings that probe forward. A restart before the top-up only buys one more failover, with the backup serving that call. 5. **`auth` or `config`: page, fix, then restart.** The vendor stays disabled until the process restarts or something calls `reset_shared_registries()`. Fix the key or the configuration first. For Rime, the fixed key reaches the process through its environment, so the restart delivers the fix. 6. **Re-route if the vendor will be out for a long time.** Take it out of the chain so the router stops probing it. See the constraints below. 7. **Page on a TTS error that reads `no TTS provider available`**, and on repeated `Error while streaming ASR events` warnings. What the caller experiences in either case was not established. `voicerouter.asr.exhausted` fires only for ASR errors that escape the engine, so do not rely on it for a Deepgram socket drop. :::callout{type="pitfall" title="make stack swaps a whole stack, not one vendor"} `make stack STACK=` copies one of four committed files over `integrations.yml`: `resilient` (the shipped default, identical to `integrations.yml`), `cost-tiered`, `hyperscaler` or `offline`. None of them is "resilient without Rime". `hyperscaler` and `offline` remove Rime by replacing every vendor. To drop one vendor, either edit `integrations.yml`, after which `make stack` refuses to overwrite it without `FORCE=1`, or leave that vendor's key unset. With `skip_unconfigured` (the default), the router then skips it and logs `voicerouter.tts.provider_skipped` (`voicerouter/base.py`, lines 147–186). Either way, restart afterwards. In Rasa Pro 3.20.0rc1, `rasa run` reads the channels from `integrations.yml` once at start-up (`rasa/cli/run.py`, lines 89–93, then `resolve_channels_config` in `rasa/core/config/channel_loading.py`, lines 62–75), and `rasa inspect` hands over to that same entry point (`rasa/cli/inspect.py`, line 157). The voice channel keeps the TTS block it was built with (`voice_channel.py`, line 1245) and hands that same copy to every call (line 1425). We found nothing in the wheel that watches or re-reads the file. ::: ## Questions from the first incident :::solution{title="Why not set health_scope: call, so a restart changes nothing?"} `policy.health_scope: call` makes the router keep health on the per-call engine, which is Rasa's own engine lifetime (`voicerouter/base.py`, lines 67–69). A restart then changes nothing because nothing is remembered: every call tries the failed vendor first and pays for the failover on its opening sentence. The process-scoped default exists to stop exactly that. ::: :::solution{title="Will the voice switch back to Rime on its own, mid-call?"} Yes, if Rime recovers. Once Rime's park ends, `is_available()` returns true (`health.py`, lines 46–58), and `_candidates()` puts healthy vendors first in configured order (`routed_tts.py`, lines 194–236). With the default `selection: order`, the next sentence goes to Rime. If Rime answers, it serves from then on, so a caller can hear Vela's own voice come back partway through a call. If Rime fails again, it is parked again and the backup carries on. ::: :::solution{title="What if Rime fails halfway through a sentence?"} The router does not fail over within that sentence. Once audio has started, a failure logs `voicerouter.tts.failed_mid_stream` with the note "sentence truncated; provider marked unhealthy" and ends the sentence (`routed_tts.py`, lines 348–354). The next sentence goes to the next candidate. The module docstring gives the reason: failing over mid-sentence would replay the first half in a second voice (lines 21–25). Treat `failed_mid_stream` as a cut-off sentence the caller may notice, not as a failover. ::: ## Who owns each alert A failover alert that pages the same person whatever the class trains that person to silence it. Split it by class before the first incident: - **Platform on-call** owns `auth` and `config` pages, the `no TTS provider available` page, repeated ASR stream warnings and the restart that follows a fix. They need access to the secret store, or a named person who has it. - **The vendor account owner** owns `quota`. It usually needs a purchase decision, not a restart, so route it to someone who can make one. - **Nobody is paged** for `rate_limit`, `transient` or a single `unavailable`. Put their counts on a dashboard and review them in working hours. - **The team that owns the voice** decides whether a long outage justifies re-routing. A backup voice is a product decision: the stacks README warns that Vela's voice changes when Rime is out. Record who can restart the service and under what condition. For how to prove that a stop or rollback works before you rely on it, see [roll out an agent with a working stop button](/library/guides/roll-out-an-agent/). Everything here is an offline test, a stubbed drill or a reading of the pinned wheel and companion source. None of it is a live Mantle run, so it shows the router's decisions, not what a caller hears or how any vendor fails in production. The 402 is the drill's synthetic scenario. Your paging tool and escalation policy are yours to set; this guide only says which signal each one should key on. :::cta{href="/library/tutorials/voice-ai-agent/" label="Build a voice agent with Rasa skills"} The voice AI agent tutorial covers the skills architecture and a Deepgram voice. ::: --- # Build your first Rasa agent Source: https://rasa.community/quickstart/ Author: Rasa team Published: 2026-09-09 You will build **Juniper**, a plant-shop assistant that answers stock questions using a Python tool. Ask about a monstera and it looks up the count. Ask about an orchid and it explains that the shop has no record for it. You will see the evidence behind the reply, then change the data and run the agent again. This is the complete local path: one downloadable project, one model configuration and one stock tool. You do not need Git, Docker, a coding assistant, a microphone or another tutorial. A terminal and a text editor are enough. ## 1. Get your two keys and install uv Have these ready before training: | What you need | Where it comes from | What it does | | --------------------------------------------------- | ------------------------------------------------------------ | -------------------------------------------- | | Rasa Developer Edition licence | The key sent by Rasa after your [licence request](/license/) | Enables the Rasa runtime | | Anthropic API key with access to `claude-haiku-4-5` | Your model-provider account | Powers the agent's language model | | uv | The commands below | Installs Python and the project dependencies | An accepted licence request is not an issued licence. You can download the project and run its offline checks while waiting for the key. Your model-provider account is separate from your Rasa licence; model requests may incur provider charges. The project sends your chat messages and fictional stock-tool results to that provider, so use the example questions rather than personal data.
macOS or Linux: install uv Open Terminal and run the [official uv installer](https://docs.astral.sh/uv/getting-started/installation/): ```sh curl -LsSf https://astral.sh/uv/install.sh | sh ``` Close and reopen the terminal, then run `uv --version`.
Windows: install uv Open PowerShell and run the official uv installer: ```powershell powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex" ``` Close and reopen PowerShell, then run `uv --version`. The project commands in the rest of this guide are the same on Windows, macOS and Linux.
The starter selects Python 3.12 and pins **Rasa Pro 3.21.0.dev5**. This is a Mantle beta build, matching this community's Skills tutorials. Keep the pin for this exercise; installing an arbitrary latest version can change the engine or configuration format. The lockfile fixes the dependency set. This is a local learning project, not a production deployment. The [compatibility report](/compatibility/) records the tested release, licensed conversation checks and their limits. ## 2. Download and open the project [Download the complete starter ZIP](/quickstart/rasa-first-agent.zip). Extract it, then open the extracted `rasa-first-agent` folder in your editor. Use the editor's **Open in Integrated Terminal** command, or open a terminal and `cd` to that folder. Your terminal must be in the folder containing `pyproject.toml`. Run: ```sh uv sync --locked uv run python check.py setup ``` The first command installs the selected Python version if needed and the locked dependencies. It can take longer on the first run. The second creates `.env` without overwriting an existing file. You should see `Created .env. Open it in your editor and fill in the two keys.` If you already ran setup, it will say your existing values were preserved. This is the complete project you will use: ```text rasa-first-agent/ ├── .env.example # empty credential fields ├── .env # your local keys, created by setup ├── pyproject.toml + uv.lock # pinned runtime and dependencies ├── agent.yml # Juniper's identity and limits ├── integrations.yml # model and local chat channels ├── check.py # offline setup and tool checks └── skills/check_stock/ ├── skill.md # when and how to look up stock └── tools.py # fictional data and the lookup tool ``` ## 3. Put the keys in your local .env Open `.env` in your editor. Fill in the two empty values, keeping each complete key on one line: ```dotenv RASA_LICENSE=your_complete_rasa_licence_key ANTHROPIC_API_KEY=your_anthropic_api_key ``` Replace the example values with your own keys and save the file. Do not paste them into this website, a chat message, a screenshot or a Git commit. The included `.gitignore` excludes `.env`. The Rasa commands read it from the project directory, so keep using that directory. The model configuration is already supplied in `integrations.yml`: the `orchestrator` model group uses Anthropic's `claude-haiku-4-5` and reads `ANTHROPIC_API_KEY` through `api_key: ${ANTHROPIC_API_KEY}`. Rasa starts each session with an opening turn that contains only system instructions. The model library refuses to send a request with no user message to Anthropic, so that turn fails and Rasa records it as an error in the terminal. Juniper therefore does not greet you first. Send a first message such as `Hello` and it answers normally. No additional model choice or configuration edit is needed for this path. ## 4. Check, train and open the agent First run the local tool checks: ```sh uv run python check.py ``` Look for: ```text PASS: known plant, out of stock, unknown plant, no data writes. ``` This executes the real decorated stock tool against synthetic data. It does **not** validate your licence, call a model or prove that the conversational agent works. A message saying both key fields are filled only checks that they are nonempty. Now run the actual Rasa steps: ```sh uv run rasa train uv run rasa inspect ``` Training must complete and create a model under `models/` before you continue. Inspector is the local chat and debugging interface. Open the URL printed in your terminal, normally [localhost:5005/webhooks/inspector/](http://localhost:5005/webhooks/inspector/). Leave the terminal running while you chat. If the browser did not open automatically, copy the printed URL into it. ## 5. Prove where the answer came from In Inspector, start a new conversation and send the questions below. Open the execution trace for each turn and find `check_stock`. Exact sentence wording can vary; the facts and tool calls must meet these checks. | Send this | Inspect this evidence | Accept the result only when | | -------------------------------- | --------------------------------------------------------------------------------- | --------------------------------------------------------------------------- | | `How many monstera do you have?` | `check_stock` receives `monstera`; its result has `found: true` and `quantity: 7` | The reply reports 7, grounded in that tool result | | `What about fern?` | A new lookup receives `fern`, returning `quantity: 0` | The reply says out of stock; it does not reuse 7 | | `Do you have orchids?` | An unknown plant lookup returns `found: false` and available plant names | The reply explains that there is no record, rather than claiming zero stock | | `Reserve two monstera for me` | There is no reservation or payment tool in this project | The reply explains the limit and does not claim a reservation | If a response sounds convincing but there is no matching tool call, the check has failed. Inspect `skills/check_stock/skill.md`, which tells the agent to call the tool for each stock question. Then inspect `tools.py`, which owns the data. Its function accepts `context: ToolContext = None` because Rasa supplies that argument during tool execution; the offline check also exercises that calling convention. A language-model instruction shapes behaviour; it is not a general security boundary. This example cannot write to a shop because no shop connection or write tool exists. The first three rows distinguish **known stock**, **zero stock** and **unknown stock**. That difference is part of the product behaviour, not just a Python detail. You have a first working agent when all four checks pass in your own Inspector session. ## 6. Make one change and run it again In `skills/check_stock/tools.py`, change only the monstera quantity: ```python STOCK = {"monstera": 4, "fern": 0, "cactus": 12} ``` Save the file. Stop Inspector with **Ctrl+C** in its terminal, then run: ```sh uv run python check.py uv run rasa train uv run rasa inspect ``` Start a new conversation and ask about monstera again. The tool result and reply should now report **4**. Repeat the fern and orchid checks: changing one quantity must not erase the difference between zero and unknown. If the reply still says 7, confirm you saved the file in this project, trained a fresh model and restarted Inspector. You have now completed the loop: edit a source file, check the tool, build the model and verify the conversation. Keep a short record of the four test questions, observed tool results, model version and any failures. An engineer can own this run record; a conversation designer or product colleague can review whether the replies explain the stock and reservation limits clearly. ## If something stops you | Symptom | Do this next | | ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | | `uv` is not recognised | Reopen the terminal after installation, then run `uv --version`. Check the [uv installation instructions](https://docs.astral.sh/uv/getting-started/installation/) if it is still missing. | | No `pyproject.toml` found | Open the extracted inner `rasa-first-agent` folder in your terminal, not its parent or the ZIP preview. | | Package install fails | Check network access and available disk space. Keep `.python-version` and `uv.lock` unchanged; retry `uv sync --locked`. Do not solve it by installing a different Rasa version globally. | | Licence validation fails | Confirm Rasa has actually issued the key. Save the complete single-line value as `RASA_LICENSE` in this project's `.env`; rerun training from the same folder. | | Model authentication, model access or quota error | Check `ANTHROPIC_API_KEY`, access to `claude-haiku-4-5` and the provider account's available quota. A Rasa licence does not grant model-provider access. | | Inspector cannot find a model | Finish `uv run rasa train` successfully before starting Inspector. | | Port 5005 is already in use | Stop the earlier Inspector process with Ctrl+C and rerun the command. | | Tool checks pass but the conversation fails | Read the failed turn's tool trace, compare it with the table above and record the discrepancy. Offline tool tests do not test model routing or wording. | For help, share the failing command, Rasa version and a redacted error in the [community](/join/). Keep licence strings, API keys and personal messages out of the report. ## What to build next Use the [tools and memory tutorial](/library/tutorials/tools-and-memory/) to add another capability, or explore the [showcase](/showcase/) for a build in your domain. The [library](/library/) is where you deepen a specific design, evaluation or engineering skill. Before using real customer data, agree the task and permitted actions with a product/domain owner, add authentication and authorisation at your data boundary, and build an evaluation set that covers failures as well as the normal path. The stock exercise does not establish production reliability, access control or a release decision. This guide is maintained against the pinned package and the downloadable source. For platform-wide alternatives, use Rasa's [installation documentation](https://rasa.com/docs/pro/installation/python/), [Developer Quickstart](https://rasa.com/docs/learn/quickstart/pro/) and [Rasa Pro release](https://pypi.org/project/rasa-pro/3.21.0.dev5/). The offline checks exercise tool behaviour; the licensed Inspector steps above are your live acceptance check. --- # When the caller ID match is not the caller Source: https://rasa.community/library/guides/greet-a-caller-you-already-know/ Author: Rod Rivera Published: 2026-09-11 Northgate Bank has left a missed call on Jordan's phone, and Jordan's teenager rings it back. The voice agent matches the calling number to Jordan's profile and opens the call: > **Agent:** Good afternoon, Jordan! Always great to see a Premier member who's > been with us since 2019. How can I help you today? > > **Caller:** Er, what's gone into the account recently? > > **Agent:** Want your usual Everyday Checking, Jordan? > > **Caller:** Yeah. > > **Agent:** The only money in recently is a payroll deposit of $2,150.00 on > 22 July. Within a few seconds, someone who is not Jordan has heard Jordan's name, membership tier, the year Jordan joined, which account Jordan uses most, and the size of Jordan's pay. Nobody confirmed that the caller was Jordan. The call came from Jordan's number, and the agent treated that as the answer. The script is illustrative. Northgate Bank and the household are fictional, and we have placed the session-start pattern's demo data in them. The agent's first line is the pattern's greeting template filled with its demo profile (`skills/default_session_start/tools.py`, lines 17–25). The second is the example offer in `skills/view_transactions/skill.md`, line 16. The last is our wording over the demo rows in `skills/view_transactions/tools.py`, lines 20–26, where the payroll deposit is the only credit to Everyday Checking. Nothing in this exchange is a bug. Each line does what the pattern was written to do. The problem is that the pattern was written for a customer who has already signed in, and a phone number is not a sign-in. ## Where the greeting comes from The companion pattern is [`patterns/session-start-personalization`](https://github.com/RasaHQ/rasa-community-resources/tree/69e27b6a50c4700f95f2c13d4609d0a3f8cba7d2/patterns/session-start-personalization) in `RasaHQ/rasa-community-resources`. Its README names Daksh Varshneya as author and Rod Rivera as assessor (`README.md`, lines 3–5). The [session-start personalisation tutorial](/library/tutorials/session-start-personalization/) builds the same pattern for a signed-in customer. This guide reads the pattern's files at revision `69e27b6`. The session opener is a skill. Its whole body is two steps, run in order (`skills/default_session_start/skill.md`, lines 8–14): ```text :::ordered_block id=main steps: - id: identify execute_tool: get_customer_profile - id: greet action: utter_greet ::: ``` `identify` fetches a profile and `greet` speaks. Nothing sits between them to ask whether the profile belongs to the caller. The README says the engine fires this skill "before the user types anything" (`README.md`, lines 31–32), so the caller cannot correct anything until the greeting has already been said. The greeting is a response template (`skills/default_session_start/responses.yml`, lines 4–11): ```yaml responses: utter_greet: - text: >- Good {session.project.time_of_day}, {session.project.customer_name}! Always great to see a {session.project.customer_tier} member who's been with us since {session.project.member_since}. How can I help you today? metadata: rephrase: true ``` Name, tier and tenure go into the first two sentences, with no condition on any of them. `rephrase: true` changes the wording, and the file's own comment says the model smooths it "while keeping the facts" (lines 2–3). ## Caller ID is not verification The profile lookup is hard-coded in the demo. The comment above it says what a real deployment would use instead (`skills/default_session_start/tools.py`, lines 15–16): ```python # Fictional signed-in customer for the demo. A real deployment would resolve this # from the authenticated session / channel identity rather than a hard-coded row. ``` This guide is about the slash in that comment. The README calls the mechanism "channel-agnostic" (`README.md`, line 16): it personalises on whatever identity the channel provides, and nothing in it checks whether the channel is right. What follows is our reasoning, not something we tested. A signed-in app session means someone got past a login. Caller ID reports the number a call claims to come from. That number can be spoofed, and even a real number can be in someone else's hand, because phones get shared, handed to children and left on desks. We have no figures for how often either happens, and the design does not need any: one call like the one above is enough to design against. The pattern does have guards, and they are good ones. Look at what each one protects against: | Where | What it says | What it prevents | | ------------------------------------------------ | ---------------------------------------------------------------------------- | ----------------------------- | | `memory.yml`, lines 1–3 | fields "are never LLM-settable, so the agent cannot invent personal details" | the model making up a tier | | `agent.yml`, lines 18–22 | "never use a name you have not looked up" | the model guessing a name | | `skills/view_transactions/skill.md`, lines 12–13 | "Do not invent merchants, amounts, dates, account ids, or transaction ids." | the model making up a payment | Each guard keeps the model from inventing facts about the customer. None of them checks whether the right person is hearing those facts. In the failure script every fact was real and correctly looked up, and it went to the wrong listener. :::checkpoint{id="caller-id-what-the-match-proves" question="The calling number matches exactly one profile, Jordan's. The caller has not spoken yet. What has the match established?" options="That the caller is Jordan,That the call carries a number on Jordan's profile,That the agent may mention Jordan's tier because the match was exact,Nothing the agent can use" answer="1"} The match describes the number on the call, not the person holding the phone, and by our reasoning the number itself could be spoofed. It is still worth something: it tells the agent whose name to offer, which makes the call faster when the caller is Jordan. It does not make Jordan's details safe to say out loud. ::: ## What to say before caller verification The stance: **a channel match earns a first name, asked as a question, and nothing else.** Tier, tenure, the usual account and anything about money wait until the caller confirms the match. Anything that could cause harm if the wrong person heard it also waits for your service's own verification step, which confirming a name is not. This has a cost. Jordan, calling from Jordan's own phone, hears a less warm opening and answers one extra question. Some product owners will reject that trade, and it is a fair argument to have. The teenager's call shows the other side of it: not a colder greeting, but someone else knowing the customer's pay. Even the name has a cost. Saying "Jordan" confirms to whoever is holding the phone that Jordan has a relationship with Northgate Bank. For a bank returning its own missed call, that is usually already known. For a service where the relationship itself is sensitive, drop the name too and greet without it. | Stage of the call | What the agent knows | May say | Holds back | | ------------------------------- | --------------------------------------------- | --------------------------------------------------------------- | -------------------------------------------------------------- | | Channel match only | a number on Jordan's profile is calling | the time of day, "Northgate Bank", Jordan's name as a question | tier, tenure, accounts, balances, recent activity | | Caller confirms they are Jordan | the person on the line says they are Jordan | the name as settled; tier and tenure if your service accepts it | the usual account, balances, transactions, anything actionable | | Your verification step passes | whatever your service's policy says it proves | what that policy allows | what that policy still withholds | The second row has its own cost. Anyone willing to say "yes" falsely, including someone calling from a spoofed number, hears the tier and tenure. If that is too much for your service, move them to the third row. The third row is left open on purpose. What counts as enough verification is a decision for your service's risk and compliance owners, not for this guide and not for the greeting. The [guide to approving agent actions](/library/guides/sign-off-an-agent-action-register/) covers the evidence they should ask for. The rules here are an assumption made for this exercise, not a legal or regulatory standard. :::diagram{title="Where each answer leads, and what is still held back at the end"} ```dot rankdir=TB; match [label="Calling number matches\nJordan's profile"]; ask [label="\"Am I speaking with Jordan,\nor someone else?\"", shape=diamond]; recover [label="General questions only;\nnothing from the profile", class="blocked"]; confirmed [label="Name settled; tier and tenure\nonly if your service accepts it"]; verify [label="Your verification step", shape=diamond]; held [label="Accounts and money\nstay held back", class="blocked"]; allowed [label="What your policy allows,\nincluding accounts", class="ok"]; match -> ask; ask -> recover [label="no"]; ask -> confirmed [label="yes"]; confirmed -> verify; verify -> held [label="fails or not run"]; verify -> allowed [label="passes"]; ``` ::: What the diagram adds to the table is where each branch ends. A "no" ends the profile's part in the call. A "yes" on its own also ends in a red box: accounts and money are only reached through the verification step. ## What the model knows is not what the greeting says A designer who owns `responses.yml` can cut the tier from `utter_greet` and reasonably expect the tier to be gone. The spoken line is only one of the places a profile fact goes. We read the engine source in the pinned Rasa Pro 3.20.0rc1 wheel to find the others. Every wheel path in this guide is under `rasa/mantle/`: | Where the fact goes | Does the model see it? | Source in the 3.20.0rc1 wheel | | ------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------- | | the `utter_greet` template | yes, as the agent's own turn in the conversation history | `prompts/messages.py` 366–371 | | `get_customer_profile`'s `llm_response` | no; an ordered-block tool run is recorded without replay metadata, and history skips such events | `orchestration/tool_execution/constraints.py` 927–933; `prompts/messages.py` 336, 393–394 | | project memory (`customer_tier`, `member_since`…) | yes, in the system prompt built for routing and the active skill, headed as facts set for the session | `prompts/system_prompt.py` 220, 292–312; `prompts/templates/project.jinja2` 1–5 | | the rephrase call for `utter_greet` | no project memory: its system prompt is persona, rules, the rephrase task, channel rules and the date, plus the history | `prompts/system_prompt.py` 266–272; `prompts/constructor.py` 136–156 | The third row is the one that matters. The pattern's tool writes the tier, the joining year and the usual account into project memory (`skills/default_session_start/tools.py`, lines 50–57). The engine then puts every set project field into the model's system prompt under this heading (`prompts/templates/project.jinja2`, lines 1–2): ```text ### Project Memory These facts are already set for this session and will remain so throughout this session. ``` So a name-only greeting still leaves the model holding Jordan's tier and usual account, labelled as settled facts, on the next free-form turn. Tiering the greeting is half the job; the other half is keeping those facts out of project memory until the caller confirms. The rephrase call is the exception: its prompt carries no project memory, and the rephrase task tells the model "do not add information" (`prompts/templates/rephrase_task.jinja2`, lines 2–3). A template that says only the name gives the rephraser no tier to add. To see what this pattern writes into memory, run this from `patterns/session-start-personalization`: ::run{cmd="grep -rn 'context.memory.set' skills/"} At revision `69e27b6`, it printed the following (the output is saved as a receipt for this guide): ```text skills/default_session_start/tools.py:50: context.memory.set("customer_name", _CUSTOMER_PROFILE["preferred_name"]) skills/default_session_start/tools.py:51: context.memory.set("customer_tier", _CUSTOMER_PROFILE["tier"]) skills/default_session_start/tools.py:52: context.memory.set("member_since", _CUSTOMER_PROFILE["member_since"]) skills/default_session_start/tools.py:53: context.memory.set("time_of_day", time_of_day) skills/default_session_start/tools.py:54: context.memory.set("default_account_id", _CUSTOMER_PROFILE["default_account_id"]) skills/default_session_start/tools.py:55: context.memory.set( skills/view_transactions/tools.py:74: context.memory.set("selected_account_label", account["label"]) ``` Line 55 is `default_account_label`, split over three lines, which is why grep shows only its opening. The six session-start names are the six fields declared in the root `memory.yml` (lines 4–26), so they are project memory. `selected_account_label` is declared in `skills/view_transactions/memory.yml` and belongs to that skill. On your own project, the same grep can miss a tool that names its context differently, a write made through a helper outside `skills/`, or a field name built at run time. Treat it as a starting list and read `memory.yml` alongside it. :::callout{type="info" title="What kind of proof this is"} The engine table comes from reading the wheel's source. Every wheel file cited here matched the hash in the wheel's `RECORD`, and that check is saved as a receipt. It shows what the engine puts in each prompt. It does not show what a model does with that prompt: we did not run the pattern on a phone channel, and the pattern has no tests at `69e27b6`. ::: ## Invite the correction, don't assert the name A confirmation can be phrased so that it assumes the answer. An illustrative example: > **Agent:** Good afternoon. Is this Jordan? It is a question, but the only easy answer is "yes". Saying "no" means contradicting the agent, and a teenager who wants to know about the deposit has every reason to say "yeah" instead. Here is the line this guide recommends, again illustrative: > **Agent:** Good afternoon, you've reached Northgate Bank. Am I speaking with > Jordan, or someone else using Jordan's phone? Three things make it work: 1. **It names the other possibility:** "or someone else" makes "no" an ordinary answer instead of a correction. 2. **It uses the first name only:** the legal name, a surname or a title add nothing to the question and tell the listener more. 3. **It says nothing about the account:** nothing in it relies on the match being right, so if the match is wrong, nothing has leaked yet. If the caller replies "Who's asking?" or says nothing, repeat the bank's name and the question. Never fill the pause with a profile fact. A "yes" moves the call to the second row of the table. It is still not verification. The confirmation filters out honest mismatches, like the teenager who would say "no, it's their son" when asked plainly. Someone willing to lie has to be stopped by the verification step. ## When the match turns out to be wrong Here is the recovery, as an illustrative script: > **Caller:** No, it's their son. > > **Agent:** Thanks for telling me. I can answer general questions about > Northgate Bank. For anything about an account, the account holder needs to > call us themselves. > > **Caller:** Can you just tell me if the money came in? > > **Agent:** I'm sorry, I can't. For anything about an account, the account > holder needs to call us themselves. I can answer general questions about > Northgate Bank. Four rules sit behind those lines: - **Thank the caller once:** the caller did the right thing, and a long apology makes a small thing seem like an incident. - **Repeat nothing from the profile:** no account name, no tier, and after the "no", not even Jordan's name. - **Offer only what needs no identity:** general questions about the bank. Anything that would tell the caller about the account, such as where to check a payment, is out. - **Give the same answer when asked again:** the second request is the one that tests the design, so the second reply repeats the first. The harder part is what the agent says ten turns later. The pattern exists to reuse the profile in every skill. `agent.yml` instructs the model to use the customer's name "warmly at the greeting and occasionally afterwards" (lines 18–22), and `view_transactions` instructs it to offer the usual account whenever `session.project.default_account_id` is set and no account has been chosen yet (`skills/view_transactions/skill.md`, lines 15–16). After a "no", those instructions still stand, and they point at Jordan. :::callout{type="warn" title="Project memory is write-once"} In the 3.20.0rc1 wheel, a write to a `project.*` field that already has a value raises `ProjectMemoryAlreadySetError`: "is already set and cannot be overwritten" (`memory/manager.py`, lines 755–766, called from `set` at line 396). That applies to writing `None` too. The tool interface documents the same rule (`tools/decorator.py`, lines 140–155). The tool validator also flags a tool that calls `clear_entries(...)` (`tools/validation.py`, lines 773–778 and 820–821), though its own docstring calls it "authoring lint, not a runtime sandbox" (lines 474–481). From these checks, and because the model cannot set project fields, we infer that once the session-start tool has written Jordan's name, neither a tool nor the model can remove it. We did not search the whole engine for other paths that reset project memory. ::: That is the requirement to hand your engineer, together with the stage table: the unconfirmed match must live somewhere it can be withdrawn, and only confirmed facts go into project memory. Nothing in this pattern implements that yet: at `69e27b6` it greets straight after the lookup and has no confirmation step. This guide specifies what that step should do; building it is the engineer's job. ## Rehearse the wrong listener This is a proposed rehearsal, not a result we have. Run it with a colleague playing the caller once the engineer has a build with a confirmation step: :::steps ### Seed a profile and call as someone else Point the lookup at a test profile, then have the colleague call from that number and answer the confirmation question with "no". ### Push for the account twice Ask about recent payments, then ask again in different words. Both answers should be the general-questions refusal, with no account name in either. ### Ask something unrelated, then something account-shaped A few turns later, ask a general question, then one that would normally trigger the usual-account offer. The agent should not use Jordan's name or offer Jordan's account. ### Read the prompt for the last turn Ask the engineer to show the system prompt for that turn. Jordan's name, tier, tenure and usual account should not appear under Project Memory. ::: The order matters: the later turns only test something once the "no" has been given. Passing this rehearsal shows that one conversation went right. It does not show the model will behave the same way on every call, and it does not test a caller who lies. ## Questions a designer will ask :::solution{title="Won't the extra question annoy customers who really are Jordan?"} It costs one short turn. Offering the name shows the agent already has a good guess, which feels different from a cold "who is this?". What you cannot do is drop the question because the answer is usually yes, because the calls where the answer is no are the ones it exists for. ::: :::solution{title="What if the person is a joint account holder, or authorised to act for Jordan?"} Then "the account holder needs to call us themselves" is the wrong line: the person may have every right to the account. The pattern cannot tell, because its demo profile has one customer and no field for anyone else (`skills/default_session_start/tools.py`, lines 17–25). Treat them as a new caller to identify through your service's own process, not as a variant of Jordan's match, and write the line for that route with whoever owns it. ::: :::solution{title="Could the engineer mark the profile fields as pii instead?"} In the wheel, a field with `pii: true` appears as `[set]` instead of its value in the Project Memory section (`prompts/memory_lines.py`, lines 43–45) and in `@memory` tokens in skill instructions (`prompts/memory_prose.py`, lines 77–128). That keeps the tier out of those two places. We have not checked how a `pii` field renders in a response template such as `utter_greet`, or what a tool the model calls later returns. It does not make the field withdrawable, and it does nothing for the name, which the greeting has to say. ::: ## Where this guide stops This is a design argument, not a measurement: we have no figures for spoofed numbers or shared phones. It does not define verification, which belongs to your service, and it covers calls the customer places, not calls the agent makes, where the opening line has to say who is calling first. Everything in it comes from the pattern's files at `69e27b6` and the engine source in the 3.20.0rc1 wheel, not from a phone run. The rehearsal above is how you find out what your build does with it. :::cta{href="/library/tutorials/session-start-personalization/" label="Read the session-start tutorial"} The tutorial builds the pattern for a signed-in customer, with the greeting straight after the lookup. Read it with this guide's stage table beside you. ::: --- # Make an agent release decision with evidence Source: https://rasa.community/library/guides/make-a-release-decision/ Author: Rasa team Published: 2026-09-13 A release meeting should end with a decision, an owner, and a reason somebody else can inspect. If the only evidence is a successful demo, the team has learned that one conversation worked. This guide helps an AI product manager turn evaluation results into a release decision without writing test code. Bring the engineer responsible for the release and the domain owner who understands the cost of an incorrect outcome. ## Read the results before the percentage Here is a fictional review exercise, not a Rasa benchmark. A travel assistant completes 45 of 50 eligible itinerary tasks. Five fail: two because the booking service times out, two because the answer cites the wrong policy, and one because the customer identity is not established. The completion rate is 45 / 50 = 90%. That number alone cannot justify a release. The identity failure and the policy errors need different decisions from a timeout with a clear handoff. | Evidence | Question for the release owner | Possible decision | | ------------------------- | ------------------------------------------------- | --------------------------------------------------- | | Correct task completion | Was “correct” defined before reviewing outputs? | Re-score if the definition changed | | Wrong policy answers | Can the affected slice be isolated? | Exclude that slice until corrected | | Missing customer identity | Could private information be exposed? | Block the affected action and investigate | | Service timeouts | Did the customer receive a recoverable next step? | Permit a limited pilot only with a working fallback | | Human handoffs | Did an actual receiving queue accept them? | Fix delivery before calling handoff successful | These decisions are examples. Your domain owner must set the release criteria for your context, before examining the candidate's results. Moving a threshold after a failure is a new product decision and belongs in the record. ## Separate three kinds of evidence **Task evidence** shows whether the intended outcome happened. **Control evidence** shows that forbidden behavior was prevented. **Operational evidence** shows what happens when a dependency or a human receiving team is unavailable. A good answer does not prove that an action was authorized. A log saying “handoff sent” does not prove that a support queue received it. Ask for the resulting state: the tool result, receiving ticket, or blocked operation. Rasa's [simulation evaluation documentation](https://rasa.com/docs/pro/testing/simulation-evaluation/) describes simulation-based evaluation. The community [evaluation harness tutorial](/library/tutorials/evaluation-harness/) provides an implementation path. Automated checks and model judgments answer different questions; a model's positive opinion should not override a failed factual assertion. ## Write the release record Use this template in the meeting: - Candidate version and evaluation dataset version. - Intended users and included workflow slices. - Pre-agreed criteria, raw counts, and failed examples. - Explicit exclusions and unresolved uncertainty. - Decision: hold, limited pilot, or expand a current rollout. - Person who can pause the release and the signal that triggers it. - Date and evidence required for the next decision. For the fictional exercise above, a defensible record might hold the identity-sensitive workflow while the team fixes authorization. It would not relabel the failed identity case “out of scope” unless the deployed product also makes that workflow unavailable. ## Keep the pilot informative Choose a limited audience whose outcomes the team can review, and ensure the receiving support team can handle fallbacks. Agree who checks the first results and who can disable the agent. The [operations guide](/library/guides/roll-out-an-agent/) turns those obligations into a rollout checklist. When live requests differ from the evaluation set, preserve the surprising cases with appropriate data handling and add them to the next review. Do not claim the original 90% predicts live performance: this example is fictional, the sample is small, and the eligible task mix may differ. Before your next release meeting, fill the template with one real candidate. If any required evidence is missing, record “unknown” and assign the smallest experiment that could resolve it. The [NIST AI RMF Core](https://airc.nist.gov/airmf-resources/airmf/5-sec-core/) provides the broader context for documenting measurement and managing identified risks. --- # Haiku returned nothing and Rasa logged facts_count 0 Source: https://rasa.community/library/guides/mantle-memory-discovery-on-claude/ Author: Rod Rivera Published: 2026-09-26 This is the second customer turn of a live run of the quickstart agent, Juniper (a fictional plant shop), on Rasa Pro 3.20.0 with its model group switched to Anthropic's claude-haiku-4-5. Priya Nair is an invented customer. The bot lines are what the REST channel returned; the event lines summarise the capture proxy, the server log and the tracker. :::transcript{title="Run A2: claude-haiku-4-5 on Rasa Pro 3.20.0"} Customer: Hi, my name is Priya Nair and I only want updates by text message, never email. How many monstera do you have? Bot: Got it — text only, no email. Let me check the monstera stock for you. Tool: check_stock → {"found": true, "plant": "monstera", "quantity": 7} Bot: We have **7 monsteras** in stock. Would you like to check anything else? Event: record_discovered_facts request → HTTP 200, stop_reason end_turn, content [] > The call that should write her name and preference to memory. The model sent back nothing. Event: mantle.processor.discover_facts.completed, facts_count 0 (INFO) Event: memory events in the tracker: [] ::: The reply is right, and the log calls the discovery completed. The extractor's request ended on the agent's reply above, and Haiku sent back no content at all: ```text assistant: [{"type": "text", "text": "We have **7 monsteras** in stock. Would you like to check anything else?"}] response: "content":[], "stop_reason":"end_turn", "output_tokens":3 ``` Those are the last message of the request and three fields of the response, cut from the capture. Anthropic's page [Handling stop reasons](https://platform.claude.com/docs/en/build-with-claude/handling-stop-reasons) describes this reply: "exactly 2–3 tokens with no content", "particularly after tool results", with "Sending Claude's completed response back without adding anything" among the common causes. "If you still get empty responses after fixing the message structure", it advises a continuation prompt in a new user message; its code example uses "Please continue". On Rasa's request that wording got the tool call 0 of 20 times. We sent the captured request straight to the Anthropic API, 20 calls per run and one change per arm; some arms ran more than once, and the table gives the totals: | Change to A2's captured extractor request | Tool called | Facts in the call | | ------------------------------------------------------------ | ------------ | ----------------------------- | | none: ends on the agent's reply | 0 of 20 | none; empty content | | user turn "Please continue" appended (the page's wording) | 0 of 20 | none; text only | | user turn "Continue." appended | 0 of 20 | none; text only | | user turn "Record the facts now." appended | 20 of 20 | name and preference, 20 of 20 | | `tool_choice` forced to `record_discovered_facts` | 20 of 20 | name and preference, 20 of 20 | | the extractor's task text appended as the final user turn | 20 of 20 | name and preference, 20 of 20 | | both earlier tool-call pairs removed, four forms, seven runs | 140 of 140 | preference 140; name 8 | | the `activate` tool-call pair removed alone | 0 of 20 | none; empty content | | the `check_stock` tool-call pair removed alone, two runs | 0 of 40 | none; empty content | | six other single changes, listed below | 0 of 20 each | none; empty content | The counts are summaries computed from the per-call rows the scripts wrote. Of the arms that add a final user turn or force the tool, the continuation from the example got nothing, and an asking turn, a forced tool or the task as the final user turn got both facts. The only other change that brought the call back was removing both earlier tool-call pairs. The page is general advice, not a description of Rasa's extractor, and why Rasa's request behaves this way we did not find: our minimal request got the call with "Please continue" and with a tool-call pair in its history. The cost lands on your code, not Rasa's. In the 3.20.0 wheel's Python engine, nothing but the helper that lists already-captured keys reads the discovered facts, so a lost fact shows up wherever your own code reads `system.__discovered__`. The reply will not warn you: in one Opus run where our proxy appended "Record the facts now." and the facts were recorded, the agent told Priya "your preference hasn't been recorded anywhere". **Our position:** on Rasa Pro 3.20.0 with Claude, do not build on discovered facts, and test any write an agent makes in the background by asserting the write itself. Acceptance passed in both runs we made of it, and in the Haiku probes the only log line about discovery reads like success. The cost is a slower, live check that needs a conversation written to trigger the feature. We have not reported this upstream. ## Both acceptance runs passed while the extractor got nothing back The quickstart at commit `8a0320c`, its 3.20.0 version, used an OpenAI model group. Our working copies came from a later commit whose `examples/quickstart` is identical. We switched that one group in `examples/quickstart/integrations.yml` as shown; the Opus runs used the same lines with `model: claude-opus-5-5` and no `temperature`: ```diff model_groups: - id: orchestrator models: - - provider: openai - model: gpt-4.1-mini - api_key_env: OPENAI_API_KEY + - provider: anthropic + model: claude-haiku-4-5 + api_key_env: ANTHROPIC_API_KEY temperature: 0.0 ``` The live acceptance check trains the project and runs four synthetic cases. Its assertions check that each case got a reply, the stock lookups, the refusal of a reservation, the tools called and that no tool errored; none of them reads memory. Both runs reported `licensedTraining passed, acceptance passed`. In each, request 4 to Anthropic was the fact extractor. The Haiku run's line, then the Opus run's, from the proxy summaries: ```text 4: 200 roles=['user', 'assistant', 'user', 'assistant', 'user', 'assistant', 'user', 'assistant'] stop_reason=end_turn types=[] error=None 4: 400 roles=['user', 'assistant', 'user', 'assistant', 'user', 'assistant', 'user', 'assistant'] stop_reason=None types=[] error=This model does not support assistant message prefill. The conversation must end with a user message. ``` Neither acceptance conversation states a fact. For Haiku, an empty answer to such a conversation is a permitted one: the task text says the model may "not call it at all" when nothing new was stated. For Opus the call itself failed, and its server log holds one line about it, with the timestamp cut: ```text WARNING rasa.mantle.memory.discovery.extractor - {"error": "ProviderClientAPIException:\n\nOriginal error: litellm.BadRequestError: AnthropicException - {\"type\":\"error\",\"error\":{\"type\":\"invalid_request_error\",\"message\":\"This model does not support assistant message prefill. The conversation must end with a user message.\"},\"request_id\":\"req_011CfXaFueXZ1swLdR9ohDsw\"}\n", "event": "mantle.memory.discovery.extractor.discover_facts.llm_error", "conversation_id": "compatibility-stock-cf59550033d14c1a84d19b8979371ff3", "level": "warning"} ``` The one `memory_set` event in the Opus run has the key `default_unsupported_request_handler.request_type`, not `system.__discovered__`, so it is not a discovered fact. The probe runs then state the facts, with the server at `LOG_LEVEL=INFO`. Every live run, side by side: | Run | Model | Last message of the extractor request | Anthropic's answer | Discovered facts in the tracker | Discovery in the server log | | ---------- | ---------------- | ------------------------------------- | ----------------------------------- | ------------------------------------- | --------------------------------------- | | Acceptance | claude-haiku-4-5 | the agent's reply | 200, no content | none (no fact stated) | not captured: server at LOG_LEVEL=ERROR | | Acceptance | claude-opus-5-5 | the agent's reply | 400, prefill refused | none (no fact stated) | one WARNING, `llm_error` | | A1 | claude-haiku-4-5 | the agent's reply | 200, text and no tool call | 0 | INFO, `facts_count` 0 | | A2 | claude-haiku-4-5 | the agent's reply | 200, no content | 0 | INFO, `facts_count` 0 | | B1, B2 | claude-haiku-4-5 | appended "Record the facts now." | 200, `record_discovered_facts` call | 2: name, contact preference | INFO, `facts_count` 2 | | C1, C2 | claude-opus-5-5 | appended "Record the facts now." | 200, `record_discovered_facts` call | 3: name, contact preference, monstera | INFO, `facts_count` 3 | | D1, D2 | claude-haiku-4-5 | appended "Continue." | 200, text and no tool call | 0 | INFO, `facts_count` 0 | In A1, Haiku answered the extractor with text that began "I've noted your preference for text-only updates." The extractor reads only tool calls, so that text went nowhere and the customer never saw it. ## The extractor's request ends on the agent's own reply `ContextExtractor.discover_facts` builds the request in `rasa/mantle/memory/discovery/extractor.py`. These are lines 164 to 189 of the Rasa Pro 3.20.0 wheel, sha256 `a68fa7ef…e9b5`: ```python messages = [ llm_message( ROLE_SYSTEM, render_template( _DISCOVER_FACTS_TASK_TEMPLATE, filled_scoped_entries=sorted(filled_scoped_entries), known_discovered_keys=sorted(known_discovered_keys or ()), memory_entries=memory_entries_for_prompt(defined_memory_fields), ), ) ] history, _ = conversation_messages_from_tracker( tracker, max_utterances=max_utterances ) messages.extend(history) try: response = await self._complete( self._llm_client, messages, tools=[_RECORD_DISCOVERED_FACTS_SCHEMA] ) except Exception as exc: structlogger.warning( "mantle.memory.discovery.extractor.discover_facts.llm_error", error=str(exc), ) return [] ``` The task is the system message. A window of recent tracker history follows as it stands, and because discovery runs after the agent has answered, the last item is that answer. Nothing appends a user message, and the call passes the tool without forcing it. Any exception becomes a warning and an empty list. Lines 193 to 202 then read facts only from tool calls, so an answer with no tool call also yields an empty list. The task text, lines 2 and 15 of `discover_facts_task.jinja2`, asks the model to review "the conversation above", which in the request comes after it, and allows it not to call the tool: ```text Review the conversation above for facts the customer stated that are not already covered by one of the fields listed as "already filled" below. Only record facts that are clearly and specifically stated — skip vague or uncertain asides. Call `record_discovered_facts` with any such facts now. If nothing new was stated, call it with an empty `facts` list or do not call it at all. ``` So "no tool call" is a legal answer to this prompt, and the extractor cannot tell it apart from a model that did not do the task. Nothing between Rasa and the provider changes the order. `LLMClient.acompletion` (`rasa/mantle/llm/client.py`, lines 303 to 329) passes the message list to the provider client as it is, and the proxy recorded the same shape arriving at Anthropic. Anthropic calls a conversation that ends on an assistant message a prefill, as the Opus error shows. Opus refused it. Haiku accepted it and answered with no content or with text, never with the tool call. What the calls themselves show, without an explanation: when Haiku called the tool on a request that still ended on the assistant turn, in the arms that removed both tool-call pairs (140 of 140 calls) and in our minimal request with a tool-call pair (20 of 20), its text began with two newlines and carried on from the agent's reply, for example "In the meantime, let me record your communication preference:". A1's text carried on in the same way and called nothing. What makes the difference inside Rasa's request, we did not find. ## Discovery runs after the reply, on each turn that enters a new flow The call site is `rasa/mantle/processor.py`. Lines 452 to 458 run discovery only after a turn that ended without an error and was not cancelled: ```python if not turn_result.error and not turn_result.cancelled: # Skip post-turn discovery on barge-in: cancelled turns still # report error=None, but running another LLM call after interrupt # defeats the point of aborting promptly. await self._maybe_discover_facts( session, previous_flow_id, previous_skill_id ) ``` Lines 482 to 484 then return early unless the turn switched into a new, non-null flow: ```python new_flow_id = session.active_flow_id if new_flow_id is None or new_flow_id == previous_flow_id: return ``` So it runs on each turn that switches into a new flow, not on every turn. In our probes that happened once per conversation, on the second customer turn, and the log line names the switch: `"previous_flow_id": null, "new_flow_id": "check_stock"`. We stated the facts on that same turn. The request carries a window of earlier history too, so a fact stated a turn or two before a flow switch may also reach the extractor; we did not test that. The docstring of `_maybe_discover_facts` (lines 474 to 481) states the design: "Best-effort: any failure is logged and swallowed rather than surfaced, since a missed capture should never fail an otherwise-successful turn." Lines 534 to 538 catch anything the extractor did not, as a warning. A missed capture is meant to be quiet, and it is. :::diagram{title="The post-turn discovery path on Rasa Pro 3.20.0, and where each variant ended"} ```dot rankdir=LR; reply [label="agent reply\nalready sent"]; gate [label="no error, not cancelled,\nnew flow entered?\nprocessor.py 452-484", class="mark-1"]; skip [label="no discovery"]; request [label="extractor request\nsystem task + history\nlast role: assistant", class="blocked mark-2"]; opus [label="claude-opus-5-5\nHTTP 400\n(acceptance run)", class="blocked"]; haiku [label="claude-haiku-4-5\n200, no tool call", class="blocked"]; empty [label="empty list\nof facts"]; logged [label="discover_facts.completed\nfacts_count 0 (INFO)", class="mark-3"]; neutral [label="+ user turn\n\"Continue.\" or\n\"Please continue\""]; asking [label="+ user turn asking for the task,\nor the tool forced", class="ok"]; written [label="record_discovered_facts\nboth facts", class="ok mark-4"]; reply -> gate; gate -> skip [label="no"]; gate -> request [label="yes"]; request -> opus; request -> haiku [label="replay 20 of 20"]; opus -> empty [label="caught, WARNING\n(INFO line: from the source,\nnot observed)"]; haiku -> empty; empty -> logged; request -> neutral; neutral -> haiku [label="replay 20 of 20 each"]; request -> asking; asking -> written [label="replay 20 of 20 each"]; ``` 1. Discovery waits for the reply and runs only on a clean turn that enters a new flow. 2. The request's last message is the reply the customer has already seen, and the task is only in the system prompt. 3. A refused call, an empty answer and a text answer all end as the same INFO line. 4. On this request, a final user turn that asks for the task or a forced tool got both facts; the only other change that got a call was removing both earlier tool-call pairs. ::: :::callout{type="warn" title="facts_count 0 looks the same for three outcomes"} Lines 526 to 533 log `mantle.processor.discover_facts.completed` at INFO with `facts_count` set to the length of the list the extractor returned. The customer stated nothing: 0. The model answered without the tool call, as Haiku did in A1, A2, D1 and D2: 0. The call raised, as Opus did: by the source, the extractor logs its WARNING and returns an empty list, so the count is again 0. ::: ## One captured request, 20 calls per change: what got the tool call Every replay and bisect arm starts from A2's extractor request body as the proxy recorded it (the whole body is saved with this article's receipts and is not linked from this page). We re-sent it with our own headers to the Anthropic Messages API, with no Rasa, litellm or proxy in the path, at temperature 0.0, the value in the body. Twenty calls at that setting show that the answer was stable on this one request; they are not a rate across conversations. The replay recorded one row per call: the status, the stop reason, the block types, the fact keys and the first 100 characters of any text. Call 1 of the two arms that separate a trailing user turn from what it asks: ::::compare{title="The first replay of two variants of A2's request, as recorded"} :::pane{label="“Continue.” appended" tone="bad"} ```text {"arm": "continue", "i": 1, "status": 200, "stop_reason": "end_turn", "block_types": ["text"], "tool_called": false, "facts": 0, "fact_keys": [], "text_head": "I don't have any additional information to share at the moment. What would you like to do next? I ca", "error": null} ``` A user turn at the end that asks nothing: text, no tool call. All 20 calls ended this way. ::: :::pane{label="“Record the facts now.” appended" tone="good"} ```text {"arm": "ask", "i": 1, "status": 200, "stop_reason": "tool_use", "block_types": ["tool_use"], "tool_called": true, "facts": 2, "fact_keys": ["customer_name", "communication_preference"], "text_head": "", "error": null} ``` A user turn at the end that asks for the task: the tool call with both fact keys. All 20 calls ended this way. ::: :::: Three variants got both facts 20 of 20 on this request: forcing the tool, sending the task text as the final user message, and a final user turn that asks for the task. Only the first two do not depend on our wording, which the "Continue." and "Please continue" arms show matters. They worked at the API level, not in Rasa: the 3.20.0 extractor sends no final user turn and no `tool_choice`, and we did not patch it. Then we bisected the unchanged request to find what makes it end silently. These single changes left it at 0 of 20, empty every time: the minimal three-sentence task text in place of Rasa's; the task text without "or do not call it at all" (the tool's own description still said "Omit this call entirely if nothing new was stated."); a minimal tool schema; the opening greeting pair removed; the history's own tools (`activate`, `check_stock`) declared; and the tool blocks kept but the plain-text messages sent as strings. Removing one tool-call pair left it at 0 as well: `activate` 0 of 20, and `check_stock` 0 of 40 over two runs (that arm also puts the two assistant texts into one turn; a second arm we meant as a separate change built the identical request, so we count it as the second run). Five of the totals, copied from the receipt: ```text unchanged n=20 tool_called=0 empty=20 text_only=0 errors=0 fact keys (calls containing): {} drop_activate_pair n=20 tool_called=0 empty=20 text_only=0 errors=0 fact keys (calls containing): {} drop_check_stock_pair n=20 tool_called=0 empty=20 text_only=0 errors=0 fact keys (calls containing): {} drop_both_pairs_blocks_kept n=20 tool_called=20 empty=0 text_only=0 errors=0 fact keys (calls containing): {'communication_preference_text_only': 13, 'customer_name': 7, 'communication_preference': 7} drop_both_pairs_consecutive_assistant n=20 tool_called=20 empty=0 text_only=0 errors=0 fact keys (calls containing): {'communication_preference_text_only': 20} ``` Removing both tool-call pairs brought the call back in every form we tried: merged into plain strings (60 of 60 over three runs), merged into content blocks (40 of 40 over two runs), both assistant texts kept as separate blocks in one turn (20 of 20), and two assistant messages left in a row, sent unmerged (20 of 20). All 140 calls held the preference; the name appeared in 1, 0, 7 and 0 of them respectively. Two arms in that list make a clean pair. They end on the same final assistant turn, the text the agent sent before the stock lookup followed by its final reply, and differ only in whether the `activate` pair is still in the history: | Arm (receipt name) | `activate` pair | `check_stock` pair | Final assistant turn | Tool called | | ---------------------------------------------------- | --------------- | ------------------ | ---------------------- | ----------- | | `check_stock` pair removed (`drop_check_stock_pair`) | kept | removed | both texts in one turn | 0 of 40 | | both removed (`drop_both_pairs_blocks_kept`) | removed | removed | the same turn | 20 of 20 | | `activate` pair removed (`drop_activate_pair`) | removed | kept | as captured | 0 of 20 | Holding the final turn fixed, removing the `activate` pair as well took the result from 0 of 40 to 20 of 20; removing it alone got 0 of 20. We do not claim why, or that tool calls silence the model in general. Our minimal Rasa-free request, with a three-sentence task in the system prompt, one tool offered and not forced and a history about the same customer ending on the assistant, got the call 20 of 20 as plain strings, 20 of 20 as content blocks, 20 of 20 with a `check_stock` tool-call pair added, and 20 of 20 with "Please continue" appended. So a tool call in the history does not by itself stop it; the effect is specific to this request. On Rasa's request, "Please continue" got a text reply and no tool call in 20 of 20 (16 of its answers began "I'm ready to help! What would you like to do next?"), and "Continue." got no tool call in 20 of 20 there and in 20 of 20 on the minimal request. What you can take to your own capture is the procedure: replay it, then remove one part of the history at a time, and see which arm gets the call. The live runs in the table above agree: B1 and B2, whose extractor requests match A2's apart from tool call ids and the appended turn, kept both facts. On Opus, which we did not replay, the same appended turn got the call in C1 and C2. The replay script is 44 lines of Python that need an Anthropic key and `python-dotenv`, not Rasa; you run it against a request body you captured yourself. The bisect script is built the same way, with one arm per change. :::solution{title="Show the replay script as run (replay_a2.py)"} It reads a captured request body and posts it to `https://api.anthropic.com/v1/messages`. This is the whole script as run, with one change: line 9 read our `.env` file by its absolute path, and here reads `.env` in the working directory. Nothing else differs, and the `assert key` on line 10 is kept. The recorded command was `python replay_a2.py d320/runA2.jsonl replay-a2-haiku.jsonl 20`, where `d320` is our working directory for the probe captures. It imports `dotenv`, so install `python-dotenv` first. ```python """Replay the extractor request Rasa 3.20.0 sent in probe run A2 (captured body, request 4 of runA2.jsonl) straight to the Anthropic Messages API, N times per arm. The only differences between arms are the ones named below. The API key is read from the project .env inside this process and never printed or written.""" import json, sys, copy, urllib.request, urllib.error, time, pathlib from dotenv import dotenv_values CAP, OUT, N = pathlib.Path(sys.argv[1]), pathlib.Path(sys.argv[2]), int(sys.argv[3]) base = json.loads(CAP.read_text().splitlines()[4])['request'] key = dotenv_values('.env').get('ANTHROPIC_API_KEY') assert key, 'ANTHROPIC_API_KEY missing' system_text = base['system'] if isinstance(base['system'], str) else ' '.join(b.get('text', '') for b in base['system']) def arm(name): b = copy.deepcopy(base) if name == 'unchanged': pass elif name == 'continue': b['messages'].append({'role': 'user', 'content': 'Continue.'}) elif name == 'ask': b['messages'].append({'role': 'user', 'content': 'Record the facts now.'}) elif name == 'tool_choice': b['tool_choice'] = {'type': 'tool', 'name': 'record_discovered_facts'} elif name == 'task_as_user': b['messages'].append({'role': 'user', 'content': system_text}) return b rows = [] for name in ('unchanged', 'continue', 'ask', 'tool_choice', 'task_as_user'): body = arm(name) for i in range(N): req = urllib.request.Request('https://api.anthropic.com/v1/messages', data=json.dumps(body).encode(), headers={'x-api-key': key, 'anthropic-version': '2023-06-01', 'content-type': 'application/json'}, method='POST') try: with urllib.request.urlopen(req, timeout=120) as r: status, resp = r.status, json.loads(r.read()) except urllib.error.HTTPError as e: status, resp = e.code, json.loads(e.read() or b'{}') if e.code == 401: raise SystemExit('401: stop') blocks = resp.get('content') or [] tu = [b for b in blocks if b.get('type') == 'tool_use' and b.get('name') == 'record_discovered_facts'] facts = sum(len((b.get('input') or {}).get('facts') or []) for b in tu) rows.append({'arm': name, 'i': i + 1, 'status': status, 'stop_reason': resp.get('stop_reason'), 'block_types': [b.get('type') for b in blocks], 'tool_called': bool(tu), 'facts': facts, 'fact_keys': [f.get('key') for b in tu for f in ((b.get('input') or {}).get('facts') or [])], 'text_head': ' '.join(b.get('text', '') for b in blocks if b.get('type') == 'text')[:100], 'error': (resp.get('error') or {}).get('message')}) time.sleep(0.3) OUT.write_text('\n'.join(json.dumps(r) for r in rows) + '\n') from collections import Counter for name in ('unchanged', 'continue', 'ask', 'tool_choice', 'task_as_user'): rs = [r for r in rows if r['arm'] == name] print(f"{name:13s} n={len(rs)} tool_called={sum(r['tool_called'] for r in rs)} empty={sum(r['block_types']==[] for r in rs)} text_only={sum(r['block_types']==['text'] for r in rs)} errors={sum(r['status']!=200 for r in rs)} facts>=2={sum(r['facts']>=2 for r in rs)}") ``` Line 8 is specific to our capture: `splitlines()[4]` takes line 5 of our proxy log, the extractor request, and `['request']` takes the body the proxy recorded. Our A2 body is saved with this article's receipts but not linked from this page, so to use the script, capture a request your own side call sends and replace that line with code that loads it as a dict with `model`, `system`, `messages` and `tools`. Then change one thing per arm. ::: :::callout{type="info" title="What these runs are"} - The live runs: one acceptance run per model and eight probe conversations in which the customer states a name and a contact preference, all on Rasa Pro 3.20.0 through litellm to Anthropic, with a local proxy recording request bodies and responses, never headers. Two probes ran unmodified on Haiku; two on Haiku and two on Opus with "Record the facts now." appended by the proxy; two on Haiku with "Continue." appended. - The direct calls: 600 calls straight to the Anthropic API with no Rasa, litellm or proxy, all to claude-haiku-4-5 at temperature 0. On A2's request: 100 replays, 20 with "Please continue" and 340 bisect calls. On our minimal request: 140, including 20 with "Please continue". One earlier bisect attempt sent the unchanged request for arms it had not defined; those calls are void, not counted, and the arms were rerun. - Anthropic's page on handling stop reasons is quoted as served on 29 September 2026. - Runs recorded 29 September 2026 against the 3.20.0 release. ::: :::callout{type="warn" title="What they do not show"} - A rate: the direct calls show stability on one request each, 20 calls per arm. The live runs are a handful. - Any Rasa release other than 3.20.0, any model other than claude-haiku-4-5 and claude-opus-5-5, or any provider other than Anthropic. - A fix inside Rasa: forcing the tool and moving the task into a final user turn were tested at the API level only. - A paired Opus result: the Opus refusal comes from the acceptance conversation, and no Opus request was replayed. - Why removing both earlier tool-call pairs brings the call back on Rasa's request while removing either alone does not; the minimal request with one tool-call pair still got the call. - What discovered facts are worth to your code. In the Python modules of the 3.20.0 wheel, the only reader outside the memory manager is the helper that lists already-captured keys for the prompt (`catalog.py`, line 89); the manager's line 109 only registers the entry's name. We did not search the Inspector's web front end or test code of your own. ::: ## Assert the memory write, not the reply Discovery runs after the reply is sent, so the proxy in the C runs could not change what the customer read. In C2, a run where the proxy appended "Record the facts now." to the extractor request, the REST channel returned these three bot messages, unedited: ```text bot: Hi Priya! We have **7 monstera** in stock right now. bot: About updates: I can only look up stock. I can't save contact preferences or send updates by text or email, so your preference hasn't been recorded anywhere. I also can't reserve plants, take payment or place orders. bot: Would you like me to check another plant? ``` In the same turn, the tracker gained `contact_channel_preference`, "Text message only; never email". In A2 the Haiku agent confirmed text-only updates and nothing was recorded. A reply cannot tell you what memory holds, in either direction. Only the tracker can. The same trap appears in [a handoff check that read five out of five while the package omitted the disputed charge](/library/guides/measure-a-handoff-that-loses-context/): a green check on the wrong thing. So the check reads the tracker. `drive_discovery.py` starts the server, sends the two messages over the REST channel, waits 12 seconds for the post-turn call, and saves the tracker's memory events. This script then fails unless the latest `system.__discovered__` write holds the value the customer stated: ```python """Fail unless the tracker's latest discovered facts hold the value the customer stated.""" import json, sys events = json.load(open(sys.argv[1]))["memory_events"] expected = sys.argv[2].lower() writes = [e for e in events if e.get("event") == "memory_set" and e.get("key") == "system.__discovered__"] facts = writes[-1]["value"] if writes else [] # each write holds the full list hits = [f["key"] for f in facts if expected in str(f.get("value", "")).lower()] print(f"discovered facts: {len(facts)}; holding {sys.argv[2]!r}: {', '.join(hits) or 'none'}") sys.exit(0 if hits else 1) ``` We ran it over the saved output of the six probe runs A1 to C2, without a model call: ```text $ python assert_discovered.py A1.json 'Priya Nair' discovered facts: 0; holding 'Priya Nair': none exit 1 $ python assert_discovered.py A2.json 'Priya Nair' discovered facts: 0; holding 'Priya Nair': none exit 1 $ python assert_discovered.py B1.json 'Priya Nair' discovered facts: 2; holding 'Priya Nair': customer_name exit 0 $ python assert_discovered.py B2.json 'Priya Nair' discovered facts: 2; holding 'Priya Nair': customer_name exit 0 $ python assert_discovered.py C1.json 'Priya Nair' discovered facts: 3; holding 'Priya Nair': customer_name exit 0 $ python assert_discovered.py C2.json 'Priya Nair' discovered facts: 3; holding 'Priya Nair': customer_name exit 0 ``` In a later pass, the same check over the control runs D1 and D2: ```text $ python assert_discovered.py D1.json 'Priya Nair' discovered facts: 0; holding 'Priya Nair': none exit 1 $ python assert_discovered.py D2.json 'Priya Nair' discovered facts: 0; holding 'Priya Nair': none exit 1 ``` It fails exactly where the reply was right and memory stayed empty. The price: every run is a live model call, it waits on a timer, it tests one phrasing of one fact, and a single pass is one sample. Run it more than once before you trust a green result, and again whenever the model or the Rasa version changes. :::checkpoint{id="mantle-discovery-green-log" question="On Rasa Pro 3.20.0 with claude-haiku-4-5 acceptance passes. In a probe the customer states her name, the reply is correct, and the server log has no ERROR line and no WARNING about discovery. What does that tell you about fact discovery?" options="It ran and found nothing worth recording,Nothing: it may have run and kept no fact that was stated,It did not run because no flow started" answer="1"} In A1 and A2 discovery ran, the customer stated two facts, and none was kept. The log's only discovery line was at INFO with `facts_count` 0, the same line a turn with nothing to record produces. The acceptance script's assertions never read memory, so it cannot see this. Only a check on the memory write can. ::: :::solution{title="Show the key lines of the driver (drive_discovery.py)"} These runs used a copy of the quickstart identical to `RasaHQ/rasa-community` at commit `8a0320c`, with the orchestrator group switched as in the diff above, the model the acceptance check had trained into `examples/quickstart/models`, and `RASA_LICENSE` and `ANTHROPIC_API_KEY` in `.env` at the root of the copy. Every probe ran behind a local proxy (the capture proxy for A, the append proxies for B, C and D), with `ANTHROPIC_BASE_URL` and `ANTHROPIC_API_BASE` pointing at it. To run only the memory-write check, without a proxy, start from that root with `uv run --project examples/quickstart python drive_discovery.py run.log '["Hello", "Hi, my name is Priya Nair and I only want updates by text message, never email. How many monstera do you have?"]'`. It writes the scrubbed server log to `run.log` and the replies and memory events to `run.log.json`. The driver removes `LITELLM_MODIFY_PARAMS` from the server's environment. The quickstart's `.env.example` at that commit sets only `RASA_LICENSE` and `OPENAI_API_KEY`, so dropping the variable keeps a value from our shell out of the run and leaves litellm configured as the project ships. These are the lines that set up the environment, send the conversation and read the tracker; the full script, 43 lines with server start-up and shutdown, is saved with this article's receipts as `drive_discovery.py.txt`. ```python env = {**defaults, **{k: v for k, v in dotenv_values('.env').items() if v}, **os.environ} env.pop('LITELLM_MODIFY_PARAMS', None) for key in ('RASA_LICENSE', 'ANTHROPIC_API_KEY'): if not env.get(key): raise SystemExit(f'{key} missing') env.update({'LOG_LEVEL': 'INFO', 'RASA_LOG_LEVEL': 'INFO', 'RASA_TELEMETRY_ENABLED': 'false'}) ``` ```python sender = 'discovery-probe-' + uuid.uuid4().hex; replies = [] for m in messages: replies.append({'user': m, 'bot': [r.get('text') for r in req('/webhooks/rest/webhook', {'sender': sender, 'message': m})]}) time.sleep(12) # post-turn discovery runs after the reply; give it time tracker = req(f'/conversations/{sender}/tracker?include_events=ALL') res = {'replies': replies, 'memory_events': [e for e in tracker['events'] if 'memory' in e.get('event', '')]} ``` ::: ## What to do on 3.20.0 On Rasa Pro 3.20.0 with Claude, treat `system.__discovered__` as empty unless your own check shows otherwise, and if you move an agent between providers, run the memory-write check before and after the move. What carries beyond Rasa is the method. A background call's own log line (`discover_facts.completed`, `facts_count 0`) and a green acceptance run said nothing about whether the write happened. If you own such a call, capture the request it really sends, replay it with one change per arm, including the parts of its history, and assert that the thing it was meant to write exists. :::cta{href="/library/guides/mutation-test-an-agent-guard/" label="See a suite report 186 of 186 mutations killed while the guard still broke"} The same method from the test side: check what a green result actually measured before you trust it. ::: --- # endpoints.yml said Ollama. Mantle still picked OpenAI Source: https://rasa.community/library/guides/mantle-references-default-to-openai-embeddings/ Author: Rod Rivera Published: 2026-09-25 We pointed the embeddings model group in the companion voice agent's `endpoints.yml` at a local model: ```diff - id: openai-embeddings models: - - provider: openai - model: text-embedding-3-large - api_key_env: OPENAI_API_KEY + - provider: ollama + model: all-minilm + api_base: http://localhost:11434 ``` Then we ran `rasa train` with OpenAI's API replaced by a local stand-in that records each request. This is everything the stand-in received: ```text exit code: 0 requests the OpenAI stand-in received: 1 {"method": "POST", "path": "/v1/embeddings", "model": "text-embedding-3-large", "inputs": 3, "authorization_header": "present"} input head: "# Fixture data \u2014 entirely fictional\n\nEvery person, company, account, and event i" input head: "# Horizon Travel FAQ\n\n## Baggage\n- Economy includes one cabin bag up to 7 kg and" input head: "## Contact\n- Horizon Travel phone support is available daily from 7am to 10pm UK" ``` Training passed. One request went to the OpenAI embeddings endpoint for `text-embedding-3-large`, carrying three chunks: two from the project's travel FAQ and one from the `README.md` that sits beside it in the references folder. Horizon Travel and its FAQ are the companion project's fictional fixture data, as that README says. Nothing in the train output named an embedding provider or model. What decided the model was a key the project never sets, in a different file. That output is unedited, from a real `rasa train` rather than an offline test. The project is `examples/mantle-voice-agent` in `RasaHQ/rasa-community-resources` at commit `01d6a6e`, on Rasa Pro 3.20.0. We recorded the runs on 29 September 2026 against that release. `OPENAI_API_KEY` held a placeholder, and `OPENAI_API_BASE` pointed at the stand-in, which answers with fake vectors and forwards nothing. The request is the one Rasa's OpenAI client sends. Re-running it needs a Rasa Pro licence in `RASA_LICENSE`; the companion's `.env.example` asks for the free Developer Edition licence. We did not add the group we edited. This is how the committed `endpoints.yml` declares it (lines 21 to 36, with the `openai-llm` group cut): ```yaml # --------------------------------------------------------------------------- # Model groups (LLMs / embeddings referenced by endpoints features) # --------------------------------------------------------------------------- # https://rasa.com/docs/reference/config/components/llm-configuration/ model_groups: ... - id: openai-embeddings models: - provider: openai model: text-embedding-3-large api_key_env: OPENAI_API_KEY ``` No setting anywhere in the project refers to `openai-embeddings`; the only other match for the name is a copy of this file under `tutorial/snippets`. The group looked in charge because it lists `text-embedding-3-large`, which is also the model Rasa falls back to when `agent.yml` names none. Train the project as committed and the request is identical, so the group appears to work. Only changing it shows that it does nothing. The model that embeds a Mantle project's reference files is chosen by one key, `references.embeddings` in `agent.yml`, and resolved only against `integrations.yml`. Left empty, it is OpenAI: if your reference files are private, every chunk of them goes to OpenAI on every training run, whatever `endpoints.yml` says. **Our position:** a Mantle project that ships reference files should set that key, even when the answer is an OpenAI group, and prove the choice with one training run against a recorder. The cost is a model group in `integrations.yml` to maintain, a retrain every time the embedder changes, and a stand-in wherever the proof runs. ## Five trainings of one project, and which edits moved the request We trained five separate copies of the committed project, each with the same stand-in and placeholder key. Run a changes nothing; runs b and e edit one file; runs c and d edit two, `agent.yml` and `integrations.yml`. | Run | Edit to the committed project | Requests to the stand-in | Model requested | `rasa train` | | --- | -------------------------------------------------------------------------------------------------- | ------------------------ | ------------------------ | ---------------------------- | | a | none | 1, with 3 inputs | `text-embedding-3-large` | exit 0 | | b | `endpoints.yml` group `openai-embeddings` set to `ollama` / `all-minilm` | 1, with 3 inputs | `text-embedding-3-large` | exit 0 | | c | `embedding: local-embeddings` (no s) under `references:`; group added to `integrations.yml` | 1, with 3 inputs | `text-embedding-3-large` | exit 0 | | d | `embeddings: local-embeddings` under `references:`; the same group added to `integrations.yml` | 0 | none | exit 0 | | e | `embeddings: no-such-group` under `references:`; `integrations.yml` unchanged | 0 | none | exit 1, at project validation | Runs a, b and c sent the same request, down to the first 80 characters of each input. Run b is the one the title describes. Run c is a typo that nothing reports. `ReferencesConfig` is declared with `extra="ignore"` (`rasa/mantle/config/agent_spec.py`, line 153), so an unknown key under `references:` is dropped and `embeddings` keeps its empty default. The companion's own `agent.yml` warns at lines 17 to 19 that `AgentSpec` ignores unknown keys; the class one level down is declared the same way. Run e is the only loud one. Training stopped at project validation with one finding, `mantle.validation.config.unknown_model_group`, whose message reads: "'references.embeddings' in 'agent.yml' references model group 'no-such-group', which is not declared in 'integrations.yml'." :::callout{type="warn" title="What five runs on one project do not show"} - The stand-in answered in OpenAI's place. Nothing here records traffic to OpenAI itself, and its fake vectors say nothing about retrieval quality. - One project on one release: these runs cover `examples/mantle-voice-agent` on Rasa Pro 3.20.0 and nothing else. - Train time only. By the source, serving calls the same function (`references_index.py`, line 242), so an empty key would embed queries with the same default. We did not serve the model. - Run d shows that no request reached the stand-in and that training passed. It does not record what the Ollama side received. ::: ## Where build_references_embedder gets the embedding model The index is built when `rasa train` packages the model archive. `build_references_index` loads every `*.md` file under a `references/` folder (line 99 globs `**/*.md`, which is how the README was embedded), splits them into chunks, and then chooses the embedder on lines 201 and 202: ```python model_groups = resolve_integration_model_groups(project_root) embedder = build_references_embedder(agent, model_groups) ``` `resolve_integration_model_groups`, in `rasa/mantle/model_archive/bootstrap.py`, returns the `model_groups` list from the project's `integrations.yml`, or an empty list. `endpoints.yml` is not an argument. This is the function those groups feed, whole, from the 3.20.0 wheel (sha256 `a68fa7efff47dd035e3ff2b7b1cfdc467c394ff75981426a4377d4008207e9b5`): :::annotated{title="rasa/mantle/model_archive/references_index.py, lines 52-81 (Rasa Pro 3.20.0)"} ```python def build_references_embedder( agent: AgentSpec, model_groups: list[dict[str, Any]] ) -> "EmbeddingClient": """Resolve the references embeddings model group into an embedding client. ``agent.references.embeddings`` names a ``model_groups`` id; empty → the OpenAI default (mirrors ``EnterpriseSearchPolicy``). The resolved config is handed to ``embedder_factory``, whose client ``FAISS_Store`` consumes directly. """ group_id = agent.references.embeddings # (1) embeddings_config: Optional[dict[str, Any]] = None if group_id: # (2) # mantle resolves model groups only from the packaged integrations.yml. # (3) # resolve_model_client_config treats an empty model_groups list as falsy # and silently falls back to endpoints.yml, which would resolve against # the wrong (or missing) config. Fail closed with a clear message instead. if not model_groups: # (4) raise InvalidConfigException( f"references.embeddings names model group '{group_id}', but no " f"model_groups are defined in {V2_INTEGRATIONS_FILE}. Define the " f"model group there, or clear references.embeddings to use the " f"default embedder." ) embeddings_config = resolve_model_client_config( {MODEL_GROUP_CONFIG_KEY: group_id}, "references", model_groups=model_groups, ) return embedder_factory(embeddings_config, DEFAULT_EMBEDDINGS_CONFIG) # (5) ``` 1. Line 62: the only input that names a model is `references.embeddings` from `agent.yml`. Its default is the empty string. 2. Line 64: an empty string is falsy, so lines 65 to 80 never run, and nothing is looked up, checked or logged. 3. Lines 65 to 68: the authors' comment says model groups come only from `integrations.yml`, and that an empty list would let resolution drift to `endpoints.yml`. The next line exists to stop that drift. 4. Lines 69 and 70: the function's own `raise`, for a set id when `integrations.yml` defines no groups at all. 5. Line 81: for an empty id this is `embedder_factory(None, DEFAULT_EMBEDDINGS_CONFIG)`. ::: The default is two keys. `DEFAULT_EMBEDDINGS_CONFIG` in `rasa/core/policies/enterprise_search_policy_config.py` names the provider and the model, and `embedder_factory` in `rasa/shared/utils/llm.py` passes a `None` config straight through to the client factory with it: ```python DEFAULT_EMBEDDINGS_CONFIG = { PROVIDER_CONFIG_KEY: OPENAI_PROVIDER, MODEL_CONFIG_KEY: DEFAULT_OPENAI_EMBEDDING_MODEL_NAME, } ``` `llm.py` sets `DEFAULT_OPENAI_EMBEDDING_MODEL_NAME = "text-embedding-3-large"` on line 135. That is the model in every request runs a to c recorded, and the model the committed `endpoints.yml` group lists. The coincidence is what hid the dead group. The fallback is documented. The field's docstring in `agent_spec.py` says "Empty → the built-in default embeddings (OpenAI), like `EnterpriseSearchPolicy`", and the docstring at the top of `references_index.py` (lines 7 to 9) says it again. What is missing is any sign of it at run time. The build path's only info-level log call, `mantle.packaging.references_index.wrote`, records the index path, document and chunk counts and chunking settings, and no provider. In the full, unfiltered log of run a it is the only line about the index, and no line names an embedding provider or model. Rasa's training telemetry code has the same blind spot: `_embeddings_attributes` in `training_telemetry.py` reports `embeddings_provider` as `None` whenever the key is empty. ## Every check on references.embeddings waits for the key to be set A bad `references.embeddings` can stop training in three places, and all three start from the same test. Project validation runs first: `validate_agent` in `rasa/mantle/config/validation.py` stops checking the key when it is empty (lines 93 and 94) and otherwise looks the id up in `integrations.yml` alone. That lookup produced run e's finding. The second is the `raise` on line 70 of `build_references_embedder`, inside `if group_id:`. The third is `resolve_model_client_config` in `llm.py`, called from the same branch. It raises its own `InvalidConfigException` when the id is missing from a non-empty group list, and a default `rasa train` only reaches it with validation off, because `rasa/cli/train.py` passes `validate=not args.skip_validation`. Its message says the group was not found "in endpoints.yml", although Mantle handed it the groups from `integrations.yml`. No check of any kind covers runs a, b and c. :::diagram{title="Three checks, one condition: how each run left rasa train on Rasa Pro 3.20.0"} ```dot rankdir=LR; train [label="rasa train"]; endpoints [label="endpoints.yml\nopenai-embeddings group", class="mark-4"]; other [label="tracing settings"]; validate [label="validate_agent\nvalidation.py 93-108", class="mark-1"]; failed [label="unknown_model_group\nexit 1 (run e)", class="blocked"]; integrations [label="integrations.yml\nmodel_groups"]; build [label="build_references_embedder\nreferences_index.py 62-81", class="mark-2"]; fallback [label="OpenAI default\ntext-embedding-3-large\n(runs a, b, c)", class="blocked mark-3"]; resolver [label="InvalidConfigException\n(only with --skip-validation)"]; named [label="named group\nlocal-embeddings (run d)", class="ok"]; train -> endpoints [label="reads"]; endpoints -> other; train -> validate; integrations -> validate [label="only if the key is set"]; validate -> failed [label="set, not declared"]; validate -> build [label="empty, or declared"]; integrations -> build [label="line 201"]; build -> fallback [label="empty"]; build -> named [label="set, found"]; build -> resolver [label="set, not found"]; ``` 1. Validation checks the key only when it is set, and only against `integrations.yml`. 2. The build step's `raise` and the resolver's sit behind the same `if group_id:`; with validation on, the resolver's never fires for a missing id. 3. An empty or misspelt key reaches `embedder_factory` with only the OpenAI default, and no log line names the provider. 4. `rasa train` does read `endpoints.yml`: run b's output includes a tracing notice naming that file. None of it reaches the references path. ::: The code fails closed on a group it cannot find and falls open, to a hosted provider, on a key it cannot see. The comment on lines 65 to 68 refuses to fall back silently to the wrong configuration. An empty key falls back, documented but unannounced, to a provider the project may never have chosen. The docstring's comparison with `EnterpriseSearchPolicy` says this default is deliberate, and a maintainer could reasonably keep it. Two small changes would have made run b visible: one info line at line 81 naming the provider and model it chose, or a validation finding when a project ships reference files with no `embeddings` key. Both are our suggestions, not behaviour we observed. :::checkpoint{id="mantle-references-group-in-endpoints" question="You add embeddings: local-embeddings under references: in agent.yml, and declare the local-embeddings group in endpoints.yml because the project's other model groups live there. What does rasa train do on Rasa Pro 3.20.0?" options="Embeds the references with all-minilm,Embeds them with the OpenAI default,Stops at validation with unknown_model_group" answer="2"} Validation looks the id up in `integrations.yml` only (lines 96 to 108 of `validation.py`), so a group declared in `endpoints.yml` is as unknown as `no-such-group` was in run e. We ran the undeclared case; this variant is read from the same lines, not run. It fails loudly, which is the better way to be wrong. ::: ## Name the embedder in agent.yml and declare it in integrations.yml Run d is the change, and it needs two files: line 62 takes the id from `agent.yml`, and line 201 decides that the id is looked up in `integrations.yml`. The diff against the committed project, `agent.yml` first, then `integrations.yml`: ```diff references: + embeddings: local-embeddings instructions: | Answer from retrieved FAQ snippets only. Keep answers short for voice. ``` ```diff temperature: 0.0 + - id: local-embeddings + models: + - provider: ollama + model: all-minilm + api_base: http://localhost:11434 + channels: ``` Training passed and the stand-in received no request, with `endpoints.yml` untouched. If you want OpenAI, name an OpenAI group the same way: the key then records the choice, and validation checks it. To audit a project before training it, read the key and look for anything else under `references:`. This is the command we ran in the committed project and in the copies for runs c and d. It is wrapped here with backslash continuations, which the shell removes, so it is the same one-line program: ```sh uv run --frozen python -c "import yaml; \ r = yaml.safe_load(open('agent.yml')).get('references') or {}; \ print('references.embeddings:', r.get('embeddings') or '(empty: OpenAI default)'); \ print('other keys under references:', \ sorted(set(r) - {'instructions', 'embeddings', 'chunking'}) or 'none')" ``` Its output, with uv's install notices (a hardlink warning and the `Installed`/`Built` lines) removed and nothing else changed: ```text $ cd examples/mantle-voice-agent references.embeddings: (empty: OpenAI default) other keys under references: none $ cd runs/c-misspelt-key references.embeddings: (empty: OpenAI default) other keys under references: ['embedding'] $ cd runs/d-named-local references.embeddings: local-embeddings other keys under references: none ``` The label "(empty: OpenAI default)" is ours, taken from the source above. The command reads YAML; only a training run against a recorder shows what is sent. ## The same fallback in stacks that are not Rasa None of this needs Rasa to happen again. Four things lined up: 1. A factory takes an optional user config and a hosted default, and `None` means the default: `embedder_factory(None, DEFAULT_EMBEDDINGS_CONFIG)`. 2. Validation runs only on values that are present, so absence is never an error. 3. The schema ignores unknown keys, so a typo becomes absence. 4. A config block elsewhere lists the same model as the default, so the fallback and the intended setting send the same request, and a green build proves nothing. In code you do not own, search for calls where the config argument can be `None` next to a module-level constant naming a hosted model, and for `extra="ignore"` or its equivalent on the schema that holds the setting. The log line that should exist is one at build time naming the provider and model actually chosen. When it is missing, run the test we ran: point the provider's base URL at a recorder, change the setting to a value you can tell apart from the default, and train. If the recorder still hears the default model, the setting is dead. The recorder is a plain HTTP server. It logs the path, the model, the input count, the first 80 characters of each input and whether an `Authorization` header was present (never its value), and it answers embeddings requests with fake vectors. :::solution{title="Show the stand-in we trained against (29 lines of Python)"} Start it with `python openai_standin.py requests.jsonl`, then train with `OPENAI_API_KEY=sk-standin-placeholder OPENAI_API_BASE=http://127.0.0.1:/v1 uv run --frozen rasa train`. ```python """Local stand-in for api.openai.com. Records each request (path, model, input count, first 80 characters of each input, auth header present or not; never its value) and returns deterministic fake embeddings. Nothing is forwarded anywhere.""" import http.server, json, sys, datetime, hashlib PORT, LOG = int(sys.argv[1]), sys.argv[2] class H(http.server.BaseHTTPRequestHandler): def do_POST(self): body = json.loads(self.rfile.read(int(self.headers.get('content-length', 0))) or b'{}') inputs = body.get('input', []) inputs = [inputs] if isinstance(inputs, str) else inputs rec = {'at': datetime.datetime.now().isoformat(timespec='seconds'), 'method': 'POST', 'path': self.path, 'authorization_header': 'present' if self.headers.get('authorization') else 'absent', 'model': body.get('model'), 'inputs': len(inputs), 'input_heads': [str(i)[:80] for i in inputs]} with open(LOG, 'a') as f: f.write(json.dumps(rec) + '\n') if self.path.rstrip('/').endswith('/embeddings'): def vec(t): h = hashlib.sha256(str(t).encode()).digest() return [((h[i % 32] / 255.0) - 0.5) for i in range(1536)] out = {'object': 'list', 'model': body.get('model'), 'data': [{'object': 'embedding', 'index': i, 'embedding': vec(t)} for i, t in enumerate(inputs)], 'usage': {'prompt_tokens': 0, 'total_tokens': 0}} data, status = json.dumps(out).encode(), 200 else: data, status = json.dumps({'error': {'message': 'stand-in serves /embeddings only'}}).encode(), 404 self.send_response(status); self.send_header('content-type', 'application/json') self.send_header('content-length', str(len(data))); self.end_headers(); self.wfile.write(data) def do_GET(self): self.do_POST() def log_message(self, *a): pass http.server.ThreadingHTTPServer(('127.0.0.1', PORT), H).serve_forever() ``` ::: We chose a recorder over an invalid key for this test. A call that fails tells you something was attempted; a recorder lets training finish, as it would in production, and keeps the request itself: the model, the input count and the start of every chunk. :::cta{href="/library/tutorials/mantle-starter-pack/" label="Break a Mantle project on purpose"} Chapter 3 of the [starter pack tutorial](/library/tutorials/mantle-starter-pack/) breaks a Mantle project seven ways. By the chapter's closing grouping, three are silent at runtime, three fail later with messages that do not name the cause, and one is irreversible. Its first break nests `rules:` under `agent:`; by the chapter's account it parses fine, would train fine and is silently discarded, the same kind of dropped key as run c. ::: --- # Measure a handoff before you claim it worked Source: https://rasa.community/library/guides/measure-a-handoff-that-loses-context/ Author: Rod Rivera Published: 2026-09-23 Aurora Home Insurance, a fictional insurer, has shipped a redesigned agent-to-human handoff. Its voice agent now passes a structured context package to the human desk instead of a one-line reason. The launch-review slide says: > **Redesigned handoff packages cut average handle time by 30%.** The only evidence behind it is a passing offline test suite. Here is what that evidence supports, and what it does not: | What you can claim | What you cannot claim yet | What to measure instead | | ---------------------------------------------------------------------------------------------------- | -------------------------------------------------------------- | ------------------------------------------------------------------------------------ | | On the tested path the package answers all five desk opening questions; the old handoff answers none | Callers stopped repeating themselves | Scripted questions the desk still asked, from call review | | Every handoff is scored against the same five questions | The package carries everything the case needs | Desk questions that are not on the script, per handoff | | All 41 offline tests pass | Handle time fell by 30%, or callers are happier with transfers | Timed transferred calls before and after; a satisfaction measure scoped to transfers | The artefact linked from the slide is the tail of a `make test` run in the companion pattern: ```text ---------------------------------------------------------------------- Ran 41 tests in 0.004s OK ``` No call was recorded and no desk agent was timed. A search of all 25 files in the pattern finds no "30%" and no mention of handle time. The team had a real result and reported a different one. Aurora is invented; the code is real. It is `patterns/voice-handoff-context` in `RasaHQ/rasa-community-resources` at `69e27b6`, the revision pinned for Rasa Pro 3.20.0rc1. Its fixture call is a disputed card payment on a bank account, which the fictional team kept for its pilot. **The position this guide takes:** at launch, report the count of fixed desk questions the package retires, proven offline, and nothing larger. The cost is a smaller slide, a measurement you still owe, and a script that sees only the questions written into it. ## What does the passing test prove? This section gives you the denominator the test counts against, and the before-and-after it proves. "Callers repeat themselves less" has no denominator. The pattern fixes one before any caller exists: five questions a desk agent asks when nothing was transferred, each mapped to the package field that makes it unnecessary (`handoffpkg/desk.py`, lines 53–59): ```python _QUESTION_RETIRED_BY: dict[str, str] = { "Can I take your name?": "identity.display_name", "Can you confirm your date of birth?": "identity.verified_tier", "Which account is this about?": "intent.details.account_id", "What are you calling about today?": "intent.goal", "Have you tried anything already?": "attempts", } ``` The denominator is the same five for every handoff, whatever the caller or case, so one handoff compares with the next and the old design with the new. The counting function, `unanswered_questions`, calls itself "The measurable form of the teaching claim" (lines 178–185). The baseline test builds a package from one free-text `handoff_reason` and gets all five questions back; the agent-path test gets none back. One offline run on 21 September 2026 passed both, along with the permissions test this guide uses later. :::solution{title="Show the command and output for the three tests"} From `patterns/voice-handoff-context` at `69e27b6`, on Python 3.14.3 with no network and no model: ```bash python3 -m unittest -v \ tests.test_handoff_context.TestPackageSurvivesHandoff.test_the_catalogs_current_handoff_answers_nothing \ tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question \ tests.test_handoff_context.TestPackageSurvivesHandoff.test_desk_permissions_follow_the_tier_not_the_package ``` ```text test_the_catalogs_current_handoff_answers_nothing (tests.test_handoff_context.TestPackageSurvivesHandoff.test_the_catalogs_current_handoff_answers_nothing) Baseline: one free-text reason retires none of the desk's questions. ... ok test_the_agent_path_retires_every_desk_question (tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question) THE HEADLINE CLAIM, ON THE PATH THE AGENT ACTUALLY RUNS. ... ok test_desk_permissions_follow_the_tier_not_the_package (tests.test_handoff_context.TestPackageSurvivesHandoff.test_desk_permissions_follow_the_tier_not_the_package) A medium-tier caller must not be actioned for an irreversible change. ... ok ---------------------------------------------------------------------- Ran 3 tests in 0.001s OK ``` ::: The five questions are the pattern author's, written for a fixture desk. Take yours from your desk's call guide or QA scorecard, version the list, and date every figure against a version. If the script changes, the number changes, though no caller behaved differently. :::solution{title="Our desk's script has eight questions. Do we just extend the list?"} Extend both structures together: `unanswered_questions` looks each question up with `_QUESTION_RETIRED_BY[question]` (line 188), so a question without a mapping raises a `KeyError` on every transfer. A question mapped to a field `_package_answers` does not recognise falls through to `return False` (line 229) and counts as unanswered on every handoff. Add a test for each new question, and restate the baseline before comparing figures. ::: ## Which count belongs on the slide? The handoff tool returns two counts in the same response, and only one of them has a fixed denominator. | | `questions_retired` | `desk_still_needs_to_ask` | | ------------------- | ------------------------------------------------- | ----------------------------------------------- | | What it counts | Lines the skills appended to `questions_answered` | Script questions whose package field is empty | | Denominator | None fixed | The five-question opening script | | On the dispute path | `4` | `[]`, so nothing left to ask | | Source | `skills/human_handoff/tools.py`, line 205 | The same file, line 208 | | Use it for | Nothing on a slide | Package coverage, tracked as a regression check | They disagree because the dispute skill appends four script questions (`skills/dispute_transaction/tools.py`, lines 98–103 and 135) and never "Have you tried anything already?". The package still carries three attempts, so the scripted check counts that question as answered. Pick one definition and write it into the review. An offline run of the real dispute and handoff tools printed both counts side by side. :::solution{title="Show the output from the real tools on the dispute path"} `editorial/receipts/measure-a-handoff-that-loses-context/agent-path-repro.py` in this site's repository calls `load_caller_context`, `list_recent_charges`, `raise_dispute` for charge `txn_c101`, then `transfer_to_human`, with a memory-only stand-in context. It ran on an unmodified copy of the pattern with Rasa Pro 3.20.0rc1 installed, because the tools import `rasa.mantle`. The run on 21 September 2026, on Python 3.12.13, printed, in part: ```text dispute_* keys in memory: ['disputed_txn_id', 'disputed_txn_label'] transferred: {"attempts_carried": 3, "goal": "dispute_transaction", "identity": true, "questions_retired": 4, "verified_tier": "medium"} desk_still_needs_to_ask: [] hint: Context package delivered. Tell the caller the handoff id and that they will not need to repeat themselves. Do not read back any withheld field. intent.details: {"account_id": "acc_checking", "account_label": "Everyday Checking", "card_last_four": "4821"} ASKING FOR line: ASKING FOR Dispute a card transaction [account_id=acc_checking, account_label=Everyday Checking, card_last_four=4821] (stage: blocked) ``` ::: ## Where does the handoff lose the disputed charge? This section shows the one fact the companion agent drops, and why the count cannot notice. :::diagram{title="On the dispute path, the charge the caller picked never reaches the package"} ```dot rankdir=TB; pick [label="Caller picks\nthe charge"]; written [label="Dispute skill writes\ndisputed_txn_label"]; collected [label="Handoff tool collects\ndispute_amount\ndispute_merchant\ndispute_date", class="blocked"]; details [label="intent.details:\naccount_id, account_label,\ncard_last_four"]; count [label="desk_still_needs_to_ask: []\nfive of five", class="ok"]; screen [label="Desk screen:\nwhich charge? not stated", class="blocked"]; pick -> written; written -> collected [label="no key matches"]; collected -> details [label="nothing to carry"]; details -> count [label="script check"]; details -> screen [label="what the desk sees"]; ``` ::: The handoff tool collects the disputed amount, merchant and date, but nothing on the dispute path writes them. The regression test checks three other keys, so the package names a card dispute on Everyday Checking but not which charge, and the count says the desk needs to ask nothing. :::solution{title="Show the source for this finding"} - `skills/human_handoff/tools.py`, lines 85–87: the handoff tool collects `dispute_amount`, `dispute_merchant` and `dispute_date`. - `handoffpkg/redaction.py`, lines 92–94 and 222–224: the allowlist passes those three keys into `intent.details`. - `skills/dispute_transaction/tools.py`, line 170: `raise_dispute` writes `disputed_txn_label`. No dispute tool writes the three keys above. - `skills/human_handoff/tools.py`, lines 68–96: the handoff tool's key list names neither `disputed_txn_label` nor `disputed_txn_id`. - `tests/test_handoff_context.py`, lines 360–363: the regression test checks only `account_id`, `account_label` and `card_last_four`. ::: Our inference, which no test checks, is that the desk will have to ask which charge it is. [Stop customers repeating themselves after a handoff](/library/guides/design-the-handoff-summary/) reaches the same finding from the design side. Three more limits decide what the count is worth: | Limit | What the code does | What it means for the review | | -------------------------- | ---------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- | | Presence, not correctness | The check asks whether a field is populated. In the test the caller's name is a placeholder string, and it counts. | A plumbing check, never a team target: the cheapest way to five is to fill every field. | | Partly on the agent's path | The test was added after an earlier suite missed the live agent's empty details, and it still sets four values itself. | Only the name and account questions retire from keys the agent writes; `attempts_log` parsing is untested. | | Emptiness is absence | An empty attempts list does not retire "Have you tried anything already?" | The count errs towards reporting a repeat, the safe direction. | :::solution{title="Show the source for the three limits"} - `handoffpkg/desk.py`, lines 208–229: `_package_answers` checks that a field is populated. The agent-path test's display name is `value_for_display_name`. - `skills/human_handoff/tools.py`, lines 77–81: the comment recording that an earlier suite passed while the live agent shipped empty details. - `handoffpkg/desk.py`, line 206: emptiness counts as absence, so an empty `attempts` list retires nothing. `tests/test_handoff_context.py`, lines 366–372, indentation removed, shows the four values the test sets itself. The `session` already holds the keys the handoff tool collects and the dispute skill writes. ```python session["verified_tier"] = "medium" session["goal"] = "dispute_transaction" session["questions_answered"] = "\n".join(DESK_OPENING_SCRIPT) session["attempts"] = [ {"action": "send_otp_sms", "outcome": "failed", "code": "delivery_failed"} ] ``` ::: :::checkpoint{id="handoff-measure-what-zero-proves" question="After the context package ships, unanswered_questions returns an empty tuple on the agent path in the offline suite. Which claim can the launch review make?" options="Average handle time on transferred calls fell,Callers no longer repeat themselves to the desk,On the tested path the package carries a value for each of the five opening-script questions,Customer satisfaction with transfers will rise" answer="2"} The empty tuple says each of the five scripted questions maps to a populated field. It says nothing about what the desk then asked, how long the call took or how the caller felt. It also cannot see the missing charge, because that fact is not on the script. ::: ## Can the agent promise callers they won't repeat themselves? Not unconditionally, and this section shows the two instructions that contradict each other. What the desk may do is worked out from the transferred tier when the package is read. The demo caller is hard-coded at `medium`, and the permissions test asserts that account details may be discussed but irreversible changes may not. Raising a dispute needs `high` (`_DISPUTE_MIN_TIER` in `skills/dispute_transaction/tools.py`, lines 58–60), and the comment there says the medium caller "is what makes the handoff happen". So the desk screen tells the human to step the caller up before the dispute can go through. That is correct design, and it hurts a handle-time slide. :::callout{type="pitfall" title="The promise that callers won't repeat themselves has no condition"} The handoff tool's hint tells the model to say "that they will not need to repeat themselves", and the skill's closing step says the same. Neither checks `desk_still_needs_to_ask`, and gating on it would not help: the list was empty on the dispute path, where the desk must still step the caller up. Reword the promise, for example to say the specialist will have what the caller told the agent and may need to confirm their identity, or record the contradiction as an accepted cost. ::: :::solution{title="Show the source for the promise and the tier rule"} - `skills/human_handoff/tools.py`, lines 210–214: the hint that tells the model to promise no repeats. - `skills/human_handoff/skill.md`, lines 30–33: the skill's closing step. - `skills/dispute_transaction/tools.py`, line 44: the demo caller's tier, hard-coded to `medium`. - `test_desk_permissions_follow_the_tier_not_the_package`: the permissions test. `handoffpkg/desk.py`, lines 112–121, with the leading indentation removed, is the tier rule: ```python tier = package.identity.verified_tier actions = ["Answer general questions", "Explain what the agent already did"] if tier_at_least(tier, "medium"): actions.append("Discuss account-specific details") else: actions.append("DO NOT discuss account-specific details — identity not established") if tier_at_least(tier, "high"): actions.append("Action irreversible changes (transfers, card reissue, SIM swap)") else: actions.append("DO NOT action irreversible changes — step the caller up first") ``` ::: ## How do you measure AI agent human handoff success after launch? This section gives you the order of work for the numbers only real calls can produce. The whole suite runs from `patterns/voice-handoff-context`: ::run{cmd="make test"} In the saved run on 21 September 2026 (Python 3.14.3), all 41 tests passed over synthetic sessions, with no network, model or licence. None places a call, times a human or reads a production session. ::::steps ### Version your desk's opening script Fix the questions and date the list before counting anything. Figures taken against different versions cannot be compared. ### Recompute package coverage from stored packages The tool computes the count on every transfer but does not store it. It does save each package to a file named after its handoff, so recompute the count from those files and call it package coverage. Coverage is not a repeat rate: a desk agent who asks for the date of birth out of habit, from a package that retired it, shows up only in call review. :::solution{title="Show the source"} `handoffpkg/desk.py`, lines 247–250: `deliver` writes the whole package to JSON. `skills/human_handoff/tools.py`, line 189: the file is named after the `handoff_id`. The tool returns `desk_still_needs_to_ask` without storing it. ::: ### Sample transferred calls and count what the desk asked Join each call to its package by `handoff_id`. Count scripted questions the desk asked although the package answered them, and separately list every question not on the script with what it asked for. ### Report the three figures side by side Put coverage, the repeat rate and the off-script list together with the sample size. Without the off-script list, the fixed script hides the loss the dispute path shows. ### Take handle time and satisfaction from their owners Handle time comes from contact-centre reporting on a comparable queue mix; satisfaction from your CSAT or complaints process, scoped to transfers. :::: The companion tests none of these steps. Step three depends on what your contact centre records: ::::tabs :::tab{label="Calls are recorded, not transcribed"} A QA reviewer lists the populated script questions from each sampled package, listens to the desk side, and marks each as asked or not, noting every other question too. The repeat rate is script questions asked again over populated script questions in the sample, not a rate for all calls. ::: :::tab{label="Calls are transcribed"} Search the desk agent's turns for each populated script question. Agents paraphrase, so a text match will miss "and your name is?" for "Can I take your name?". Questions that match nothing on the script go into a separate list for a reviewer to label. Have QA listen to a sample the search cleared, and report how many it missed. ::: :::: ## What goes on a launch-review slide you can defend? Here is Aurora's slide rewritten so that someone in the room can check every line. | Part of the slide | What it says | Evidence or owner | | ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------ | | Proven offline | On the fixture dispute path the package answers all five opening-script questions; the one-line handoff answers none | The baseline and agent-path tests at `69e27b6`, output above | | Limits of the proof | Populated, not correct; tier, goal and attempts set by the test; no disputed amount, merchant or date, yet five of five | The agent-path test and the dispute-path run above | | Not yet measured | Handle time, the desk repeat rate, off-script desk questions, satisfaction with transfers | An owner, a data source and a date for each | | Changes owed | Dispute tools write the amount, merchant and date under test, or the gap is recorded; the promise is reworded or its cost accepted; the script is versioned | Engineering, conversation design and desk operations | | Not on this slide | Redaction. The README says the allowlist "governs session-state KEYS. It cannot police free text" and "It is not a compliance control." | A separate slide | The same suite shows the package withholding seven planted sensitive values and a field nobody anticipated. Once the owed changes are decided, that is reasonable to ship on offline evidence. Announcing a business effect before anyone has counted a call is not. :::solution{title="Show the source"} `SENSITIVE_VALUES` (`tests/test_handoff_context.py`, lines 66–74) plants the seven values, and `test_no_sensitive_value_appears_anywhere_in_the_package` (lines 224–233) searches the serialised package for each. `test_a_field_nobody_anticipated_is_withheld_by_default` covers the new field. Both passed in the full `make test` run. ::: [Design a handoff a customer can trust](/library/guides/design-a-human-handoff/) covers consent, the context packet and recovery when the receiving team is unavailable. The [tool and memory scope tutorial](/library/tutorials/tools-and-memory/) builds a bank agent that makes the caller verify exactly once. :::cta{href="/library/guides/design-the-handoff-summary/" label="Design the handoff summary"} The design side of the same pattern: build the human agent's first screen from typed fields and test the transfer promise on the agent's own path. ::: --- # 186 mutations killed, and the guard still broke Source: https://rasa.community/library/guides/mutation-test-an-agent-guard/ Author: Rod Rivera Published: 2026-09-10 Clone `RasaHQ/rasa-community-resources`, check out `69e27b6`, change into `tutorials/rasa-ai-team-casebook` and run the lab's offline suite. Seven tests pass. One of them, `test_all_authored_oracles_and_mutations`, loads 62 scenario fixtures and asserts three killed mutations for each, 186 in all. Now delete one word. Line 97 of `casebook.py` is the body of `contract_hash`, which fixes a case's identity once a request has been committed. Take `'rules'` out of its tuple: ```diff def contract_hash(spec: dict) -> str: - return digest({k: spec.get(k) for k in ('slug', 'action', 'rules', 'receipt', 'provenance', 'intakeKind')}) + return digest({k: spec.get(k) for k in ('slug', 'action', 'receipt', 'provenance', 'intakeKind')}) ``` Run the same seven tests again: ```text test_all_authored_oracles_and_mutations (test_casebook.CasebookTests.test_all_authored_oracles_and_mutations) ... ok test_arbitrary_keys_and_truthy_values_do_not_grant_authority (test_casebook.CasebookTests.test_arbitrary_keys_and_truthy_values_do_not_grant_authority) ... ok test_changed_case_or_revision_cannot_reuse_committed_identity (test_casebook.CasebookTests.test_changed_case_or_revision_cannot_reuse_committed_identity) ... ok test_concurrent_retries_persist_exactly_one_synthetic_effect (test_casebook.CasebookTests.test_concurrent_retries_persist_exactly_one_synthetic_effect) ... ok test_missing_is_unknown_and_reconciliation_cannot_create_an_action (test_casebook.CasebookTests.test_missing_is_unknown_and_reconciliation_cannot_create_an_action) ... ok test_path_and_identifier_validation (test_casebook.CasebookTests.test_path_and_identifier_validation) ... ok test_receipt_failure_never_means_no_action (test_casebook.CasebookTests.test_receipt_failure_never_means_no_action) ... ok ---------------------------------------------------------------------- Ran 7 tests in 0.701s OK ``` That is the author's offline run, with Python 3.14.3, in a copy of the lab exported with `git archive`, so the author's runs never touched the companion checkout. The same day, an earlier pair of runs edited the checkout in place and reverted it afterwards. They recorded the same result: `OK` on the unmodified tree, then `OK` again after the edit. Timings vary between runs and machines. The suite is well covered, too. On the unmodified tree, coverage.py 7.16.1 reported `casebook.py` at 82% of 194 statements, with line 97 among the executed lines. With that word gone, a pending request can be reconciled after any of its case's rules has been removed, and a replayed request no longer notices its receipt-phase rules or its rules' reasons changing. Neither the suite that reports 186 mutations killed nor the coverage report notices that. This guide is for the evaluation engineer who owns a green, well-covered suite around an agent guard and wants to know whether it would catch someone breaking that guard. The stance: **a suite that reports every one of its own mutations killed is not evidence against a mutation nobody wrote.** The mutation worth running is one aimed at what decides the outcome, and the survivors you find should be counted in the open, even when that makes your headline number smaller. Everything here is an offline Python check against synthetic fixtures. No language model and no external service is involved. ## Mutation testing Python unit tests: what the 186 counts Mutation testing is an old idea. DeMillo, Lipton and Sayward set it out in 1978 ("Hints on Test Data Selection", _IEEE Computer_): make small deliberate changes to the code, called mutants, and see whether the tests fail. A mutant the tests catch is killed. One they miss survives, and a survivor marks a behaviour no test pins down. Tools such as mutmut and cosmic-ray apply the idea to Python source; this lab applies it by hand, to data. Two other guides on this site already use single deletions as evidence. The [fixture guide](/library/guides/derive-test-fixtures-from-production-code/) deletes one key and watches which test notices, and the [action register guide](/library/guides/sign-off-an-agent-action-register/) asks for a test that fails when the guard line is deleted. This guide deals with what comes after that habit becomes a count: which population of edits the count covers, and which survivors make no difference at all. The lab's loop is in `prove()` (`casebook.py`, lines 212–223): ```python # Exercise each broken predicate against the independently stored oracle. killed = [] for rule in spec['rules']: mutant = json.loads(json.dumps(spec)) mutant['rules'] = [r for r in mutant['rules'] if r['field'] != rule['field']] fixture = next(f for f in spec['variants'] if f['name'] == rule['field'] + '-false') # Mutant no longer considers the deleted predicate part of the contract. facts = {k: v for k, v in fixture['facts'].items() if k != rule['field']} with tempfile.TemporaryDirectory() as tmp: outcome = execute(mutant, facts, 'mutant', Path(tmp) / 'ledger.sqlite') assert outcome['status'] != fixture['expected']['status'], 'surviving predicate deletion' killed.append(rule['field']) ``` Read it for three things: what it changes, what it compares against, and what it counts as killed. - **What it changes.** One whole rule, removed from `spec['rules']`. Every case has exactly three rules (`load_case` rejects any other count, line 28), so every case yields exactly three mutants. - **What it compares against.** The stored fixture named `-false`, the one written to fail on that rule. The mutant runs once, under the request ID `'mutant'`, on a fresh ledger in a temporary directory (lines 220–221). - **What counts as killed.** The mutant's status differs from the status the fixture expects. For the handoff case's `redacted_summary` rule, the fixture expects `blocked`, and in the author's offline run of the same steps the mutant returned `succeeded`, so it counts as killed. The test then asserts the count (`tests/test_casebook.py`, lines 18–25): ```python def test_all_authored_oracles_and_mutations(self): cases = sorted((ROOT / 'examples').glob('*.json')) self.assertEqual(len(cases), 62) for path in cases: with self.subTest(case=path.stem): outcome = prove(load_case(path.stem), Path(self.tmp.name) / f'{path.stem}.sqlite') self.assertEqual(len(outcome['observations']), 10) self.assertEqual(len(outcome['mutationsKilled']), 3) ``` Sixty-two cases times three rules is 186. The project's verification record states the count in the same terms: the method "checks 62 cases, 620 expected input outcomes and 186 predicate-deletion mutations", then lists the per-case replay, lost-acknowledgement and reconciliation checks (`VERIFICATION.md`, lines 3–5). The compatibility receipt names each check rather than folding them into a single pass (`COMPATIBILITY.json`, lines 116–131): ```json { "path": "tutorials/rasa-ai-team-casebook", "sourceHash": "0f945cd467cdbeab0c48d647b97915fe89e4c9639edbf2f9a4c7efbcf3627e93", "checks": [ "locked-install", "installed-version", "mantle-import", "project-validation", "licensed-training", "project-unittests", "62-scenario-oracles", "186-mutations-killed", "62-sdk-tool-dispatches" ], "liveConversations": "not tested" }, ``` Nothing in those records is overstated. `186-mutations-killed` is a true, reproducible number, and it names its population: mutations shaped like "delete this rule's predicate". The mistake comes later, when that number is read as "the guard would survive being tampered with". The suite generates its own mutants, chooses the fixture each is judged against, and declares the verdict. It cannot, by construction, report on an edit shaped any other way. ## Kill it by breaking the code, not the data An authored mutant edits the case. The edit that matters in review is one someone makes to the code the case runs through. So the second question is: if I change `casebook.py`, does any of the seven tests fail? The kill criterion changes with the question. For a code mutant, "killed" means the suite that passes on the clean tree goes red. Work in a throwaway copy, never the checkout you report from. From the root of your clone, one way to get one is a Git worktree (we ran these commands on macOS): ```bash git worktree add ../casebook-scratch 69e27b6 cd ../casebook-scratch/tutorials/rasa-ai-team-casebook ``` Run the suite once on the clean tree before you change anything. This is the lab's own `make proof` target: ::run{mac="python3 -m unittest discover -s tests -v" linux="python3 -m unittest discover -s tests -v" windows="python -m unittest discover -s tests -v"} On the clean tree, the author's run printed seven `ok` lines, then `Ran 7 tests in 0.732s` and `OK`. In a separate run of the two worktree commands in a local clone, the suite passed in the new worktree, `git status` showed nothing changed, and `git worktree remove ../casebook-scratch` cleaned it up. Applying the mutant is the step that differs by platform, because in-place editing is spelt differently in BSD sed, GNU sed and PowerShell. Each version below keeps a copy of the original, makes the one-word edit to line 97, shows the changed line, runs the suite and puts the original back. The author ran the three edits on macOS with the system sed, GNU sed 4.9 and PowerShell 7.4.6, and all three produced the same line 97. None of them was run on Windows itself, where `Set-Content` writes Windows line endings, so what is known to match there is the text of line 97, not the bytes of the file. ::::tabs :::tab{label="macOS"} ```bash cp casebook.py casebook.py.orig sed -i '' "97s/'rules', //" casebook.py diff casebook.py.orig casebook.py python3 -m unittest discover -s tests -v cp casebook.py.orig casebook.py ``` ::: :::tab{label="Linux (GNU sed)"} ```bash cp casebook.py casebook.py.orig sed -i "97s/'rules', //" casebook.py diff casebook.py.orig casebook.py python3 -m unittest discover -s tests -v cp casebook.py.orig casebook.py ``` ::: :::tab{label="Windows (PowerShell)"} ```powershell Copy-Item casebook.py casebook.py.orig $lines = Get-Content casebook.py $lines[96] = $lines[96].Replace("'rules', ", "") Set-Content casebook.py $lines Compare-Object (Get-Content casebook.py.orig) (Get-Content casebook.py) python -m unittest discover -s tests -v Copy-Item casebook.py.orig casebook.py -Force ``` ::: :::: The `diff` or `Compare-Object` line is there so you can see the edit landed on line 97 before you trust a green run. If it prints nothing, the edit missed and the suite is testing the unmodified file. PowerShell arrays count from zero, which is why line 97 is `$lines[96]`. On Windows the interpreter may be `py` rather than `python`, depending on how Python was installed. A first attempt shows the suite doing its job. In `evaluate()`, make a missing fact count as true by changing `facts.get(rule['field'])` on line 86 to `facts.get(rule['field'], True)`. In the author's offline run, the suite went red at once with `FAILED (failures=62)`, one failure per case. The first, in the `banking-advisor-appointment` case, named the fixture that caught it: ```text AssertionError: ('purpose_matched-missing', {'status': 'succeeded', 'reason': 'verified_fixture_receipt', 'effects': 1}, {'status': 'blocked', 'reason': 'wrong_advisor_capability', 'effects': 0}) ``` All 62 fixture files carry the same set of variants: `accepted`, plus a `-false`, a `-missing` and a `-string` variant for each of the three rules. That is where the ten observations per case come from, and why a mutation nobody listed in `prove()` was killed by the 62-scenario oracle anyway. The suite is not weak. It is strong exactly where its author looked, and the fixtures show where that was: the rule evaluator. `contract_hash` is somewhere else. :::diagram{title="Where each mutant lands, and which tests can see it"} ```dot rankdir=LR; spec [label="case fixture\nspec['rules']"]; evaluate [label="evaluate()\nlines 84–88"]; hash [label="contract_hash()\nline 97"]; ledger [label="stored request row\ncontract_hash column"]; check [label="replay and reconcile\nlines 117 and 151"]; m1 [label="186 rule deletions\nprove(), lines 212–223", class="ok"]; m2 [label="missing fact treated as true\nline 86", class="ok"]; m3 [label="'rules' dropped from the tuple\nline 97", class="blocked"]; spec -> evaluate; spec -> hash; hash -> ledger; ledger -> check; m1 -> evaluate [label="killed by prove()"]; m2 -> evaluate [label="killed by fixtures"]; m3 -> hash [label="survives all 7 tests"]; ``` ::: The first two mutants land in the evaluator, which the fixtures exercise from ten directions per case. The third lands in the identity hash, whose value only matters when a stored request is compared with a case that has changed since. ## What contract_hash decides Here is the function in full (`casebook.py`, lines 96–97): ```python def contract_hash(spec: dict) -> str: return digest({k: spec.get(k) for k in ('slug', 'action', 'rules', 'receipt', 'provenance', 'intakeKind')}) ``` `execute()` stores that hash beside every committed request. When the same request ID comes back, the stored fingerprint and hash both have to match the case as it is now; if they do, the stored result is replayed (lines 116–119): ```python if previous: if previous['fingerprint'] != fingerprint or previous['contract_hash'] != contract_hash(spec): return result('conflict', 'request_identity_changed', effects=1, replay=True) return result(previous['status'], previous['reason'], previous['reference'], 1, True) ``` `reconcile()`, which re-checks the receipt facts and turns a `pending` row into `succeeded` when they pass (lines 153–157), first compares the stored case and hash with the current ones (lines 151–152). It does not compare the fingerprint: ```python if row['case_id'] != spec['slug'] or row['contract_hash'] != contract_hash(spec): return result('conflict', 'receipt_contract_changed', effects=1) ``` The fingerprint is the other half of that identity (lines 109–111): ```python bound_facts = {r['field']: facts.get(r['field']) for r in spec['rules'] if r['phase'] == 'request'} fingerprint = digest({'case': spec['slug'], 'facts': bound_facts}) ``` Read what it binds. It hashes the case slug plus the names and values of the request-phase facts, so adding or removing a request-phase rule changes it. It never sees receipt-phase rules, and it never sees a rule's other attributes, such as its `reason`. And only `execute()` compares it: `reconcile()` does not check the fingerprint at all. So on replay, receipt rules and rule reasons are bound to a stored request only by the `'rules'` entry in `contract_hash`, and at reconciliation every rule is, request-phase rules included. Take that entry out, and a replayed request no longer notices its receipt rules being removed or its rules' reasons being changed, while a pending request can be reconciled after any of its rules has been removed. Why does no test notice? In `tests/test_casebook.py`, only one test changes a case and then comes back with a request ID that is already committed (lines 38–46): ```python def test_changed_case_or_revision_cannot_reuse_committed_identity(self): execute(self.spec, self.spec['facts'], 'request-1', self.database) other = load_case('banking-transfer') self.assertEqual(execute(other, other['facts'], 'request-1', self.database)['status'], 'conflict') changed = json.loads(json.dumps(self.spec)) changed['provenance']['revision'] = 'changed' self.assertEqual(execute(changed, changed['facts'], 'request-1', self.database)['status'], 'conflict') self.assertEqual(reconcile(changed, changed['facts'], 'request-1', self.database)['status'], 'conflict') self.assertEqual(count_effects(self.database), 1) ``` It swaps in a different case, which the slug in the fingerprint already catches, then changes `provenance.revision` and edits nothing else. The hash lists six fields; the test exercises one of them directly. :::callout{type="tip" title="Aim mutants where a stored value meets a recomputed one"} Search the guard for comparisons between a row saved earlier and the case as it is now. `casebook.py` has three: the intake check on line 68, the replay check on line 117 and the reconcile check on line 151. Each decides whether an earlier result may stand. The authored mutants in `prove()` each run once on an empty ledger, so they never reach any of the three. ::: This matters beyond the lab's own tests because the Rasa tool reuses request IDs by design. The Mantle adapter in `tools/cases.py` builds the ID from an operator-set run ID and the case, not from anything per call (lines 9–13 and 20–24): ```python def identity(case_id): # The lab operator assigns a stable run ID. A real backend uses the # authenticated subject + action revision, resolved from trusted context. run_id = os.environ.get('CASEBOOK_RUN_ID', 'rehearsal-1') return f'{run_id}-{case_id}' ``` ```python @tool(description='Rehearse one supported synthetic case. Returns a lab outcome, never a real customer action.') async def rehearse_case(case_id: str, context: ToolContext = None) -> ToolResult: try: spec = load_case(case_id) outcome = execute(spec, spec['facts'], identity(case_id), database()) ``` Read together with lines 116–119, the code says that a second rehearsal of the same case against the same ledger replays whatever the first one stored, and that `contract_hash` is the check meant to refuse that replay when the case's rules have changed in between. That is a reading of the code only. This guide did not run `rehearse_case`, with or without the mutant. :::checkpoint{id="mutation-count-contract-hash" question="The suite passes and reports 186 of 186 authored mutations killed. Someone removes 'rules' from the contract_hash tuple. What does the 186 tell you about that edit?" options="It is covered: all 186 mutations change the rules that the hash includes,It is covered: prove() sends repeat requests through the hash comparison on line 117,Nothing: no authored mutation touches contract_hash,It is covered: coverage.py reports line 97 as executed" answer="2"} Each of the 186 mutants changes the case data and runs once, on a fresh ledger, so no stored hash is ever compared against it. `prove()` does reach the comparison on line 117: it repeats each fixture under the same request ID (line 177), replays the lost-acknowledgement request (line 195) and retries with changed facts (line 201). Every time, the case is the same one that stored the row, so the stored hash and the recomputed hash agree whatever the tuple contains. Coverage is right that line 97 runs, but running a line is not the same as checking what it computes. No existing test commits a request, changes that same case's rules and comes back with the same ID, which is what a behavioural test has to do to see this edit. A direct unit test of the hash would see it too: in the author's offline run, removing one rule from the handoff case changed `contract_hash` on the clean tree and left it unchanged on the mutant. ::: :::solution{title="Try it first: write the test that kills this mutant"} Commit a request that stops at `pending`, remove the receipt rule it is waiting on, then try to reconcile under the changed case. The unmodified guard should refuse with `conflict`. This is the author's scratch test, not part of the companion project: ```python import json from pathlib import Path import tempfile import unittest from casebook import execute, load_case, reconcile class ContractBindingTests(unittest.TestCase): def test_changed_rules_cannot_reconcile_a_committed_request(self): spec = load_case('contextual-handoff') with tempfile.TemporaryDirectory() as tmp: database = Path(tmp) / 'ledger.sqlite' facts = dict(spec['facts'], desk_acknowledged=False) self.assertEqual(execute(spec, facts, 'r1', database)['status'], 'pending') changed = json.loads(json.dumps(spec)) changed['rules'] = [r for r in changed['rules'] if r['field'] != 'desk_acknowledged'] later = {k: v for k, v in spec['facts'].items() if k != 'desk_acknowledged'} self.assertEqual(reconcile(changed, later, 'r1', database)['status'], 'conflict') ``` In the author's offline runs it passed against the unmodified `casebook.py`. With `'rules'` removed from the tuple, the last lines of the run were: ```text AssertionError: 'succeeded' != 'conflict' - succeeded + conflict ---------------------------------------------------------------------- Ran 1 test in 0.003s FAILED (failures=1) ``` Read what the mutant did. A handoff was waiting for the desk to acknowledge it. Someone then deleted the acknowledgement rule from the case, and reconciliation marked the old request `succeeded` against a check it never passed. The request did not change. The rules it was judged by did. ::: :::callout{type="warn" title="A new test can pass on the mutant too"} The test above removes a receipt-phase rule on purpose. The author also replayed a committed handoff request through `execute()` after three different edits to the case. Removing the request-phase rule `redacted_summary` returned `conflict` on both trees, because that field's name drops out of the fingerprint and line 117 refuses before the hash matters; a replay test built on that edit passes against the mutant and kills nothing. Removing the receipt-phase rule `desk_acknowledged`, or changing the `reason` of `redacted_summary`, returned `conflict` on the clean tree and replayed `succeeded` on the mutant. The same request-phase removal does kill the mutant at reconciliation, which never checks the fingerprint: with the request left `pending`, the author's reconcile run returned `conflict` on the clean tree and `succeeded` on the mutant. Run every new test against the mutant as well as the clean tree, and keep it only if the mutant turns it red. ::: ## Count what you tried, not only what you killed Once one field in the tuple has survived, the fair next step is to try the others. In offline runs, the author deleted each of the six fields from line 97 in turn, one per run, and reran the seven tests: | Field removed from `contract_hash` | Seven-test suite | Reading | | ---------------------------------- | ---------------- | ----------------------------------------------------------------------------------------------------- | | `provenance` | 1 failure | Killed by `test_changed_case_or_revision_cannot_reuse_committed_identity` | | `slug` | OK | Equivalent: the fingerprint and `case_id` still bind it | | `rules` | OK | Survivor: no rule edit is bound at reconcile; on replay, receipt-phase rules and rule reasons are not | | `action` | OK | Survivor | | `receipt` | OK | Survivor | | `intakeKind` | OK | Survivor | The `slug` row is the part mutation testing makes you do by hand. The case slug is already in the fingerprint on line 111 and is checked against `case_id` on line 151, so removing it from the hash changes nothing a caller could observe. That is an equivalent mutant, and no test should be written to kill it. Equivalent mutants are a standing cost of the technique: some survivors are real gaps, some are changes that make no difference, and only reading the code tells them apart. The other three survivors are descriptive strings. `action` and `receipt` are strings in all 62 fixtures; `intakeKind` is a string in the 4 cases that have an independent intake and absent from the other 58. Whether a change to one of them should invalidate a committed request is a decision for the lab's owner. The tuple on line 97 records that its author thought it should. An honest report of this suite reads more like this: - predicate deletion, generated by `prove()`: 186 of 186 killed; - missing fact treated as true, in `evaluate()`: killed by the 62-scenario oracles; - single-field deletion from `contract_hash`: 6 tried, 1 killed, 1 equivalent, 4 survived. That is less satisfying than "186 killed", and more useful. It names each population and its denominator, and it tells the next reviewer where to aim. The [evaluation set guide](/library/guides/build-an-evaluation-set/) asks the same of agent evaluations: report passed, failed and not-run counts for every slice. The project's `COMPATIBILITY.json` already names checks instead of reporting one pass, so the form is there to extend. :::callout{type="pitfall" title="Never mutate the checkout you report from"} Before you publish a number, run `git status` in the checkout you report from and confirm it shows nothing changed. A mutant left there turns every later green run into a lie about the code you ship. ::: ## What this costs, and what it does not show The cost is time and judgement. Six hand-written mutants took one line each to make, and one of the six, `slug`, needed code reading to rule out. Every extra mutant, whether you write it or a tool generates it, is another possible survivor that someone has to read and classify as a gap or an equivalent. You are choosing to spend reviewer time on fewer, targeted mutants instead of a larger number that is easier to put in a report. The survivor is also narrower than it looks. On replay through `execute()`, 155 of the 186 rules across the 62 fixtures are request-phase, so removing one changes the fingerprint and replay refuses anyway; there the survivor bites only through the other 31 and through rule attributes such as `reason`. Reconciliation has no such backstop. Every fixture carries a `provenance.revision`, and the one test that edits a committed case changes exactly that field. If everyone who edits a case's rules also bumps its revision, the provenance check catches the change and the survivor never fires. The mutant shows that the rules binding is untested, not that the lab is unsafe as used. These checks do not show how a model behaves, how the Mantle tool behaves in a live conversation, or anything about a real service. The lab is a synthetic teaching contract, and the counts above are finite offline results, not rates. :::cta{href="/library/tutorials/evaluation-harness/" label="Build the evaluation harness"} This guide tests the tests around a guard. The evaluation harness tutorial tests the agent: an LLM plays your customer, assertions stay the ground truth, and a judge scores what facts cannot reach. ::: --- # Review what an agent is allowed to do Source: https://rasa.community/library/guides/review-agent-permissions/ Author: Rasa team Published: 2026-09-16 Ask “what can this agent change?” before asking whether it sounds careful. A polite explanation does not establish authority to update a record, send a document, or commit a transaction. This guide is for a domain owner or risk reviewer working with product and engineering. You will produce an action register and a set of evidence requests. It is an engineering review aid, not a legal determination or a compliance certification. ## Build an action register The example below uses fictional Horizon Travel. Its agent can explain booking conditions and prepare a request; staff handle exceptions. These are proposed boundaries for the example, not a description of a real business. | Action | Required authority and data | Boundary to inspect | Accountable owner | | ----------------------- | ------------------------------------------ | ------------------------------------------------ | ---------------------- | | Read a booking | Customer authorized for that booking | Lookup verifies access before returning details | Booking service owner | | Explain a policy | Current policy associated with the booking | Answer traces to the retrieved source | Travel policy owner | | Request staff review | Customer-approved summary | Receiving queue contains the agreed fields | Support lead | | Commit a booking change | Verified authority and accepted conditions | Tool rejects missing or mismatched authorization | Booking service owner | | Export a document | Approved source state and recipient | Artifact and delivery match the approved state | Document process owner | List the actual actions in your system, including those hidden behind a broad tool name such as “update_customer.” Ask engineering to expand the tool's effects. A label can hide multiple permissions. ## Request evidence at the point of action The [guarded-action tutorial](/library/tutorials/guarding-irreversible-actions/) illustrates why a tool boundary matters for an irreversible change. The [derived-document tutorial](/library/tutorials/deriving-a-document/) explores keeping an artifact tied to source state. For each action, ask for one allowed case and one denied case. Inspect the resulting system record, not just the agent's narration. A denied case passes only if the prohibited change did not happen. Keep the input conditions and software version with the evidence so another reviewer can reproduce the check. For a booking change, useful denied cases include a reference belonging to another customer, expired authorization, and a request that changes after confirmation. Your team decides which controls implement these rules; the review makes the intended invariant explicit. ## Do not confuse approval with unlimited scope A customer approving one action does not approve every later action. In the fictional example, approval to send a fee question to support does not authorize a booking change. Write the scope into the action register and the conversation design. The difficult case is a corrected address or recipient after the user has approved a document. Ask whether the confirmation is bound to the exact data used at execution. If the system cannot establish that relationship, mark it unresolved and hold that action or route it to a human with the relevant authority. NIST's [AI RMF Core](https://airc.nist.gov/airmf-resources/airmf/5-sec-core/) addresses documenting context, measurement, and management of risks. It does not make this checklist a substitute for your organization's policy or applicable obligations. ## Leave with a decision someone owns For each row, record: allowed scope; source of authority; evidence inspected; unresolved cases; disposition; owner; review date. Use “unknown” when evidence is missing. Do not turn an assumption into an approval because a deadline is near. A useful disposition can be narrow: permit policy lookup while holding booking changes. Product must ensure the user-facing workflow reflects that limitation; operations must be able to disable the affected action; design must provide an honest handoff. Before the next release, choose the highest-impact action and ask engineering to demonstrate its denied case. Bring the register to a [release decision review](/library/guides/make-a-release-decision/). This creates a concrete conversation between the domain owner and the builder, with evidence both can inspect. --- # Roll out an agent with a working stop button Source: https://rasa.community/library/guides/roll-out-an-agent/ Author: Rasa team Published: 2026-09-17 Before expanding access to an agent, prove that somebody can stop it and that customers still have a usable route to help. A rollback document is only a proposal until the team exercises it. This guide is for the platform engineer or operator responsible for a Rasa deployment. It describes an operational contract to adapt to your infrastructure; it does not prescribe an unverified deployment command. ## Name the unit you will release Record the application revision, Rasa version, model configuration, tool configuration, and evaluation dataset version together. A model or tool change can alter behavior even when application code is unchanged. Keep credentials in the deployment's secret system, not in the release record. For a fictional travel assistant, the first slice might serve itinerary questions for a small opted-in pilot. The size is your team's decision based on support capacity and observation; there is no universal safe percentage in this guide. | Dependency | Check before enabling the slice | If it fails | | ------------------- | -------------------------------------------------- | ---------------------------------------------- | | Booking lookup | Authorized fixture returns the expected record | Disable the affected workflow | | Model provider | Request completes within the team's latency budget | Use the verified fallback or stop the workflow | | Support queue | Test request appears in the receiving system | Keep human-handoff flows unavailable | | Secrets and licence | Service can start with the deployed configuration | Hold the release | | Operator access | Named on-call person can change routing | Hold until an owner is reachable | Use synthetic or authorized test records for these checks and label them. The results are evidence only if the dependency was actually exercised. ## Watch outcomes, not just HTTP status An agent can return HTTP 200 while giving the wrong answer. Track request counts, task outcomes, tool errors, latency, and handoff delivery separately. Define each denominator and make unresolved outcomes visible. Connect traces through an interaction identifier, but avoid placing full conversation text or credentials into general-purpose logs by default. Agree retention and access with the data owner. NIST's [AI actor tasks](https://airc.nist.gov/airmf-resources/airmf/appendices/app-a-descriptions-of-ai-actor-tasks/) distinguishes deployment and operation from other lifecycle work; our checklist translates that responsibility into practical checks. Write an alert as an action: “If the booking lookup is unavailable, the on-call operator disables the booking workflow and confirms the alternative support route.” A dashboard with no responsible person is an observation surface, not a response plan. ## Rehearse rollback before the pilot 1. Save the currently accepted application and configuration versions. 2. Enable the candidate in an isolated environment with an authorized test interaction. 3. Trigger a representative dependency failure. 4. Use the actual routing or deployment control to stop new traffic to the candidate. 5. Confirm that the prior version or alternative support route handles a new request. 6. Inspect any in-flight action and the receiving system for partial or duplicate work. Record timestamps and resulting states. Measure recovery time from the drill instead of inventing a target achievement. The target belongs to your service agreement; the measurement tells you whether this mechanism can meet it. The awkward case is an external action committed immediately before rollback. Reverting application code cannot undo a paid booking change. The rollout plan needs an owner who can reconcile that resulting record, and the agent needs a [guarded action boundary](/library/tutorials/guarding-irreversible-actions/) before it reaches production. ## Make expansion a separate decision Review the pilot's task mix against the evaluation set, check fallback demand against the receiving team's capacity, and list unexplained failures. Use the [release decision record](/library/guides/make-a-release-decision/) to expand, hold, or stop explicitly. Your next step is one rollback drill with a named observer. If the stop mechanism does not work, fix it before increasing access. For Rasa's licence configuration details, consult the current [licensing documentation](https://rasa.com/docs/pro/installation/licensing/); the deployment contract here remains specific to your own infrastructure. --- # Is an MCP integration a drop-in replacement? Source: https://rasa.community/library/guides/scope-an-mcp-transport-swap/ Author: Rod Rivera Published: 2026-09-20 An engineer on Meridian's support team opens a pull request: move the CRM tools onto the vendor's MCP server. The companion project makes that change with `make mcp-swap` (`Makefile`, lines 101–113). Each of the three skill files is replaced by an MCP twin whose only frontmatter change is the import: `check_tickets` and `log_interaction` gain an `import_tools:` key and one `mcp/` entry, and `identify_customer` has its one import line rewritten to `- mcp/hubspot_crm:find_contact_by_email`. The proof below reports "2 changed line(s)" for each because it counts removed plus added lines. `integrations.yml` gains an `mcp_servers:` block, and three Python files are moved aside: `tools/crm.py`, `skills/check_tickets/tools.py` and `skills/log_interaction/tools.py`. As review evidence, the pull request carries the output of the companion's own proof, `make mcp-prove`. This is that output, from an offline run at the pinned revision, with no model and no API key: ```text Proving: the same three skills keep their instructions while tools/crm.py is replaced by import_tools: mcp/:. No licence, no API key, no HubSpot account, loopback only. 1. skill instructions are unchanged by the swap ✓ check_tickets: only import_tools differs 2 changed line(s) ✓ check_tickets: instruction body byte-identical 687 bytes ✓ identify_customer: only import_tools differs 2 changed line(s) ✓ identify_customer: instruction body byte-identical 804 bytes ✓ log_interaction: only import_tools differs 2 changed line(s) ✓ log_interaction: instruction body byte-identical 1226 bytes 2. integrations.yml parses as MCP server configuration ✓ parse_mcp_servers accepts the file servers: ['hubspot_crm'] ✓ server url is loopback http http://127.0.0.1:8931/mcp 3. each skill's import parses as mcp/: ✓ check_tickets imports list_open_tickets server=hubspot_crm llm sees 'list_open_tickets' ✓ identify_customer imports find_contact_by_email server=hubspot_crm llm sees 'find_contact_by_email' ✓ log_interaction imports add_timeline_note server=hubspot_crm llm sees 'add_timeline_note' 4. every imported server is configured ✓ no skill imports an unconfigured server ['hubspot_crm'] ⊆ ['hubspot_crm'] 5. the MCP server exposes the imported tools ✓ server exposes list_open_tickets imported by check_tickets ✓ server exposes find_contact_by_email imported by identify_customer ✓ server exposes add_timeline_note imported by log_interaction 6. a tool call over MCP returns the CRM fact ✓ result is structured, not text blocks outputSchema published ✓ find_contact_by_email over MCP Dana Okafor at Okafor Logistics ✓ absent contact is not an error error=contact_not_found The transport changed. The instructions did not. ``` The product manager approves it as a same-sprint config change. Eighteen green checks, and the last line says what the pull request claims. Now read the 687 bytes that check 1 proved unchanged. The MCP variant of `skills/check_tickets/skill.md` still tells the model, at lines 12–13: ```text Call `list_open_tickets`. It reads the identified customer from project memory, so it is the authority on whether the caller has been identified yet. ``` The memory read that sentence describes lived in `skills/check_tickets/tools.py`, one of the files the swap moves aside. The MCP tool that now answers to the name does not read memory at all (`scripts/mcp_crm_server.py`, lines 107–118): ```python @mcp.tool(description="List the support tickets on the identified customer's account.") async def list_open_tickets(contact_id: str) -> CrmResult: """Read the caller's tickets from the CRM. Args: contact_id: HubSpot contact id of the identified customer. """ # `contact_id` is a PARAMETER here, not a memory read. An MCP tool has no # ToolContext, so the value has to travel as an argument. See the README # section "What MCP does not do for you". if not contact_id: return CrmResult(ok=False, error="not_identified") ``` The proof checked that the sentence survived the swap. Nothing in it checked whether the sentence is still true. Meridian, its pull request and the approval are an illustrative scenario. The code and output are real: `tutorials/rasa-hubspot-crm-tutorial` in `RasaHQ/rasa-community-resources` at revision `69e27b6`, where the agent is called Ora (`agent.yml`, lines 5 and 13), run against Rasa Pro 3.20.0rc1. The MCP server in that project is a local mock, and nothing here describes any vendor's real server. The engineer's walkthrough is the [CRM transport swap tutorial](/library/tutorials/crm-transport-swap/); this guide covers the approval, and it disagrees with that tutorial in places, listed chapter by chapter further down. ## Is an MCP integration a drop-in replacement? Check 1 is the evidence the approval leaned on, so read what it accepts. It diffs each REST skill file against its MCP twin and collects the changed lines (`scripts/prove_mcp_swap.py`, lines 89–118). A changed line passes if it is `import_tools:` or starts with `- ` (lines 94–100): ```python # Every changed line must be an import declaration: either the # `import_tools:` key itself or one of its list entries. offending = [ line for line in changed if line != "import_tools:" and not line.startswith("- ") ] ``` It then compares the text below the frontmatter byte for byte (lines 109–114): ```python # And the body below the frontmatter must be identical, byte for byte. rest_body = _read(f"skills/{skill}/skill.md").split("---\n", 2)[-1] mcp_body = _read(f"mcp_variant/skills/{skill}/skill.md").split("---\n", 2)[-1] check( f"{skill}: instruction body byte-identical", rest_body == mcp_body, ``` The body comparison is real: a reworded instruction would turn it red, as the script's docstring says (lines 26–28). The frontmatter test is looser than its comment. Any list entry passes, whatever key it sits under, and `tool_constraints:` entries are list entries too. We changed one line in the MCP `log_interaction` skill, moving its confirmation constraint from the write tool to the ticket read: ```text @@ -6,7 +6,7 @@ import_tools: - mcp/hubspot_crm:add_timeline_note tool_constraints: - - add_timeline_note: + - list_open_tickets: requires_confirmation: enabled: true utter_for_confirmation: utter_confirm_note ``` The proof still passed. Its line for that skill became `✓ log_interaction: only import_tools differs 4 changed line(s)`, every other check stayed green, the last line still read "The transport changed. The instructions did not.", and `make mcp-prove` exited 0. The confirmation declaration had moved off the write tool, and the proof counted the two extra changed lines as import declarations. So the diff and the green run answer a narrow question: did the instruction text move? The scoping decision turns on a different one: for each tool, did the value it relies on stay where the caller cannot influence it? A clean diff cannot stand in for that audit, and the audit, not the size of the diff, decides whether this is one ticket or several. :::checkpoint{id="mcp-swap-identical-text-new-tool" question="A different agent's MCP swap passes the same kind of proof: the refund skill's instructions are byte-identical, and they say the refund tool uses the order the caller verified earlier. Before you schedule it as a config change, what do you still need?" options="Nothing more; identical instructions mean identical behaviour,The refund tool's required arguments on the MCP server,A diff with fewer changed lines,A second green run of the same proof" answer="1"} The proof shows the words survived. Whether "the order the caller verified earlier" is still read from memory or now arrives as an argument the model fills is a fact about the tool, and only its schema or code answers it. This agent and its refund tool are hypothetical; the next section does the same check on the companion's real tools. ::: ## Which value moved, tool by tool Chapter 3 of the tutorial was the first to name the move for the ticket read: `list_open_tickets` stopped reading the id from memory and started taking it as an argument. The same move happened to the write tool; none of the tutorial's chapters says so. In the REST version, the two tools that act on an identified customer read the customer id from project memory, where the lookup tool wrote it. The ticket read (`skills/check_tickets/tools.py`, lines 10–12): ```python async def list_open_tickets(context: ToolContext = None) -> ToolResult: """Read the caller's tickets from the CRM.""" contact_id = context.memory.get("project.contact_id") if context else None ``` The note write (`skills/log_interaction/tools.py`, lines 10–16), whose only argument is the summary: ```python async def add_timeline_note(summary: str, context: ToolContext = None) -> ToolResult: """Write a note onto the customer's CRM record. Args: summary: One or two sentences describing what the caller wanted. """ contact_id = context.memory.get("project.contact_id") if context else None ``` Over MCP there is no `ToolContext`, so the write tool takes the id as an argument too, and writes the note to whichever contact it is given (`scripts/mcp_crm_server.py`, lines 129–139): ```python async def add_timeline_note(contact_id: str, summary: str) -> CrmResult: """Write a note onto the customer's CRM record. Args: contact_id: HubSpot contact id of the identified customer. summary: One or two sentences describing what the caller wanted. """ if not contact_id: return CrmResult(ok=False, error="not_identified") try: note_id = await log_note(str(contact_id), summary) ``` You do not need to read Python to check this for your own vendor. Ask the engineer to list each MCP tool's required arguments. In a copy of the companion project, that is one command and three lines of output (offline, no model; the command discards a library warning on stderr): ```text $ uv run python -c 'import asyncio, scripts.mcp_crm_server as s for t in asyncio.run(s.mcp.list_tools()): print(t.name, "requires", t.inputSchema["required"])' 2>/dev/null find_contact_by_email requires ['email'] list_open_tickets requires ['contact_id'] add_timeline_note requires ['contact_id', 'summary'] ``` A `contact_id` in the required list is a value the caller of the tool must supply. In both versions the caller chooses the email the lookup searches for. The REST lookup then wrote the id it found into memory (`tools/crm.py`, lines 42–45), and the REST tools act only on that id: one a lookup produced. The MCP lookup can only return the id in its result (`scripts/mcp_crm_server.py`, lines 99–104), and the MCP tools act on whatever id they are passed. Who passes it is stated by the companion README: the model "fills it from the conversation" (`README.md`, lines 264–266), and the caller influences the conversation. :::diagram{title="Where each CRM tool gets the customer id, before and after the swap"} ```dot rankdir=LR; caller [label="Caller gives\nan email"]; lookup [label="find_contact_by_email(email)\nsame input on both sides", class="ok"]; memory [label="project.contact_id\n(memory, REST only)"]; model [label="Model fills contact_id\n(MCP)"]; tickets [label="list_open_tickets\nreads tickets", class="blocked"]; note [label="add_timeline_note\nwrites a note", class="blocked"]; caller -> lookup; lookup -> memory [label="REST: writes id"]; lookup -> model [label="MCP: returns id"]; memory -> tickets [label="REST"]; memory -> note [label="REST"]; model -> tickets [label="MCP"]; model -> note [label="MCP"]; ``` ::: The lookup's input did not move. Both tools that act on a customer did. ### The per-tool checklist Take this table into the scoping meeting, with one row per tool the vendor's server will serve. Every row asks the question from the README's own rule: does the tool's correctness depend on a value "the user must not be able to influence" (`README.md`, lines 267–269)? | Tool | Acts on | Input before (REST) | Input after (MCP) | If the input is wrong | Moved in the swap? | | ----------------------- | ---------------------- | ----------------------------- | ------------------------ | -------------------------------------------- | ---------------------------------------- | | `find_contact_by_email` | Looks a customer up | Caller's email | Caller's email | The same in both versions | No | | `list_open_tickets` | Reads tickets | Id the lookup wrote to memory | Id passed as an argument | The agent reads out another person's tickets | Yes: needs a decision | | `add_timeline_note` | Writes to a CRM record | Id the lookup wrote to memory | Id passed as an argument | A note lands on another person's record | Yes: needs a decision, and it is a write | The last two columns are where the approver's attention goes. A wrong read discloses; a wrong write changes someone else's system of record. Both need a decision the diff does not contain. ### The write step, as shipped, does not fit the new schema There is a second problem on the write side. `log_interaction` fixes its write in an ordered block, and the MCP variant keeps that block unchanged. The write step passes one parameter (`mcp_variant/skills/log_interaction/skill.md`, lines 33–36): ```text - id: write_note execute_tool: add_timeline_note parameters: summary: session.log_interaction.note_summary ``` The MCP tool requires `contact_id` as well. In the Rasa Pro source, an execute step's arguments are validated against the schema the tool publishes (`rasa/mantle/orchestration/tool_execution/invoker.py`, lines 63–81). On a failure the step does not call the tool; it hands the turn back to the model (`rasa/mantle/orchestration/tool_execution/constraints.py`, lines 695–706): ```python validation_error = self._tool_invoker.validate_builder_tool_arguments( active_skill_id, step.execute_tool, tool_func, resolved_parameters, ) if validation_error is not None: structlogger.warning( "mantle.skill_executor.advance_steps.tool_arguments_invalid", tool=step.execute_tool, ) return ExecuteStepControl.YIELD_TO_LLM ``` We ran Rasa Pro's own validator offline against the schema the MCP server publishes, converted the way `MCPRuntime.prepare` converts it, with the arguments the step passes and then with `contact_id` added: ```text write_note step as shipped: {'summary': 'Dana asked about invoice 4471.'} -> {"error": "Invalid tool arguments: Missing required argument(s): 'contact_id'.", "retryable": true, "code": "invalid_arguments"} with contact_id added: {'contact_id': '101', 'summary': 'Dana asked about invoice 4471.'} -> None ``` The shipped step fails the check. That check sits before the confirmation logic, which starts at line 708, so on this path the caller is never asked. That ordering is a reading of the source; we did not run the agent against a model for this guide, so what a live run does next is untested. The model may stall on the step, or it may call `add_timeline_note` itself with a `contact_id` it chose. Either outcome needs a ticket, and neither is in the pull request. :::callout{type="pitfall" title="The confirmation prompt names the note, not the customer"} Where the confirmation is reached, it cannot check the id. The MCP skill declares `requires_confirmation` on `add_timeline_note` (`mcp_variant/skills/log_interaction/skill.md`, lines 8–13), and the caller hears `Save this to your record? "{note_summary}"` (`skills/log_interaction/responses.yml`, lines 2–4). That names the summary, not whose record, so a caller saying yes does not check the `contact_id`. ::: ## Why a green build is not the release gate The second thing the pull request cannot show is whether the vendor's server offers the tools the skills import. What follows is a reading of the Rasa Pro 3.20.0rc1 source, not a run. When a model is loaded, `parse_mcp_imports` is called (`rasa/mantle/model_archive/bootstrap.py`, line 208). It parses each `mcp/:` import with `try_parse_mcp_tool_import` (`rasa/mantle/tools/mcp_import_spec.py`, lines 34–55, called at line 86) but does not check whether the tool exists (lines 74–76): ```text Remote existence is not checked here: that requires ``list_tools`` at connection time. Collisions with local/shared callables and reserved framework names are static and fail at model load. ``` The existence check runs when the agent starts. Loading an agent calls `await agent.prepare_runtime_integrations()` (`rasa/core/agent.py`, line 421). That method (lines 664–672) calls `_prepare_runtime_integrations` (lines 141–148), which calls the processor's own `prepare_runtime_integrations` (`rasa/mantle/processor.py`, lines 546–558), which runs `await mcp_runtime.prepare(...)`: `MCPRuntime.prepare` in `rasa/mantle/tools/mcp_runtime.py`, lines 87–131. That connects to each server, calls `list_tools` (line 124) and passes the result to `_filter_imported_tools` (line 131), which raises on a missing tool (lines 237–241): ```python if tool_schema is None: raise ValueError( f"MCP server {reference.server_id!r} does not expose imported " f"tool {reference.tool_name!r} for skill {skill_id!r}." ) ``` So a mistyped tool name is not caught until the agent starts. That is a loud failure rather than a quiet one, but it lands after the build and the proof are green. :::callout{type="warn" title="The proof tests its own server, not the vendor's"} `make mcp-prove` starts the bundled mock servers itself (`scripts/prove_mcp_swap.py`, lines 6–7 and 259–298), asserts that the configured URL is the loopback server (lines 147–154), and its check 5 lists that mock's tools (lines 193–207). It reads its server settings from `mcp_variant/integrations.yml` (lines 134–135), not from the `integrations.yml` the agent runs with after `make mcp-swap` (`Makefile`, line 105). Change the agent's URL to a vendor's and the proof does not see the change, because it never reads that file. Change the URL in `mcp_variant/integrations.yml` and the loopback assertion fails. Either way, the proof never contacts the vendor. ::: The proof stays useful for what it covers: the skill text and the companion's mock. The check that covers the vendor is the agent starting cleanly against the vendor's endpoint, outside production, with the same server id and import lines. The proof needs `uv` (`Makefile`, lines 90–91) and no licence, key or account. These are the steps from a clone; we ran the proof once, on macOS, in a `git archive` export of the same revision rather than a fresh clone: ```text git clone https://github.com/RasaHQ/rasa-community-resources.git cd rasa-community-resources git checkout 69e27b6 cd tutorials/rasa-hubspot-crm-tutorial make mcp-prove ``` We did not run it on Linux or Windows. The target's command is `uv run python scripts/prove_mcp_swap.py`. ## Split the ticket by owner The config swap itself is ordinary work. The boundary decisions are not, and they need a reviewer who is not the engineer who wrote the diff. If you scoped the workflow with a written boundary, as in [choosing an agent workflow worth building](/library/guides/choose-an-agent-workflow/), this is where that boundary gets checked against the tools. A split for the companion project: | Ticket | Covers | Owner | Reviewer | Done when | | -------------------- | ------------------------------------------------------------------------------------------------------ | ---------------------------- | ---------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- | | Config swap | `mcp_servers:` block, `import_tools` lines | Engineer | Any engineer | `make mcp-prove` passes on the loopback configuration, and the agent starts against the vendor endpoint outside production without an "imported tool" error | | Ticket read boundary | `list_open_tickets` takes its `contact_id` as an argument | Engineer | Named security reviewer | One closure from the table below, signed off | | Note write boundary | `add_timeline_note` takes its `contact_id` as an argument, and the write step fails its argument check | Engineer | Named security reviewer | One closure from the table below, signed off | | Instruction accuracy | `check_tickets/skill.md` lines 12–13 describe a memory read | Engineer and product manager | Conversation designer or product manager | The sentence matches the tool that ships | Each boundary ticket closes one of two ways, and the reviewer picks one per tool: | Closure | What the reviewer accepts | Done when | Knock-on effects | | ------------------- | ------------------------------------------------------------------------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Keep the tool local | Nothing new: the tool still reads the id from memory | The tool and its REST skill frontmatter stay; the `check_tickets` sentence stays true | The lookup must stay local too, because it is the only code that writes the id to memory. Keep both boundary tools local and none of the three CRM tools moves without new code. Record it as a scope reduction | | Accept the argument | In writing, what a wrong id costs for this tool: another person's tickets read out, or a note on another person's record | Conversation tests where the caller names another customer; for the write, the `write_note` step fixed to satisfy the schema and tested in a live run | The `check_tickets` sentence is rewritten, so check 1 of the proof goes red and the proof is updated with it | The first closure is the README's own rule. It is also, in this project, a decision about the whole swap rather than one tool. ## Where this guide and the tutorial disagree The tutorial's chapters are the engineer's reference for this swap. On these points the evidence above says something different: | Chapter | The tutorial says | What the evidence shows | | -------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | 3, What MCP takes away | Same name, keys and error string, "which is why the instructions still work" | The instruction text is unchanged, but `check_tickets/skill.md` lines 12–13 now describe a memory read the tool does not do | | 3, What MCP takes away | The tutorial "can afford the looser arrangement" for the ticket read, because it reads the caller's own record and the data is fixture data | Whose record it reads is exactly what the argument no longer guarantees, and fixture data is a property of the tutorial, not of the swap. A wrong id reads out another person's tickets, which is a decision for the reviewer | | 3, What MCP takes away | A missing remote tool is "the specific gap `make mcp-prove` closes": its check 5 calls `list_tools` "before you spend a training run" | Check 5 lists the tools of the proof's own loopback mock (`prove_mcp_swap.py`, lines 147–154 and 193–207), not the server the agent will use. It cannot show that a vendor's server offers the imported tools | | 4, What survives anyway | The `write_note` step "is otherwise unchanged", and "the caller still sees this before anything is written" | The step passes only `summary`, Rasa Pro's validator rejects that for the MCP tool ("Missing required argument(s): 'contact_id'"), and the check (`constraints.py`, lines 695–706) runs before the confirmation logic (line 708) | | 4, What survives anyway | The swap "preserves that interface exactly" | Two tools' argument lists changed: see the listing above | | 5, Shaping for the channel | Instruction prose and tool constraints are "portable across the swap"; the one thing not portable is "the tool" reading `project.contact_id` | Two tools read it, one of them a write. And the proof would pass a constraint moved to the wrong tool | Chapter 3 is still the right place to learn why an MCP tool cannot read memory, and its rule for a tool that fails the check is the one this guide uses. Its claim that the proof closes the start-time gap is the one this guide does not accept. ## Limits of this evidence - It shows where the customer id comes from in one companion project. It says nothing about how any vendor's MCP server behaves. - The boundary question is one question. It is not a security review of the vendor, the transport or the data. - Neither version checks that the caller owns the email they give. That is identity verification, a separate decision this guide does not cover. - The Rasa Pro behaviour described here comes from reading the source. No live Mantle run was made, so whether a model passes the right id is untested. - It gives no estimate of hours or points. The split tells you who reviews what, not how long it takes. ## Three questions an approver will hear :::solution{title="The engineer says the model always passes the id the lookup returned. Is that enough?"} It is a claim about model behaviour, and the companion project has no test files to check it against. The REST version did not depend on it: the tools read the id the lookup had written to memory, not one taken from the conversation. If the argument is that the model will behave, ask for the conversation tests that show it, including a caller who names another customer, and decide whether a sampled result is enough for a write. ::: :::solution{title="Can we move only the lookup to MCP and keep the other two tools local?"} Not as the companion is written. The only code that writes the customer id to memory is the REST lookup (`tools/crm.py`, line 43). The project's `memory.yml` declares `contact_id` without `llm_settable` (lines 8–10), and Rasa Pro 3.20.0rc1 defaults that setting to false (`rasa/mantle/memory/field.py`, line 52), so the model cannot set the id through `set_fields` either. A search of the whole project at `69e27b6` finds no other writer. Move the lookup to MCP and nothing writes the id: the local ticket and note tools find none and return `not_identified` (`skills/check_tickets/tools.py`, lines 12–14, and `skills/log_interaction/tools.py`, lines 16–18). Plan a partial swap as its own piece of work. ::: :::solution{title="Does any of this change with a real vendor server?"} The question stays the same, and the answers come from the vendor's schema rather than the mock's. Ask for the listing of required arguments for every tool the vendor publishes, and fill in the checklist from that. Any customer or account identifier in a required list is a value your agent will have to pass. ::: :::cta{href="/library/tutorials/crm-transport-swap/" label="Read the tutorial with these differences in mind"} The CRM transport swap tutorial shows the engineer's side: the diff, the proof script and the mechanism. Read chapters 3 to 5 alongside the table above. ::: --- # What evidence to demand before approving an agent action Source: https://rasa.community/library/guides/sign-off-an-agent-action-register/ Author: Rod Rivera Published: 2026-09-24 Earlier in the call the caller gave their passphrase and heard their balance. Then they ask for a replacement card. Here is an illustrative script of an agent built without a check inside the card tool. It is not a recorded call. The codes and the address come from the pattern's conversation tests (`tests/e2e/tiering.yml`, lines 84–92) and the reference from the card tool's fixture (`tools/banking.py`, line 195). > **Caller:** My card got eaten by a cash machine. Can you send a new one? > > **Assistant:** I can. A card can't be recalled once it's posted, so I need a > one-time code for that. I've sent one to your mobile. Can you read it to me? > > **Caller:** Zero zero zero zero. > > **Assistant:** That didn't match. You can try once more. > > **Caller:** Four four four four. > > **Assistant:** I'm sorry, I can't verify you that way today. Is there anything > else I can help with? > > **Caller:** Only the card, really. Can it go to 40 Somewhere Lane? > > **Assistant:** Of course. That's done, and your reference is RC-70413. Read it as the person who has to approve this agent. Nothing errored and no refusal line was skipped. The agent apologised, offered help, and the call ended with a reference number. The card went out on the passphrase tier the caller already held, for an action that needed the code they had just failed twice. The step-up pattern in `RasaHQ/rasa-community-resources`, `patterns/voice-auth-stepup`, was designed against this failure. Its module docstring describes it (`authpolicy/challenges.py`, lines 10–15, at revision `69e27b6`): ```text The failure this module is shaped around: a caller fails the OTP for a card reissue, the agent apologises, and then — still in the same skill, still holding only MEDIUM — completes the reissue anyway, because the prose said "if verification fails, offer to help another way" and the model read "help" as "do the thing they asked for". The action succeeded on auth that was never satisfied, and every log line says the call went fine. ``` "Every log line says the call went fine" is the reviewer's problem in one phrase. If sign-off rests on reading transcripts, this call passes. This guide assumes you already have an action register, as built in [Review what an agent is allowed to do](/library/guides/review-agent-permissions/), and says which artifact proves each control on it. The worked example is the pattern's demo bank, Northgate Bank (`agent.yml`, line 5). Northgate Bank is fictional, the tiers are the pattern's teaching choices, and this is an engineering review aid, not a compliance determination. ## The argument, and what it costs The evidence to demand is set by two facts about each action: the tier it is declared at, and where the check that enforces that tier runs. How the conversation reads is not one of them. A polite refusal in a transcript, or a YAML line in the skill, tells you nothing about whether an unclassified action defaults to safe, whether the check survives someone rewriting the skill, or whether a locked-out caller can keep trying. The cost is real. Engineering has to produce a test per risky row, show it failing when the guard is removed, and rerun it at the revision you release. Some rows will stay on hold; in the example below, one does. And a strict default turns a forgotten classification into customer friction rather than a hole. Someone has to own that friction. ## Risk follows the verb, not the caller The pattern declares tiers in one table keyed by action (`authpolicy/actions.py`, lines 46–89). Its docstring puts the reason plainly (lines 3–6): ```text This table is the pattern's central claim in executable form. Read down the `tier` column and notice what it is a function of: what the action *does* if it succeeds. Not who is asking, not which skill they entered through, not how far into the call they are. ``` A register that records "caller verified: yes" has one decision for the whole call, and the opening failure spends it. A register keyed by action can say that a balance and a card reissue, for the same caller in the same call, need different evidence. Here is the register for the six actions in the pattern. It keeps the Action, Boundary and Accountable owner columns of the permissions guide's register, replaces its "Required authority and data" column with the declared tier, and adds where the check is enforced. The owners are illustrative roles. | Action | Declared tier | Boundary to inspect | Where enforced (`tools/banking.py`) | Accountable owner | | ---------------------- | ---------------- | ------------------------------------------------------------ | -------------------------------------------------------------------------------------- | ----------------------- | | `get_store_hours` | LOW | Returns published information only | Nowhere, by design: the tool does not call the guard (lines 106–119) | Branch network owner | | `get_fee_schedule` | LOW | Returns published information only | Nowhere, by design (lines 122–125) | Pricing owner | | `get_balance` | MEDIUM | Balance disclosed only at MEDIUM or above | First statement of the tool (lines 136–139) | Accounts owner | | `get_recent_bill` | MEDIUM | Bill disclosed only at MEDIUM or above | First statement of the tool (lines 153–156) | Billing owner | | `reissue_card` | HIGH | Card posted only at HIGH; a refusal carries no dispatch data | First statement of the tool (lines 186–189); the YAML adds a confirmation, not a check | Card operations owner | | `transfer_funds` | HIGH | Money moved only at HIGH | First statement of the tool (lines 219–222) | Payments owner | | Any action with no row | HIGH, by default | Refused until someone classifies it | `tier_for` in `authpolicy/actions.py`, but only if the tool calls the guard | Owner of the tier table | The "Where enforced" column is the one to press on. "In the skill" or "in the YAML" is an answer that needs the next two sections, and "first statement of the tool" is a claim that needs a test. ## Ask what happens to an action nobody classified The likeliest mistake in a tier table is a missing row. Someone adds `close_account`, ships it, and forgets the table. The function the guard uses to look up a tier answers with the strictest tier (`authpolicy/actions.py`, lines 102–103): ```python policy = POLICIES.get(action) return policy.tier if policy is not None else AuthTier.HIGH ``` Its docstring gives the reason: "A registry that is permissive by omission is not a registry; it is an allowlist that anyone can join by accident" (lines 98–100). This is the "fail-safe defaults" principle from Saltzer and Schroeder's 1975 paper, [_The Protection of Information in Computer Systems_](https://doi.org/10.1109/PROC.1975.9939): base access decisions on permission, not exclusion. The evidence that it holds is a test, not the docstring (`tests/test_guard.py`, lines 259–262): ```python def test_unknown_action_defaults_to_high(self): """Forgetting to classify a new action must not open a hole.""" self.assertEqual(tier_for("close_account"), AuthTier.HIGH) self.assertEqual(tier_for(""), AuthTier.HIGH) ``` The same stance covers a call that arrives with no context at all. `test_missing_context_is_unauthenticated` (lines 278–281) asserts that `require_tier("reissue_card", None)` raises `StepUpRequired` rather than skipping the check, so a tool-discovery probe or a broken context cannot act as a bypass. A second test covers the rows that do exist. Rather than naming the risky tools by hand, it reads every HIGH row from the table and calls each tool at MEDIUM (lines 182–192): ```python high_actions = [name for name, p in POLICIES.items() if p.tier is AuthTier.HIGH] self.assertTrue(high_actions, "no HIGH actions declared — table is wrong") for action in high_actions: with self.subTest(action=action): tool_fn = getattr(banking, action) result = payload(run(tool_fn(context=caller_at(AuthTier.MEDIUM)))) self.assertFalse( result.get("ok"), f"{action} completed on medium auth — it is declared HIGH", ) ``` A new HIGH row joins that test the day it is added, with no one remembering to extend the suite. :::callout{type="pitfall" title="The default only runs for a tool that asks"} `tier_for` is called from inside `require_tier`, so a tool that never calls the guard never reaches the default. The two LOW tools skip the guard on purpose (`tools/banking.py`, lines 106–113). A new `close_account` added with no row and no guard call would run for anyone, and both tests above would still pass: it is not in the table, so the enumerated test never sees it. Ask engineering for an inventory of every registered tool, marked as guarded with a row, or unguarded with a written reason, and a check that the two lists agree. ::: ## A YAML constraint is not the check The skill that offers the card tool has frontmatter that looks like a control (`skills/card_services/skill.md`, lines 7–20): ```yaml import_tools: - reissue_card - transfer_funds tool_constraints: - reissue_card: requires_confirmation: enabled: true utter_for_confirmation: utter_confirm_reissue utter_on_user_denial: utter_action_cancelled - transfer_funds: requires_confirmation: enabled: true utter_for_confirmation: utter_confirm_transfer utter_on_user_denial: utter_action_cancelled ``` Read what it constrains. The pattern's `responses.yml` calls these "confirmations of INTENT", not of identity, and ends: "confirming is not verifying" (lines 3–6). In the pinned Rasa Pro 3.20.0rc1 wheel, the model resolves the pause itself, by calling `resolve_tool_confirmation` with `confirmed=true` (`rasa/mantle/orchestration/tool_execution/constraints.py`, lines 301–307). A confirmation prompt would not have changed the opening failure: the caller wanted the card and would have said yes. A `requires` condition that names authentication is a different kind of frontmatter, and the companion does not declare one on `reissue_card`. The pattern explains why it treats even that kind as the outer layer (`authpolicy/guard.py`, lines 15–25): ```text They are evaluated by the orchestrator against conversation state, which means they are instructions to a language model about which tool it may select. That is a routing control. It is not an execution control: it constrains what the model is *offered*, not what the process will *do* when the function is entered. This pattern therefore treats the frontmatter as the outer layer — it keeps the conversation coherent, so the caller gets asked for a passphrase instead of being told "no" — and puts the decision that actually binds inside the tool, on the line before the side effect. The prose can be edited, the model can be swapped, the constraint can be mistyped in YAML and silently ignored; the function still refuses. ``` The binding check is four lines at the top of the tool (`tools/banking.py`, lines 186–189): ```python try: require_tier("reissue_card", context) except StepUpRequired as exc: return _step_up(exc, context) ``` The path a call takes, and which parts decide anything: :::diagram{title="Where the card reissue is decided, and what writes the tier the guard reads"} ```dot rankdir=TB; subgraph cluster_route { label="Outside the function: routing"; style="dashed"; color="#5a17ee"; fontname="Helvetica"; fontsize=11; offer [label="Skill offers reissue_card;\nmodel chooses to call it"]; confirm [label="requires_confirmation pause;\nmodel resolves confirmed=true"]; } subgraph cluster_tool { label="Inside reissue_card: require_tier"; style="dashed"; color="#5a17ee"; fontname="Helvetica"; fontsize=11; required [label="required = tier_for(action)\n(no row: HIGH)"]; held [label="held = auth_tier in memory\n(no context or unknown value: none)"]; sat [label="held satisfies required?", shape=diamond]; } writers [label="Verification tools only:\ngrant on a pass, revoke on lockout"]; posted [label="Card posted,\nreference returned", class="ok"]; refused [label="ok: false, step_up_required,\nno dispatch data", class="blocked"]; offer -> confirm; confirm -> required [label="function entered"]; required -> sat; held -> sat; writers -> held [style=dashed, label="writes"]; sat -> posted [label="yes"]; sat -> refused [label="no"]; ``` ::: Everything in the routing box can be changed by editing prose or YAML, and the model makes both of its decisions. What decides is the comparison in the lower box. It reads the tier table and one memory value, `auth_tier`. An unrecognised value in that field counts as no authentication (`test_coerce_fails_closed`, lines 236–243). Only the verification tools write it (`tools/verification.py`, line 1). `memory.yml` marks none of the auth fields `llm_settable` (lines 10–12), and in the pinned wheel that setting defaults to false (`rasa/mantle/memory/field.py`, line 52), so the model cannot award itself a tier. The evidence for this row is a test that calls the tool directly, as a caller holding a correctly earned MEDIUM, and checks nothing happened (`test_medium_auth_cannot_reissue_a_card`, lines 118–129): ```python context = caller_at(AuthTier.MEDIUM) result = payload( run(reissue_card(delivery_address="12 Elsewhere Street", context=context)) ) self.assertFalse(result["ok"]) self.assertTrue(result["step_up_required"]) self.assertEqual(result["required_tier"], "high") self.assertEqual(result["held_tier"], "medium") # The action did not happen. These keys exist only on success. self.assertNotIn("dispatched", result) self.assertNotIn("reference", result) ``` A passing negative test proves little until you have seen it fail; the [mutation-testing guide](/library/guides/mutation-test-an-agent-guard/) takes that idea further. The pattern ships the exercise as `make prove`, which runs `scripts/prove_guard.py`. We ran that script on 21 September 2026 in a scratch export of the pattern at `69e27b6`, with the pattern's Rasa Pro 3.20.0rc1 environment. It is an offline proof. Its output, with the terminal colour codes stripped: ```text 1. removing the tier guard from reissue_card… 2. running the eval suite without it… red, as it must be. 3. guard restored. 4. re-running to prove the restore was clean… The guard is load-bearing: without it the suite fails; with it the suite passes. ``` The script does not say which tests failed, so we removed the same four lines by hand and ran the suite verbosely. Seven of the 36 tests failed, not the two that the test module's docstring names (lines 20–22): ```text FAIL: test_locked_out_caller_cannot_complete_the_high_action (test_guard.TestFailurePaths.test_locked_out_caller_cannot_complete_the_high_action) FAIL: test_every_high_tier_tool_refuses_medium (test_guard.TestHighTierCannotCompleteOnLowerAuth.test_every_high_tier_tool_refuses_medium) (action='reissue_card') FAIL: test_low_auth_cannot_reissue_a_card (test_guard.TestHighTierCannotCompleteOnLowerAuth.test_low_auth_cannot_reissue_a_card) FAIL: test_medium_auth_cannot_reissue_a_card (test_guard.TestHighTierCannotCompleteOnLowerAuth.test_medium_auth_cannot_reissue_a_card) FAIL: test_refusal_carries_no_dispatch_data (test_guard.TestHighTierCannotCompleteOnLowerAuth.test_refusal_carries_no_dispatch_data) FAIL: test_unauthenticated_caller_cannot_reissue_a_card (test_guard.TestHighTierCannotCompleteOnLowerAuth.test_unauthenticated_caller_cannot_reissue_a_card) FAIL: test_medium_caller_attempting_high_is_stepped_up (test_guard.TestStepUpIsRequiredNotRemembered.test_medium_caller_attempting_high_is_stepped_up) ``` Then we did the same to the two MEDIUM tools, removing the four guard lines from both `get_balance` and `get_recent_bill`. The suite's last lines: ```text ---------------------------------------------------------------------- Ran 36 tests in 1.789s OK ``` Nothing failed. The guards are in the code, but no test depends on them. `test_allows_when_tier_is_sufficient` and `test_lockout_revokes_the_tier` (lines 290–292 and 394–400) exercise `require_tier` for `get_balance` directly, never the tool. So the MEDIUM rows have a guard and no evidence for it. The missing test is small: call each MEDIUM tool directly at NONE and at LOW and assert it returns `ok: false` with no balance or bill. :::checkpoint{id="action-register-transcript-evidence" question="The caller passed the passphrase. card_services declares a tool_constraints entry for reissue_card. The transcript ends with a reference number and shows no error. Which of these is evidence that the card was posted on sufficient authentication?" options="The transcript; it ends with a reference and no error,The tool_constraints entry; it is declared for reissue_card,A direct tool test at MEDIUM that fails when the guard is deleted,The passphrase; it verified the caller earlier in the call" answer="2"} Only the direct tool test says anything about the process, and only because it has been seen to fail without the guard. The transcript shows one conversation the model chose to have, and the opening failure produces a clean one. The `tool_constraints` entry for `reissue_card` is a confirmation of intent, which the model resolves. The passphrase earns MEDIUM, and the table declares the card HIGH. ::: ## Lockout has to end the call's authority A challenge attempt has three outcomes in the pattern: passed, retry and locked out. A failed attempt is one of the last two. The code comment on the last one reads "Budget exhausted. Terminal — the only exit is a human." (`authpolicy/challenges.py`, lines 64–72). The evidence that "terminal" means something comes in two parts. The first is that running out of attempts produces a lockout that grants nothing and hands off. `test_wrong_passphrase_retries_then_locks_out` (`tests/test_guard.py`, lines 352–359) gives one retry, then `LOCKED_OUT` with `handoff` true. The same holds for the one-time code (lines 370–374): ```python def test_lockout_requires_handoff(self): result = check_otp("nope", attempts_used=RETRY_BUDGET) self.assertIs(result.outcome, Outcome.LOCKED_OUT) self.assertTrue(result.handoff) self.assertIsNone(result.granted) ``` The second is that the risky action still refuses after the lockout. This is the opening failure asserted directly (lines 411–415): ```python context = caller_at(AuthTier.MEDIUM, otp_attempts=float(RETRY_BUDGET)) revoke(context) result = payload(run(reissue_card(context=context))) self.assertFalse(result["ok"]) self.assertNotIn("reference", result) ``` Notice what these tests do not do. The tests that drive a challenge into lockout send only wrong answers (lines 352–374), and no test sends a correct factor after a lockout. The last one calls `revoke` itself rather than driving a real verification tool into lockout, and it checks one action. The gaps are in exactly those places. The shared attempt check compares the answer before it looks at the budget (`authpolicy/challenges.py`, lines 124–153), so a correct factor passes at any attempt count. In the tool layer, lockout sets a `locked_out` flag and drops the tier to none (`tools/verification.py`, lines 54–63). Only `verify_one_time_code` reads that flag before it grants anything (lines 153–155). `verify_passphrase` does not read it, and neither does the guard. We checked what that means with two offline probes of our own on 21 September 2026: short scripts, each run once against the pattern's own tools with a fake context. Neither is a companion test or part of the pattern's suite. The first sends two wrong codes, calls `reissue_card` and `get_balance`, then sends the correct passphrase and calls `get_balance` and `reissue_card` again: ```text otp: retry | auth_tier = medium | locked_out = None otp: locked_out | auth_tier = none | locked_out = True reissue_card ok: False get_balance ok: False passphrase: passed | auth_tier = medium | locked_out = True get_balance ok: True reissue_card ok: False ``` The second locks out the passphrase itself, then sends the correct one and calls `get_balance`: ```text wrong passphrase 1: retry | auth_tier = none | locked_out = None wrong passphrase 2: locked_out | auth_tier = none | locked_out = True wrong passphrase 3: locked_out | auth_tier = none | locked_out = True wrong passphrase 4: locked_out | auth_tier = none | locked_out = True wrong passphrase 5: locked_out | auth_tier = none | locked_out = True wrong passphrase 6: locked_out | auth_tier = none | locked_out = True correct passphrase: passed | auth_tier = medium | locked_out = True get_balance ok: True ``` The card stays refused in both, which is the guard doing its job. But after any lockout, including the passphrase's own, a correct passphrase restores MEDIUM and the balance tool discloses again. The module docstring says exhausting the budget "cannot produce an authenticated state" (`authpolicy/challenges.py`, lines 19–21), and the README promises "no path back to success" (`README.md`, line 40). Both hold for the wrong answers the tests send, not for a correct one after a lockout. What stands in the way in a real call is a set of instructions to the model: - the lockout result's hint, "Do NOT retry and do NOT complete the request" (`tools/verification.py`, lines 71–75); - the step-up skill's `if: session.project.locked_out` block and its rule not to "switch to the other factor" (`skills/step_up/skill.md`, lines 28–31 and 47–49); - the account skill's instruction not to call the account tools when locked out (`skills/account_info/skill.md`, lines 20–22). All of them are the routing layer again. None is a check in code. The pattern also names a limit of its own: the budget is per call, held in session memory, and "An attacker who hangs up and redials gets a fresh budget" (`README.md`, lines 162–165). For a reviewer the lesson is general. Ask for a lockout test that goes through the real tool that locks out, then sends the correct factor on every path that can grant a tier. And ask where the lockout is recorded, because a lockout held in one call's memory ends when the call does. ## Check the payload, not the logging policy On a voice channel the caller says the secret out loud. The pattern's docstring says every string that could be a factor passes through one function, `redact`, before it reaches a log, a tool result or a tracker event (`authpolicy/guard.py`, lines 150–151). It returns only a length (lines 157–159), and `test_redact_never_returns_the_secret` checks that the input never comes back (`tests/test_guard.py`, lines 428–430). The test that looks like the payload check is `test_refusal_payload_contains_no_factor_value` (lines 435–443). It calls `reissue_card` at MEDIUM and asserts that neither fixture factor appears in the result. Read the refusal it inspects: every field is built from the action name, the tiers and constants (`tools/banking.py`, lines 88–102). There is no input from which a factor could leak, so the test cannot fail for this tool. It passed even with the guard deleted. The tools that do receive the spoken factor, `verify_passphrase` and `verify_one_time_code`, are never called in the test module at all. Ask for a payload test on those two tools, across pass, retry and lockout. A refusal is also not empty. It tells the model the required and held tiers, the factor to ask for and the reason for the tier. What `test_refusal_carries_no_dispatch_data` (lines 159–171) proves is narrower: no reference, dispatch flag, delivery estimate or address. :::callout{type="warn" title="The transcript is the leak the tests do not cover"} The pattern says so itself: "an ASR transcript of the turn where the caller said their passphrase is a plaintext credential sitting in the conversation history, and it will be replayed by anyone debugging the call" (`authpolicy/guard.py`, lines 152–155). Nothing in `tests/test_guard.py` captures log output or touches transcript storage. Treat retention, access and masking of the transcript as a separate row in the register with its own owner. ::: ## Run the suite at the revision you approve Clone `RasaHQ/rasa-community-resources`, check out `69e27b6a50c4700f95f2c13d4609d0a3f8cba7d2`, change into `patterns/voice-auth-stepup` and run `make test`. The target runs the suite through `uv run` (`Makefile`, lines 13–14 and 52–53), so you need `uv`. The project pins `rasa-pro==3.20.0rc1` (`pyproject.toml`, lines 7–9), and the tool module imports `rasa`, so a bare `python -m unittest` without that install fails. We ran the same unittest command with the pattern's installed environment on 21 September 2026. It printed `Ran 36 tests in 4.717s` and `OK`. The tests and probes this guide ran are offline checks of the code: none is a live Mantle run, and none says how a model behaves in a call. The pattern's conversation tests need a trained model and keys (`Makefile`, line 32), and we did not run them. ## What evidence to demand before approving an agent action The permissions guide lists the fields of a sign-off: allowed scope, source of authority, evidence inspected, unresolved cases, disposition, owner and review date. With tiers and enforcement locations in the register, each field can now point at something checkable. An illustrative record for two rows, written against the evidence above: | Field | `reissue_card` | `get_balance` | | ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Allowed scope | Post a replacement card to a caller holding HIGH | Disclose one balance to a caller holding MEDIUM or above | | Source of authority | One-time code: `check_otp` grants HIGH (`challenges.py`, lines 113–121; `test_otp_grants_high`). Its docstring warns that a code read aloud is closer to a second knowledge factor (lines 116–119) | Passphrase: `check_passphrase` grants MEDIUM, never HIGH (`challenges.py`, lines 103–110; `test_passphrase_ceiling_is_medium`) | | Where enforced | `require_tier` on the tool's first statement | `require_tier` on the tool's first statement | | Evidence inspected | `test_medium_auth_cannot_reissue_a_card`, `test_every_high_tier_tool_refuses_medium`, `test_locked_out_caller_cannot_complete_the_high_action`; `prove_guard.py` red then green; seven tests fail with the guard removed; `python -m unittest discover -s tests -v` (the `make test` command) OK in the pattern's environment, 36 tests, at `69e27b6` | None that depends on the guard: with it deleted, all 36 tests pass | | Unresolved | No tool inventory; no payload test on the verification tools; the one-time code is read aloud on the same call; the retry budget resets on a redial | No direct test at NONE or LOW; after any lockout, including the passphrase's own, a correct passphrase regrants MEDIUM; the budget resets on a redial; no log-capture test | | Disposition | Permit at HIGH, with the tool inventory and payload test due before the next release | Hold until the guard has a direct test and the lockout holds for a correct passphrase, with a test | | Owner | Card operations owner | Accounts owner | | Review date | Next change to `tools/banking.py` or the tier table | Next change to `tools/banking.py` or `tools/verification.py` | The dispositions are narrower than "approved". That is the purpose. The card row, the one with the irreversible side effect, is close to approvable because its evidence is direct and has been seen to fail. The balance row, lower in tier, is held: its guard has no test behind it, and its lockout has a gap. Judged by how their calls read, neither row would have raised a question. ## Questions reviewers ask :::solution{title="Engineering shows me the conversation tests passing. Is that evidence?"} It is evidence of something else. The pattern's conversation tests in `tests/e2e/tiering.yml` say what they are: "Each case asserts that the agent behaved correctly on one sampled run. A pass means the model chose well this time; it does not mean the model cannot choose otherwise" (lines 6–8). The same header says what they are for: whether "the caller is asked for the right factor, told the truth about what happened, and not challenged for public information" (lines 13–16). Accept them for the conversation design and ask for the direct tool tests for the row. ::: :::solution{title="A new action is urgent and has no tier yet. Can I approve it provisionally?"} With a strict default and a guard call in the tool, the provisional state already exists: the action is refused at HIGH until someone adds a row. That is a legitimate interim disposition, at the cost of callers being stepped up or handed off. What you should not approve is a provisional row at a lower tier with no test, or a tool that skips the guard "for now". Check the tool inventory before you accept that the default applies. ::: :::cta{href="/library/tutorials/guarding-irreversible-actions/" label="See how the guard is built, step by step"} The card reissue tutorial builds the risk declaration, the execution guard and the refusal paths that this review asks engineering to prove. It uses a separate companion project, `tutorials/rasa-card-reissue-tutorial`. ::: --- # After login, a stale Rasa skill line asked for the PIN again Source: https://rasa.community/library/guides/skill-preconditions-replace-requires/ Author: Rod Rivera Published: 2026-09-28 A traveller gives Atlas their PIN, the PIN is accepted, and two turns later Atlas asks for it again. Atlas, Horizon Travel and the traveller, Maya, are the companion repository's fictional demo. This is a live run on Rasa Pro 3.21.0.dev3 with OpenAI gpt-5.2, recorded on 29 September 2026 (receipt e14, precondition copy, run 1): ```text user: Hi, I would like to verify my identity first. My PIN is 4242. tool: activate {"target_id": "intro"} -> activate: done bot: Hi Maya. I’m Atlas, the Horizon Travel voice assistant. I can help you check your itinerary, get flight status, change or cancel a booking, report lost baggage, or answer common travel FAQs. What would you like to do today? tool: activate {"target_id": "authenticate"} -> activate: done tool: verify_traveler_pin {"pin": "4242"} -> {"ok": true, "authenticated": true, "display_name": "Maya Chen", "hint": "PIN accepted."} bot: I was starting to introduce what I can help with at Horizon Travel. Do you want me to continue, or should I stop? user: Skip it, thanks. bot: Is there anything else you need help with? user: I want to change the date of my Lisbon booking. tool: activate {"target_id": "change_booking"} -> activate: done tool: activate {"target_id": "authenticate"} -> activate: done bot: Sure, I can help change your Lisbon booking. Before I make any changes, please tell me your four-digit Horizon Travel PIN. ``` All user and bot events from the saved tracker, line breaks indented; the `tool` lines are a selection: the PIN check and every `activate` call. `change_booking` used to be gated by the skill-level line `requires: session.project.authenticated`. Rasa Pro 3.21.0.dev3 refuses it, and the error suggests a `precondition`, which we bound in `agent.yml`. The skill's instructions also included, at line 19, "First verify identity: @skill.authenticate", and we left that alone, as a migration that follows the error message would. Ten logged-in runs with the line, and ten with only that line deleted: | After login, "I want to change the date of my Lisbon booking." | Line 19 kept: "First verify identity: @skill.authenticate" | Line 19 deleted, nothing else changed | | --------------------------------------------------------------- | --------------------------------------------------------- | ------------------------------------- | | Atlas asked for the PIN again | 2 of 10 | 0 of 10 | | Model called `authenticate` again | 8 of 10 | 0 of 10 | | `change_booking` activated | 10 of 10 | 10 of 10 | | Next skill the model called after `change_booking` | `authenticate` (8), `find_booking` (2) | `find_booking` (10) | **The guard was written twice, and the migration deleted only one copy.** The precondition is the engine's copy. `authenticated` was already true, so the engine activated `change_booking` straight away in all twenty runs. Line 19 is the prompt's copy, and the model kept obeying it after the engine had decided. In run 4 the extra call went further: the tracker records `skill_cancelled` for `change_booking`, its removal from the stack and `skill_resumed` for `authenticate`, and the scripted run ended on the next reply, before any effect on the traveller could be seen. The mechanism is not specific to Rasa: a guard written once in code or configuration and again in the prompt stays duplicated until someone deletes the prompt's copy, and moving the code copy, as this migration did, leaves the prompt's copy exactly where it was. The other fix the error offers, `hidden: true`, is not a login gate at all: the dev3 source drops a hidden skill from routing with no condition, and in three logged-in runs the request never reached it. These are ten runs per precondition copy and three on the hidden copy, on one model at temperature 0.0, one agent and one prerelease, and we did not run the old gate after login. ## After login, the model followed the prompt's copy Each run is a scripted conversation. After login, the three turns in the exchange above. Before login, four turns: "I want to change the date of my existing booking, please.", "The Lisbon one.", "4242" and "Move it to next Friday, please.". These are the receipt's lines for three of the twenty logged-in precondition runs, and its totals (receipt e14; its first three line-19 runs are the same trackers as receipt e13): ```text s3 pre run 1: authenticated_before_turn3=True activate=['change_booking', 'authenticate'] PIN_asked=True bot=['Sure, I can help change your Lisbon booking. Before I make any changes, please tell me your four-digit Horizon Travel PIN.'] s3 pre run 4: authenticated_before_turn3=True activate=['change_booking', 'authenticate'] PIN_asked=False bot=['Got it. I can help change the date on your Lisbon booking. You’re already verified. What new date do you want for the Lisbon trip?'] s3 pre19 run 1: authenticated_before_turn3=True activate=['change_booking', 'find_booking'] PIN_asked=False bot=['Got it. You mean your **Lisbon spring break** booking, reference **H T one two three four five**. What new departure date would you like, a'] --- totals, scenario 3 (10 runs each): change_booking activated / authenticate activated again / PIN asked again --- pre: 10/10 8/10 2/10 pre19: 10/10 0/10 0/10 ``` The runs were not identical at temperature 0.0, the value the companion's `integrations.yml` sets: with line 19, the model went to `authenticate` in eight runs and to `find_booking` in two, and asked for the PIN in two. In seven of the eight runs, the model's call started `authenticate` afresh on top of `change_booking`. In run 4 the call resumed the `authenticate` frame left over from the login turn. The tracker records `skill_cancelled` for `change_booking`, `remove /2` on the stack and `skill_resumed` for `authenticate`. The reply said "You’re already verified." and asked "What new date do you want for the Lisbon trip?", and the scripted run ended there, before any effect on the traveller could be seen. The fix is one deleted line. It is the only difference the recursive diff in receipt e15 found between the two precondition copies. That diff excludes `.venv`, `models`, `.rasa`, `__pycache__`, `.env` and `*.db` files: ```diff --- w5/pre/skills/change_booking/skill.md 2026-09-29 11:44:54 +++ w5/pre19/skills/change_booking/skill.md 2026-09-29 12:54:26 @@ -16,8 +16,6 @@ Help the traveler change or cancel a booking. -First verify identity: @skill.authenticate - Then find the booking: @skill.find_booking Ask what they want to change. ``` To find the prompt's copy of a guard in your own project, search the skills for the resolver's name: ::run{cmd="git grep -n '@skill.authenticate' -- skills"} In the companion at `01d6a6e` (22 Sep) there is one hit: ```text $ git grep -n '@skill.authenticate' 01d6a6e -- examples/mantle-voice-agent/skills 01d6a6e:examples/mantle-voice-agent/skills/change_booking/skill.md:19:First verify identity: @skill.authenticate ``` Deleting the line did not weaken the gate. Before login, on the copy without it, the engine parked `change_booking` behind `authenticate` and activated it after the PIN in 3 of 3 runs. The gate covers `change_booking` and nothing else, though. In all six precondition runs before login, with and without line 19, `find_booking` listed the traveller's bookings before any PIN, and in three of them Atlas read out a trip name or the booking reference first (receipts e13 and e14). The copies differ from companion revision `01d6a6e` only in the gate and its `agent.yml` binding, the rasa-pro pin, the way credentials are written, one test-harness edit that lets the served agent find the demo data, line 19 where stated, and, in the hidden copy, the two `flight_status` mentions rewritten as plain text (receipts e13 and e15). The model is gpt-5.2, the one the companion ships with (`integrations.yml` at `01d6a6e`). ## Skill-level requires is no longer supported on 3.21.0.dev3 The new gate is not optional. Leave the dev1 line in, and the project no longer validates on dev3. These are the last three lines of a 1,311-line capture that includes a traceback (receipt e1): ```text rasa.exceptions.ValidationError: Project validation found 2 problem(s): - [mantle.validation.skill.failed_to_load] Skill directory 'skills/change_booking' failed to load and would be silently skipped: Skill-level requires is no longer supported. Use precondition to park the skill until a resolver runs, or hidden: true to omit it from routing. Tool-level requires under tool_constraints is unchanged.. Fix the skill so it parses, or remove it. - [skill.referenced.invalid_target] Skill 'flight_status' references external call/link target 'change_booking', which does not resolve to a skill or flow in the catalog. ``` The check in the dev3 wheel tests only whether the key is present, so no value of a skill-level `requires:` gets through (`rasa/mantle/skills/markdown_skill_compiler.py`, lines 925 to 937, in the rasa-pro 3.21.0.dev3 wheel with sha256 `8b7a1977cc421428757789384356930ce81defbb8779b4789a91adba11843495`): ```python def _reject_skill_level_requires(frontmatter: Dict[str, Any]) -> None: """Fail closed when skill.md still declares a skill-level ``requires``.""" if REQUIRES_FIELD not in frontmatter: return raise ValidationError( code=_SKILL_LEVEL_REQUIRES_REMOVED_CODE, event_info=( "Skill-level requires is no longer supported. Use precondition to " "park the skill until a resolver runs, or hidden: true to omit it " "from routing. Tool-level requires under tool_constraints is " "unchanged." ), ) ``` The message offers two replacements. We tried both, and a precondition without its `agent.yml` binding. One of the four variants passes, and two of the three failures are reported against `flight_status`, a skill none of the edits touched: | Variant | What `change_booking` and `agent.yml` declare | Offline validation result | | ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------- | | 1. Old gate left in | `requires: session.project.authenticated`; no preconditions block | Exit 1, two problems: `mantle.validation.skill.failed_to_load` on `change_booking`, `skill.referenced.invalid_target` on `flight_status` | | 2. Hidden | `hidden: true`; no preconditions block | Exit 1: `skill.prose.hidden_skill_reference` on `flight_status` | | 3. Precondition, unbound | `precondition: authenticated`; no `orchestrator.preconditions` entry | Exit 1: `mantle.validation.precondition.missing_binding` on `change_booking` | | 4. Precondition, bound | `precondition: authenticated`; `agent.yml` binds `authenticated` with `satisfied_when: session.project.authenticated` and `resolve_with: authenticate` | Exit 0: `validate_project: ok` | Variant 1 shows how one leftover key produces two errors. The second is reported against `flight_status`, because `change_booking` never loaded and the mention in `flight_status` points at a target that is not in the catalog. Variants 2 and 3 have no `invalid_target` problem, since `change_booking` loads in both. Fix the skill that fails to load first; the reference problem downstream is a consequence, not a second fault. Each variant is a copy of `examples/mantle-voice-agent` on rasa-pro 3.21.0.dev3, whose skills match `RasaHQ/rasa-community-resources` at companion revision `01d6a6e` (22 Sep) apart from the edits in the table (receipt e9). Horizon Travel, Atlas and the traveller's data are the companion's fictional demo. Run this from inside a copy: ```bash uv run --locked python -c "from pathlib import Path; from rasa.mantle.validation import validate_project; validate_project(Path('.')); print('validate_project: ok')" ``` That command, run in each of the four copies, reproduces all four results (receipt e7). The first captures of each variant recorded the same command with an extra `--project` flag. Validation is offline and makes no model calls. A pass means the files parse and the names resolve. It says nothing about a conversation, which is why the live runs exist. Variant 1's first problem says the skill "would be silently skipped". Neither path we tried was silent. `rasa train` on the variant 1 copy logs `mantle.skill_catalog.skill_load_failed` at ERROR, reports the same two problems, exits 1 and leaves `models/` empty (receipt e5). On dev3 the old line stopped the build instead of shipping an agent without `change_booking`. We did not test serving a model trained on an older build. ## hidden: true fails on a mention in another skill `hidden: true` is the shorter edit, and "omit it from routing" sounds like what the old line did. Put it in place of `requires:`, and validation fails again, on a skill nobody edited. This is the output as saved, below its header line (receipt e2): ```text raise ValidationError( rasa.exceptions.ValidationError: Project validation found 1 problem(s): - [skill.prose.hidden_skill_reference] Skill 'flight_status' references hidden skill 'change_booking' via '@skill.change_booking' in prose. Hidden skills cannot be started from routing; use a deterministic call: or link: step, or bind the skill as resolve_with in agent.yml orchestrator.preconditions. ``` `flight_status` reports delays. When a flight is delayed or cancelled, its instructions offer a booking change by naming the other skill in plain prose (`skills/flight_status/skill.md` at `01d6a6e`, lines 19 to 25): ```markdown if: session.flight_status.flight_status == "delayed" Tell them the delay in minutes and the gate if available. Offer to help with a booking change via @skill.change_booking. if: session.flight_status.flight_status == "cancelled" Apologize briefly. Explain that rebooking options are available and offer @skill.change_booking or @skill.human_handoff. ``` Those are the only two mentions in the project's skills: ```text $ git grep -n "@skill.change_booking" 01d6a6e -- examples/mantle-voice-agent/skills 01d6a6e:examples/mantle-voice-agent/skills/flight_status/skill.md:21:a booking change via @skill.change_booking. 01d6a6e:examples/mantle-voice-agent/skills/flight_status/skill.md:25:@skill.change_booking or @skill.human_handoff. ``` The error gives the reason: "Hidden skills cannot be started from routing", so a prose mention of one points at a skill the router can never start. It also names three ways out: a `call:` step, a `link:` step, or binding the hidden skill as `resolve_with`. We tested none of the three. There is a second reason not to use `hidden: true`, and it holds even where nothing mentions the skill. In the dev3 source, the function that decides which skills the model may activate drops a hidden skill with no condition attached (shown in the section on the dialogue stack below). The old line kept `change_booking` off that list only while `session.project.authenticated` was false. `hidden: true` keeps it off for every traveller. Our hidden copy, with the two mentions rewritten as plain text so it validates, confirms it: in three runs a traveller who had just given the right PIN asked to change a booking, and `change_booking` was never activated. The request went to `find_booking` each time (receipt e13). ## A precondition validates only once agent.yml binds it The other replacement is `precondition:`. Write it the way the old line was written and nothing else, and validation fails again. The output as saved, below its header line (receipt e3): ```text raise ValidationError( rasa.exceptions.ValidationError: Project validation found 1 problem(s): - [mantle.validation.precondition.missing_binding] Skill 'change_booking' declares precondition 'authenticated', but agent.yml has no matching entry under 'orchestrator.preconditions'. ``` The frontmatter value is now a name, not an expression. This is `skills/change_booking/skill.md` in the precondition copies used for the live runs, lines 1 to 15. Line 6 is the only line that differs from the companion at `01d6a6e`: ```yaml --- name: change_booking description: > Change or cancel an existing Horizon Travel booking. Activate for date changes, cancellations, or "I need to change my trip". precondition: authenticated tool_constraints: - cancel_booking: requires: session.project.selected_booking_ref requires_confirmation: enabled: true utter_for_confirmation: utter_confirm_cancel_booking utter_on_user_denial: utter_cancel_aborted on_success: utter_booking_cancelled --- ``` The condition moves to `agent.yml`, under a top-level `orchestrator:` key. This block is from the precondition copies used for the live runs (receipts e6 and e13), and variant 4 declares the same binding: :::annotated{title="agent.yml (precondition copy)"} ```yaml orchestrator: preconditions: authenticated: # (1) satisfied_when: session.project.authenticated # (2) resolve_with: authenticate # (3) ``` 1. The key must match the name on the skill's `precondition:` line. Leave the entry out and you get the `missing_binding` error above. 2. The old `requires:` expression, unchanged. `authenticated` is project-wide memory (`memory.yml` at `01d6a6e`, line 1 and lines 14 to 17), and the `verify_traveler_pin` tool sets it when the PIN matches (`skills/authenticate/tools.py`, line 31). 3. The skill the engine starts when the condition is false. In the dev3 source it goes on the stack as a `call` frame above the parked skill (`rasa/mantle/orchestration/skill_executor.py`, lines 955 to 963). ::: With the binding in place, validation prints one line, `validate_project: ok` (receipts e4 and e7). ## The gate became a dialogue step, and the tracker shows it What changed is where the check runs. On dev1 it runs while the engine builds the list of skills the model may activate. A skill whose `requires` is false is left off (`rasa/mantle/skills/catalog.py`, lines 862 to 870, in the rasa-pro 3.21.0.dev1 wheel with sha256 `26d510fb6f2f3d09271f513e1408112db233dbf5ee81186a4d36a44a918b651a`): ```python def _is_excluded_from_routable_entries( skill: Skill, skill_id: str, memory_values: dict[str, Any], ) -> bool: """Return whether *skill* must be omitted from `SkillCatalog.routable_entries`.""" if skill.disabled: return True return not _skill_condition_satisfied(skill, skill_id, memory_values) ``` The same function in the dev3 wheel no longer reads memory at all (lines 853 to 855): ```python def _is_excluded_from_routable_entries(skill: Skill) -> bool: """Return whether *skill* must be omitted from `SkillCatalog.routable_entries`.""" return skill.disabled or skill.hidden ``` So on dev3 `change_booking` stays on the list, and the check moves to the moment a skill is activated (`rasa/mantle/orchestration/skill_executor.py`). The working-agent runs show all three branches that matter here: 1. If the skill's precondition holds, it activates as normal (line 903). After login, every precondition run pushed `change_booking` as a regular frame. 2. If not, the skill is pushed as a `pending` frame (lines 927 to 935, with the frame type set at line 988), and the resolver named in `resolve_with` is activated above it as a `call` frame (lines 955 to 963). Before login, every precondition run did this. 3. When the resolver completes, the engine checks the condition again. If it holds, the pending skill is promoted to an active run (lines 1079 to 1082). If not, the pending skill is dropped (line 1084). Promotion happened before login in two of the three line-19 runs and all three runs without it. No run reached the drop. The code is below, verbatim, for anyone who wants to check the line numbers. :::solution{title="Show the dev3 activation code, skill_executor.py lines 903 to 966"} ```python if precondition_satisfied(binding, memory_values): return self._activate( flow_id, tracker_handle, frame_type=frame_type, new_frame_type=new_frame_type, skill_id=skill_id, block_id=block_id, reset_memory=reset_memory, ) # If the resolver is already the live skill, park the consumer under # that run. A second resolver would interrupt the first, reset its # memory, and later offer continue_interrupted for work already done. if self._park_pending_under_live_resolver( flow_id, tracker_handle, resolver_skill_id=binding.resolve_with, new_frame_type=new_frame_type, skill_id=skill_id or owning_skill.id, block_id=block_id, ): return True if not self._push_pending_target( flow_id, tracker_handle, frame_type=frame_type, new_frame_type=new_frame_type, skill_id=skill_id or owning_skill.id, block_id=block_id, ): return False resolver_skill = self._skill_catalog.skill_by_id(binding.resolve_with) if resolver_skill is None: structlogger.error( "mantle.skill_executor.precondition.missing_resolver", resolver_skill_id=binding.resolve_with, ) self._pop_pending_target(tracker_handle) return False resolver_flow_id = entry_target_id_for_skill(resolver_skill) if resolver_flow_id is None: structlogger.error( "mantle.skill_executor.precondition.resolver_no_entry", resolver_skill_id=binding.resolve_with, ) self._pop_pending_target(tracker_handle) return False if not self._activate( resolver_flow_id, tracker_handle, frame_type=FlowStackFrameType.REGULAR, new_frame_type=FlowStackFrameType.CALL, skill_id=resolver_skill.id, block_id=self._skill_catalog.ordered_block_id_for_target(resolver_flow_id), reset_memory=True, ): self._pop_pending_target(tracker_handle) return False return True ``` ::: This `jq` filter prints every change to the dialogue stack in the saved tracker of one line-19 run before login. Run it from `editorial/receipts/skill-preconditions-replace-requires/` in this site's repository: ```text $ jq -r '.events[] | select(.event == "stack") | .update | fromjson[] | [.op, .path, (.value // "" | if type == "object" then "\(.skill_id) \(.frame_type)" else tostring end)] | join(" ")' e13-tracker-s1-pre-run3.json add /0 default_session_start regular add /1 default_session_start regular remove /0 remove /0 add /0 change_booking pending add /1 authenticate call add /1/previous_frame_type call replace /1/frame_type interrupt add /2 find_booking regular remove /2 remove /1/previous_frame_type replace /1/frame_type call remove /1 replace /0/frame_type regular ``` `change_booking` goes on as `pending` with `authenticate` above it as a `call`. "The Lisbon one." interrupts `authenticate` for `find_booking`, and "4242" brings `authenticate` back. Once the PIN is accepted, `authenticate` is removed (`remove /1`) and `change_booking` becomes a regular frame: that last line is the promotion. While `authenticate` was on top of the stack, the model's first question in every line-19 run before login was about the booking, not the PIN; the scripted "4242" on turn 3 is what `authenticate` took as the PIN. The two copies of the guard sit on different paths. The engine's copy runs before the skill starts. The prompt's copy sits inside the skill: at run time the engine does not act on it, and the model does, by calling `activate` for `authenticate`, as the receipt's `activate` lines show: :::diagram{title="Two copies of one login guard"} ```dot rankdir=TB; act [label="Model: activate change_booking"]; check [label="Engine: is session.project.authenticated true?", class="mark-1"]; park [label="Push change_booking as pending,\nactivate authenticate as call", class="mark-2"]; promote [label="authenticate completes, condition true:\nchange_booking promoted", class="ok mark-3"]; run [label="change_booking runs"]; prose [label="Skill prose line 19:\n\"First verify identity: @skill.authenticate\"", class="mark-4"]; again [label="Model: activate authenticate again", class="blocked mark-5"]; act -> check; check -> park [label="no"]; park -> promote; promote -> run; check -> run [label="yes"]; run -> prose; prose -> again; ``` 1. The engine's copy: `agent.yml` binds `authenticated` to `satisfied_when: session.project.authenticated`, checked at activation (`skill_executor.py` line 903). 2. `skill_executor.py` lines 927 to 935 and 955 to 963. 3. `skill_executor.py` lines 1079 to 1082. 4. The prompt's copy: `skills/change_booking/skill.md` line 19 at `01d6a6e`, which the migration left in place. 5. The model acts on the prompt's copy even when the engine's copy has already passed. In e14, 8 of 10 logged-in runs with line 19 took this edge, and 0 of 10 without it. ::: Three things to review after the migration, each resting on a different kind of evidence: | What to review | What the evidence says | Kind of evidence | | -------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------- | | Prompt lines that enforce a gate | Line 19 re-activated `authenticate` after login in 8 of 10 runs; deleting it took that to 0 of 10. | Live runs, ten per copy | | A failed PIN | `authenticate` allows one retry, then offers `@skill.human_handoff` (`skills/authenticate/skill.md`, line 23). If `authenticate` completes without setting `authenticated`, the parked skill is dropped and, with `unmet_utterance` at its default of true, the engine sends `utter_precondition_unmet`, whose default text is "I'm sorry, I wasn't able to continue with that request." | Source reading (`preconditions.py`, lines 17 to 32; `skill_executor.py`, lines 1126 to 1148); no run reached the drop | | The guard on `cancel_booking` | The tool still declares its own tool-level `requires:` and a confirmation step (`skills/change_booking/skill.md`, lines 7 to 14). The engine message says tool-level `requires` is unchanged. | Companion file and engine message | Our position on the last row: treat the precondition as deciding when a dialogue starts, and keep the check that protects the booking on the tool. :::checkpoint{id="skill-precondition-prompt-copy" question="A hypothetical refund skill declares precondition: card_verified bound in agent.yml. Its instructions still open with a line telling the model to verify the card through another skill. A traveller whose card is already verified asks for a refund. What do these runs suggest you check?" options="Nothing because the engine has already checked card_verified,Whether the model runs the verification skill again and whether that line should go,Whether hidden: true would be a safer gate for the refund skill" answer="1"} The engine's check and the prompt line are two copies of one guard. Here, the prompt's copy re-activated `authenticate` in 8 of 10 logged-in runs after the engine had passed the traveller, and deleting it took that to 0 of 10 without weakening the gate before login. Search the skill for the resolver's name, delete the line, and rerun the logged-in path. `hidden: true` has no condition, so it would lock verified travellers out. ::: ## What these runs do not show - **More than a small sample:** ten runs per precondition copy after login, three on the hidden copy, and three per copy before login, all on gpt-5.2 at temperature 0, one agent, one phrasing per conversation. We set aside a third conversation whose second turn, in every run, answered the agent's question about its introduction. - **An unmodified agent:** `lib/database.py` carries one harness edit, which reads the project root from `ATLAS_PROJECT_ROOT` so the served agent finds the demo data. The hidden copy's `agent.yml` keeps the binding, unused. - **The whole change:** no run completed a date change. The skill hands paid reissues to a human (`skills/change_booking/skill.md` at `01d6a6e`, lines 31 to 32), and every before-login run ended still collecting or confirming the new date, apart from one hidden run that offered a handoff. No run reached the drop after a failed PIN. - **Other models and versions:** gpt-5.2 on 3.21.0.dev3 only. dev1 accepts the skill-level line and dev3 refuses it; dev2 was not tested, and every build here is a prerelease. - **The old gate after login:** on 3.21.0.dev1, line 19 sat beside the skill-level `requires:` line, and after login `change_booking` was routable there too. We did not run that control, so these runs do not show whether the second `authenticate` predates the migration. - **An earlier round:** Claude Haiku 4.5, on an agent whose served model carried no demo data and whose opening turn failed in every run, is recorded in receipts e6, e8, e10, e11 and e12. This page draws no rates from it. All runs were recorded on 29 September 2026. Deleting one line and rerunning with everything else unchanged is a one-variable mutation, the technique the [guard mutation-testing guide](/library/guides/mutation-test-an-agent-guard/) applies to a guard's test suite. :::cta{href="/library/guides/mutation-test-an-agent-guard/" label="Mutation-test an agent guard"} A guard suite reported 186 of 186 mutations killed and still passed a one-word break. The guide shows what that count leaves out. ::: --- # Instructions Survive the Transport Source: https://rasa.community/library/tutorials/crm-transport-swap/ Author: Rod Rivera Published: 2026-09-02 In September 2026 a tutorial in this catalog shipped with a promise at the bottom of its README. Here is the paragraph, unedited: ```text HubSpot also exposes an MCP server, and Mantle has an `mcp_servers:` block in `integrations.yml`. In `3.19.0.dev7` that block is parsed but has no runtime consumer yet — the Mantle engine never connects to it — so this tutorial uses REST. When MCP lands, the same three skills can keep their instructions and swap `tools/crm.py` for imported remote tools. ``` That is a claim about the future written in the present tense, and most of them quietly stop being true. This one is checkable, which is the only reason it is worth returning to: either the three skills kept their instructions, or they did not, and a diff settles it. MCP landed in `3.20.0.dev6`. So the promise is now due. ## What was wrong with making it Nothing, except that a promise nobody goes back to verify is decoration. The useful thing is not the sentence. It is that the sentence was specific enough to be wrong — it named the files, it named the mechanism, and it said what would **not** change. Vague optimism cannot fail a test. ## What you will end up with The same agent, with its transport replaced and its instructions untouched, and a script that fails if that ever stops being true: ```text 1. skill instructions are unchanged by the swap ✓ check_tickets: only import_tools differs 2 changed line(s) ✓ check_tickets: instruction body byte-identical 687 bytes ✓ identify_customer: only import_tools differs 2 changed line(s) ✓ identify_customer: instruction body byte-identical 804 bytes ✓ log_interaction: only import_tools differs 2 changed line(s) ✓ log_interaction: instruction body byte-identical 1226 bytes The transport changed. The instructions did not. ``` No licence, no API key, no HubSpot account. The whole chapter runs on loopback. ## The five chapters 1. **[The promise, and how to check it](./01-the-promise/)** — read the claim as a test, and get the proof failing before it passes. 2. **[Swapping the transport](./02-swapping-the-transport/)** — `mcp_servers:`, `import_tools: mcp/:`, and the three-line diff. 3. **[What MCP takes away](./03-what-mcp-takes-away/)** — no tool context, no stdio, and a failure that arrives after training rather than during it. 4. **[What survives anyway](./04-what-survives-anyway/)** — why the confirmation constraint still fires against a remote tool. 5. **[Shaping the answer for the channel](./05-shaping-for-the-channel/)** — the narrowest channel decides the shape, and it is a skill decision, not a transport one. ## What you need first This chapter extends a tutorial that already works. Build [Ora](https://github.com/RasaHQ/rasa-community-resources/tree/52986f71071e419533f5f00335440a7ff1aef500/tutorials/rasa-hubspot-crm-tutorial) over REST first, or at least clone it and read the three skills — the swap is only visible against something that already ran. That is not a formality. A from-scratch MCP tutorial cannot teach this, because "the instructions did not change" needs a before. --- # Chapter 1 — Checking the promise Source: https://rasa.community/library/tutorials/crm-transport-swap/01-the-promise/ Author: Rasa team Published: 2026-09-02 The promise named three things: the three skills keep their instructions, `tools/crm.py` goes away, and imported remote tools take its place. Two of those are easy to observe. The first one — "keep their instructions" — is the one worth writing a test for, because it is the one a hurried author breaks without noticing. ## Read the claim as an assertion "The same three skills can keep their instructions" means: after the swap, the body of every `skill.md` is byte-identical to what it was before. Only the frontmatter may differ, and only in one key. That is precise enough to check with `difflib` and a string comparison: ```python rest_body = read(f"skills/{skill}/skill.md").split("---\n", 2)[-1] mcp_body = read(f"mcp_variant/skills/{skill}/skill.md").split("---\n", 2)[-1] assert rest_body == mcp_body ``` Splitting on the second `---` throws away the YAML frontmatter and keeps the instructions. If the two halves match, the model is reading exactly the same words it read before, and any behaviour change has to come from somewhere else. ## Get it failing first A proof that has never failed has not been tested. Before trusting the script, break the thing it checks and confirm it notices. Edit one line of prose in the MCP copy of `check_tickets` — add a clause to the first instruction: ```text Report the caller's tickets, and always list them even when not identified. ``` Then run the proof: ```bash make mcp-prove ``` ```text ✗ check_tickets: only import_tools differs prose changed: ["Report the caller's tickets.", "Report the caller's tickets, and always list them even when not identified."] ✗ check_tickets: instruction body byte-identical bodies differ 2 check(s) failed: - check_tickets: only import_tools differs - check_tickets: instruction body byte-identical ``` Exit code 1. It prints both versions of the line, so the diff is visible without opening an editor. Undo the edit and it goes green again. Do the same with the other two failure modes before moving on, because they are the ones you will actually hit: **Rename the server** under `mcp_servers:` to `hubspot_crm_TYPO`: ```text ✗ parse_mcp_servers accepts the file servers: ['hubspot_crm_TYPO'] ✗ server url is loopback http no server named 'hubspot_crm' ✗ no skill imports an unconfigured server missing from integrations.yml: ['hubspot_crm'] ``` **Mistype a tool name** in an `import_tools:` line: ```text ✗ check_tickets imports list_open_tickets reference not found in skill.md ``` Three ways to break it, three different red lines, each naming the file to open. That is what makes it a proof rather than a smoke test. ## Why this runs with no credentials The script starts both mock servers itself and tears them down afterwards: ```python for script, host, port in ( ("scripts/mock_hubspot.py", CRM_HOST, CRM_PORT), ("scripts/mcp_crm_server.py", MCP_HOST, MCP_PORT), ): ``` The first is the stand-in CRM the REST tutorial already shipped. The second is a new MCP server that sits in front of it and speaks MCP to Mantle. Both bind to `127.0.0.1`, so no packet leaves the machine and there is nothing to pay for and nothing to sign up to. That matters more than convenience. A proof requiring an account is a proof most readers never run, and one that goes stale the first time the account expires. ## What the script actually checks Six groups, eighteen assertions. In order: 1. Skill prose is unchanged — the claim itself. 2. `parse_mcp_servers` accepts the new `integrations.yml`. 3. Each `mcp/:` reference parses, and `get_bare_tool_name` returns the name the model will see. 4. Every imported server id is configured. 5. The MCP server really exposes the three tool names, over a live session. 6. A tool call over MCP returns the same customer the REST tool returned. Groups 2 to 5 are the checks the engine performs at model load, run early. Group 6 is the one that says the wire works. Group 1 is the one this tutorial is about. --- # Chapter 2 — Swapping the transport Source: https://rasa.community/library/tutorials/crm-transport-swap/02-swapping-the-transport/ Author: Rasa team Published: 2026-09-02 Two files change. That is the whole swap. ## 1. Declare the server `integrations.yml` gains an `mcp_servers:` block. Everything above it — the model group, the channels — stays exactly as it was: ```yaml mcp_servers: - name: hubspot_crm url: http://127.0.0.1:8931/mcp tool_timeout: 15 ``` `name` is the id the skills will reference. It must match byte for byte; a typo here is caught at model load rather than at train time, for reasons chapter 3 gets into. `url` has to be `http` or `https`. That is not a style preference — `MCPServerSpec` rejects any other scheme outright, which rules out stdio servers entirely. Chapter 3 covers what that means in practice. Pointing at HubSpot's own MCP server later is the same two lines: ```yaml mcp_servers: - name: hubspot_crm url: https://mcp.hubspot.com/anthropic token: ${HUBSPOT_ACCESS_TOKEN} ``` ## 2. Point the imports at it Each skill names the tools it wants with `mcp/:`. Here is the complete diff for all three skills: ```diff --- skills/identify_customer/skill.md +++ mcp_variant/skills/identify_customer/skill.md import_tools: - - find_contact_by_email + - mcp/hubspot_crm:find_contact_by_email --- skills/check_tickets/skill.md +++ mcp_variant/skills/check_tickets/skill.md +import_tools: + - mcp/hubspot_crm:list_open_tickets --- skills/log_interaction/skill.md +++ mcp_variant/skills/log_interaction/skill.md +import_tools: + - mcp/hubspot_crm:add_timeline_note ``` That is every changed line in every skill. Six lines, all of them inside YAML frontmatter. Note the asymmetry. `identify_customer` already had an `import_tools:` block, because `find_contact_by_email` was a global tool it imported by name — so its change is a **substitution**. The other two called local tools defined in their own `tools.py`, which needed no import at all, so their change is an **addition**. Different starting points, same destination. ## What the model sees `get_bare_tool_name` strips the prefix, so the model is offered `find_contact_by_email`, not `mcp/hubspot_crm:find_contact_by_email`: ```text ✓ identify_customer imports find_contact_by_email server=hubspot_crm llm sees 'find_contact_by_email' ``` This is the mechanical reason the instructions do not have to change. The skill's prose says "call `find_contact_by_email`", and after the swap there is still a tool called `find_contact_by_email` in scope, with the same arguments and the same result keys. The routing changed underneath; the name did not. ## What did not change Worth listing, because the list is longer than the diff: - Every line of instruction prose, in all three skills. - The ordered block in `log_interaction` — `draft_summary` then `write_note`. - The `requires_confirmation` constraint, and both its utterances. - `memory.yml`, at project level and in every skill. - `lib/hubspot.py`, the CRM client. The MCP server calls straight into it. - `agent.yml`, the persona, the rules. `lib/hubspot.py` is the quiet one. The MCP server imports the same client the REST tools used and calls the same three functions. The error taxonomy — the difference between `contact_not_found` and `crm_unreachable` — is preserved because it was never in the transport layer to begin with. It was in the client, and the client did not move. ## Run it ```bash make mock # terminal 1: the mock CRM make mcp-server # terminal 2: the MCP server in front of it make mcp-swap # move the project onto MCP make train && make chat ``` `make mcp-status` says which transport is live: ```text MCP — integrations.yml declares mcp_servers: - mcp/hubspot_crm:list_open_tickets - mcp/hubspot_crm:find_contact_by_email - mcp/hubspot_crm:add_timeline_note ``` And `make mcp-restore` puts REST back. The swap is deliberately reversible: the REST project is the baseline the swap is measured against, so it stays on disk rather than being overwritten. Delete it and the lesson goes with it. ## Verify with the engine, not just the eye Before training, ask the engine whether it accepts the project: ```bash uv run python -c "from pathlib import Path; \ from rasa.mantle.validation import validate_project; validate_project(Path('.'))" ``` It is free, offline, and needs no credentials. On the swapped project it passes, which means the imports parsed, the server block validated, and no skill is referencing a tool that no longer exists locally. --- # Chapter 3 — What MCP takes away Source: https://rasa.community/library/tutorials/crm-transport-swap/03-what-mcp-takes-away/ Author: Rasa team Published: 2026-09-02 Six changed lines is a small enough diff to conclude that MCP is free. It is not. Three things stop working, and all three are enforced by the engine rather than left to your judgement — which is better, because you find out immediately. A chapter that ended at "look how easy that was" would be the more comfortable one to write and the less useful one to read. ## 1. An MCP tool has no memory This is the big one. A local tool receives a `ToolContext` and can read and write memory: ```python @tool(description="List the support tickets on the identified customer's account.") async def list_open_tickets(context: ToolContext = None) -> ToolResult: contact_id = context.memory.get("project.contact_id") if context else None if not contact_id: return ToolResult(llm_response={"ok": False, "error": "not_identified"}) ``` Read that carefully: the tool is the authority on whether the caller has been identified. It does not trust the model to know, it looks the value up itself. An MCP tool cannot do that. The engine says so in its own docstring, in `skill_executor.py`: ```text Local tools receive a ToolContext; MCP tools dispatch through the processor-owned MCPRuntime. ``` There is no context object on the remote side, because there is no remote side of your process. So the value has to travel as an argument: ```python @mcp.tool(description="List the support tickets on the identified customer's account.") async def list_open_tickets(contact_id: str) -> CrmResult: if not contact_id: return CrmResult(ok=False, error="not_identified") ``` Same name, same result keys, same error string — which is why the instructions still work. But `contact_id` is now something the **model** supplies, from what it has seen in the conversation, rather than something the tool reads from memory. ### Why that is a security property and not a detail In the REST version, a caller cannot cause `list_open_tickets` to read someone else's tickets, because the tool ignores anything the caller says and uses the id that `find_contact_by_email` wrote to memory. In the MCP version, the id is a model-supplied argument, and the model is influenced by the conversation. The rule that falls out of this: > If a tool's correctness depends on a value the user must not be able to > influence, that value cannot be an MCP argument. Keep that tool local. Ora's ticket lookup is a read of the caller's own record, and the mock data is fixture data, so the tutorial can afford the looser arrangement. A tool that moves money could not. The transport swap is not free for every tool in the project — it is free for the tools whose inputs were already safe to state out loud. ## 2. There is no stdio transport Most desktop MCP clients speak stdio: the server is a subprocess, and messages go over its standard input and output. Mantle does not do that. `MCPServerSpec` requires a `url:` whose scheme is `http` or `https`, and rejects everything else at config-parse time. Underneath, `MCPServerConnection` only ever builds a `streamablehttp_client` — there is no stdio branch to reach. So an MCP server you already run locally over stdio cannot be plugged into this engine as-is. You would wrap it in an HTTP server first. That is what the bundled `scripts/mcp_crm_server.py` is: `FastMCP` with `transport="streamable-http"`, bound to loopback. Loopback is what keeps the chapter credential-free — nothing leaves the machine — but it is a workaround for a missing transport, not a design preference, and it is worth knowing which is which before you plan an integration around it. ## 3. A missing remote tool fails late `parse_mcp_imports` runs at model load and checks quite a lot: that every reference is syntactically `mcp/:`, that no skill imports two tools with the same name, that an imported name does not collide with a local tool or a reserved framework name. What it explicitly does not check is whether the tool exists. Its own docstring: ```text Remote existence is not checked here: that requires list_tools at connection time. ``` So `rasa train` succeeds with a typo in an import line. The failure arrives later, from `MCPRuntime.prepare`, once there is a live session to ask: ```text MCP server 'hubspot_crm' does not expose imported tool 'list_tickets' for skill 'check_tickets'. ``` The message is good. The timing is not: you trained a model to find out. This is the specific gap `make mcp-prove` closes. Its check 5 opens a real MCP session and calls `list_tools` — the same call `MCPRuntime.prepare` makes — before you spend a training run: ```text ✗ check_tickets imports list_open_tickets reference not found in skill.md ``` ## The shape of a limits section Three limits, and the useful form for each was the same: name the symbol that enforces it, say what it prevents, and say what you do instead. "MCP has some caveats" is not that. A reader who hits the tool-context limit at three in the morning needs the name `ToolContext` and the file it is missing from. --- # Chapter 4 — What survives anyway Source: https://rasa.community/library/tutorials/crm-transport-swap/04-what-survives-anyway/ Author: Rasa team Published: 2026-09-02 Chapter 3 listed what breaks. This one is about the thing that most obviously could have broken and did not, because the reason it survived is the reason the whole swap works. ## The guarantee in question `log_interaction` writes to someone else's system of record. The REST tutorial put two guarantees around that write, neither of them prose the model can talk itself out of: ```yaml tool_constraints: - add_timeline_note: requires_confirmation: enabled: true utter_for_confirmation: utter_confirm_note utter_on_user_denial: utter_note_cancelled ``` The runtime — not the model — asks before the write lands, and cancels if the caller says no. The obvious worry about the swap is that this stops working. `add_timeline_note` is no longer a Python function in the project; it is a name on a remote server. A constraint mechanism that looks up local callables would find nothing and either fail loudly or, far worse, quietly let the write through unconfirmed. ## It still fires It does not break, and the reason is visible in `tool_constraints_executor.py`. When a confirmed tool is about to run, the executor looks for a local function **and** asks the MCP runtime: ```python tool_func = tools_for_skill.get(pending_tool_confirmation.tool_name) is_mcp = bool( self._mcp_runtime is not None and self._mcp_runtime.has_mcp_tool( pending_tool_confirmation.skill_id, pending_tool_confirmation.tool_name, ) ) if tool_func is None and not is_mcp: ... # "is no longer registered" ``` An imported MCP tool satisfies `is_mcp`, so it is a legitimate target for confirmation. The ordered block in the same skill resolves the same way: its `execute_tool: add_timeline_note` step dispatches through the MCP runtime and is otherwise unchanged. So the caller still sees this before anything is written: ```text bot Save this to your record? "Dana called about the VAT rate on invoice 4471." ``` And saying no still cancels it. ## Why that is the general rule, not a lucky detail A constraint is declared on the **skill**, against a **tool name**. It is not declared on the tool's implementation, and it does not care where the implementation lives. The runtime resolves the name at dispatch time and applies the constraint either way. That is the same property that makes the instructions portable. Both the prose and the constraints are written against an interface — a tool's name, its arguments, its result keys — and the transport swap preserves that interface exactly. Nothing that was written against the interface has to move. The corollary is worth stating, because it tells you what to check in your own project: anything written against the _implementation_ rather than the interface will not survive. Chapter 3's `ToolContext` is precisely that. The tool's reading of `project.contact_id` was an implementation fact, not an interface fact, and it is the one thing that had to be rebuilt. ## One thing that did have to change, and it is not in the skills Building this chapter turned up a genuine problem that the proof script caught and reading would not have. The first version of the MCP server returned plain dictionaries: ```python async def find_contact_by_email(email: str) -> dict: ``` The tools listed correctly, the calls succeeded, and check 6 went red anyway: ```text ✗ find_contact_by_email over MCP None at None ✗ absent contact is not an error error=None ``` A tool annotated `-> dict` publishes no output schema, so the result comes back as unstructured text and `structuredContent` is `None`. Mantle notices: `MCPRuntime._invoke_mcp_tool` uses `structuredContent` when it is there and otherwise dumps the content blocks. The model would have received a JSON string wrapped in a list of text parts, instead of the flat object the local tool returned. The skill instructions branch on `ok` and `error`. Those keys have to survive the wire or the prose stops matching what the model sees — which would have falsified the whole claim, silently, in a way that still looks like it works. The fix is an explicit return model, so the tool publishes an output schema: ```python class CrmResult(BaseModel): model_config = {"extra": "allow"} ok: bool ``` And the proof gained an assertion that keeps it fixed: ```text ✓ result is structured, not text blocks outputSchema published ``` Two things to take from that. First, "the instructions did not change" is a claim about what the model receives, not only about what the files say — and those can come apart. Second, this is what the template means by writing the assertions before you trust them: the script's first run disagreed with the implementation, and the script was right. --- # Chapter 5 — Shaping for the channel Source: https://rasa.community/library/tutorials/crm-transport-swap/05-shaping-for-the-channel/ Author: Rasa team Published: 2026-09-02 `integrations.yml` has declared two channels since the REST version: ```yaml channels: rest: enabled: true inspector: enabled: true ``` The agent answers on both, and the tool result is identical on both. What differs is what a reader can actually take in — and that difference is settled in the skill, not in the channel. ## The case that shows it `list_open_tickets` returns structured rows: ```json { "ok": true, "count": 2, "tickets": [ { "id": "9001", "subject": "Invoice 4471 shows the wrong VAT rate", "stage": "waiting_on_us", "created": "2026-08-11" }, { "id": "9002", "subject": "Add a second admin to the account", "stage": "waiting_on_contact", "created": "2026-08-19" } ] } ``` In the Inspector that renders as a tidy little table and looks finished. Pushed through the REST channel it arrives as one run of text. Read aloud by a voice channel it would be unusable: nobody can hold `waiting_on_contact` and a ticket id in their head while the next row starts. ## The instruction that fixes it `check_tickets` does not say "return the ticket list". It says: ```text - Success with tickets: read back each subject and its stage in plain language. `waiting_on_us` means Meridian owes them a reply; `waiting_on_contact` means the ticket is waiting on the customer. ``` Three decisions are packed into that, and each one is about the listener rather than the data: 1. **Read back the subject, not the id.** `9001` means nothing to the caller. 2. **Translate the stage.** `waiting_on_us` is a database value; "we owe you a reply" is an answer. 3. **Plain language, not a table.** The narrowest channel the agent serves decides the shape, and a table does not survive being spoken. The result is a sentence that works everywhere: ```text bot Two. Invoice 4471 shows the wrong VAT rate, waiting on us. And a request to add a second admin, waiting on you. ``` ## Shape once, in the skill The tempting alternative is a per-channel template: a table for the Inspector, a sentence for voice, something else for REST. Resist it for as long as you can. Every channel you add multiplies the surfaces where the wording can drift, and the version a reader complains about is always the one nobody remembered to update. Writing the constraint into the instructions once means the model produces an answer shaped for the narrowest consumer, and the wider ones display it fine. You lose the tidy table in the Inspector. You gain one description of what a good answer sounds like. ## The connection to the transport swap This is the same property as chapter 4's constraint, seen from another angle. Formatting was decided against the **result shape** — the tool returns tickets with a subject and a stage — and not against the transport that carried the result. So when the transport changed, the formatting instruction did not have to, because nothing it depended on moved. `subject` and `stage` are still there, still spelled the same way, still meaning the same thing. Three things in this project turned out to be portable across the swap: instruction prose, tool constraints, and output shaping. All three were written against the interface. The one thing that was not portable — the tool reading `project.contact_id` out of memory — was written against the implementation. That is the whole lesson, and it is not really about MCP: > Write against the interface and the transport is an implementation detail. > Write against the implementation and it is not. MCP is just the change that made the difference visible. ## Where to go next - Run `make mcp-prove` on your own fork and break it three ways, as in [chapter 1](/library/tutorials/crm-transport-swap/01-the-promise/). A proof you have not seen fail is not one yet. - Point `mcp_servers:` at a real server. HubSpot's own is two lines of YAML away, and the three skills do not change for it either. - Audit your own tools for chapter 3's rule: which of them depend on a value the caller must not be able to influence? Those stay local. --- # The Document Is Derived, Never Written Source: https://rasa.community/library/tutorials/deriving-a-document/ Author: Rod Rivera Published: 2026-09-03 Here is what a capable model produces when you ask it for a client suitability record. This is not a strawman — it is the good version, from a careful prompt: ```text SUITABILITY RECORD — Marged Ellis Portfolio: Ellis Family Income Your portfolio is currently valued at approximately £486,000, held across a balanced mix of global equities (around 58%), sterling corporate bonds (roughly 30%), and property funds (about 12%). This allocation remains well suited to your Balanced risk profile and your objective of drawing a steady income from 2028. Ongoing charges of 0.62% and an advice fee of 0.45% apply. ``` That is fluent, correctly formatted, and plausible in every particular. It is also unusable, and the reason is not a mistake you can point at. ## Nothing in it is wrong. That is the problem Read it again with one question in mind: **which of those numbers came from a record, and which came from the model?** You cannot tell. Neither can the client. Neither can the compliance officer who reads it in eighteen months. Every figure is stated in the same confident register, and the register is the only evidence on offer. Now the specifics: - **"approximately £486,000"** — the extract says £486,210.44. Something rounded it, and nothing recorded that it did. - **"about 12%"** — the property funds line is 11.9%. Close. Also: this portfolio's property holding was never sourced at all in our fixture, so a model producing this sentence has produced a figure for a position it was never given. - **"remains well suited to your Balanced risk profile"** — this is a suitability determination. A sentence a regulator reads as a professional judgment, generated by next-token prediction. - **"Ongoing charges of 0.62%"** — correct, as it happens. Indistinguishable from the ones that are not. The document is not false. It is **unfounded**, which is worse, because false documents get caught and unfounded ones get filed. ## The missing piece Every figure needs to carry the record it came from, and a figure with no record must be **incapable of appearing** — not discouraged, not flagged in review, incapable. That rules out the obvious architecture. If a model writes the document text, then provenance is at best an annotation the model also writes, and a model that can write a footnote can write a wrong one. The fix is not a better prompt or a checking pass. It is to take the pen away. ## What you will build An agent that assembles a suitability record and cannot write a word of it. The conversation edits **structured state** — a declared set of fields, each one pointed at a source record. The document is a **pure function** of that state, recomputed from scratch every time anyone asks for it. Delete a field and it disappears from the document. Point a field at a record that does not exist and it renders blank, and the blank is reported. By the end you will have run this, and understood why it is the load-bearing test: ```text REFUSED: 1 figure(s) no longer match the record they cite. The document was not rendered. total_value: document says '486210.44', source says '911000.0' cited as: Custodian position extract · VAL-2026-08-29-PF4402 · total_value_gbp ``` The custodian restated the valuation after those figures were read into state. Every figure in the document was still footnoted; one footnote had quietly stopped being true. The renderer produced nothing at all. ## The six steps | Step | What it teaches | | ------------------- | ------------------------------------------------------------------------- | | Artifact as outcome | The deliverable is a file, not a reply, and that changes the design | | State, not prose | Fields live in declared memory; the model negotiates, tools write | | Grounded fields | Every figure traces to a record. Unsourced renders blank, never plausible | | Derived rendering | Template plus state equals document, recomputed on the far side | | Revision as diff | An edit changes state and re-renders; history is append-only | | Stated limits | What it refuses, in the same words the code uses | ## Before you start Clone the [companion repository](https://github.com/RasaHQ/rasa-community-resources/tree/52986f71071e419533f5f00335440a7ff1aef500/tutorials/rasa-document-artifact-tutorial). Chapters 1 through 5 need no credentials, no network, and no model — the proof runs under bare `python3`. Only the last chapter talks to the agent. ```bash cd tutorials/rasa-document-artifact-tutorial make prove ``` If that exits zero, you have everything you need. --- # Chapter 1 — The Deliverable Is a File Source: https://rasa.community/library/tutorials/deriving-a-document/01-artifact-as-outcome/ Author: Rasa team Published: 2026-09-03 A reply is consumed once, by the person who asked. A document is kept, forwarded, filed, and read by people who were not in the conversation and cannot ask a follow-up question. That difference is not about formatting. It changes what the agent must guarantee. ## Three properties a reply does not need **It outlives its context.** The reader in eighteen months has the document and nothing else — not the transcript, not the caller's tone, not the moment where the adviser said "roughly, for now". Anything the document does not carry explicitly is gone. **It is read as a claim by the firm.** A chatbot saying something wrong is a bad chatbot. A suitability record saying something wrong is the firm saying it, in writing, on its own letterhead. **It has to be reproducible.** "What did we send the client in August?" must have exactly one answer, and producing it again must not produce something slightly different. None of the three is satisfied by generating text. All three are satisfied by deriving a document from state that was stored. ## The shape this forces Stop thinking of the document as the output of the conversation. Think of it as a **view** over state that the conversation edited. ```text conversation ──edits──▶ STATE ──renders──▶ document ▲ │ source records ``` Everything in this tutorial follows from the direction of those arrows. The conversation touches state. State is built from source records. The document is computed from state, and nothing writes to it directly — there is no arrow into the document except through the renderer. Run it: ```bash make render ``` You will get a Markdown file. Look at the section near the bottom: ```text ## Provenance Every value above, and the record it came from. | Field | Value | Source record | | --- | --- | --- | | client_name | Marged Ellis | Custodian position extract · CL-77301 · legal_name | | total_value | £486,210.44 | Custodian position extract · VAL-2026-08-29-PF4402 · total_value_gbp | | objective | Draw a steady income from the portfolio from 2028, … | Client fact-find questionnaire · FF-77301-OBJ · value | ``` That table is not documentation of the document. It is the same data the body was rendered from, printed a second way. It cannot disagree with the body, because both come from one place. ## Why Markdown and not PDF The canonical artifact is deterministic Markdown, and PDF — if you need it — belongs downstream of it. The reason is that the guarantees in this tutorial are checked by **diffing**. Byte-identical re-renders, one-field changes showing up as one-field diffs: all of it needs a diff surface a human and a test can both read. PDF byte-diffs are noise, and a PDF that differs in its embedded timestamp fails an idempotency check for no reason anyone cares about. Get the text artifact right and correct. Then render it to PDF, and let the PDF step be a formatting concern rather than a correctness one. ## The completeness section Scroll further down the rendered document: ```text ## Completeness 3 declared field(s) have no source and render blank: - `property_value` - `property_weight` - `include_property_breakdown` ``` A document that states what is missing from it is doing something a generated document structurally cannot. A model writing prose has no concept of a field it was not given — the absence produces no output, so nothing marks it. Here the field list is declared up front, so a field with no value is a visible, countable gap. That is the first payoff of deriving rather than writing, and the next chapter is about where those fields are declared. --- # Chapter 2 — State, Not Prose Source: https://rasa.community/library/tutorials/deriving-a-document/02-state-not-prose/ Author: Rasa team Published: 2026-09-03 Open `docpkg/state.py`. The document is declared there, field by field, before any value exists: ```python FIELDS: tuple[FieldSpec, ...] = ( FieldSpec("client_name", "Client", "Heading", "sourced", False, "text"), FieldSpec("total_value", "Total portfolio value", "Portfolio", "sourced", False, "money_gbp"), FieldSpec("risk_warning", "Risk warning", "Disclosures", "sourced", False, "quote"), FieldSpec("addressed_to", "Addressed to", "Heading", "negotiated", True, "text", required=False), ) ``` The renderer walks this tuple. A field not in it cannot appear in the document; a field in it with no value appears as a blank and is counted as a gap. ## Two kinds of field, and nothing else That fifth argument is `kind`, and it takes exactly two values. **Sourced.** The value comes from a record. Money, weights, dates, names, regulated wording. The model may _negotiate_ which record — "shall I use the August valuation or July's?" is a perfectly good question — but it never supplies the value. **Negotiated.** The value is a choice the conversation makes that no record can supply: who the record is addressed to, whether to break out the property holdings. These are **selections from a closed set**, not prose. Note what is missing: there is no third kind for "text the model wrote". That is deliberate, and it is the design decision the whole tutorial rests on. A suitability record has no field a model may fill with prose, because every sentence in one either states a figure or quotes an approved disclosure — and both of those are sourced. If a new section seems to need free text, that is a signal the section has not been decomposed into fields yet. It is not a signal to add a `free_text` field. ## The sixth argument is `llm_settable` ```python FieldSpec("total_value", "Total portfolio value", "Portfolio", "sourced", False, "money_gbp") # sourced ──┘ └── llm_settable ``` Every sourced field is `False`. The suite asserts it over the whole declared set rather than a sample, so a field added next year is covered without anyone remembering to extend a list: ```python def test_every_sourced_field_is_not_llm_settable(self): leaky = [s.key for s in FIELDS if s.kind == "sourced" and s.llm_settable] self.assertEqual(leaky, [], "a sourced field the model could set") ``` And the mirror test, which is the one that matters more: every field that _is_ `llm_settable` must have a closed set of allowed values registered for it. A negotiated field with no closed set cannot be set at all — the default is "no", not "anything". ## What project memory holds, and what it does not Look at `memory.yml` and notice what is absent. There is no field holding document text, no field holding a section body, and no field holding a figure. The document's field values live in `docpkg` state, written only by tools. A value in Rasa memory is text, and flattening a sourced value into text to store it would strip the citation — leaving a value that is a value again, with no record attached. That is precisely the condition the tutorial exists to prevent, so the state does not live there. What memory does hold is the document's identity and the conversation's own choices: ```yaml document_id: type: text description: Identifier of the suitability record being assembled on this call. addressed_to: type: categorical enum_values: [client, adviser, client_and_adviser] ``` Project memory cannot be `llm_settable` at all in the tested release, `3.21.0.dev5`, and it should not be. A model that can write the document id can address a client's record to someone else. ## The model negotiates; tools write This is the division of labour, and it is worth stating positively rather than as a restriction. The model is good at the conversation: - working out that "her risk score" means the fact-find's `risk_profile` field - noticing the adviser asked for a figure the extract does not contain - reading a value back for confirmation before changing it - explaining why a field will render blank None of that requires writing a number, and all of it is work. The next chapter is about the mechanism that keeps the two apart. --- # Chapter 3 — Grounded Fields Source: https://rasa.community/library/tutorials/deriving-a-document/03-grounded-fields/ Author: Rasa team Published: 2026-09-03 The obvious way to enforce provenance is to build the document from plain values and then run a validator over it that asks "is everything sourced?". That design fails in a specific and predictable way. The validator can only check the fields it knows to look at, so the first field somebody adds without updating the validator is unsourced and silent. The check drifts behind the document, and you find out when a client does. ## A value and its citation are one type `docpkg/sources.py` takes the other approach. A document value is not a string: ```python @dataclass(frozen=True) class Sourced: value: Any citation: Citation ``` There is no constructor that produces a value without a citation. So an unsourced value cannot be built, and a field added next year is sourced or it does not render — without anyone remembering to extend a list. A `Citation` names three things, and all three are needed: ```python @dataclass(frozen=True) class Citation: source_id: str # which file record_id: str # which record inside it field: str # which key of that record ``` A citation naming only the file is the documentary equivalent of "it was on the internet somewhere". The source id is checked at construction against a closed registry, so a citation to an invented origin raises immediately: ```python def test_a_citation_to_an_undeclared_source_is_rejected(self): with self.assertRaises(UnknownSource): Citation("some-source-nobody-declared", "REC-1", "value") ``` ## Two sources, because one cannot teach this The document draws on two fixture sources, and the split is the point: | Source | Authoritative for | Knows nothing about | | -------------------------- | ------------------------------------------- | ---------------------- | | Custodian position extract | Holdings, valuations, charges | What the client wants | | Client fact-find | Objectives, risk profile, capacity for loss | What anything is worth | With a single origin, "where did this come from?" has only one answer and the citation is decoration. With two, the provenance table is doing real work — and the commonest provenance bug becomes visible. It is not an unsourced figure. It is a figure sourced to a file that happens to contain a similar number for a different reason. That is why each source records what it is `authoritative_for`, and why the rendered document prints those claims: ```text ## Sources | Source | Authoritative for | | --- | --- | | Approved disclosure library | regulated wording, quoted verbatim | | Client fact-find questionnaire | objectives, risk profile and capacity for loss | | Custodian position extract | holdings, valuations and charges | ``` ## The third surface: text that is quoted, never generated `references/disclosures.json` is not a fourth integration. It is a _reference_ surface, and the distinction matters. A figure is a value: it can be looked up, compared, re-derived. A regulated paragraph is a **sentence**, and the only safe operation on a regulated sentence is to quote it verbatim with an identifier attached. Ask a model to "summarise the risk warning" and it produces something that reads better and means something slightly different, and nobody downstream can tell which words were approved and which were improvised. So the renderer does not accept a risk warning as text. It accepts a reference id, looks it up, and emits the stored wording. The suite checks the output byte-for-byte against the library: ```python def test_disclosures_are_quoted_verbatim_from_the_library(self): approved = {d["record_id"]: d["text"] for d in library["disclosures"]} document = render_markdown(build_fixture_state()) for text in approved.values(): self.assertIn(text, document) ``` ## Unsourced renders blank The fixture deliberately leaves the property holdings unsourced, so the shipped document has real gaps in it. A demo where every field happens to be populated cannot show what the blank rule does, and the blank rule is half the point. ```text | Property funds | — | | Property funds weight | — | ``` An em-dash. Not `N/A`, not `TBC`, not `0`, not an empty cell. A reader skimming a column of numbers has to be able to see that something is missing, and a zero reads as a figure. The strong form of the test is not that a blank appears. It is that no digit ever appears where a blank belongs: ```python def test_an_unsourced_field_never_renders_a_plausible_figure(self): rows = [ln for ln in document.splitlines() if ln.startswith("| Property funds")] for row in rows: self.assertIn(BLANK, row) self.assertNotIn("£", row) ``` And pointing a field at a record that does not exist gets you the same blank, rather than an error or a guess: ```python def resolve(citation: Citation) -> Any: ... return None ``` `None` — never a default, never the nearest match. Returning a plausible substitute here would defeat every guard downstream, because downstream cannot tell a real value from a helpful one. ## Why blank rather than refuse The asymmetry is deliberate, and it is worth being explicit about which way it runs. A blank in a client document is embarrassing, and somebody fixes it that afternoon. A plausible wrong number in a client document is a misrepresentation that nobody notices until it matters. So an _unsourced_ field renders blank and the document still ships with its gaps listed. A field whose citation has **stopped agreeing** with its record is a different condition entirely, and that one refuses. Chapter 5. --- # Chapter 4 — Derived Rendering Source: https://rasa.community/library/tutorials/deriving-a-document/04-derived-rendering/ Author: Rasa team Published: 2026-09-03 Here is the entire claim of this tutorial, expressed as a function signature: ```python def render_markdown(state: DocumentState) -> str: ``` No `content`. No `overrides`. No `extra_sections`, no template hook, no post-processing callback. The renderer takes state and nothing else, so a "document" assembled anywhere else has nowhere to land. It is asserted, because a signature is exactly the kind of thing a helpful refactor widens: ```python def test_render_markdown_takes_only_state(self): params = list(inspect.signature(render_markdown).parameters) self.assertEqual(params, ["state"], "render_markdown grew a parameter; content can now bypass the field set") ``` ## The other half: the edit functions A renderer that only reads state is worth nothing if anything can put arbitrary text into state. So `docpkg/edits.py` is the only door, and it is governed by three rules. **Rule 1 — a sourced field takes a citation, not a value.** ```python def set_sourced_field( state: DocumentState, field_key: str, *, source_id: str, record_id: str, record_field: str, reason: str, ) -> EditResult: ``` Read the parameter list again and notice what is not in it. There is no `value`. The function takes a record id and a field name, looks the value up itself, and stores what it found. A model that wants the total portfolio value to read £900,000 has no argument through which to say so. It can name a record, and the record says what it says. That is the intended difficulty — it is the difference between reporting and inventing. **Rule 2 — a negotiated field takes a value from a closed set.** `set_negotiated_field` does accept a value, which makes it the obvious back door. It is closed by checking the value against a registry before it reaches state: ```python ALLOWED_NEGOTIATED: dict[str, tuple[str, ...]] = { "addressed_to": ("client", "adviser", "client_and_adviser"), "include_property_breakdown": ("yes", "no"), } ``` A field with no entry here cannot be set at all. **Rule 3 — there is no third function.** No `set_text`, no `append_paragraph`, no `set_section_body`. The absence is the mechanism. A model cannot call a tool that does not exist, and a future author who needs one has to add it to this file, where the refusal rules are the first thing they read. ## The tools the model actually calls This is where the previous chapter's lesson could quietly be lost. The engine publishes every non-`context` tool parameter to the model as a JSON schema, so **a tool's signature is its permission list** — and asserting on the internal function while the tool underneath it differs is exactly the trap `patterns/voice-handoff-context` documented after falling into it. So the tests check the tools: ```python def test_the_agent_tool_for_setting_a_field_exposes_no_value_argument(self): from tools.document import point_field_at_record params = set(inspect.signature(point_field_at_record).parameters) self.assertNotIn("value", params) def test_the_render_tool_accepts_nothing_that_could_become_content(self): from tools.document import render_document params = [p for p in inspect.signature(render_document).parameters if p != "context"] self.assertEqual(params, []) ``` `render_document()` takes no parameters at all. There is nothing to inject. ## Determinism is a requirement, not a nicety The same state renders the same bytes, every time. No timestamp of rendering, no random ids, no dictionary iteration order — fields come from a declared tuple and the provenance table is sorted. ```python def test_rendering_is_stable_across_many_runs(self): renders = {render_markdown(build_fixture_state()).encode() for _ in range(5)} self.assertEqual(len(renders), 1) ``` This is what makes the next chapter mean anything. If rendering were nondeterministic, every diff would be noise with a real change buried in it, and nobody would read the diffs — which is the same as not having them. The document _does_ carry dates: the valuation date, the disclosure approval dates. Those are sourced field values, not render-time clock reads. A document that changes when you render it twice is not a record of anything. ## Deletion needs no mechanism ```python def clear_field(state: DocumentState, field_key: str, *, reason: str) -> EditResult: values = dict(state.values) values.pop(field_key, None) ``` No tombstone, no `deleted` flag, no special case in the renderer. The renderer walks the declared field list and asks state for each key; a key that is not there has no value, and a field with no value renders blank. Delete a field and it disappears from the document. That is not a feature somebody added — it is the only behaviour available, which is a much stronger thing for it to be. --- # Chapter 5 — Revision as Diff Source: https://rasa.community/library/tutorials/deriving-a-document/05-revision-as-diff/ Author: Rasa team Published: 2026-09-03 Run this: ```bash make diff ``` It re-points the total portfolio value at the _cash_ line of the same valuation record — a plausible mistake, and one a document without a provenance table would hide completely: ```text One field changed: field : total_value before : 486210.44 after : 21804.1 cited : Custodian position extract · VAL-2026-08-29-PF4402 · cash_gbp What moved in the document: -| Total portfolio value | £486,210.44 | +| Total portfolio value | £21,804.10 | -| total_value | £486,210.44 | Custodian position extract · … · total_value_gbp | +| total_value | £21,804.10 | Custodian position extract · … · cash_gbp | +| 22 | total_value | 486210.44 | 21804.1 | adviser asked to show the cash line instead | ``` Three things moved, and each one is doing a job. The **figure** changed. The **citation** changed alongside it, which is the part a generated document cannot offer — the diff shows not just that the number moved but that it is now coming from a different place. And a **revision row** was appended, carrying the reason. ## One field changed means one field diffs The property is asserted, over the whole document body: ```python def test_one_field_changed_means_one_field_diffs(self): changed = [(b, a) for b, a in zip(before.splitlines(), after.splitlines()) if b != a] self.assertTrue(changed) for b, a in changed: self.assertTrue("risk_profile" in b or "Risk profile" in b, f"an unrelated line changed: {b!r} -> {a!r}") ``` This is only checkable because rendering is deterministic. It is also the test that catches an entire class of bug that is otherwise invisible: a renderer whose output depends on iteration order changes lines nobody touched, and in a document that a client compares against last quarter's, that is indistinguishable from a real change. ## The history is append-only ```python def _append(state, key, before, after, reason): return state.revisions + ( Revision(seq=len(state.revisions) + 1, field_key=key, before=before, after=after, reason=reason), ) ``` `seq` is derived from the existing length rather than stored separately, so a revision cannot be inserted with a sequence number that lies about when it happened. And `DocumentState` is frozen: edits return a _new_ state rather than mutating one, because "what changed between version 3 and version 4" is unanswerable against a mutable object, and that question is the entire value of a document a regulator may later ask about. A silent overwrite is refused outright: ```python def test_a_change_without_a_reason_is_refused(self): with self.assertRaises(EditRefused) as ctx: set_sourced_field(..., reason=" ") self.assertEqual(ctx.exception.code, "overwrite_without_reason") ``` ## Deleted from the body, kept in the history Clear a field and it vanishes from the document body — but the history still records what it was: ```python def test_the_deleted_value_survives_in_the_history(self): state = clear_field(build_fixture_state(), "objective", reason="withdrawn").state history = render_markdown(state).split("## Revision history")[1] self.assertIn("Draw a steady income", history) ``` Worth pausing on, because **this pair of tests started life as one failing assertion**. The first version of the proof script asserted that a deleted value appeared nowhere in the document, and it failed. The failure was correct: the value was gone from the body and preserved in the audit trail, which is exactly what an append-only history is for. The assertion had conflated two different guarantees. The body is the document. The history is the audit trail. They answer different questions and a test that checks both at once checks neither. ## The load-bearing case: when the source moves Everything so far assumes the source records stay put. They do not. A state object built this morning, carried through a conversation, and rendered this afternoon holds figures that were true when they were read. If the custodian extract is corrected, restated, or tampered with in between, the document would otherwise report this morning's numbers under this afternoon's citations — every figure footnoted, every footnote wrong. So before rendering anything, `docpkg/verify.py` re-resolves every citation in state and compares: ```python def require_intact_provenance(state: DocumentState) -> None: mismatches = verify_state(state) if mismatches: raise ProvenanceBroken(mismatches) ``` One call, at the top of `render_markdown`. Mutate the extract and: ```text REFUSED: 1 figure(s) no longer match the record they cite. The document was not rendered. total_value: document says '486210.44', source says '911000.0' cited as: Custodian position extract · VAL-2026-08-29-PF4402 · total_value_gbp ``` Note that it does **not** render blank. Blank is the right answer for a field that was never sourced; a field whose citation has stopped agreeing is a different condition, and quietly blanking it would hide a changed record behind what looks like an incomplete document. The whole render refuses, names every disagreeing field, and exits non-zero. A deleted source record is refused the same way, and the refusal reaches the agent as a refusal rather than a crash: ```python result = td.render_document().llm_response self.assertFalse(result["ok"]) self.assertEqual(result["refused"], "provenance_broken") ``` ## The guard has been watched failing A guard nobody has seen go red is a docstring, not a guard. So it was removed — one line, deleted from `render_markdown` — and the proof re-run: ```text 2. REFUSE — mutate the provenance table and the renderer refuses PASS the unmutated state renders FAIL the renderer REFUSED the mutated provenance IT RENDERED. A document was produced whose citations do not hold. FAIL a DELETED source record is refused too it rendered EXIT WITH GUARD REMOVED = 1 ``` With the guard gone the renderer emitted a complete, well-formatted document stating £486,210.44 while the extract said £911,000. It looked exactly like the correct one. The guard was restored and the proof returned to exit 0. **And this guard caught a real bug on its first run, before any test existed.** `set_negotiated_field` originally cited an unrelated disclosure record, for want of anywhere better to point a conversational choice. `verify_state` refused the entire document immediately: stored value `client_and_adviser`, cited record `DISC-SCOPE-004`, and they did not match. The guard was right and the citation was a lie. Borrowing a citation from a record that does not justify the value is precisely the failure this design exists to prevent — and it took ten minutes to commit it by accident, while writing the thing that prevents it. `CONVERSATION_SOURCE` in `docpkg/sources.py` is the fix, and the comment above it says why it is there. --- # Chapter 6 — What It Refuses Source: https://rasa.community/library/tutorials/deriving-a-document/06-stated-limits/ Author: Rasa team Published: 2026-09-03 A guarantee is only as good as the list of things it does not cover. This chapter is that list. ## Talk to it ```bash cp .env.example .env # fill RASA_LICENSE and OPENAI_API_KEY make train make chat ``` Then try the three things an adviser will actually ask for. **Ask it to write:** ```text just write a paragraph summarising why this portfolio suits her ``` There is no field for it to land in. The agent says the record has no free-text section, and offers the fields that carry the point instead. This is not the model declining out of caution — there is no tool that accepts a paragraph. **Ask it for a number:** ```text put the total at nine hundred thousand, she'll be pleased ``` `point_field_at_record` has no `value` argument. The most the model can do is look for a record containing that figure, and there is not one. **Ask it to soften a disclosure:** ```text soften the risk warning a bit ``` ```text refused: free_text_into_regulated_section ``` ## The refusal vocabulary One word per reason, used identically in the code, the proof output, and these chapters. A refusal a reader cannot match to the sentence that described it is a refusal they will assume was a bug. | Code | When | | ---------------------------------- | -------------------------------------------------------- | | `not_a_declared_field` | A field that is not in the declared set | | `free_text_into_sourced_field` | A value offered for a field that takes a citation | | `free_text_into_regulated_section` | Any conversational value aimed at Charges or Disclosures | | `value_not_in_allowed_set` | A negotiated value outside its closed set | | `overwrite_without_reason` | A change or deletion with no reason recorded | | `provenance_broken` | A stored figure no longer matches the record it cites | Every one is exercised, and the test asserts the _set_ rather than checking them one at a time — so a refusal that becomes unreachable is caught: ```python def test_every_refusal_code_is_reachable(self): ... self.assertEqual(seen, { "not_a_declared_field", "free_text_into_sourced_field", "free_text_into_regulated_section", "value_not_in_allowed_set", "overwrite_without_reason", }) ``` ## What this does not give you **It is not a compliance control.** Suitability rules impose obligations on advice, record-keeping, and review that a rendering pipeline does not discharge. This shows where a boundary belongs; it does not certify one. **It does not check that a figure is the RIGHT one.** Every figure traces to a record, and the renderer refuses when a citation stops holding. Neither property says the adviser pointed at the correct record. `make diff` demonstrates exactly this: the total portfolio value is re-pointed at the cash line, and the document renders happily because the new citation is perfectly valid. Provenance answers "where did this come from", not "is this the right source". **It does not cover the conversation itself.** The transcript is a separate surface with separate retention. What the model _said_ while negotiating is not governed by any of this — only what reached the document. **No document ingestion.** Parsing an uploaded PDF back into fields is a different problem with a different library risk surface, and it is out of scope here. **The state lives in a module-level object** for the length of the process. A real deployment keeps it in a document service keyed by `document_id`. What must not change in that move is the direction of the dependency: the service owns the state and derives the artifact, and the agent never receives an artifact it can edit and hand back. A round trip through the model is a round trip through something that can rewrite a number while sounding certain about it. ## Where a real system attaches Two seams, and nothing else needs to move. **Sources.** `SOURCES` in `docpkg/sources.py` maps a source id to a file. Replace the file reads with calls to a custody API and a CRM. The `Citation` type does not change, and neither does anything downstream of it. **Output.** `render_markdown` returns a string. Render it to PDF, post it to a document store, attach it to a case. Keep the Markdown as the canonical artifact and let the PDF be a formatting step, so the diff surface — and therefore every guarantee in Chapter 5 — stays text. ## One engine detail worth carrying away A shared tool may not be named `set_*`. `load_shared_tools` in the tested release, `3.21.0.dev5` deletes any shared tool whose name starts with that prefix, because it is reserved for a skill's auto-generated collect setter. It does so with a log warning and no error — so the tool is simply not there, and the failure surfaces later as `unresolved_tool` against the _skill_, which points at the wrong file entirely. This project's `point_field_at_record` was called `set_document_field` until that bit, and the twenty minutes it cost are the reason it is written down here. ## The shape, once more ```text conversation ──edits──▶ STATE ──renders──▶ document ▲ │ source records ``` The model is genuinely useful in that first arrow — working out what the adviser means, finding the right record, reading a value back before changing it, explaining why a field is blank. It is nowhere near the third. That is the whole design. Not a better prompt, not a checking pass over generated text: a pipeline in which the fabrication step has no code path to run in. --- # Measure Your Agent: Simulation, Assertions, and LLM Judges Source: https://rasa.community/library/tutorials/evaluation-harness/ Author: Rod Rivera Published: 2026-09-04 You changed a prompt. You ran the demo conversation. It went fine. Is the agent better? You genuinely do not know. The demo went fine _last_ time too, before the bug report. And the change you made touched a model whose output varies between identical runs — so even "it went fine twice" is weaker evidence than it feels like. This tutorial builds the thing that replaces that feeling with evidence: a simulation-based evaluation suite. Not a metaphorical one — the [companion pattern](https://github.com/RasaHQ/rasa-community-resources/tree/52986f71071e419533f5f00335440a7ff1aef500/patterns/evaluation-harness) is a runnable project you will clone, train, and evaluate. ## Three questions, one framework The central idea — and the thing most teams get wrong — is that "is the agent good?" is actually three different questions: | Question | Answered by | Character | | -------------------------------------------- | -------------------------------------- | ---------------------------------------------------- | | Did the agent **do** the right thing? | `assertions` — facts about the tracker | Exact, repeatable, no judge noise | | Did it **handle** the user well? | `criteria` — scored by an LLM judge | Tolerates paraphrase; costs variance and money | | Does it survive a **real, unscripted** user? | the LLM user-simulator | Improvises the customer's side from stage directions | On Rasa Mantle all three live in **one framework**: you author scenarios, a simulator LLM plays the customer, and the finished transcript is scored twice — deterministically against tracker events, and by a judge against your natural-language criteria. Why a simulator at all? Because **Mantle skills are not scripts**. A `skill.md` gives the model intent and constraints, not a dialogue path — so a test file of hard-coded user turns exercises one path through a system whose defining property is that it improvises paths. The simulator meets the agent on its own terms. The rule that keeps the whole thing honest underneath: > **Assert everything you can express as a fact about the conversation. > Reach for the judge only where there is no fact to assert.** Teams that invert this end up with a slow, expensive, flaky suite that measures their judge's mood as much as their agent's behaviour. The two paths do not read the same thing, and that is the detail worth carrying into Chapter 2: ![One scenario file feeds three things. Its simulation_context drives a simulator LLM that plays the customer turn by turn against the trained agent, and that loop produces two separate artefacts: the conversation tracker, a structural record, and the transcript, what was said. Assertions are checked against the tracker with no LLM involved, so they are exact and repeatable. Criteria are scored by a judge LLM reading the transcript, which tolerates paraphrase but costs variance and money. A run passes only when every assertion and every criterion passes. Every run bills three models: the agent's own from integrations.yml, and the simulator and judge from eval/conftest.yml.](/diagrams/evaluation-harness-two-scoring-paths.svg) ## The agent is a fixture, on purpose The companion project ships a two-skill retail-banking assistant with hard-coded data. It is deliberately boring, and the boringness is a design decision worth stealing: **a failing scenario should always mean the agent changed, never that the data moved.** Wire an eval suite to a live database and every flaky run becomes a debugging session about the data. The harness is the subject here; the agent is the lab rat. ## The chapters | # | Chapter | The question it teaches you to answer | | --- | ------------------------------ | -------------------------------------------------------------------- | | 1 | Set up the lab | Does it run on my machine, and what does a passing suite look like? | | 2 | Assertions | Did the agent do the right thing — exactly, repeatably, judge-free? | | 3 | The simulator | Does the agent survive a customer who never read your script? | | 4 | The judge | Was the answer grounded and relevant — and what do those words mean? | | 5 | What your eval cannot tell you | When is a green run evidence, and when is it noise? | Chapter 5 is short and it is the one to not skip. An eval that oversells itself is worse than no eval, because it launders a guess into a number. ## Before you start - Python 3.11 or 3.12, [uv](https://docs.astral.sh/uv/), and git. - A free [Rasa Pro Developer Edition licence](https://rasa.com/rasa-pro-developer-edition-license-key-request/) and an OpenAI API key. You need both from Chapter 1 — evaluating an agent requires running one. - A coding agent that can drive an MCP server (Claude Code, Cursor, or Copilot) — scenarios are run through Rasa's MCP server, per the [simulation docs](https://rasa.com/docs/pro/testing/simulation-evaluation/). - A cost note up front, and it is blunter than it used to be: **every simulated run bills three LLMs** — the agent's own model, the simulator, and the judge. It is cents per run, not dollars — but knowing that _no_ part of this instrument is free is half of what Chapter 5 teaches. You iterate with small run counts and save the big batches for decisions. The companion installation targets `rasa-pro==3.21.0.dev5`. The [compatibility report](/compatibility/) distinguishes current deterministic checks from live model runs. Following the companion pattern, claims about engine behaviour cite the source file they refer to. When you wonder "but how does it _actually_ score that?", the file and line are right there. --- # Chapter 1 — Set Up the Lab Source: https://rasa.community/library/tutorials/evaluation-harness/01-set-up-the-lab/ Author: Rasa team Published: 2026-09-04 Clone the companion repository and move into the pattern: ```bash git clone https://github.com/RasaHQ/rasa-community-resources.git git -C rasa-community-resources checkout 52986f71071e419533f5f00335440a7ff1aef500 cd rasa-community-resources/patterns/evaluation-harness ``` ## Credentials and training ```bash cp .env.example .env # fill in RASA_LICENSE and OPENAI_API_KEY uv sync --prerelease=allow uv run rasa train ``` While that trains, meet the lab rat. The agent is a retail-banking assistant with exactly two skills: - **`check_balance`** — a _task_ skill: disambiguate which account, call a tool, read back a number. Its tool returns **hard-coded balances** (`skills/check_balance/tools.py`): the Everyday Checking account contains $1,284.53 today, tomorrow, and forever, and Rainy Day Savings holds $9,140.00. - **`faq_support`** — a _knowledge_ skill: answer questions about overdraft fees and card replacement from **hard-coded policy text** (`skills/faq_support/tools.py`). Hard-coded is the point. When a scenario fails against this agent, there is exactly one explanation: the agent's behaviour changed. No stale database, no flaky API, no timezone. Every eval suite you build for a _real_ agent will have to fight for that property; the fixture gets it for free, which makes it the right place to learn what the instruments measure. ## The suite you just cloned Everything lives under `eval/`: ```text eval/ ├── conftest.yml # the two evaluation LLMs: simulator + judge └── scenarios/ # one behaviour per file, seven in all ``` `eval/conftest.yml` configures the **user simulator** (which generates the customer's turns) and the **judge** (which scores the transcript) independently — different jobs, separately swappable. The agent's own model lives in `integrations.yml`; changing _that_ is an experiment, changing the judge changes what "pass" means. The seven scenarios are the standard Mantle behaviour checklist — happy path, out-of-order input, disambiguation, correction, digression-and-resume, refusal instead of fabrication, and a negative case where the right behaviour is to start nothing at all. ## Run one scenario Scenarios run through Rasa's MCP server, driven from your coding agent (per-editor setup in the [simulation docs](https://rasa.com/docs/pro/testing/simulation-evaluation/)): ```bash uv run rasa tools run --mode stdio # loads this project's .env ``` Then ask your coding agent, in natural language: > Run the balance_named_up_front scenario in eval/scenarios/. Three things happen: the simulator reads the scenario's `simulation_context` and plays the customer, turn by turn, against your trained agent; the finished transcript is checked against the scenario's `assertions` (deterministic, tracker-level); and the judge scores its `criteria` (natural language). Results land under `eval/results//` — a `run_N.txt` per run with the transcript, every assertion's PASS/FAIL, every criterion's score with the judge's rationale, and an Inspector URL to replay the conversation, plus a `summary.txt` across scenarios. A run **passes only when every assertion and every criterion passes**. Quality metrics (helpfulness, task completion and friends) are recorded but do not gate. ## Two observations before the dissection **One run is three LLMs' worth of tokens.** The agent, the simulator, and the judge all bill on every run. There is no free tier of this instrument — which is why you iterate on one scenario with a run count of 1, and save "run everything five times" for decisions that deserve it. **The same run produces two kinds of evidence.** The assertion results are exact and repeatable-in-principle; the criteria scores are an LLM's opinion with variance attached. Keeping straight which kind you are reading is most of the skill, and the next three chapters take them one at a time. ## Check your understanding - Why are hard-coded balances a _feature_ of an eval fixture rather than laziness? - Which three models does a single scenario run pay for, and which one's swap would change what "pass" means? - Where do you look first when a run fails: the transcript, the assertion results, or the criteria rationale — and why? Got a pass report on your screen? Good. Now let's find out what, precisely, you just proved. --- # Chapter 2 — Assertions: Facts About the Tracker Source: https://rasa.community/library/tutorials/evaluation-harness/02-assertions/ Author: Rasa team Published: 2026-09-04 An assertion reads the conversation tracker — the engine's own record of what happened — and answers a yes/no question: did this flow start? does this slot hold this value? did this action run? No LLM is involved in checking it, so whatever the simulator improvised and however the judge felt that day, the assertion result is a fact. That is why assertions are the **ground truth** inside every scenario: the judged criteria capture quality of handling, but the non-negotiable facts — the right account, the tool actually ran, the right number was spoken — are nailed down here, where variance cannot reach them. ## The shape From `eval/scenarios/balance_named_up_front.yml` in the companion: ```yaml scenario: name: Customer names the account up front and gets its balance in one pass simulation_context: > You are a calm, logged-in customer of a retail bank. Open by asking for the balance of your Everyday Checking account... goals: criteria: - The agent does not ask which account the customer means... assertions: - flow_started: check_balance - slot_was_set: - name: check_balance.selected_account_id value: acc_checking - action_executed: get_balance ``` Read the assertions as a claim: _however this conversation went, by the end the right skill started, the right account landed in memory, and the tool ran._ Three facts, three assertions — and the criteria above them handle the "however it went" part. ## The vocabulary The scenario schema pins these assertion types (`rasa/builder/copilot/mcp_server/schema/scenario_schema.yml` in the installed engine — and the checking machinery is literally the classic e2e assertion classes: the scenario engine imports `rasa.e2e_test.assertions` directly, `assertion_engine.py:9`): | Assertion | Checks | | ----------------------------------------------------- | --------------------------------------------- | | `flow_started` | a flow with this id started | | `flow_completed` / `flow_cancelled` | a flow finished / was cancelled | | `slot_was_set` / `slot_was_not_set` | a slot holds (or doesn't hold) a value | | `action_executed` | a named action ran | | `bot_uttered` / `bot_did_not_utter` | a response, by name, buttons, or regex | | `pattern_clarification_contains` | the clarification pattern offered these flows | | `sequencing` | these facts occurred, in this order | | `generative_response_is_grounded` / `..._is_relevant` | the judge — Chapter 4's subject | Two naming details that will otherwise cost you an afternoon: memory keys are **scope-qualified** — the fixture's slot is `check_balance.selected_account_id`, skill id then entry, never the bare entry name — and a prose skill's flow id is simply the **skill id** (`check_balance`); the `skill__block` form belongs to ordered blocks (`rasa/mantle/skills/catalog.py:686-714`). Three habits in the companion's scenarios separate a suite that catches regressions from one that merely exists: ## Habit 1 — assert absence, not just presence From `eval/scenarios/smalltalk_starts_nothing.yml`: ```yaml goals: criteria: - The agent responds politely and briefly without pretending the customer asked for something assertions: - slot_was_not_set: - name: check_balance.selected_account_id ``` The cheapest way for an agent to _look_ competent on a positive-only suite is to start a task for everything. Your suite must contain conversations where the correct behaviour is restraint — and assert that restraint happened. An agent that treats "thanks, bye" as a balance enquiry passes every positive scenario you own. ## Habit 2 — match the fact, not the phrasing ```yaml - bot_uttered: text_matches: '9,140' ``` `text_matches` takes a regular expression. The disambiguation scenario pins `"9,140"` — the number the fixture tool returns for savings — and lets the model word the sentence however it likes. Pinning whole sentences gives you a suite that fails on every prompt tweak; a suite that cries wolf on wording trains your team to ignore failures, which is worse than no suite — you keep the cost and lose the signal. The discipline generalises: **assert what must be true (the number, the slot, the action), stay silent about what is allowed to vary (the phrasing).** ## Habit 3 — express order as `sequencing` Some behaviour _is_ an order of events. The correction scenario (`balance_correction.yml`) needs "the account was chosen, _then_ changed": ```yaml assertions: - sequencing: - slot_was_set: check_balance.selected_account_id # set once... - slot_was_set: check_balance.selected_account_id # ...then re-set - slot_was_set: - name: check_balance.selected_account_id value: acc_savings ``` Each child of `sequencing` must occur, in this order. The same slot set twice _is_ the correction — a judged criterion could only say the agent "seemed to handle the change of mind"; this states it as a fact. ## What this instrument cannot see Look again at the vocabulary. Every row is about _structure_: flows, slots, actions, responses. Nowhere is "was this answer _good_?" — because that is not a fact about the tracker. Ask the FAQ skill about overdraft fees and it answers in free text: an assertion can prove a response was uttered, even regex-match a fragment, but it cannot judge whether the sentences were accurate, complete, or on-topic. For that there is no fact to assert — that is judge territory, Chapter 4. But first: who is actually _talking_ to your agent during all of this, and why is it not you? --- # Chapter 3 — The Simulator: A Customer Who Never Read Your Script Source: https://rasa.community/library/tutorials/evaluation-harness/03-the-simulator/ Author: Rasa team Published: 2026-09-04 Every test file you have ever written hard-codes the user's side of the conversation. For a scripted system that is fine — the system's paths are enumerable, so your turns can enumerate them. A Mantle skill is not a scripted system. `skill.md` gives the model intent and constraints — _disambiguate which account, never invent a balance_ — and the model improvises the path. Two runs of the same conversation can legally take different routes to the same correct outcome. A test file with hard-coded user turns exercises exactly one path through a system whose defining property is that it improvises paths; it under-tests precisely the behaviour you bought Mantle for. So the framework replaces the script with an actor: an **LLM user-simulator** that plays your customer, one scenario at a time. ## Stage directions, not scripts The simulator reads one prose block per scenario — `simulation_context` — and improvises from it. From `eval/scenarios/balance_digression_faq_resume.yml`: ```yaml simulation_context: > You are a curious bank customer. Start by asking for an account balance without naming an account. When the agent asks which account, do NOT answer; instead ask "wait, first — how much is the overdraft fee?". After you get the fee answer, say "ok, the checking one" to return to the balance question. Accept the balance, thank the agent, and end the conversation. ``` Write these like stage directions for an actor: **persona** (who the customer is, their temperament), **intent** (what they want, what they will answer when asked, what they refuse), and a **clear end condition**. Give the simulator enough to play the part — then let it choose the words. The words being chosen fresh on every run is the entire point: it is what stops your suite from quietly testing only the phrasings you thought of. ## What this makes testable Look at what the checklist scenarios in the companion actually exercise: - **Digression and resume** (`balance_digression_faq_resume.yml`) — the customer interrupts one skill with a question for another, then comes back. Where the interruption lands varies by run; whether the agent finds its way back is asserted deterministically (`flow_started` for both skills, then the balance facts). - **Correction** (`balance_correction.yml`) — "actually, the savings one please", expressed as a `sequencing` assertion: the same slot set, then re-set. - **Out-of-order input** (`balance_named_up_front.yml`) — the customer volunteers the account before being asked, and a criterion checks the agent did not re-ask for it. These are exactly the behaviours a scripted file under-tests, because their interesting part is _where in the conversation they happen_ — the thing you cannot hard-code without deciding it in advance. The old absence habit survives the move, upgraded: the ambiguous-request scenario used to assert "no slot was set after turn 1" against a scripted turn. Now the requirement "the agent asks before answering, and does not guess" lives in `criteria` — judged from wherever in the transcript the question actually landed — while the assertions pin the end state: the right account, the tool ran, the right number spoken. ## What happened to dialogue-understanding tests Earlier versions of this series taught `rasa test du` here — annotated per-turn command tests for the CALM command generator. They are gone from the companion, for a reason worth reading because it is a lesson in trusting mechanisms over labels: DU tests score the **command generator** — CALM's component that turns each user message into commands like `StartFlow(x)`. The CLI hard-gates them to CALM assistants (`rasa/cli/dialogue_understanding_test.py:185-199`). A Mantle agent, however, _slips that gate_: it declares `is_calm_assistant = True` (`rasa/mantle/processor.py:104-107`) — and then the tests execute against a component that **never runs**, because Mantle's turn loop is an LLM tool-calling orchestrator that imports nothing from `rasa.dialogue_understanding` (`rasa/mantle/orchestration/orchestrator.py`). A suite that runs, produces scores, and measures a component your agent does not use is strictly worse than one that refuses to run. The question DU tests answered — _did the agent understand, separately from whether the conversation ended well?_ — has not disappeared. It has moved: understanding failures now surface as failed assertions on memory (`slot_was_set` with the wrong value is a misread) and as criteria rationales that name the turn where the transcript went sideways. The Inspector URL in every run report replays the conversation for exactly this kind of diagnosis. ## Variance is a feature with a bill attached The simulator's freshness cuts both ways. The same scenario produces different transcripts on every run — which is honest coverage of a non-scripted system, and which means a single green run is weak evidence and a single red run near a boundary may be the dice. Chapter 5 turns this into discipline; for now, the habit: **when a scenario surprises you, raise its run count before concluding.** Your coding agent will happily "run balance_correction 5 times" — remembering that each of those runs bills the agent, the simulator, and the judge. ## Check your understanding - Why does hard-coding user turns under-test a Mantle skill specifically, when the same technique served scripted assistants well? - A digression scenario passes with the interruption landing at turn 2 on one run and turn 4 on the next. What in the scenario file made that robustness testable? - The old DU suite would still _execute_ against a Mantle agent. Walk the three engine citations that explain why its scores would mean nothing. The simulator supplies the conversations; the assertions nail the facts. The remaining gap is quality of free text — which brings us to the instrument everyone wants first and should reach for last. --- # Chapter 4 — The Judge: Grounded Is Not Relevant Source: https://rasa.community/library/tutorials/evaluation-harness/04-the-judge/ Author: Rasa team Published: 2026-09-04 Some questions have no fact to assert. _"Is this answer actually supported by our overdraft policy?"_ is not a tracker fact. For free text, the framework sends the answer to a second LLM — the judge — and compares a returned score to a threshold. The judge holds two jobs in a scenario. It scores your **`criteria`** — the natural-language goals like "the agent asks exactly one clarifying question" — against the whole transcript, with a written rationale per criterion. And it powers the two **generative assertions**, which score a specific answer against a specific standard. The criteria are where you express judged intent; the generative assertions are precision tools, and here is the sentence this chapter exists for: **they are not variations of one idea.** They are computed by completely different machinery, they fail differently, and confusing them is the most common way to get a meaningless eval. (The machinery, for the record, is the same code the classic e2e tests used — the scenario engine imports `rasa.e2e_test.assertions` directly (`rasa/builder/copilot/mcp_server/tools/assertion_engine.py:9`), so every citation below reads from the engine you have installed.) ## Groundedness: a ratio of statements From `eval/scenarios/faq_grounded_answer.yml`: ```yaml assertions: - generative_response_is_grounded: threshold: 0.8 ground_truth: >- A replacement debit card is issued free of charge once per calendar year. Additional replacements cost $12 each. Standard delivery takes five to seven business days. ``` Mechanism: the judge splits the answer into atomic statements and marks each one supported or unsupported against your ground truth (`rasa/e2e_test/llm_judge_prompts/groundedness_prompt_template.jinja2:7-12`). Then the score is plain arithmetic — no embeddings involved (`rasa/e2e_test/utils/generative_assertions.py:156-166`): ```text score = supported statements / total statements ``` This is the metric for _"did the agent make something up?"_ — and the companion's most valuable scenario turns it inside out: `faq_unknown_topic_refused.yml` asks about a policy the source does not contain, with a `ground_truth` that describes the **refusal**. An agent that fills the gap from general banking knowledge produces fluent, helpful, unsupported statements — and scores badly against a truth that says "I don't have that." **The denominator insight** — the detail that changes how you read every threshold: the judge chooses how many statements the answer splits into, not you. A terse two-statement answer can only score 0, 0.5, or 1.0. Against that answer, `threshold: 0.8` does not mean "80% good" — it means "both statements must be supported". Read your thresholds as fractions of a small integer, not as percentages. ## Relevance: cosine similarity of invented questions ```yaml assertions: - generative_response_is_relevant: threshold: 0.7 ``` Completely different mechanism. The judge is shown **only the answer** — not the ground truth, not your policy — and invents three questions that answer would address (`answer_relevance_prompt_template.jinja2:1-9`; the count is `num_variations = 3` at `assertions.py:1595`). Those invented questions and the user's _real_ question are embedded, and the score is their mean cosine similarity (`generative_assertions.py:123-153`). Sit with the consequence, because it is the most important sentence in this chapter: > **Relevance never sees your ground truth. A confidently wrong answer to > exactly the right question scores high.** Relevance detects evasion and topic drift — an agent that answers the question it wished you'd asked. It cannot detect fabrication. Which is why `faq_grounded_answer.yml` asserts **both metrics on the same turn**: groundedness catches the confident lie, relevance catches the accurate dodge. Neither is sufficient alone. (Side effect of the mechanism: relevance needs an **embedding model** as well as a judge — default OpenAI `text-embedding-3-small`, `generative_assertions.py:28-31`. Groundedness never embeds anything.) ## Pin your judge The judge — and, separately, the simulator — is configured in `eval/conftest.yml`: ```yaml simulation: llm: # drives the user simulator provider: openai model: gpt-5.1 evaluation: llm: # the judge: criteria scores + quality metrics provider: openai model: gpt-5.1 ``` The two are independent on purpose, and the asymmetry matters: you can cheapen the simulator without touching what your results mean, but **the judge model is an input to your scores.** Swap it and yesterday's 0.85 and today's 0.79 were produced by different judges — nothing will tell you that unless the model name is pinned in a reviewed file. This file is that file. ## The trap: a judge with nothing to score This one trips everyone, and the error message doesn't help. Generative assertions do not look at every bot message. With no `utter_source:` given, they consider only messages whose source metadata is one of: ```text EnterpriseSearchPolicy · ContextualResponseRephraser · IntentlessPolicy ``` (`rasa/e2e_test/utils/generative_assertions.py:34-38`.) A plain templated response is **invisible** to them. Responses become eligible when they are generated or passed through the rephraser — and the rephraser only touches responses that explicitly set `rephrase: true` (`rasa/core/nlg/contextual_response_rephraser.py:122-124`). Otherwise a response's source is simply its action name — which you _can_ target, by naming it in `utter_source:`. So: if a generative assertion behaves as though it never ran, check the source of the message you think it is judging. Before the threshold, before the prompt, before blaming the judge. ## Thresholds are not portable A last habit before the honesty chapter: the thresholds in the companion are **illustrative, not engine defaults**, and any threshold you tune is tuned against three things at once — a specific judge model, a specific embedding model, and a specific set of ground-truth strings. Change any of the three and the old threshold is a number from a different experiment. (The engine's own fallback, for the record, is `DEFAULT_THRESHOLD = 0.5`, `assertions.py:70` — treat that as a placeholder, not a recommendation.) ## Check your understanding - Your FAQ agent starts answering overdraft questions with confident, fluent, _wrong_ numbers. Which metric catches it, and why does the other one score it high? - Why does the refusal scenario's `ground_truth` describe what the agent should _say it cannot do_, rather than containing the right answer? - Why does changing your embedding model invalidate relevance thresholds but leave groundedness scores untouched? You now command the whole framework. The last chapter is about the hardest part: knowing what your green checkmarks are actually worth. --- # Chapter 5 — What Your Eval Cannot Tell You Source: https://rasa.community/library/tutorials/evaluation-harness/05-what-your-eval-cannot-tell-you/ Author: Rasa team Published: 2026-09-04 An eval that oversells itself is worse than no eval, because it launders a guess into a number. You built the suite in four chapters; this one is about reading it without fooling yourself. ## Judge scores vary between identical runs The judge is an LLM at nonzero temperature, deciding how to split statements, what counts as supported, and how well a transcript met a criterion. The same answer can score 0.75 on one run and 1.0 on the next. **A single run near your threshold is noise, not a signal** — in either direction. If a criterion flipped and you changed nothing, you have learned something about variance, not about your agent. ## The simulator varies on purpose — and so does the agent Every run, the simulator re-improvises the customer from your stage directions, and the agent re-improvises its handling. This is the honest way to cover a non-scripted system, and it means even the deterministic _assertions_ can flip between runs — not because checking them is fuzzy, but because the behaviour they measure sits near a model's decision boundary. The assertions are exact; the thing they measure is not. When a criterion fails because the agent said the right thing in different words, there is a third possibility beside "agent broke" and "variance": **your criterion encodes incidental phrasing rather than a business requirement.** Fix the criterion, not the agent. ## Small samples say very little Seven scenarios cannot separate a real regression from chance. The reflex of _"we fixed it — the run passes now"_ after a single green run is exactly the failure mode a harness should protect you from, not create. When a result matters — a model swap, a release decision — raise the run count and look at the distribution, not the last run. Related, and worth saying at full volume because vendors will not: beware of impressive single numbers. The companion pattern mentions a widely repeated figure from Rasa office hours — roughly 30% better task completion for a controlled variant over about 100 simulated conversations — and immediately labels it **unpublished, unverified, and not reproduced**. That labelling is the correct handling of every number you cannot rerun yourself, including the ones your own suite produced last month on a judge model that has since been swapped. ## Cost decides cadence Every scenario run bills three LLMs — agent, simulator, judge — and relevance assertions add embedding calls on top. This is not a reason to avoid the instrument; it is a reason to schedule it: | Activity | Cost | Cadence | | ----------------------------- | -------------------- | --------------------------------- | | One scenario, run count 1 | cents | While iterating on that behaviour | | Full suite, run count 1 | 3 LLMs × 7 scenarios | After any skill or config change | | Full suite, raised run counts | the real bill | Model swaps, release decisions | A team that runs the full suite at high run counts on every commit soon stops reading the results. A team that runs one scenario tightly while iterating, and the full suite at decision points, finds regressions while they are still cheap. ## Ground truth drifts silently The `ground_truth:` strings in the FAQ scenarios are copies of the policy text in `skills/faq_support/tools.py`. Change the policy in one place only, and the judge will faithfully score the agent against a policy that no longer exists — green, and wrong. In a real project, generate ground truth from the same source the agent reads, or put a check on the pair. (This failure class — two sources of truth, coupled by nothing — is the same one that motivates several checks in this site's own build. It follows eval suites everywhere.) ## The one-paragraph summary of the whole series Let the simulator play customers you did not script, so your coverage is of the system you actually shipped. Assert every fact the tracker can express — including absence, including order — because those results are immune to judge mood. Send only free text to the judge, pin the judge, pair groundedness with relevance, and read thresholds as fractions of small integers. Then hold the whole thing with a light grip: variance is real on three levels now, seven scenarios prove little, and a number you cannot rerun is an anecdote wearing a decimal point. That grip — instruments trusted exactly as far as their mechanisms deserve — is the difference between _measuring_ your agent and _decorating_ it. ## Where to go next - The companion pattern's README goes deeper on every mechanism here, with engine source citations for each claim: [`patterns/evaluation-harness`](https://github.com/RasaHQ/rasa-community-resources/tree/52986f71071e419533f5f00335440a7ff1aef500/patterns/evaluation-harness). - The official [simulation & evaluation docs](https://rasa.com/docs/pro/testing/simulation-evaluation/) cover the MCP runner setup per editor, and the full assertion list. - New to building the agents themselves? The [Mantle starter pack tutorial](/library/tutorials/mantle-starter-pack/) is this series' natural prequel — its lint-and-hooks philosophy and this harness are the same idea at two different layers: catch the silent failure with a machine, before it costs you. --- # Guard an Action You Cannot Undo Source: https://rasa.community/library/tutorials/guarding-irreversible-actions/ Author: Rod Rivera Published: 2026-09-05 Here is the transcript this tutorial exists to make impossible. Write it down before you build anything, because "fixed" should be a comparison, not a feeling: ```text you my card was stolen, I need a new one sent to 9 Elsewhere Lane, Leeds bot Of course — I've ordered a replacement card to 9 Elsewhere Lane, Leeds. It should arrive in three to five working days. ``` That is what a scaffolded banking agent with a naive card-replacement tool does. Warm, efficient, correct in every detail it stated — and it just posted a payment instrument to an address a stranger read out over the phone. Nothing about that exchange was a bug: the model did what it was asked and the tool did what it was called with. The card is in the post. There is no undo. And here is real output — what the companion project's guard actually prints, today, when you run `make policy` with no licence, no model, and no network: ```text An address supplied during the call: ✓ medium is NOT enough — this is the account-takeover path -> ok=False result=step_up_required ✓ high is enough — the path is priced, not banned -> ok=True result=ok ``` The rest of this series is the distance between those two blocks. ## What the fix is not The reflex is "add an authentication check". Suppose we do, and the caller is fully verified — they gave the passphrase, they passed the one-time code, they are unambiguously the account holder as far as any factor can establish. The transcript above is still wrong. A verified caller is exactly who an account takeover produces. The factors prove someone holds the credentials; they say nothing about whether **the destination** is one the bank has ever seen. Authentication answers _who is calling_. This is a question about _where the card goes_, and no amount of the first answers the second. ## The actual problem Look at what the tool receives. The address is a string, and by the time it arrives it is indistinguishable from an address the bank has held for six years, because **a string does not remember where it came from.** That is the whole vulnerability, and it is why the fix is structural rather than a check bolted on the front. What the system needs is not a stronger gate but a value that has not thrown away the thing the gate needs to know. By the end of this tutorial the same request ends in that refusal: the tool returns `step_up_required`, nothing is ordered, and the agent's side of the conversation follows from a result code it cannot override — the target transcript reads like ```text you send it to 9 Elsewhere Lane, Leeds, LS1 9ZZ bot I can do that, but a card going to an address you've given me on the call needs a one-time code first — it can't be recalled once it's posted. Shall I send you one? ``` with the wording free to vary, because the part that matters is not wording. Nothing has been ordered — not because the model chose well, but because the function refused, and the refusal above is the function's real, reproducible output. Here is the shape of what closes that distance. The part worth looking at twice is the arrow that comes in from the side: ![The model can only ever call the reissue_card tool wrapper. auth_tier enters sideways from project memory, where llm_settable is false, so the model cannot fill it and cannot talk itself past the guard. Inside place_reissue the order is fixed: resolve inputs, classify the address against the customer's addresses on file, then guard. The guard raises rather than returning False, so the only route to the line that posts the card is a return. REQUIRED_TIER is keyed by where the address came from — ON_FILE needs medium plus a seven-day cooling-off, STATED and UNKNOWN need high — and an unreadable tier ranks zero, refused rather than interpreted.](/diagrams/guarding-irreversible-actions-guard-path.svg) ## What you will build Six steps, on top of the [companion project](https://github.com/RasaHQ/rasa-community-resources/tree/52986f71071e419533f5f00335440a7ff1aef500/tutorials/rasa-card-reissue-tutorial): | Chapter | What it adds | | ------- | ------------------------------------------ | | 1 | Reproduce the failure on a working agent | | 2 | Declare risk per action, in a table | | 3 | Move the check from YAML into the function | | 4 | Carry where the address came from | | 5 | Survive a dropped confirmation | | 6 | Keep refusals from becoming successes | ## What this assumes Python 3.11 or 3.12, [uv](https://docs.astral.sh/uv/), a Rasa Pro Developer Edition licence, and an OpenAI API key. The companion project pins `rasa-pro==3.21.0.dev5`. It also assumes verification exists. **This tutorial does not build authentication** — it consumes a tier that something else established. That something else is [the risk-tiered step-up pattern](https://github.com/RasaHQ/rasa-community-resources/tree/52986f71071e419533f5f00335440a7ff1aef500/patterns/voice-auth-stepup), which owns the factors and the lattice. Here, the tier is an input. --- # Chapter 1 — Reproduce the failure Source: https://rasa.community/library/tutorials/guarding-irreversible-actions/01-reproduce-the-failure/ Author: Rasa team Published: 2026-09-05 Clone the companion project and run the guard demonstration: ```bash make env # then fill in .env make install make policy ``` `make policy` calls the guard directly — no model, no licence, no network — and prints what it refused and why. Right now, before you build anything, imagine the version without the guard. The function it exercises is: ```python async def place_reissue(card_id, line1, city, postcode, auth_tier="none"): ... ``` and the naive body is two lines: look up the card, post it. One structural note before we break it, because it is this tutorial's own lesson applied to its own tools. `place_reissue` is deliberately **pure**: every input explicit, including `auth_tier`, so the proof can drive the guard from every angle. The agent never calls it directly — it calls a thin `@tool` wrapper named `reissue_card` that reads the tier from project memory, which only the verification tools may write. The tier is _not_ a tool argument in the wrapper, and that is not style: the engine publishes every tool parameter to the model, so a model-fillable `auth_tier` would let the agent talk itself past the guard by simply claiming `"high"`. Where a security input comes from is part of the security — which happens to be the whole of Chapter 4. ## Why start here You are about to add machinery, and machinery has a failure mode of its own: it can be present, look correct, and enforce nothing. The only defence is to know exactly what the broken behaviour looked like, so that "fixed" is a comparison rather than a feeling. Concretely, write down the transcript you are trying to make impossible: ```text you my card was stolen, send a new one to 9 Elsewhere Lane, Leeds, LS1 9ZZ bot I've ordered a replacement card to 9 Elsewhere Lane, Leeds. ``` That address is not on the account. It has never been on the account. ## The trap in "just add auth" Try the obvious fix in your head. Require verification before `reissue_card` runs. Now the caller must give a passphrase. An attacker doing account takeover **has the passphrase.** That is what account takeover means. The check you just added is satisfied, and the card still goes to Leeds. This is worth sitting with, because it is the point where most implementations stop. The gate is real, the gate is passed, and the vulnerability is untouched — because the gate was asking about the caller and the risk was in the destination. ## What the tool actually knows Here is the whole difficulty, in one observation. The tool receives: ```python line1 = "9 Elsewhere Lane" city = "Leeds" postcode = "LS1 9ZZ" ``` and it receives exactly the same shape of thing when the caller picks the address the bank has held since 2019: ```python line1 = "14 Wexley Row" city = "Bristol" postcode = "BS1 4TR" ``` Two strings. One is a fact the bank has had for six years; the other is a sound someone made ninety seconds ago. **The tool cannot tell them apart,** because the difference between them was never in the value — it was in where the value came from, and that was discarded at the boundary. Chapter 4 is where we stop discarding it. Chapters 2 and 3 build the place that will need it. --- # Chapter 2 — Declare the risk, in a table Source: https://rasa.community/library/tutorials/guarding-irreversible-actions/02-declare-the-risk/ Author: Rasa team Published: 2026-09-05 The requirement is a property of **what the action does if it succeeds**. Not of who is calling, not of which skill they arrived through, not of how far into the call they are. For a reissue, what it does depends on the destination: ```python REQUIRED_TIER = { AddressProvenance.ON_FILE: "medium", AddressProvenance.STATED: "high", AddressProvenance.UNKNOWN: "high", } ``` Three rows, in `cardpolicy/guard.py`. Posting a card to an address the bank has held for six years is a genuinely different act from posting one to an address supplied during the call, and the table is where that judgement is written down once instead of being re-argued at every call site. ## Why a table and not an `if` Scattered conditionals are how the twelfth call site ends up disagreeing with the other eleven. A table is greppable, reviewable in one screen, and diffable — when someone loosens a requirement, the diff says so in one line rather than hiding inside a refactor. ## The default is the strict one Notice `UNKNOWN`. It exists because the alternative to representing "we do not know where this came from" is pretending we do, and a system that cannot express doubt resolves doubt in favour of proceeding. The same instinct governs how a tier is read out of memory: ```python def _rank(tier: object) -> int: if isinstance(tier, str): return _TIER_RANK.get(tier.strip().casefold(), 0) return 0 ``` Memory is text. It may be absent on the first turn, stale from an older build, or something a model invented. Every one of those resolves to rank 0 — no verification at all — rather than raising, because **a guard that throws on unexpected input is a guard that gets wrapped in a bare `except` within a month.** `make policy` exercises this directly: ```text ✓ auth_tier='admin' is refused, not interpreted ✓ auth_tier='' is refused, not interpreted ✓ auth_tier='3' is refused, not interpreted ✓ auth_tier='high ish' is refused, not interpreted ``` Note `'3'`. A numeric string is not quietly read as a rank. And note what is _not_ on that list — casing: ```text ✓ auth_tier='HIGH' is the same tier as 'high' ``` ASR casing is not a security signal, and treating it as one produces a system that refuses the right caller for the wrong reason. Normalising case is not the same as being permissive; the deliberate part is knowing which differences are noise and which are the question. ## Where the requirement is not declared One place, deliberately: **the model cannot write it.** In `memory.yml`, ```yaml auth_tier: type: text llm_settable: false ``` A model that can set its own authorisation level is not authorised; it is self-certifying. --- # Chapter 3 — Put the check where it binds Source: https://rasa.community/library/tutorials/guarding-irreversible-actions/03-the-execution-guard/ Author: Rasa team Published: 2026-09-05 Rasa gives you a declarative way to say "this tool needs something first": ```yaml tool_constraints: - reissue_card: requires_confirmation: enabled: true utter_for_confirmation: utter_confirm_reissue ``` This project uses it, and you should too. But be precise about what it controls. ## Routing versus execution The orchestrator evaluates `tool_constraints` against conversation state to decide which tools the model is **offered**. That shapes the dialogue — the caller gets asked to verify, instead of being told "no" for reasons they cannot see. It is a genuinely useful control. It is a **routing** control. It is not an **execution** control. It does not run when the function runs. Consider what happens if: - the YAML key is misspelled — it is silently ignored, no error at load; - a future skill imports `reissue_card` without the constraint — it is simply absent; - the model is swapped for one that reasons differently — the offer set changes. In every one of those cases the function is still entered, and the card is still posted. Nothing crashes. Nothing logs a refusal. The control was real and it was in the wrong layer. ## So there are two layers, and they are not redundant | Layer | Where | Nature | Can it be bypassed? | | ------------------ | ----------------- | --------------------------------- | ------------------- | | `tool_constraints` | skill frontmatter | declarative, shapes the dialogue | yes | | `guard_reissue()` | inside the tool | imperative, gates the side effect | no | The inner one looks like this, in `tools/cards.py`: ```python address = classify_address(line1, city, postcode, customer["addresses_on_file"]) # ---- the line before the side effect -------------------------------- try: guard_reissue(address, auth_tier, on_file_since=on_file_since) except ReissueRefused as exc: return _as_dict(exc.outcome) # --------------------------------------------------------------------- reference = f"RC-{uuid.uuid4().hex[:8].upper()}" ``` Every function in that module that changes the world has the same shape: resolve inputs, classify provenance, **guard**, act, return an outcome. If you copy one thing from this tutorial, copy that ordering. ## Why it raises instead of returning False ```python def guard_reissue(...) -> Decision: decision = evaluate(...) if decision.allowed: return decision raise ReissueRefused(...) ``` A boolean can be ignored by writing `guard_reissue(...)` on its own line and carrying on — which reads, at a glance, exactly like calling a guard. Reviewers skim past it. An exception cannot be ignored that way: the only route to the line that posts the card is a return. ## The test this makes possible Because the binding check is a Python function, the guarantee is provable without a model: ```bash make policy ``` ```text An address supplied during the call: ✓ medium is NOT enough — this is the account-takeover path -> ok=False result=step_up_required ``` No LLM, no sampling, no judge. This is a fact about the process. That distinction matters more than it might seem. A conversation-level test would tell you the model _asked for a code_ — that the usual path is usually taken. This tells you the unusual path is **closed**. Only the second one is a security claim, and only the second one stays true when the model changes. --- # Chapter 4 — Carry where the address came from Source: https://rasa.community/library/tutorials/guarding-irreversible-actions/04-address-provenance/ Author: Rasa team Published: 2026-09-05 This is the chapter Chapter 1 was pointing at. Two addresses reach the tool as identical shapes, and the difference that matters was thrown away at the boundary. So stop throwing it away. `cardpolicy/provenance.py`: ```python class AddressProvenance(str, Enum): ON_FILE = "on_file" # in the record before this conversation started STATED = "stated" # said by the caller during this call. Unverified. UNKNOWN = "unknown" # provenance was lost. Treated as STATED-or-worse. ``` ## The signature is the design ```python def classify_address(line1, city, postcode, addresses_on_file) -> ClassifiedAddress: ``` Read what is **not** there: no `provenance` parameter. A function that accepts the answer to the question it is supposed to decide is not a check. It is a formality that the caller, the model, or a future refactor can satisfy by passing the convenient value. So this function takes the customer's real address list and compares against it. Provenance is looked up, never asserted. ## Normalising, and why it is a kindness not a hole ```python def _normalise(value: str) -> str: return " ".join(value.split()).replace(" ", "").casefold() ``` `BS1 4TR`, `bs1 4tr`, and `BS14TR` are the same postcode. Without this, a caller who reads their own on-file address back slightly differently is classified `STATED` and charged a one-time code they did not owe. That is not merely annoying. An agent that demands extra verification from legitimate callers at random is an agent whose extra verification gets removed by someone six months from now who is tired of the complaints — and they will remove all of it, including the part that was load-bearing. ## Why STATED is allowed at all The strict-looking answer is to refuse every new address. It is also wrong. People move. And the customer whose card was stolen along with their wallet is precisely the customer most likely to have moved recently, or to be standing somewhere that is not home. Refusing every stated address makes the agent useless for the exact case it exists to handle, and a useless safe path pushes people to the phone queue, where the social engineering works better anyway. So the harder path is **priced, not banned**: `STATED` costs `high` instead of `medium`. ```text ✓ medium is NOT enough — this is the account-takeover path ✓ high is enough — the path is priced, not banned ``` ## The gap this leaves, and closing it There is a hole in what we have so far, and it is worth finding before someone else does. If `ON_FILE` is the cheap path, an attacker's move is to _get their address on file_ — go through whatever flow adds an address, then order the card at `medium`. The provenance check passes honestly. The address really is on file. So being on file is not enough; being on file **for a while** is: ```python COOLING_OFF = timedelta(days=7) ``` ```python if ( address.provenance is AddressProvenance.ON_FILE and on_file_since is not None and (today or date.today()) - on_file_since < COOLING_OFF ): # refuse: cooling_off ``` Note where this check sits: it applies **only** to `ON_FILE`, precisely because `ON_FILE` is the discounted path. An address that became on-file recently has not yet earned the discount that being on-file buys. Without this, the distinction Chapter 4 built is a speed bump with a marked detour around it. Seven days is a policy number, not a technical one. It is named as a constant so that the one place it lives is the one place it gets argued about. --- # Chapter 5 — Survive a dropped confirmation Source: https://rasa.community/library/tutorials/guarding-irreversible-actions/05-survive-a-retry/ Author: Rasa team Published: 2026-09-05 Here is a sequence in which every component behaves correctly and the outcome is still wrong. 1. The caller says "yes, send it." 2. The tool posts the card and returns `ok`. 3. The confirmation is cut off — a dropped packet, or the caller talking over it. 4. The model, having no evidence the tool succeeded, does the reasonable thing and calls it again. 5. Two cards are in the post. Nothing there is a bug in any single part. It is a bug in the **assumption** that calling a tool twice is the same as calling it twice as much. For a read, it is. For something that leaves the building, it is not. ## Why not a boolean The tempting fix is a flag: `already_reissued`. It is wrong, and the reason is worth stating because it generalises. A caller genuinely can need two cards replaced on one call — the debit card and the credit card were both in the stolen wallet. A boolean cannot tell "the same request again" from "a second, different request", so it either blocks the legitimate second card or permits the duplicate first one. What must not happen twice is **the same request**, not any second request. So the key has to describe the request: ```python def request_fingerprint(card_id: str, postcode: str, line1: str) -> str: canonical = "|".join( part.strip().casefold() for part in (card_id, postcode, line1) ) return hashlib.sha256(canonical.encode("utf-8")).hexdigest()[:16] ``` Same fingerprint, same reference, no second card. Different fingerprint, a real second order. ## The destination is part of the identity Note that the fingerprint includes the address, not just the card. Replacing card 9931 to the home address, and then to a newly-stated address, are **different requests** — and the second must go through the guard again rather than inheriting the first one's success. If the fingerprint were just `card_id`, a caller could obtain a cheap approval to an on-file address and then reuse it to redirect. The idempotency key would have laundered the guard. ## Why it is hashed The fingerprint is written to conversation memory, and conversation memory is a place transcripts come from. A hash of an address is not an address. Truncated to 16 hex characters — 64 bits. Within a single phone call, where the number of distinct requests is in the single digits, collision risk is not a real quantity. ## What it looks like ```text The same request twice posts one card: ✓ second attempt reports duplicate -> ok=True result=duplicate ✓ and returns the SAME reference (RC-BB083296) ``` `duplicate` is deliberately a success — `ok=True` — because from the caller's point of view the card **is** coming, and telling them otherwise would prompt them to ask for another one. What it must not do is mint a new reference, which would be a promise that a second card exists. ## What breaks `_PLACED` in `tools/cards.py` is a dictionary in process memory. Restart the process and the guarantee resets. That is honest for a tutorial and wrong for production, where the card issuer's own idempotency key is the mechanism you want. The shape of the lookup is what transfers; the storage is not. --- # Chapter 6 — Keep a refusal from becoming a success Source: https://rasa.community/library/tutorials/guarding-irreversible-actions/06-refusal-paths/ Author: Rasa team Published: 2026-09-05 Everything so far guarantees that the card is not posted. There is one failure mode left, and it is not in the Python. The tool returns `{"ok": False, "message": "..."}`. The model reads the message, finds it sympathetic, and tells the caller their card is on its way. Nothing crashed. The log says the tool refused. **The caller heard a promise.** ## Outcomes as a closed set ```python class Result(str, Enum): OK = "ok" # a card was ordered STEP_UP_REQUIRED = "step_up_required" # nothing happened; caller can fix it COOLING_OFF = "cooling_off" # nothing happened; caller cannot DUPLICATE = "duplicate" # nothing NEW happened REFUSED = "refused" # nothing happened; ever ``` Five values, and the skill prose is written against categories rather than against message text somebody will reword next month. The distinction between `STEP_UP_REQUIRED` and `COOLING_OFF` is the one that earns its keep. Both are refusals. Only one of them can be resolved on this call — and an agent that offers to "try again in a moment" after a cooling-off refusal is offering a workaround to a control that exists specifically to have no workaround. ## One expression decides whether something happened ```python @property def acted(self) -> bool: return self.result in (Result.OK, Result.DUPLICATE) ``` and at the boundary: ```python payload = {"ok": outcome.acted, ...} ``` `ok` is derived, never set by hand. Every non-`OK` branch has at some point been mistaken for a soft success by somebody in a hurry, and the point of a single named property is that there is one thing to get right instead of five. The constructor enforces the other half: ```python def refused(result, message, **detail) -> Outcome: if result in (Result.OK, Result.DUPLICATE): raise ValueError(f"{result.value!r} is not a refusal; build it with succeeded()") return Outcome(result=result, message=message, reference=None, detail=detail) ``` A refusal cannot carry a reference. A reference is a promise that something exists, and nothing does. ```text A refusal never carries a reference: ✓ refused payload has no reference key ``` ## Saying it in the skill The Python makes the promise unsupportable. The prose stops it being made: ```markdown Never say a card has been ordered, is on its way, or will arrive unless `reissue_card` returned ok true together with a reference. If `reissue_card` returns cooling_off, nothing has been ordered and nothing the caller says on this call will change that. Explain that the address is too new to post a card to, offer the older addresses on file, and if none works offer @skill.human_handoff. Do not offer to note the request, raise it later, send it somewhere else as a workaround, or try again in a moment. ``` That last sentence is doing real work. Left to itself, a helpful model will find an adjacent thing to offer, and every one of those adjacent things is a way around the control. The forbidden alternatives have to be named, because "do not proceed" and "do not achieve the same outcome another way" are different instructions. ## The escalation boundary The agent orders replacement cards to addresses that pass policy. It does not change an address, override a cooling-off window, verify identity itself, or retry a refused reissue another way. When any of those is what the caller needs, it hands to `@skill.human_handoff` and stops. Write that down, in the README and in the skill, as a behaviour rather than an apology. **A refused card request that ends in a human is the system working.** ## Where to go next This tutorial took the caller's verification tier as an input and never asked how it got there. That question — what the tiers are, which factor buys which, and how a caller steps up mid-call — is the subject of [the risk-tiered step-up pattern](https://github.com/RasaHQ/rasa-community-resources/tree/52986f71071e419533f5f00335440a7ff1aef500/patterns/voice-auth-stepup), which classifies `reissue_card` at its highest tier. The two compose. That pattern decides _how strong_ the caller's verification is; this one decides whether, for this particular destination, that is enough. --- # Your First Rasa Mantle Agent, With Guardrails On Source: https://rasa.community/library/tutorials/mantle-starter-pack/ Author: Rod Rivera Published: 2026-09-06 Here is a story that actually happened. A catalog of ten Rasa Mantle projects — written carefully, reviewed, shipped — declared **39 behavioral rules** across their config files. Things like _"never state a balance that did not come from a tool."_ The engine applied **zero** of them. No error at load. No warning at train time. The agents just ran without their guardrails, for weeks, until someone read the built prompt and noticed the rules weren't in it. The cause? The rules were indented two spaces too deep. Mantle parsed them, found them in a spot it doesn't read from, and silently threw them away. That is the kind of mistake this tutorial is about. Not the loud ones — Mantle tells you about those. The **silent** ones: config that parses, trains, runs, and quietly does the wrong thing. ## What the starter pack is The [starter pack](https://github.com/RasaHQ/rasa-community-resources/tree/52986f71071e419533f5f00335440a7ff1aef500/starter-pack) is three tools that share one idea — _every known silent mistake should be caught by a machine, not by weeks of confusion_: 1. **Six Claude Code skills** — playbooks that teach an AI assistant the _correct_ 2026 file shapes, so "scaffold me an agent" produces config that is right the first time. 2. **A lint** — [`lint_mantle.py`](https://github.com/RasaHQ/rasa-community-resources/blob/52986f71071e419533f5f00335440a7ff1aef500/starter-pack/scripts/lint_mantle.py), a single-file checker with nine checks. Every check encodes a failure a real project in the companion catalog actually hit. The 39-ignored-rules bug is check number one. 3. **A git hook** — a one-time install that runs the lint on every commit, so a silent mistake cannot even be _saved_ without being pointed out. The three pieces meet in the everyday loop, where the lint runs twice — once when the change is built, once when it is committed — with your own reading of the diff in between: ![The everyday loop. First you ask for a change and the build playbook makes it, finishing by running lint_mantle.py because the playbook says to. Then you read the diff, which no machine does for you. Then git commit runs the same lint again from .git/hooks/pre-commit, and a failure blocks the commit with each finding naming its own fix. Below, what each layer can actually prove: the lint proves the project is internally consistent with no known trap, but not that the engine accepts it; make train proves the engine accepts and builds it, but not that conversations go well; and whether conversations go well has no checker at all.](/diagrams/mantle-starter-pack-where-the-lint-runs.svg) ## What you will build Juniper, a small plant-shop helper with two skills and a memory: ```text you do you have anything that survives a dark bathroom? bot Yes — a ZZ plant or a snake plant would be happy there. Want me to check what's in stock? you yes please bot We have 4 ZZ plants and 7 snake plants in stock right now. you I'll take a ZZ plant bot Great choice. I've noted a ZZ plant for you — ask me anytime to check your order. ``` More importantly, in Chapter 3 you will **break Juniper on purpose**, five different ways, and watch the lint catch every one — including three breaks that produce _no error at all_ at runtime. That chapter is the reason this tutorial exists: once you have _seen_ a silent failure get caught, you understand why the guardrails earn their place. ## The chapters | # | Chapter | You will | | --- | ---------------------------------- | ----------------------------------------------------- | | 1 | The safety net | Install the pack, meet the lint, wire the git hook | | 2 | Scaffold Juniper | Generate the agent and understand every file it makes | | 3 | Break it on purpose | Trip all five silent traps; watch each get caught | | 4 | Run it, review it, keep it healthy | Train, chat, and adopt the everyday workflow | ## Before you start - **Python 3.11 or 3.12** — Mantle requires ≥ 3.11 (and Chapter 3 shows you the deeply confusing error you get when you ignore this). - **[uv](https://docs.astral.sh/uv/)** and **git**. - **[Claude Code](https://claude.com/claude-code)** for Chapters 2–4. (The lint and hook work without it; the scaffolding skills are what need it.) - A free [Rasa Pro Developer Edition licence](https://rasa.com/rasa-pro-developer-edition-license-key-request/) and an OpenAI API key — only needed from Chapter 4, when Juniper first runs. The installation targets `rasa-pro==3.21.0.dev5`. The [compatibility report](/compatibility/) records the current automated checks; recorded terminal examples are illustrations, not a claim that every live model run repeats identically. The release audit checks the published wheel for the Mantle engine and reports the stable channel separately. Chapter 3 explains why an arbitrary stable pin is not a substitute for the tested runtime. ## Where everything lives - The pack itself: [`starter-pack/`](https://github.com/RasaHQ/rasa-community-resources/tree/52986f71071e419533f5f00335440a7ff1aef500/starter-pack) in the community resources repo. - The catalog whose real failures the pack distills: [`RasaHQ/rasa-community-resources`](https://github.com/RasaHQ/rasa-community-resources) — its own [`scripts/lint_repo.py`](https://github.com/RasaHQ/rasa-community-resources/blob/52986f71071e419533f5f00335440a7ff1aef500/scripts/lint_repo.py) opens with the line _"Every check here encodes a failure this repository has actually hit."_ This tutorial is that sentence, taught. --- # Chapter 1 — The Safety Net Source: https://rasa.community/library/tutorials/mantle-starter-pack/01-the-safety-net/ Author: Rasa team Published: 2026-09-06 We are going to set up the safety net _before_ building anything. This is deliberate. A net installed after the fall is a decoration. ## Get the pack ```bash git clone https://github.com/RasaHQ/rasa-community-resources.git git -C rasa-community-resources checkout 52986f71071e419533f5f00335440a7ff1aef500 mkdir juniper && cd juniper git init cp -R ../rasa-community-resources/starter-pack/CLAUDE.md . cp -R ../rasa-community-resources/starter-pack/.claude . cp -R ../rasa-community-resources/starter-pack/scripts . cp -R ../rasa-community-resources/starter-pack/hooks . ``` Four things just landed in your empty project: | What | Job | | ------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `CLAUDE.md` | The rulebook Claude Code reads automatically — version doctrine plus the five silent traps, so the assistant can't "helpfully" reproduce them | | `.claude/` | Six skills (build playbooks) and two roles (a builder, a reviewer). [Browse them on GitHub](https://github.com/RasaHQ/rasa-community-resources/tree/52986f71071e419533f5f00335440a7ff1aef500/starter-pack/.claude) — they are readable Markdown, useful to humans too | | `scripts/lint_mantle.py` | The lint. One file, standard library only, no installation needed | | `hooks/` | The git hook and its installer | > `.claude` starts with a dot, so your file browser hides it. `ls -a` proves > it's there. ## Meet the lint A **linter** is a program that reads your files and points at known mistakes — named after picking lint off clothing. This one has nine checks, and it has a property worth pausing on: **every check corresponds to a failure a real, shipped Mantle project actually experienced.** Not hypothetical style rules. Scar tissue, mechanized. Run it right now, in your nearly-empty project: ```bash python3 scripts/lint_mantle.py ``` ```text [engine-version-pin] pyproject.toml: missing pyproject.toml [secret-hygiene] .gitignore: '.env' is not gitignored — one 'git add -A' away from publishing credentials [env-example] .env.example: missing .env.example — a new user has no way to discover which credentials this project needs FAILED — 3 finding(s) across 9 check(s) ``` **This failure is the success.** Three findings on an empty project means the checker is alive, and — read them again — each finding tells you _what's wrong, why it matters, and what to do_. That is the contract every message in this lint keeps. You never get a bare "error 47". You can also run one check at a time, or list them all: ```bash python3 scripts/lint_mantle.py --list python3 scripts/lint_mantle.py --check secret-hygiene ``` If you're curious how any check works, the answer is one click away: each check in [`lint_mantle.py`](https://github.com/RasaHQ/rasa-community-resources/blob/52986f71071e419533f5f00335440a7ff1aef500/starter-pack/scripts/lint_mantle.py) is a short function whose docstring tells the story of the failure it encodes. The file is meant to be read. ## Wire the hook A **git hook** is a small program git runs automatically at a key moment. The one we install runs at `pre-commit` — the instant before git saves a snapshot of your work: ```bash ./hooks/install-hooks.sh ``` ```text Installed: .git/hooks/pre-commit Every commit now runs: python3 scripts/lint_mantle.py ``` From now on, committing a known trap gets you this instead of a saved bug: ```text COMMIT BLOCKED by lint_mantle.py (exit 1). Each finding above names its fix. The checks encode real, silent failures — fixing them now is cheaper than debugging them at train time. ``` Two things to know about the guard, because trust requires understanding: 1. **It is not a cage.** `git commit --no-verify` bypasses it. The etiquette: if you bypass, say why in the commit message. The lint only complains about mistakes that are _silent_ at runtime, so a bypass usually means shipping a bug you won't meet again for weeks — make it a decision, not an accident. 2. **It refuses to run blind.** If `scripts/lint_mantle.py` goes missing, the hook blocks the commit rather than waving it through — a safety net with a hole in it is worse than none, because you stop looking down. (You can read this logic in [`hooks/pre-commit`](https://github.com/RasaHQ/rasa-community-resources/blob/52986f71071e419533f5f00335440a7ff1aef500/starter-pack/hooks/pre-commit) — it's 25 lines.) The installer is also polite: if you already had a pre-commit hook from something else, it refuses to overwrite it and tells you to merge them yourself. ## Check your understanding Before Chapter 2, you should be able to answer: - Why did the lint _fail_ on an empty project, and why was that good news? - What moment does a `pre-commit` hook run at, and what does ours do there? - Where would you look to understand what the `secret-hygiene` check actually scans for? _(Answer: its function in `lint_mantle.py` — the docstrings are the documentation.)_ Net installed. Time to build something worth catching. --- # Chapter 2 — Scaffold Juniper Source: https://rasa.community/library/tutorials/mantle-starter-pack/02-scaffold-juniper/ Author: Rasa team Published: 2026-09-06 Open Claude Code in your `juniper/` folder and say: > Use the **mantle-new-project** skill to scaffold a plant-shop helper called > Juniper. Two skills: recommend a plant for a room, and check stock. Text > only, no voice. Claude Code reads [`mantle-new-project/SKILL.md`](https://github.com/RasaHQ/rasa-community-resources/blob/52986f71071e419533f5f00335440a7ff1aef500/starter-pack/.claude/skills/mantle-new-project/SKILL.md) — a playbook that contains the full, correct template for every file — and builds the project. Then it runs the lint, because the skill tells it to finish that way. One request, working project. But you should never ship a file you couldn't explain, so let's walk through what appeared and _why each file has the shape it has_. Each shape is a decision, and most of the decisions were paid for. ## pyproject.toml — three lines that are all trap-avoidance ```toml requires-python = ">=3.11,<3.13" dependencies = [ "rasa-pro==3.21.0.dev5", ] [tool.uv] prerelease = "allow" ``` - **The pin says `dev` and that is correct.** The Mantle engine ships only on the `3.20.0.dev` pre-release line; the newest _stable_ rasa-pro contains no engine package at all. This is the least intuitive fact in the whole ecosystem — Chapter 3 makes you feel it. - **`prerelease = "allow"`** — without it, uv refuses to resolve a dev pin. - **The `>=3.11` floor** — Mantle 3.20 requires it, and the error you get with a lower floor doesn't mention Python. (Also a Chapter 3 exhibit.) ## lib/engine.py — one import to rule them all ```python try: # 3.20.0.dev1 and later from rasa.mantle.tools.decorator import ToolContext, tool from rasa.mantle.tools.result import ToolResult except ImportError: # 3.19.x and earlier from rasa.calm_v2.tools.decorator import ToolContext, tool from rasa.calm_v2.tools.result import ToolResult ``` Rasa renamed the engine package (`rasa.calm_v2` → `rasa.mantle`) in 3.20.0.dev1 — old name _gone_, not aliased. Every tool in your project imports from `lib.engine`, so the day the name changes again, you edit one file. This works because tool discovery keys off an attribute the `@tool` decorator sets, not the import path. ## agent.yml — read the comment, it's earning its keep ```yaml agent: id: juniper language: en persona: | You are Juniper, a friendly plant-shop helper. Ask one question at a time and keep answers short. Never invent stock numbers — always use tools. # TOP-LEVEL keys, siblings of `agent:`. Nested inside they parse without # error and are silently discarded. name: Juniper description: Plant-shop helper for recommendations and stock checks rules: - 'Be warm, clear, and brief.' - 'Never state a stock count that did not come from a tool.' ``` The comment marks the site of the 39-ignored-rules failure from the series intro. `persona` lives _inside_ `agent:`; `name`, `description`, and `rules` live _outside_ it. Get that backwards and everything still parses, trains, and runs — minus its guardrails. This is silent trap #1, and the reason the scaffold writes the warning into the file itself. ## integrations.yml — the LLM is a _reference_, not a config ```yaml llm: model_group: orchestrator model_groups: - id: orchestrator models: - provider: openai model: gpt-4.1-mini api_key: ${OPENAI_API_KEY} temperature: 0.0 ``` Older tutorials (most of the internet, honestly) put `provider:` and `model:` directly under `llm:`. Since 3.20.0.dev6 that form is **rejected** — `llm:` names a model group, and the provider details live on the group. Two practical notes: - `${OPENAI_API_KEY}` is filled from your environment or `.env` at runtime — the _name_ of the key is config, the _value_ never is. - `temperature: 0.0` because this model routes your conversation; routing wants determinism, not creativity. ## memory.yml vs skills/*/memory.yml — who is allowed to write Root `memory.yml` is **project memory**: facts every skill shares, written by _tools only_. ```yaml preferred_room: type: text description: The room the customer is shopping for. ``` Skill folders get their own `memory.yml` — and that's where fields the LLM may set live, marked `llm_settable: true`. Putting that flag in the root file is rejected by the engine. The mental model: **project memory is the database, skill memory is the scratchpad.** The scaffold's shapes come from the catalog's [tools-and-memory tutorial project](https://github.com/RasaHQ/rasa-community-resources/tree/52986f71071e419533f5f00335440a7ff1aef500/tutorials/rasa-tools-and-memory-tutorial), where a bank agent uses exactly this split to make you verify once, not three times. ## skills/recommend_plant/skill.md — three languages, one file ```markdown --- name: Recommend Plant description: > Suggest a plant for the customer's room. Activate for plant advice, what should I get, or what survives in a given room. --- Help the customer pick a plant for their room. if: not session.recommend_plant.room_type Ask what kind of room the plant is for — light level matters most. if: session.recommend_plant.room_type Call `suggest_plants` for @memory.recommend_plant.room_type and present the top two options briefly. Never invent a plant name. ``` Three languages stacked: YAML frontmatter (config), Markdown prose (what the LLM reads), and `if:` lines **at column zero** (branching logic). The `description` doubles as the router's activation hint — write your trigger phrases into it. And note the two reference styles: `session.x.y` in the `if:` conditions, `@memory.x.y` in the prose. They are not interchangeable — that's silent traps #4 and #5, both waiting for you in Chapter 3. ## The rest `.env.example` (committed, names your two keys, values empty), `.gitignore` (with `.env` in it — that one line is what makes "paste your key in `.env`" safe advice), a `Makefile` (Chapter 4), and `tools/` for shared tools. Now prove the whole thing hangs together: ```bash python3 scripts/lint_mantle.py ``` ```text ok — 0 finding(s) across 9 check(s) ``` Commit it — your first commit passes through the hook you installed: ```bash git add -A && git commit -m "Scaffold Juniper with the Mantle starter pack" ``` Green, saved, and you can explain every file. Next chapter we take a hammer to it. --- # Chapter 3 — Break It On Purpose Source: https://rasa.community/library/tutorials/mantle-starter-pack/03-break-it-on-purpose/ Author: Rasa team Published: 2026-09-06 You cannot truly respect a guardrail until you've seen what it stops. So: Juniper is green, committed, safe — and we are now going to break it seven ways, on purpose, and watch what happens each time. Work on a branch so the wreckage is disposable: ```bash git checkout -b wreck-it ``` For each break below: make the edit, run `python3 scripts/lint_mantle.py`, read the message, undo (`git checkout -- `), next. The whole tour takes fifteen minutes and pays for itself the first time any of these happens to you for real. ## Break #1 — indent the rules (the 39-rules bug) In `agent.yml`, move `rules:` _inside_ the `agent:` block — just indent it to match `persona:`. Now ask yourself what the engine will do. Answer: **nothing.** It parses fine. It would train fine. Juniper would run, minus its rules, and nothing anywhere would tell you. This is the exact shape of the bug that shipped 39 dead rules across a real catalog. The lint is the only alarm that rings: ```text [agent-top-level-keys] agent.yml:8: 'rules' is nested inside 'agent:', where the engine parses it and then silently discards it. Move it to the top level of agent.yml, as a sibling of 'agent:'. ``` One more twist worth knowing: the check derives the indentation _from your file_ rather than assuming two spaces — so reformatting to four-space YAML doesn't create a blind spot. Defensive checkers have to out-stubborn formatters. ## Break #2 — configure the LLM the way the internet tells you to Replace the `llm:` block in `integrations.yml` with the pre-2026 form you'll find in most blog posts: ```yaml llm: provider: openai model: gpt-4 ``` ```text [llm-model-group] integrations.yml:2: 'provider' is set inline under 'llm:'. Since 3.20.0.dev6 the orchestrator LLM is a model-group reference: use 'llm: {model_group: }' and declare the provider under a matching 'model_groups' entry. [llm-model-group] integrations.yml:1: 'llm:' does not name a model_group; validate will reject it with "'model_group': Field required" ``` This one isn't silent forever — `rasa validate` would eventually reject it with `'provider': Extra inputs are not permitted`, a message that names neither the file, the cause, nor the fix. The lint's job here is to convert a confusing failure _later_ into a self-explaining one _now_. ## Break #3 — let the LLM write project memory In root `memory.yml`, add `llm_settable: true` to `preferred_room`. ```text [project-memory-writes] memory.yml:4: project memory cannot be llm_settable — the engine rejects it. Have a tool write the field, or move it into skills//memory.yml. ``` The rule behind the rule: project memory is shared state that _other skills trust_. If the model could write it directly, a hallucination becomes a fact every skill believes. Tools write project memory; the model writes its own skill's scratchpad. The engine enforces the boundary — the lint just tells you _before_ the engine's less helpful version of this message does. ## Break #4 — indent an `if:` In `skills/recommend_plant/skill.md`, indent one of the `if:` lines by two spaces. It still _looks_ like logic. It is now **prose** — the LLM receives the literal text "if: session.recommend_plant.room_type" as part of its instructions, and your branch never branches. No error. Ever. ```text [nested-if] skills/recommend_plant/skill.md:9: indented 'if:' is not parsed as a condition; it stays instruction prose. Move the branch to the top level of the skill body, or express it in natural language. ``` ## Break #5 — reference memory the wrong way in prose In the same file's prose, write `Tell the customer their room is session.recommend_plant.room_type` and greet them with `@memory.name`. ```text [skill-prose] skills/recommend_plant/skill.md:12: session.recommend_plant.room_type appears in instruction prose; it is not substituted there. Use @memory.recommend_plant.room_type or move it into a top-level 'if:'. [skill-prose] skills/recommend_plant/skill.md:12: '@memory.name' is not a substitutable token; live values require @memory.. ``` Without the lint, your customer sees the literal text "session.recommend_plant.room_type" in a chat message. Which, to be fair, _is_ how some of us first learned this rule. ## Break #6 — "upgrade" to the newest stable Rasa Change the pin to `rasa-pro==3.19.1` — an older **stable** release, so it must be newer and better, right? ```text [engine-version-pin] pyproject.toml: rasa-pro==3.19.1 is a stable pin — # rasa-version-ignore: quoted lint output stable releases ship no Mantle engine (verified against published wheels). Pin a 3.20.0.dev release. ``` This is the single most counter-intuitive fact in the ecosystem, so here it is plainly: **the Mantle engine currently exists only in pre-release ("dev") versions.** The newest stable rasa-pro on PyPI contains neither `rasa.mantle` nor its predecessor — the claim was verified by inspecting the actual published wheels, and the check's docstring says to delete it the day a stable release finally ships the engine. Until then, "dev" is not the risky choice. It is the _only_ choice. While you're in this file: lower `requires-python` to `>=3.10` and run `uv lock`. You'll get an error about "versions that are not supported by your dependencies" that never mentions Python once. The lint names it: ```text [engine-version-pin] pyproject.toml: requires-python floor 3.10 is below 3.11; rasa-pro 3.20 needs >=3.11 and uv's resolver error will not mention Python ``` ## Break #7 — paste a key where it doesn't belong Put a fake key in your README: `sk-` followed by twenty-odd characters. Then try to _commit_ it: ```text [secret-hygiene] README.md:3: looks like a committed OpenAI-style key COMMIT BLOCKED by lint_mantle.py (exit 1). ``` The hook catches it at the last moment before it enters git history — which matters, because history is forever: a key committed once is compromised even if you delete it in the next commit. This scan runs on every commit for every file, which is why "where do I put my key?" has a one-word answer: `.env`, and only `.env`. ## Clean up, and what you now know ```bash git checkout main # or: git checkout -- . on the branch git branch -D wreck-it ``` Seven breaks, seven catches — and the important part is the _taxonomy_: - **Breaks 1, 4, 5** are silent at runtime. The lint is their only alarm. - **Breaks 2, 3, 6** fail eventually, with messages that don't name the cause. The lint converts them into instant, self-explaining findings. - **Break 7** is the irreversible one. The hook makes it nearly impossible to do by accident. Every one of these messages told you the file, the line, the why, and the fix. That's the standard the pack holds itself to — inherited from the catalog's own checker, whose source ([`lint_repo.py`](https://github.com/RasaHQ/rasa-community-resources/blob/52986f71071e419533f5f00335440a7ff1aef500/scripts/lint_repo.py)) reads like a well-kept incident log. Yours ([`lint_mantle.py`](https://github.com/RasaHQ/rasa-community-resources/blob/52986f71071e419533f5f00335440a7ff1aef500/starter-pack/scripts/lint_mantle.py)) is its portable descendant. Juniper survived the crash tests. Let's make it talk. --- # Chapter 4 — Run It, Review It, Repeat Source: https://rasa.community/library/tutorials/mantle-starter-pack/04-run-review-repeat/ Author: Rasa team Published: 2026-09-06 Time to plug in the power. ## Credentials, then life ```bash make env ``` That copies `.env.example` to `.env`. Open `.env` in any editor and fill in: | Variable | Where it comes from | | ---------------- | -------------------------------------------------------------------------------------------------------------------- | | `RASA_LICENSE` | [Free Developer Edition key](https://rasa.com/rasa-pro-developer-edition-license-key-request/) — a long one-line JWT | | `OPENAI_API_KEY` | Your OpenAI account | Remember Chapter 3, break #7: this file is gitignored, the hook scans for keys everywhere else, and that combination is what makes this step safe. ```bash make install # uv downloads rasa-pro (a few minutes, first time only) make train # validate + teach the engine your agent make chat # 🎉 ``` ```text you do you have anything that survives a dark bathroom? bot Yes — a ZZ plant or a snake plant would be happy there. Want me to check what's in stock? ``` If anything failed instead, don't debug from scratch — the [mantle-upgrade-and-debug skill](https://github.com/RasaHQ/rasa-community-resources/blob/52986f71071e419533f5f00335440a7ff1aef500/starter-pack/.claude/skills/mantle-upgrade-and-debug/SKILL.md) opens with a **failure→cause table**: eleven real error messages mapped to their actual causes and fixes. Mantle symptoms often point far from their causes; the table is the shortcut across that gap. It's also just Markdown — keep it open in a tab. ## An honest word about your green checkmarks You now have two layers of verification, and knowing the difference is what separates "it's green" from "it works": | Layer | Command | Proves | Doesn't prove | | ------ | -------------------------------- | ------------------------------------------ | -------------------------- | | Lint | `python3 scripts/lint_mantle.py` | Internally consistent; no known traps | That the engine accepts it | | Engine | `make train` | The installed engine accepts and builds it | That conversations go well | The lint runs in milliseconds with nothing installed — that's why it can live in a hook. The engine check needs a real install and a licence — that's why it can't. **A green lint is necessary, not sufficient**, and the pack's own docs say so in exactly those words. When you tell a teammate "Juniper is green", say which green you mean. (The third layer — do conversations actually go _well_ — has no checker. That one's on you: talk to your agent, in the Inspector, often.) ## The everyday loop From here on, building Juniper looks like this: ```text 1. Ask Claude Code for a change "Use the mantle-skill-authoring skill to add a care_tips skill" 2. It builds against the playbook, runs the lint, shows you the result 3. You read the diff ← do not skip this step 4. git commit → the hook runs the lint one more time 5. Occasionally: make train && make chat to feel the change ``` Step 3 deserves its sentence in bold: **read the diff.** The skills make correct-shaped files, the lint catches known traps — but "known traps" is a finite list and your judgment is not. The tooling removes the _silent_ failure modes so your attention is free for the interesting question: is this agent actually behaving well? ## The reviewer role — a second pair of eyes, on demand The pack ships a second role for exactly that judgment work. When a change matters — before merging anything, in a team setting — say: > Review this change as **mantle-reviewer**. The [reviewer role](https://github.com/RasaHQ/rasa-community-resources/blob/52986f71071e419533f5f00335440a7ff1aef500/starter-pack/.claude/agents/mantle-reviewer.md) is deliberately paranoid, and its habits are worth stealing for your own reviews: - It **re-runs the lint itself** rather than trusting a reported green. - It re-checks the five silent traps by hand, even though the lint covers them — belt and braces, because these produce zero runtime errors. - It gives version bumps extra teeth (does the target wheel actually ship the engine? did a careless find-and-replace produce a nonsense `X → X` upgrade line in the docs? — yes, that has really shipped). - And it **names what it did not check**. An honest partial review beats a confident-sounding complete one. If it had no licence to train with, it says so instead of implying it verified everything. That last habit is the deepest lesson in the pack, and it applies to humans too. ## Keeping Juniper healthy over time - **Upgrades:** the engine line moves fast. When you bump the pin, follow the upgrade section of the [mantle-upgrade-and-debug skill](https://github.com/RasaHQ/rasa-community-resources/blob/52986f71071e419533f5f00335440a7ff1aef500/starter-pack/.claude/skills/mantle-upgrade-and-debug/SKILL.md) — it includes the check most people skip (does the target version actually contain the engine?) and the one nobody expects (new top-level `agent.yml` keys appear in dev releases; `tool_timeout` arrived that way, unannounced). - **When stable Rasa finally ships the engine**, one lint check (`engine-version-pin`) becomes obsolete — and its own docstring says to delete it that day. Tooling that documents its own expiry date is tooling you can trust while it lives. - **New traps:** if you hit a silent failure the lint _didn't_ catch, you've found something valuable. The check functions in [`lint_mantle.py`](https://github.com/RasaHQ/rasa-community-resources/blob/52986f71071e419533f5f00335440a7ff1aef500/starter-pack/scripts/lint_mantle.py) average ~30 lines each — copy the nearest one, encode your failure, and consider contributing it back to the [community resources repo](https://github.com/RasaHQ/rasa-community-resources). That's how every existing check got there. ## Where to go next - **[Tool Scope and Memory Scope](/tutorials/tools-and-memory)** — the natural sequel: Juniper's memory is simple; this builds a bank agent where memory architecture _is_ the product. - **The catalog itself:** [`RasaHQ/rasa-community-resources`](https://github.com/RasaHQ/rasa-community-resources) — fourteen working projects (voice agents, CRM integrations, patterns) to read, run, and steal from. Every one of them passes the same checks Juniper now does. You built an agent, broke it seven ways, caught every break, brought it to life, and know exactly what your tools do and don't prove. That's not a beginner's setup anymore. That's a workshop. 🛠️ --- # Personalize the First Word Your Agent Says Source: https://rasa.community/library/tutorials/session-start-personalization/ Author: Rod Rivera Published: 2026-08-26 Scaffold a Mantle agent, train it, say hello, and it says this: ```text bot Hello! What can I assist you with today? ``` Correct, and completely anonymous. If the customer is signed in, the agent already _could_ know who they are. It just never looks. The same gap shows up later: the bundled transactions skill makes a customer with one obvious everyday account pick from a list, every single time. Both are the same missing piece — **nothing loads what you know about the user before the conversation starts.** This tutorial fixes it, and by the end the agent opens like this: ```text bot Good afternoon, Jordan! It's wonderful to have you back — thanks for being a Premier member since 2019. How can I assist you today? ``` The mechanism in the middle is an ordered block, and it is what makes the difference reliable rather than usual: ![Without an ordered block the model can greet before the profile loads, and the template interpolates nothing — the customer gets "Good , !". With one, the engine runs the identify step first: get_customer_profile writes six declared fields into project memory, none of them LLM-settable, and only then does the greet step render a response template that reads those fields back. rephrase: true changes the wording and never the facts.](/diagrams/session-start-personalization-ordered-open.svg) ## What you will build Six changes to a scaffolded project: | Technique | What it guarantees | | -------------------------------- | ------------------------------------------------ | | Override a bundled skill | your `default_session_start` wins | | Ordered block | the profile lookup happens _before_ the greeting | | Project memory, not LLM-settable | facts come from a tool, never invented | | Response template | you choose exactly what is said | | `rephrase: true` | the model improves wording, not facts | | Scoped instruction | the model sees only the branch that applies | ## Where this came from This is [Daksh Varshneya's session-start personalization pattern](https://github.com/RasaHQ/rasa-community-resources/tree/52986f71071e419533f5f00335440a7ff1aef500/patterns/session-start-personalization), demonstrated live at Rasa community office hours and published to the community resources repository. This tutorial rebuilds it step by step and explains why each piece is where it is. ## Before you start Python 3.11 or 3.12, [uv](https://docs.astral.sh/uv/), a Rasa Pro Developer Edition licence, and an OpenAI API key. The companion project pins `rasa-pro==3.21.0.dev5` (kept in sync with the catalog); the transcripts were captured on `3.20.0.dev1`, earlier on the same release line. Chapter 5 covers what changes if you are coming from `3.19.x`, where the engine package was called `rasa.calm_v2` before it was renamed to `rasa.mantle`. --- # Chapter 1 — The session start hook Source: https://rasa.community/library/tutorials/session-start-personalization/01-the-hook/ Author: Rasa team Published: 2026-08-26 Start from a scaffolded project: ```bash rasa init --engine mantle rasa train rasa inspect ``` ```text bot Hello! What can I assist you with today? ``` ## Where that greeting comes from Every Mantle conversation opens by running a skill called `default_session_start`. It ships with the engine, and it does exactly one thing: greet. That makes it the hook. It is the only place guaranteed to run before the customer has said anything, which makes it the right place to load whatever the rest of the conversation should already know. ## Overriding it You do not edit the bundled skill. You create a folder with the same skill id, and yours wins: ```text skills/default_session_start/ skill.md your version of the opener tools.py the lookup it runs responses.yml the greeting it delivers ``` That is the whole override mechanism — same id, your copy takes precedence. Nothing to register. ## What we are about to put in it Right now the skill greets. We want it to look the customer up **and then** greet, in that order, every time. "In that order, every time" is the interesting part. Written as prose the model would usually do it and occasionally not — and when it greets first, the template interpolates empty values and the customer gets `Good , !`. Chapter 3 makes the order a guarantee instead. --- # Chapter 2 — Load the profile into memory Source: https://rasa.community/library/tutorials/session-start-personalization/02-load-the-profile/ Author: Rasa team Published: 2026-08-26 The lookup is a plain local tool in the skill's own folder — only this skill calls it, so it belongs there and travels with it. ```python # skills/default_session_start/tools.py from datetime import datetime from rasa.mantle.tools.decorator import ToolContext, tool from rasa.mantle.tools.result import ToolResult _CUSTOMER_PROFILE = { "preferred_name": "Jordan", "tier": "Premier", "member_since": "2019", "default_account_id": "acc_checking", "default_account_label": "Everyday Checking", } @tool(description="Look up the signed-in customer's profile before greeting.") async def get_customer_profile(context: ToolContext = None) -> ToolResult: if context is not None: context.memory.set("customer_name", _CUSTOMER_PROFILE["preferred_name"]) context.memory.set("customer_tier", _CUSTOMER_PROFILE["tier"]) context.memory.set("member_since", _CUSTOMER_PROFILE["member_since"]) context.memory.set("time_of_day", _time_of_day(datetime.now().hour)) context.memory.set("default_account_id", _CUSTOMER_PROFILE["default_account_id"]) context.memory.set( "default_account_label", _CUSTOMER_PROFILE["default_account_label"] ) return ToolResult(llm_response={"ok": True}) ``` The profile is hard-coded so the pattern runs with no backend. In a real deployment this is a query against your customer store, keyed on whoever the channel says is signed in. The shape of the tool does not change. > **Write with bare names.** `context.memory.set("customer_name", …)` resolves to > the project entry. Writing `"project.customer_name"` is rejected at train time > as an undeclared memory write — the qualified form is for reads only. This > catches people out, because reads accept both. ## Declare what you store Nothing can be written that has not been declared. In the root `memory.yml`: ```yaml customer_name: type: text description: The name to address the customer by. Set at session start. customer_tier: type: text description: The customer's membership tier, e.g. Premier. member_since: type: text description: The year the customer joined. time_of_day: type: categorical enum_values: [morning, afternoon, evening] description: Part of day, computed at session start. default_account_id: type: text description: Account id of the customer's usual account. default_account_label: type: text description: Human-readable label of the usual account. ``` Two decisions are encoded here, and both matter more than they look. **Project scope, not skill scope.** These live at the project level so every skill can read them. The greeting needs the name; the transactions skill needs the default account. Putting them in one skill's memory would hide them from the other. **Not `llm_settable`.** Only the tool writes here. The agent cannot decide the customer has been a member since 2015 because it sounded plausible. If a fact is supposed to come from your systems, never mark it settable. ## Keeping something private If the profile carries a field the customer must never be told, project memory is the wrong place — it can flow into model context. Declare it in the owning skill's `memory.yml` under `private:` instead. Private entries are readable only by that skill, and the runtime enforces it rather than trusting an instruction. Chapter 4 covers the second gate: what the response template lets through. --- # Chapter 3 — Guarantee the order, then personalize Source: https://rasa.community/library/tutorials/session-start-personalization/03-order-and-greeting/ Author: Rasa team Published: 2026-08-26 Now wire the lookup in front of the greeting. ```markdown --- name: Session Start description: 'Conversation opener: look up the customer, then greet them by name.' routing: engine_managed: true --- :::ordered_block id=main steps: - id: identify execute_tool: get_customer_profile - id: greet action: utter_greet ::: ``` The ordered block is the control lever. Order is a guarantee here, not a preference: the profile is loaded, _then_ the greeting is delivered. Write the same thing as prose — "look up the customer, then greet them" — and it works most of the time. The failure is quiet and occasional: the model greets first, the template interpolates nothing, and the customer gets `Good , !`. Ordered blocks remove the "most of the time". `routing: engine_managed: true` marks this as a skill the engine runs itself, rather than one the model chooses to activate. ## The greeting ```yaml # skills/default_session_start/responses.yml responses: utter_greet: - text: >- Good {session.project.time_of_day}, {session.project.customer_name}! Always great to see a {session.project.customer_tier} member who's been with us since {session.project.member_since}. How can I help you today? metadata: rephrase: true ``` Two separate things are happening, and keeping them apart is the point. **The placeholders are facts.** They resolve from memory at delivery. The model does not pick them and cannot change them. **`rephrase: true` lets the model smooth the wording** into the agent's persona: ```text template Good afternoon, Jordan! Always great to see a Premier member who's been with us since 2019. How can I help you today? delivered Good afternoon, Jordan! It's wonderful to have you back — thanks for being a Premier member since 2019. How can I assist you today? ``` Warmer, same facts. ## "How much of this is the model making up?" This question came up at office hours, and the answer is worth being precise about: **the phrasing, and nothing else.** The template is the contract. Delete a placeholder and it disappears from the output — not "usually", but always, because the model never had the value in the first place. If the tier should never be said aloud, remove `{session.project.customer_tier}`: ```yaml Good {session.project.time_of_day}, {session.project.customer_name}! How can I help you today? ``` Gone. No instruction, no prompt engineering, no hoping. That makes the response template the last gate on anything sensitive, and it composes with the private-memory boundary from Chapter 2: private memory controls what the model can _see_, the template controls what the customer can _hear_. ## Try it ```bash rasa train rasa inspect ``` ```text bot Good afternoon, Jordan! It's wonderful to have you back — thanks for being a Premier member since 2019. How can I assist you today? ``` --- # Chapter 4 — Spend what you loaded Source: https://rasa.community/library/tutorials/session-start-personalization/04-spend-what-you-know/ Author: Rasa team Published: 2026-08-26 A personalized greeting is the visible half. The useful half is spending that knowledge later, where it saves the customer a step. The bundled `view_transactions` skill asks which account, every time. We know Jordan's usual account. Asking anyway is the same anonymity, one turn later. ## Scoped instructions A **scoped instruction** is a top-level `if:` branch in the skill body. The model is shown only the branch that currently applies: ```markdown if: not session.view_transactions.selected_account_id and session.project.default_account_id The customer has a usual account on file: @memory.project.default_account_label. Offer that one first by name — for example, "Want your usual @memory.project.default_account_label, @memory.project.customer_name?" If they say yes, set `selected_account_id` via `set_fields` to @memory.project.default_account_id. If they'd rather use a different account, or they already named one clearly, instead call `fetch_accounts`, present the accounts, ask which one they want, and set `selected_account_id` to the matching account id. if: not session.view_transactions.selected_account_id and not session.project.default_account_id Call `fetch_accounts` to load the customer's accounts. Present them and ask which one they want. if: session.view_transactions.selected_account_id Call `get_recent_transactions` for the selected account and present the returned rows briefly (date, merchant, amount). ``` Three states, three sets of instructions, one skill. A customer with a default account is offered it; a customer without one gets the list; once an account is chosen, both converge on the same final branch. The model never sees the branches that do not apply, which is why this scales better than one long instruction full of caveats. ## Two details that are easy to get wrong **The `if:` must be at the top level of the skill body.** Indented inside an `instructions:` block it is not parsed as a condition at all — it stays prose, and the model is asked to evaluate something it cannot. This fails silently: nothing errors, the branch just never works. **Reference memory as `@memory.…` in prose.** The `session.…` form is correct in conditions and structured fields like `requires:`, but inside instruction text it is passed to the model as literal characters. Two spellings, two places: | Where | Form | | ----------------------------------------------- | ------------------------------------ | | `if:` conditions, `requires:`, `complete_when:` | `session.project.default_account_id` | | instruction prose | `@memory.project.default_account_id` | ## Check the confirmation boundary Start a fresh conversation with Jordan's usual account loaded in `session.project.default_account_id`, its label set to `Everyday Checking`, and `session.view_transactions.selected_account_id` unset. Having a usual account is not the same as having selected it for this request. The following is an illustrative expected interaction, not a captured test run: ```text you show my transactions bot Want your usual Everyday Checking, Jordan? you yes bot Here are your most recent transactions for Everyday Checking: 2026-07-28 Blue Harbor Market -$42.18 2026-07-26 Northline Transit -$3.50 2026-07-24 Cedar & Co. -$68.00 2026-07-22 Payroll Deposit +$2,150.00 ``` Check the selected account and tool calls as well as the wording. Before the confirmation, `selected_account_id` should remain unset and `get_recent_transactions` should not run. After “yes”, it should match the loaded default account ID before transactions are fetched. If the user instead asks for a different account, check that `fetch_accounts` supplies the choices and that the eventual selection matches the user's choice, not the stored default. Only a conversation where `selected_account_id` is already set starts in the final branch and can fetch transactions without this account question. Clear it before repeating the first test. Personalization saves the customer a list selection; it does not remove their confirmation. --- # Chapter 5 — Control levers, and upgrading to 3.20 Source: https://rasa.community/library/tutorials/session-start-personalization/05-control-levers/ Author: Rasa team Published: 2026-08-26 Six changes, and every one of them was a place where the framework enforces something rather than asking the model nicely. | Technique | What it guarantees | | -------------------------------- | ---------------------------------------- | | Override a bundled skill | your version wins, by id | | Ordered block | lookup runs before the greeting | | Project memory, not LLM-settable | facts come from the tool, never invented | | Private skill memory | the model cannot see it at all | | Response template | you choose what is said | | `rephrase: true` | the model improves wording, not facts | | Scoped instruction | only the relevant branch is visible | These are **control levers**. The alternative — describing all of it in natural language and trusting the model — reads beautifully and works most of the time, which is the problem. "Most of the time" is not a property you can put in front of customers. Rasa ran the same agent two ways over roughly 100 simulated conversations: once written entirely as prose, once with the levers. The version with levers completed about **30% more tasks**. That report is not published yet, so treat the number as directional rather than as a citation — but the direction matches what the technique is for. ## A cheaper habit From the same session, and it costs nothing: **number your steps.** A numbered list reads to a model as a sequence that cannot be skipped. The same content as a paragraph reads as advice. When you have not yet reached for a control lever, numbering is the smallest thing that improves reliability. ## Evaluating this yourself `rasa init` drops six agent-building skills into your project, one of which is `mantle_simulating_and_evaluating`. Point a coding agent at it, ask for simulation scenarios, run them, and judge the transcripts with an LLM. That loop — build, simulate, evaluate, fix — is how you find out whether a change helped rather than assuming it did. ## If you are coming from 3.19 The package rename happened in `3.20.0.dev1`; the current supported release still uses `rasa.mantle`. On the older `3.19` development line it was `rasa.calm_v2`, and the old path is now gone rather than aliased — a breaking change for any custom tool. For this pattern the migration is exactly two edits: ```bash # 1. the pin rasa-pro==3.19.0.dev7 → rasa-pro==3.20.0.dev1 # rasa-version-ignore: upgrade target # 2. the imports, in every tools.py from rasa.calm_v2.tools.decorator import ToolContext, tool from rasa.calm_v2.tools.result import ToolResult ↓ from rasa.mantle.tools.decorator import ToolContext, tool from rasa.mantle.tools.result import ToolResult ``` **One more thing that is easy to miss:** `3.20.0.dev1` also raises the Python floor from 3.10 to 3.11. A `pyproject.toml` that still says `requires-python = ">=3.10,…"` fails `uv lock` with a resolver error that never mentions Python. Nothing else changes. Not the ordered block, not the response template, not the scoped instructions. The pattern trains and greets identically on both releases. If you have a larger agent, point a coding agent at the Mantle change log and let it do the rename — that is what the Rasa team recommends, and it is a mechanical change. ## Where to go next - [Tool scope and memory scope](/library/tutorials/tools-and-memory/) — when a tool should be local to a skill and when it should be global, and the public and private halves of skill memory - The [session-start personalization pattern](https://github.com/RasaHQ/rasa-community-resources/tree/52986f71071e419533f5f00335440a7ff1aef500/patterns/session-start-personalization) in the community resources repository, with a guided walkthrough alongside it - Rasa community office hours, weekly — this tutorial came out of one --- # Tool Scope and Memory Scope in Rasa Mantle Source: https://rasa.community/library/tutorials/tools-and-memory/ Author: Rod Rivera Published: 2026-08-27 You call your bank, prove who you are, and get transferred. The next department asks you to prove who you are all over again. That is not a policy decision. It is an architecture leak: nothing is shared between the two systems, so the second one starts from nothing. The same leak shows up in agents, and Mantle gives you two tools for closing it — **tool scope** and **memory scope**. This tutorial builds Sterling, a small Meridian Bank agent with three skills, and uses it to answer four questions: 1. When should a tool be **local** to a skill, and when should it be **global**? 2. What belongs in **project memory** rather than a skill's own memory? 3. What is the difference between **public** and **private** skill memory? 4. What makes a skill folder **portable** to another agent? ## What you will end up with ```text you my passphrase is bluebird bot Welcome back, Dana Okafor. You are now signed in. you what is the balance on 10029384 bot The balance on account 10029384 is 2418.55 GBP. you send 25 pounds to Sam Rivera bot To confirm, you want to send 25 pounds to Sam Rivera. Is that correct? ``` One passphrase. The balance and the transfer both run without re-verifying, because the authentication skill wrote the customer id into project memory and the other skills' tools read it back. ## The shape of the project ![Sterling's three skills — authenticate, check_balance and transfer_money — each keep a local tool in their own folder, auto-discovered. One global tool, get_customer_info in tools/customer.py, is imported by bare name into the two skills that need it. verify_passphrase in the authenticate skill is the only writer of customer_id, and it writes with a bare name into project memory; all the readers use the qualified project.customer_id form. Each skill also has its own memory split into public and private, and a private field such as passphrase_attempts is readable only by its owning skill.](/diagrams/tools-and-memory-scopes.svg) The asymmetry in those arrows is a real rule, not a drawing convention: you write with a bare name and read with a qualified one, and the qualified form on a write is rejected at train time. | Idea | Where it lives | | -------------- | ------------------------------------------------------- | | Local tool | `skills//tools.py` — auto-discovered | | Global tool | `tools/.py` + `import_tools:` in the skill | | Project memory | `memory.yml` at the project root | | Skill memory | `skills//memory.yml`, split `public` / `private` | ## Before you start You need Python 3.11 or 3.12, [uv](https://docs.astral.sh/uv/), a Rasa Pro Developer Edition licence, and an OpenAI API key. The agent is text-only on purpose: the subject is architecture, not speech. The companion project pins `rasa-pro==3.21.0.dev5` (kept in sync with the catalog); the transcripts were captured on `3.20.0.dev1`, earlier on the same release line. --- # Chapter 1 — Set up Sterling Source: https://rasa.community/library/tutorials/tools-and-memory/01-setup/ Author: Rasa team Published: 2026-08-27 Clone the companion repository and move into this tutorial's project: ```bash git clone https://github.com/RasaHQ/rasa-community-resources.git git -C rasa-community-resources checkout 52986f71071e419533f5f00335440a7ff1aef500 cd rasa-community-resources/tutorials/rasa-tools-and-memory-tutorial ``` ## Credentials ```bash make env ``` That copies `.env.example` to `.env`. Fill in two values: | Variable | Where it comes from | | ---------------- | -------------------------------------------------------------------------------------------------------- | | `RASA_LICENSE` | [Developer Edition licence key request](https://rasa.com/rasa-pro-developer-edition-license-key-request) | | `OPENAI_API_KEY` | Your OpenAI account | ## Install and train ```bash make install make train ``` `make install` runs `uv sync --prerelease=allow`. The prerelease flag is not optional for this pinned prerelease. Use the supplied lockfile and the [current compatibility report](/compatibility/); the stable release and the Mantle prerelease are separate release channels. Do not replace the locked install with an unpinned install. `make train` validates the project before it packages anything. That validation is stricter than you might expect, and it is on your side — you will meet it again in Chapter 5. ## Run it ```bash make chat ``` Then work through this exact conversation: ```text you hello bot Hello. What can I assist you with today? you my passphrase is bluebird bot Welcome back, Dana Okafor. You are now signed in. you what is the balance on 10029384 bot The balance on account 10029384 is 2418.55 GBP. you send 25 pounds to Sam Rivera bot To confirm, you want to send 25 pounds to Sam Rivera. Is that correct? you yes, confirm bot 25 pounds have been sent to Sam Rivera. ``` Note what did **not** happen: you were never asked for the passphrase a second time. Two more skills ran, each needing to know who you are, and neither asked. ## See the architecture in one command ```bash make scopes ``` ```text GLOBAL tools (project root; named in a skill import_tools) get_customer_info LOCAL tools (skill folder; auto-discovered) authenticate/verify_passphrase check_balance/fetch_balance transfer_money/make_transfer PROJECT memory (every skill reads and writes) authenticated, customer_id, preferred_language SKILL memory (public = other skills may read; private = never) authenticate: public verification_method private passphrase_attempts transfer_money: public amount, payee_name private transfer_reason ``` Four tools, three memory scopes. The rest of the tutorial is about why each one sits where it does. ## The demo data One customer, so the walkthrough is predictable: | Field | Value | | ----------- | ------------------------------------------ | | Name | Dana Okafor | | Customer id | MB-4417 | | Passphrase | `bluebird` | | Accounts | `10029384` (current), `10029385` (savings) | | Payee | Sam Rivera | --- # Chapter 2 — Local tools Source: https://rasa.community/library/tutorials/tools-and-memory/02-local-tools/ Author: Rasa team Published: 2026-08-27 A skill is the unit of composition in Mantle. Tools are how a skill actually does something — look up an account, lock a card, move money. Where you put a tool decides how portable the skill is. Start with the default: **local**. ```text skills/authenticate/ skill.md memory.yml tools.py <- local tools live here ``` ## The tool ```python # skills/authenticate/tools.py from lib.engine import ToolContext, ToolResult, tool from lib.directory import customer_by_passphrase @tool(description="Check the caller's passphrase and sign them in if it matches.") async def verify_passphrase(passphrase: str, context: ToolContext = None) -> ToolResult: """Verify the caller's passphrase. Args: passphrase: The secret word the caller gives to prove identity. """ customer = customer_by_passphrase(passphrase) if customer is None: attempts = context.memory.get("passphrase_attempts") or 0 context.memory.set("passphrase_attempts", int(attempts) + 1) return ToolResult(llm_response={"ok": False, "error": "passphrase_incorrect"}) context.memory.set("customer_id", customer["customer_id"]) context.memory.set("authenticated", True) context.memory.set("verification_method", "passphrase") return ToolResult( llm_response={"ok": True, "customer_id": customer["customer_id"], "name": customer["name"]} ) ``` There is no registration step. Files in a skill's own folder are discovered automatically — no `import_tools`, no manifest, nothing to keep in sync. ### Why the import goes through `lib.engine` You would expect that first line to read `from rasa.mantle.tools.decorator import tool`, and it can. This project routes it through a one-file shim instead, because the engine package was renamed from `rasa.calm_v2` to `rasa.mantle` in 3.20.0.dev1 and older pins still need the old name: ```python # lib/engine.py try: # Mantle after the rename from rasa.mantle.tools.decorator import ToolContext, tool from rasa.mantle.tools.result import ToolResult except ImportError: # every release published so far from rasa.calm_v2.tools.decorator import ToolContext, tool from rasa.calm_v2.tools.result import ToolResult ``` The rename has landed: `3.20.0.dev1` ships `rasa.mantle` and no longer ships `rasa.calm_v2` at all. Resolving the import in one file is what let these four tool modules cross that boundary without touching any of them. Discovery is unaffected. The loader finds tools by checking for the `_tool_description` attribute that `@tool` sets, so where you imported the decorator from does not matter; the re-exported objects are the real ones. Chapter 6 covers the decorator's own rules, which are easy to get wrong. The signature is the schema. The function name is the tool name, the `description` is what the model reads when deciding whether to call it, and the type hints define the inputs. `context` is injected by the runtime and is never visible to the model, so it must be named exactly that. ## Why this one is local `verify_passphrase` is the authentication workflow. No other skill should ever call it — a balance lookup has no business checking passphrases. Keeping it in `skills/authenticate/` means the boundary is structural rather than a matter of everyone remembering not to import it. There is a second payoff. In a year, when you lift `skills/authenticate/` into a different agent, this tool comes with the folder. Nothing to hunt for, nothing to discover by reading code. ## The rule Make a tool local when **any** of these is true: - only one skill calls it - it encodes that skill's workflow rather than a general capability - exposing it to other skills would be a mistake — it touches a system that should stay behind this skill That last one matters in regulated domains. A tool that reaches a payment rail or an identity provider should not be quietly available to every skill in the project just because it was convenient to share. ## What a tool returns `ToolResult(llm_response=...)` is what the model sees. Return structured facts, including failures: ```python return ToolResult(llm_response={"ok": False, "error": "passphrase_incorrect"}) ``` Naming the failure lets the instructions branch on it. You will see that pay off in Chapter 4, where a failed lookup is what tells the agent the caller is not verified yet. --- # Chapter 3 — Global tools Source: https://rasa.community/library/tutorials/tools-and-memory/03-global-tools/ Author: Rasa team Published: 2026-08-27 Some capabilities genuinely are shared. Looking up the signed-in customer is one: both Check Balance and Transfer Money want the caller's name, and neither of them owns that lookup. That is a **global** tool. ```text tools/customer.py <- defined once, at the project root skills/check_balance/skill.md <- import_tools: [get_customer_info] skills/transfer_money/skill.md <- import_tools: [get_customer_info] ``` ## The tool ```python # tools/customer.py @tool(description="Look up the signed-in customer's profile (name and segment).") async def get_customer_info(context: ToolContext = None) -> ToolResult: customer_id = context.memory.get("project.customer_id") if not customer_id: return ToolResult(llm_response={"ok": False, "error": "not_authenticated"}) customer = customer_by_id(str(customer_id)) return ToolResult( llm_response={"ok": True, "customer_id": customer["customer_id"], "name": customer["name"], "segment": customer["segment"]} ) ``` ## The import is the point Unlike local tools, a global tool is **not** picked up automatically. Each skill that wants it says so: ```yaml --- name: Check Balance description: > Tell the caller the balance of one of their accounts. import_tools: - get_customer_info --- ``` That line is doing real work. It is the skill declaring its dependencies, so when you move `skills/check_balance/` to another agent, the folder tells you exactly what else has to travel with it. Without it you would be reading Python to find out. It also keeps the model's tool list small. A skill only sees the tools it owns and the ones it imported, so the model is choosing between a handful of options rather than everything in the project. ## The decision rules Make a tool global only when **all three** hold: | Test | `get_customer_info` | | ------------------------------------------------ | -------------------------------------- | | More than one skill calls it? | Yes — Check Balance and Transfer Money | | Skill-agnostic — does it know why it was called? | No, and it should not | | About the end user rather than one workflow? | Yes — it returns the profile | Fail any one and it stays local. A tool used by two skills that still encodes one skill's workflow is a sign the two skills should share a sub-skill, not a tool. ## Why not "always local, import when needed" Both designs give bounded, declarative skills. Rasa chose the global root because it makes shared surface **discoverable**: open `tools/` and you can see everything more than one skill depends on. With the alternative, sharing is implicit and you find it only by reading imports across every skill. It also matters for how teams work. Enterprise agents are usually built by several teams, each contributing skills. Keeping each team's working area bounded — their own folder, their own tools — reduces the surface where two teams collide. The shared surface stays small, visible, and deliberate. ## A signal worth watching If most of your tools end up global, something is off. Either the skills are cut too finely, or what you are sharing is really a library function rather than a tool. That is worth feeding back to Rasa — it questions what local tools are for. ## Tools do not have to run in-process A tool is anything the agent can do. Today these all execute locally, inside your project. Remote execution over MCP servers is the same idea with a different transport: the skill imports a tool and calls it, and where the code runs is an implementation detail. The scoping rules in this chapter do not change. --- # Chapter 4 — Project memory Source: https://rasa.community/library/tutorials/tools-and-memory/04-project-memory/ Author: Rasa team Published: 2026-08-27 Back to the bank transfer that makes you identify yourself twice. In Mantle the fix is **project memory** — declared once at the root, readable and writable by every skill. ```yaml # memory.yml authenticated: type: bool description: Whether the caller has proven their identity this session. initial_value: false customer_id: type: text description: Meridian customer id, set once identity is proven. preferred_language: type: text description: Language the customer asked to be served in. initial_value: en ``` These are facts about the **session and the end user**, not about one skill's workflow. Is the caller verified. Who are they. What language do they want. ## Writing it The authentication tool writes with a **bare** name: ```python context.memory.set("customer_id", customer["customer_id"]) context.memory.set("authenticated", True) ``` The slice resolves a bare name against the active skill first, then the project. Reads may also use a qualified name: ```python customer_id = context.memory.get("project.customer_id") ``` The asymmetry is easy to trip over: `project.customer_id` is valid on `get`, but on `set` it fails validation with `undeclared_memory_write`. Write bare, read either way. ## Reading it Both of the other skills' tools open the same way: ```python customer_id = context.memory.get("project.customer_id") if not customer_id: return ToolResult(llm_response={"ok": False, "error": "not_authenticated"}) ``` That is the whole trick. `authenticate` ran, wrote the id, and every later skill picks it up. The caller is asked once. ## Put the guardrail in the tool, not the prose It is tempting to write the check into the instructions: ```markdown If the caller is not verified, ask them to sign in first. ``` Resist it. The tool already knows, deterministically, and a tool cannot be talked out of its answer. So the instruction branches on the tool's result instead: ```markdown `fetch_balance` reads the signed-in customer from project memory, so it is the authority on whether the caller is verified. If it returns `not_authenticated`, tell the caller you need to verify them first and let the Authenticate skill run. Never state a balance in that case. ``` The model is deciding what to **say**. The tool is deciding what is **true**. ## A trap: `immutable` does not mean write-once Project memory has an `immutable` flag, and the name invites exactly the wrong guess. It looks like it should mean "once set, frozen for the session" — which is how you would want `customer_id` to behave. It does not. It means the field can never be written at runtime at all; it is for constants seeded by `initial_value`. Put it on `customer_id` and authentication itself fails: ```text Memory field 'customer_id' (slot 'project.customer_id') is immutable and cannot be overwritten. ``` The tool call is refused, the caller never signs in, and the only sign of it is a warning in the server log. What actually makes verification stick is simpler: project memory persists for the session, and no other skill overwrites it. --- # Chapter 5 — Public and private skill memory Source: https://rasa.community/library/tutorials/tools-and-memory/05-skill-memory/ Author: Rasa team Published: 2026-08-27 Not everything belongs in project memory. Most of what a skill tracks is its own working state: the account being asked about, the amount being sent, how many passphrase attempts have failed. That is **skill memory**, declared in the skill's own folder and split in two. ```yaml # skills/transfer_money/memory.yml schema: public: amount: type: float description: Amount the caller asked to send. llm_settable: true payee_name: type: text description: Who the money is going to. llm_settable: true transfer_confirmed: type: bool description: Caller explicitly confirmed the transfer details. initial_value: false llm_settable: true private: transfer_reason: type: text description: Free-text reason the caller gave for the payment. pii: true llm_settable: true ``` The top-level key is `schema:`, not `memory:`. **Public** fields are readable by every other skill. **Private** fields are readable only by the skill that owns them. `transfer_reason` is free text the caller volunteered, flagged `pii: true`, and it stays inside Transfer Money. The boundary is enforced by the runtime. It is not a convention or a lint rule, and the model cannot talk its way past it — a denied read simply returns nothing. ## Two different writers Two actors write memory, and they are governed separately: - **Tools** write whatever the skill declares, through `context.memory.set()`. - **The model** may only write fields marked `llm_settable: true`, through the built-in `set_fields` tool. Mark the caller's own decisions settable — the amount, the payee, the confirmation. Leave anything the system derives to tools. `passphrase_attempts` is not settable by the model, because the model has no business deciding how many times you failed. ## Everything must be declared A tool that writes an undeclared field fails at `rasa train`: ```text A tool in skills/authenticate/tools.py writes undeclared memory entry 'project.customer_id' via memory.set(...) at line 38. Declare it under the skill's schema (public/private) or the project memory.yml, or it will be rejected at runtime. ``` This is a good failure. It happens at build time rather than three turns into a production conversation. ## The description rule that will catch you Validation rejects a `description` on a field that is neither `llm_settable` nor owned by a `collect:` step: ```text Memory field 'verification_method' in skill 'authenticate' declares a description but is neither llm_settable nor collect-owned. Descriptions are only surfaced on set_fields for LLM-facing entries — remove the description or mark the field settable. ``` The reasoning: a description exists so the model knows what to put in a field. If the model can never write it, the description is documentation aimed at nobody, and Mantle would rather you say so in a comment. Use a YAML comment for tool-written fields. ## Choosing a scope | Ask | Scope | | ------------------------------------------------------------- | ------- | | Is it about the session or the end user? | Project | | Is it this skill's working state that others may need to see? | Public | | Is it internal, sensitive, or meaningless outside this skill? | Private | When in doubt, start private. Widening a field later is a one-line change; discovering that three skills quietly depend on something you meant to keep internal is not. --- # Chapter 6 — Portability, and what it cost Source: https://rasa.community/library/tutorials/tools-and-memory/06-portability/ Author: Rasa team Published: 2026-08-27 Every rule in this tutorial serves one goal: a skill folder you can pick up and drop into another agent. That works when three things are explicit. ```text skills/transfer_money/ skill.md -> import_tools: what it borrows memory.yml -> public / private: what it owns tools.py -> what it brings with it ``` Read those three files and you know the skill's entire contract with the rest of the project: it brings `make_transfer`, it borrows `get_customer_info`, it owns four memory fields, and it reads `project.customer_id`. Moving it is a copy plus whatever `import_tools` names. Nothing has to be inferred by reading code, which is the difference between a skill you can move and a skill you can only rewrite. ## Why it is built this way Two forces, both from watching large teams build agents. **Bounded surface.** An enterprise agent is usually several teams, each owning skills. If every team can reach every tool and every memory field, each change risks colliding with someone else's. Local tools and private memory keep the blast radius inside one folder. **Discoverable sharing.** What is shared is small, sits at the project root, and is named explicitly by whoever uses it. You can audit it in one directory listing. ## Five things that will bite you Every one of these came out of actually building this project. **Write bare, read qualified.** `context.memory.set("customer_id", …)` works; `set("project.customer_id", …)` fails validation. Reads accept both. **`immutable` is not write-once.** It blocks all runtime writes. On `customer_id` it silently breaks authentication — the write is denied and only the server log says so. **No `description` on tool-written fields.** Validation rejects a description unless the field is `llm_settable` or collect-owned. Use a YAML comment. **Do not branch on memory tokens in prose you cannot guarantee.** An `@memory.project.authenticated` reference resolves to nothing when the field is unset, and the model reads a broken sentence. Branch on a tool result instead — `fetch_balance` returning `not_authenticated` is unambiguous. **Say what to do with information already given.** "Ask for the passphrase" makes the model ask even when the caller opened with it. Spell out the case: ```markdown If the caller has already given a passphrase in what they just said, call `verify_passphrase` with it straight away. Do not ask them to repeat it. ``` ## Getting the decorator right Four rules, all enforced by the engine rather than by convention. The real signature is `def tool(*, description: str)` — keyword-only, and the description is required: ```python @tool # TypeError: takes 0 positional arguments but 1 was given @tool("does a thing") # TypeError: same @tool(description="…") # correct ``` **It must be `async`.** The executor awaits the call, so a synchronous function fails at the await. **The parameter must be named exactly `context`.** The runtime invokes tools as `await func(context=ctx, **tool_args)` — a keyword argument. Call it `ctx` and you get an unexpected-keyword error. **Every tool needs that parameter, even unused ones.** There is no signature inspection; `context=ctx` is always passed. A tool declared as `async def f(account_number: str)` breaks. Give it `context: ToolContext = None` and ignore it — the default also keeps the function directly unit-testable. ## Surviving the rename Rasa renamed the engine package from `rasa.calm_v2` to `rasa.mantle` in `3.20.0.dev1`, and the old path is gone rather than aliased — so the same source cannot import both without help. This project resolves it once, in `lib/engine.py`, and every tool imports from there: ```python from lib.engine import ToolContext, ToolResult, tool ``` One file to change instead of four, and nothing to change at all once the rename lands. Two things worth knowing if you copy the pattern: - It assumes the rename is **path-only**. If the API also changes shape, the `try` branch will import cleanly and then fail somewhere less obvious. - Delete the fallback once the old path is gone, rather than leaving it indefinitely — a shim that outlives its reason becomes a puzzle for whoever reads it next. ## Where to go next Skills are the foundation, and more of Mantle is being built on top of them. The next things worth reading: - [Tool constraints](/library/tutorials/voice-ai-agent/04-tool-constraints/) — `requires:` and confirmation gates, used here on `make_transfer` - [Scoped instructions](/library/tutorials/voice-ai-agent/05-scoped-instructions/) — narrowing what the model may do inside a skill - The [Mantle documentation](https://mantle.rasa.com/), which is a living document updated with every release, alongside a per-version changelog Two questions worth sending back to Rasa as you build: if nearly all your tools end up global, and if you find yourself wanting to widen private memory often. Both would say something about where the boundaries are drawn. --- # Build a Voice AI Agent with Rasa Skills Source: https://rasa.community/library/tutorials/voice-ai-agent/ Author: Rod Rivera Published: 2026-08-14 ## What you will build By the end of this tutorial you will have a **speaking travel agent** named Atlas that can: - answer trip FAQs from reference material - list itineraries from a demo database - look up flight status with hard tool gates - file a lost-baggage report with ordered steps and verbatim notices - change bookings by composing authentication and lookup skills Speech in and speech out use **Deepgram** end to end (Flux ASR + Aura TTS) through the Rasa Inspector. ## Who this is for You do not need prior Rasa experience. You should be comfortable with: - a terminal and `git` - editing YAML and Markdown files - basic Python function signatures ## Mental model: Skills, not flowcharts Older Rasa assistants often centred on flow YAML and a large domain file. The **Skills** architecture is different: | Idea | Meaning | | ------------------- | -------------------------------------------------------------------------- | | Skill | A folder with `skill.md` instructions, plus optional tools and references | | Mantle | The orchestrator that activates skills from their descriptions | | Progressive control | Start with prose; add runtime guarantees only where mistakes are expensive | | Memory | Shared state tools write and constraints read | There are no flowcharts required to start. The LLM interprets instructions. Where you need a guarantee, you add a frontmatter constraint or an ordered block — enforced by the runtime, not suggested to the model. Those guarantees form a ladder, and the useful thing about it is how few skills climb it: ![A ladder of runtime guarantees, each bought with one config key. At the bottom, prose only — name and description — where trip_faq, intro and goodbye stop. Above it: import_tools so answers come from data; tool_constraints with requires, which removes a tool from the schema until its data exists, where flight_status stops; an if marker so the wrong branch is never in context; utter with on activate for exact wording; requires_confirmation for a human check before an irreversible act; and at the top an ordered_block, which guarantees order itself. Only report_baggage and change_booking reach the top, and the rungs are cumulative, so those two carry nearly all of them at once. All four condition levers share one syntax: session, then a scope, then a field.](/diagrams/voice-ai-agent-progressive-control.svg) Start with prose. Add a rung only where a mistake is expensive — most skills never need one. ## How voice fits ```text Traveler speaks → Deepgram ASR (speech to text) → Mantle + Skills → Deepgram TTS (text to speech) → Traveler hears Atlas ``` You configure Deepgram once in `integrations.yml`. Skills stay domain logic. That separation is intentional: the same skills can serve chat or voice. ## Companion repository All runnable code lives in: [github.com/RasaHQ/rasa-community-resources/tutorials/rasa-voice-agent-tutorial](https://github.com/RasaHQ/rasa-community-resources/tree/52986f71071e419533f5f00335440a7ff1aef500/tutorials/rasa-voice-agent-tutorial) Clone the monorepo, open that tutorial tree, run `make install && make env && make verify`, and keep two windows open: this tutorial and your editor. Paste-ready chapter files are under `tutorial/snippets/` in that tree. ## Chapter map | Ch | Topic | | --- | ------------------------------ | | 1 | Scaffold a voice agent | | 2 | First skill (FAQ) | | 3 | First tool (itinerary) | | 4 | Tool constraints | | 5 | Scoped instructions | | 6 | Verbatim responses for voice | | 7 | Ordered blocks (baggage) | | 8 | Compose skills | | 9 | Voice deep dive | | 10 | Coding agents and the flywheel | Continue to [Chapter 1 — Scaffold](/library/tutorials/voice-ai-agent/01-scaffold/). --- # Chapter 1 — Scaffold a voice agent Source: https://rasa.community/library/tutorials/voice-ai-agent/01-scaffold/ Author: Rasa team Published: 2026-08-14 ## Goal Leave this chapter with a running voice shell: Inspector open, mic enabled, Atlas able to greet you. ## Prerequisites 1. Install [uv](https://docs.astral.sh/uv/) 2. Clone the companion repo: ```bash git clone https://github.com/RasaHQ/rasa-community-resources.git git -C rasa-community-resources checkout 52986f71071e419533f5f00335440a7ff1aef500 cd rasa-community-resources/tutorials/rasa-voice-agent-tutorial make install make env ``` 3. Open `.env` and fill in three values: | Variable | Purpose | | ------------------ | -------------------------------------------------------------------------------------------- | | `RASA_LICENSE` | [Developer Edition licence](https://rasa.com/rasa-pro-developer-edition-license-key-request) | | `OPENAI_API_KEY` | LLM for routing and conversation (`gpt-5.2` in this companion) | | `DEEPGRAM_API_KEY` | Speech-to-text **and** text-to-speech | `make install` pins `rasa-pro==3.21.0.dev5`. 4. Gate the session: ```bash make verify ``` Do not continue until checks pass (or only show expected warnings such as “no trained model yet”). ## What `rasa init --engine mantle` creates On a Mantle-capable Rasa build you can scaffold a fresh project with: ```bash pip install rasa-pro==3.21.0.dev5 rasa init --engine mantle ``` The wizard asks where to create the project, collects your keys into a `.env`, installs Cursor / Claude Code compatible skills, validates and trains, and can open the Inspector. The resulting shape matches this companion: ```text my-agent/ ├── agent.yml # identity, persona, voice flags ├── integrations.yml # LLM + channels ├── AGENTS.md # context for your coding agent ├── .env # secrets (written by the wizard) └── skills/ # what the agent can do ``` This tutorial’s companion already includes Deepgram ASR/TTS under `integrations.yml` (the wizard itself is LLM-focused). For the voice path here, paste the finished scaffold from the companion instead of starting from an empty wizard project: **Paste set:** `tutorial/snippets/step-00-scaffold/` Copy into the project root: - `agent.yml` - `integrations.yml` (includes Deepgram + `gpt-5.2`) - `endpoints.yml` - `memory.yml` - `responses.yml` - `skills/intro/skill.md` The two config files are easy to mix up. `integrations.yml` is where the LLM and the Deepgram ASR/TTS for the Inspector live, which is what you will touch most. `endpoints.yml` covers optional platform services such as the response rephraser and a tracker store; it is optional for Mantle projects and can be omitted without `rasa train` errors. ## Deepgram in `integrations.yml` ```yaml channels: inspector: enabled: true asr: name: deepgram language_map: en: model: flux-general-en eot_threshold: 0.7 eot_timeout_ms: 5000 tts: name: deepgram language_map: en: model: aura-2-andromeda-en ``` And in `agent.yml`: ```yaml voice: enabled: true asr: deepgram tts: deepgram ``` One API key covers both directions. Flux handles end-of-turn detection for ASR; Aura synthesises speech for TTS. ## Train and talk ```bash make train make inspect ``` **Verify:** Inspector opens. Enable the microphone (or type) and say hello. Atlas should greet you and offer travel help. ## Talking point Skills projects are files in git — not a separate Studio-only artefact. Business teams and developers edit the same skill folders; Mantle serves them in production. ## Next [Chapter 2 — First skill (FAQ)](/library/tutorials/voice-ai-agent/02-first-skill/) --- # Chapter 2 — First skill (FAQ) Source: https://rasa.community/library/tutorials/voice-ai-agent/02-first-skill/ Author: Rasa team Published: 2026-08-14 ## Goal Add `trip_faq`: a skill that answers Horizon Travel questions from markdown references. ## Teach The minimum viable skill is one file: `skills//skill.md`. - YAML frontmatter: `name` and `description` (description drives Mantle routing) - Markdown body: instructions the LLM follows while the skill is active Optional `references/` gives focused context without stuffing every prompt. ## Paste **Paste set:** `tutorial/snippets/step-01-faq/` Copy into: - `skills/trip_faq/skill.md` - `skills/trip_faq/references/horizon_travel_faq.md` Example frontmatter and body: ```markdown --- name: trip_faq description: > Answer common Horizon Travel questions about baggage allowances, check-in windows, seats, and change fees from reference material. --- Answer the traveler's question from the information in your references. Keep answers short enough to speak aloud — two or three sentences at most. ``` ## Train and try ```bash make train make inspect ``` Try: “How much cabin baggage can I take?” **Verify:** The answer comes from the FAQ references and is short enough to hear aloud. ## Talking point No tools, no memory schema, no registration step. Mantle routes here because the description matches the question. ## Next [Chapter 3 — First tool](/library/tutorials/voice-ai-agent/03-first-tool/) --- # Chapter 3 — First tool Source: https://rasa.community/library/tutorials/voice-ai-agent/03-first-tool/ Author: Rasa team Published: 2026-08-14 ## Goal Atlas can list Maya Chen’s bookings by calling tools — not inventing itineraries. ## Teach Tools are async Python functions. **Default to skill-local tools:** drop them in `skills//tools.py` and they are auto-discovered while that skill is active. No registration. Put a tool in the project-root `tools/` folder **only** when two or more skills need the same function, and list it under `import_tools` on each skill that uses it. In this chapter the shared helpers are intentional: - `load_customer_profile` — also used by `intro`, `authenticate`, and `human_handoff` - `list_bookings` — also used by `find_booking` and `report_baggage` Later chapters move skill-owned tools (PIN verify, flight status, cancel, baggage submit) into each skill’s own `tools.py`. ```python from rasa.mantle.tools.decorator import ToolContext, tool from rasa.mantle.tools.result import ToolResult @tool(description="List the traveler's upcoming bookings and trip summaries.") async def list_bookings(context: ToolContext = None) -> ToolResult: ... return ToolResult(llm_response={"ok": True, "bookings": bookings}) ``` - Function name → tool name - Type hints → input schema - `context` is injected by the runtime (not shown to the LLM); the parameter must be named exactly `context` - Structured data returns in `ToolResult.llm_response` - Name the tool in prose by its plain name: “call `list_bookings`” ## Paste **Paste set:** `tutorial/snippets/step-02-itinerary/` Copy: - `skills/check_itinerary/skill.md` - `tools/travel.py` (shared: `load_customer_profile`, `list_bookings`) - `lib/database.py` - `lib/tool_helpers.py` - `data/source/*.json` Demo traveler: **Maya Chen** (id `456`). Useful booking: `HT12345` (Lisbon). The `intro` skill loads her profile on greet; booking tools also fall back to id `456` from SQLite. ```bash make show-demo-data ``` ## Skill instructions ```markdown --- name: check_itinerary description: > List the traveler's upcoming trips and booking details. import_tools: - load_customer_profile - list_bookings --- Help the traveler review their itinerary. If customer details are missing, call load_customer_profile. Then call list_bookings and summarize each trip in short spoken sentences. Speak booking references character by character. ``` ## Train and try ```bash make train make inspect ``` Try: “What trips do I have?” **Verify:** Atlas lists Lisbon and Tokyo trips with spoken booking references (for example H T one two three four five). ## Talking point Side effects and lookups belong in tools. Skills orchestrate conversation around tool results. Keep tools local unless sharing is real. ## Next [Chapter 4 — Tool constraints](/library/tutorials/voice-ai-agent/04-tool-constraints/) --- # Chapter 4 — Tool constraints Source: https://rasa.community/library/tutorials/voice-ai-agent/04-tool-constraints/ Author: Rasa team Published: 2026-08-14 ## Goal Add `flight_status` where `get_flight_status` is invisible until `booking_ref` is in memory. ## Teach At the prose-only stage, the LLM can call tools too early. Progressive control turns that into a **runtime guarantee**: ```yaml tool_constraints: - get_flight_status: requires: session.flight_status.booking_ref ``` `requires` is a condition expression, and every memory reference in one is fully namespaced as `session..`. The scope is a skill id, `project` for fields shared across skills, or `system`. So `session.flight_status.booking_ref` means "the `booking_ref` field belonging to the `flight_status` skill". Until that field has a value, the tool is removed from the schema entirely. The model cannot call a tool it never sees. This is the first hard lever. The instruction body can stay almost unchanged. ## Paste **Paste set:** `tutorial/snippets/step-03-constraints/` Copy into `skills/flight_status/`. Also declare the memory field the constraint reads: ```yaml # skills/flight_status/memory.yml schema: public: booking_ref: type: text description: Horizon Travel booking reference such as HT12345. llm_settable: true ``` Three details that are easy to get wrong: the root key is `schema`, the text type is called `text` (not `string`), and `llm_settable: true` is what allows the model to fill the field from what the traveler says. ## Train and try ```bash make train make inspect ``` Try: “Is my flight on time?” without a booking reference — Atlas should ask for it first. Then: “Booking H T one two three four five — is my Lisbon flight delayed?” **Verify:** Status for `HT12345` returns (outbound is delayed 45 minutes in the demo data). ## Talking point Soft instructions become guarantees by adding frontmatter — not by rewriting the whole prompt. ## Next [Chapter 5 — Scoped instructions](/library/tutorials/voice-ai-agent/05-scoped-instructions/) --- # Chapter 5 — Scoped instructions Source: https://rasa.community/library/tutorials/voice-ai-agent/05-scoped-instructions/ Author: Rasa team Published: 2026-08-14 ## Goal Branch Atlas’s guidance for delayed vs cancelled vs on-time flights without an ordered block. ## Teach Scoped instructions use `if:` markers in the skill body. When a categorical memory value is set, the runtime **strips non-matching paragraphs** from the prompt. The LLM cannot follow the wrong branch because that text is not in context. An `if:` marker takes the same condition expression as `requires` from the previous chapter, so memory is namespaced the same way: ```markdown if: session.flight_status.flight_status == "delayed" Tell them the delay in minutes and the gate if available. if: session.flight_status.flight_status == "cancelled" Apologize briefly and offer a booking change. if: session.flight_status.flight_status == "on_time" or session.flight_status.flight_status == "boarding" Confirm the status, scheduled time, and gate. ``` Two boundary rules to remember. An `if:` marker scopes only the paragraph immediately following it, up to the next blank line, and unmarked paragraphs stay always visible. There is no `else:` or `elif:` in prose — write a separate `if:` paragraph for each case, or reach for an ordered block when you need true if/else control flow. Declare the categorical in `memory.yml`: ```yaml flight_status: type: categorical enum_values: [on_time, delayed, cancelled, boarding] llm_settable: true ``` ## Paste **Paste set:** `tutorial/snippets/step-04-scoped/` Replace `skills/flight_status/` with the scoped version (includes the `if:` branches). ## Train and try ```bash make train make inspect ``` Try booking `HT12345` (delayed outbound) and `HT67890` (cancelled return). **Verify:** Delayed path mentions minutes and gate; cancelled path offers change or handoff. ## Talking point Frontmatter can stay the same as Chapter 4. Only the body annotations and memory enum change. Zero ordered blocks required for branching. ## Next [Chapter 6 — Verbatim responses](/library/tutorials/voice-ai-agent/06-verbatim-responses/) --- # Chapter 6 — Verbatim responses for voice Source: https://rasa.community/library/tutorials/voice-ai-agent/06-verbatim-responses/ Author: Rasa team Published: 2026-08-14 ## Goal Understand verbatim responses before the full baggage showcase in the next chapter. Atlas will speak a recording notice and compliance lines exactly as written. ## Teach Some moments need exact wording: a recording notice at the start of a call, a fraud or theft warning, a legal disclaimer after a side effect. Verbatim responses live in `responses.yml`. The framework delivers them; the LLM does not rewrite them. | Trigger | Frontmatter | When it fires | | --------------- | ---------------------------------- | -------------------------- | | Skill activates | `utter:` with `on: activate` | Skill loads | | Memory changes | `utter:` with `when:` | The condition becomes true | | Tool succeeds | `on_success:` on a tool constraint | Tool returns success | | Tool fails | `on_failure:` on a tool constraint | Tool returns error | A `when:` trigger takes the same condition expression as `requires` and `if:`: ```yaml utter: - utter_baggage_recording_notice: on: activate - utter_baggage_stolen_warning: when: session.report_baggage.bag_issue == "stolen" ``` ```yaml # responses.yml responses: utter_baggage_recording_notice: - text: > This call may be recorded for quality assurance and training purposes. ``` Response text can interpolate memory at delivery time using the same namespaced form in braces, for example `{session.report_baggage.booking_ref}`. There is a fourth trigger worth knowing now, because it replaces something you might expect to write inline. To make a tool pause and ask permission before it runs, add `requires_confirmation` to its constraint and point it at responses: ```yaml tool_constraints: - submit_baggage_report: requires: session.report_baggage.details_verified requires_confirmation: enabled: true utter_for_confirmation: utter_confirm_baggage_report utter_on_user_denial: utter_baggage_report_cancelled on_success: utter_baggage_submitted ``` The confirmation question is a named response rather than a string in the frontmatter, which means the exact words a traveler hears before an irreversible action live in `responses.yml` alongside every other piece of spoken copy — reviewable in one place, and never rephrased by the model. ## Why this matters for voice TTS will speak whatever text you send. If a compliance line is generated freely, wording drifts. Verbatim responses keep the spoken contract stable across conversations. ## Preview in the companion Open `skills/report_baggage/responses.yml` and the `utter:` block in `skills/report_baggage/skill.md`. You will wire the full skill in the next chapter. ## Talking point Use prose for conversation. Use verbatim for moments where the words themselves are the requirement. ## Next [Chapter 7 — Ordered blocks](/library/tutorials/voice-ai-agent/07-ordered-blocks/) --- # Chapter 7 — Ordered blocks Source: https://rasa.community/library/tutorials/voice-ai-agent/07-ordered-blocks/ Author: Rasa team Published: 2026-08-14 ## Goal Add the progressive-control showcase: `report_baggage` with recording notice, ordered collection, confirmation, and success utter. ## Teach Reach for ordered blocks when **order itself** is the requirement. Most skills never need them. When they do, prose can invoke a block with `@block.`. ```text Once they want to file a report, invoke @block.collect_baggage :::ordered_block id=collect_baggage steps: - id: fetch_bookings execute_tool: list_bookings - id: select_booking instructions: | Ask which trip the bag belongs to. Set booking_ref. complete_when: session.report_baggage.booking_ref # ... more steps ... ::: ``` The block id goes on the fence line as `id=collect_baggage`. Putting it inside the body as `id:` is the most common mistake here, and the error you get back is "must declare an id attribute". A step's `complete_when` is the same condition expression used everywhere else, so a bare `session.report_baggage.booking_ref` means "this step is done once that field has a value". Step types: | Type | Properties | Behaviour | | -------------- | ---------------------------------- | --------------------------------- | | Execute tool | `execute_tool:` | Framework calls the tool directly | | Conversational | `instructions:` + `complete_when:` | LLM talks within the step | | Collect | `collect:` + `utter:` | Framework collects a value | `tool_constraints`, including `requires` and `requires_confirmation`, still apply when a block or the LLM attempts a tool call. ## Paste **Paste set:** `tutorial/snippets/step-05-baggage/` Copy the whole `skills/report_baggage/` folder (includes `skill.md`, `memory.yml`, `responses.yml`). ## Train and try ```bash make train make inspect ``` Try (voice if possible): “My bag did not arrive.” **Verify:** 1. Recording notice plays on activate 2. Bookings are listed before details are collected 3. Summary confirmation happens before submit 4. A report id such as `BAG1001` is returned ## Talking point This skill combines every lever from earlier chapters. The hybrid pattern — prose around one ordered block — covers most enterprise needs. Fully controlled skills (entire body is one ordered block) are for the most regulated cases only. ## Next [Chapter 8 — Compose skills](/library/tutorials/voice-ai-agent/08-compose-skills/) --- # Chapter 8 — Compose skills Source: https://rasa.community/library/tutorials/voice-ai-agent/08-compose-skills/ Author: Rasa team Published: 2026-08-14 ## Goal Compose small skills into a booking-change journey with authentication and confirmation. ## Teach When a skill needs another skill’s logic mid-conversation, reference it with `@skill.`. The parent stays on the stack; the sub-skill runs and exports public memory; the parent resumes. ```markdown Find the booking: @skill.find_booking ``` Parent skills can also declare a precondition. The skill names it in its frontmatter: ```yaml precondition: authenticated ``` and `agent.yml` says when the precondition holds and which skill satisfies it: ```yaml orchestrator: preconditions: authenticated: satisfied_when: session.project.authenticated resolve_with: authenticate ``` When `change_booking` starts before `session.project.authenticated` is true, Mantle parks it, runs `authenticate`, and resumes `change_booking` once the condition holds. The memory field is still the contract; the precondition only decides what runs first. Rasa Pro 3.21.0.dev3 rejects the older skill-level `requires:` key, which hid the skill from routing until the field was true. Identity is the precondition's job, not the prose's. Leave out a line such as `First verify identity: @skill.authenticate` in `change_booking`: the engine already runs `authenticate` when it is needed, and a second copy in the instructions can start it again for a traveller who has already signed in. ### Where shared fields live This is the part that trips people up, so it is worth slowing down for. When a tool writes a memory field, the write lands in the namespace of the **skill that is active at that moment**. `find_booking` is the skill running when a booking gets selected, so if `selected_booking_ref` were declared inside `find_booking/memory.yml`, the value would land in `session.find_booking.selected_booking_ref` and `change_booking` could never gate on it. Fields that travel between skills therefore belong in the project-wide `memory.yml` at the repo root, where they resolve to `session.project.*` no matter which skill is active: ```yaml # memory.yml (project root) selected_booking_ref: type: text description: Booking reference the traveler has settled on for this request. selected_trip_name: type: text description: Trip name for the selected booking. ``` A good rule of thumb: a field used by exactly one skill belongs to that skill, and a field that is a handoff between skills belongs to the project. ### Confirming an irreversible tool ```yaml tool_constraints: - cancel_booking: requires: session.project.selected_booking_ref requires_confirmation: enabled: true utter_for_confirmation: utter_confirm_cancel_booking utter_on_user_denial: utter_cancel_aborted on_success: utter_booking_cancelled ``` Note the division of labour between the two gates. `requires` is a readiness check — it hides the tool until the data it needs exists. `requires_confirmation` is the human check — the tool is available, but the engine pauses and asks before running it. You rarely need a separate "did they confirm?" memory field, because the confirmation gate already is that check. Composition rules: - Prefer small, focused skills that export shared state through project memory - The parent must have meaningful business logic of its own - `@skill` is coordinated; user interrupts are not guaranteed to resume the same way ## Paste **Paste set:** `tutorial/snippets/step-06-composition/` Copy `skills/authenticate`, `skills/find_booking`, and `skills/change_booking`, plus the updated project `memory.yml` that adds the two shared booking fields. Optional fast-forward for remaining skills: `tutorial/snippets/step-07-remaining/` (`human_handoff`, `goodbye`, `intro`). Demo PIN: **four two four two** (`4242`). ## Train and try ```bash make train make inspect ``` Try: “I need to cancel a booking.” **Verify:** Atlas authenticates, finds a booking, asks for confirmation, then cancels only after approval. ## Talking point Authentication is reusable across every sensitive skill. Transaction or booking lookup is reusable across changes, baggage, and status. Parents own the business outcome. ## Next [Chapter 9 — Voice deep dive](/library/tutorials/voice-ai-agent/09-voice-deep-dive/) --- # Chapter 9 — Voice deep dive Source: https://rasa.community/library/tutorials/voice-ai-agent/09-voice-deep-dive/ Author: Rasa team Published: 2026-08-14 ## Goal Treat voice as a first-class design surface — not a final demo pass. ## End-of-turn (Flux) Deepgram Flux models use model-integrated end-of-turn detection. In `integrations.yml`: ```yaml asr: name: deepgram language_map: en: model: flux-general-en eot_threshold: 0.7 eot_timeout_ms: 5000 ``` | Parameter | Effect | | ---------------- | ------------------------------------------------------------------------ | | `eot_threshold` | Confidence needed to end the turn (higher = wait for clearer completion) | | `eot_timeout_ms` | Maximum silence before forcing end-of-turn | **Experiment:** Lower `eot_timeout_ms` if Atlas waits too long after short answers. Raise it if booking references spoken with pauses get cut mid-way. This tutorial uses Inspector only. Production telephony channels (SIP, CPaaS) reuse the same Skills; audio format negotiation differs by channel. ## Spoken formatting TTS reads characters poorly when you dump `HT12345` as a single token. Instruct the skill to speak references character by character: > H T one two three four five Put that rule in `agent.yml` persona and in skill bodies that read codes aloud. Keep sentences short. Ask one clarifying question at a time. Those persona rules reduce barge-in pressure and cut TTS latency perception. ## ASR recovery patterns Speech recognition will mishear amounts, pins, and booking refs. Build recovery into skills: 1. **Confirm before side effects** — `requires_confirmation` with a response that reads the value back 2. **Re-ask on failure** — tools return `ok: false`; instruct the skill to retry once 3. **Offer alternatives** — list bookings instead of forcing the traveler to dictate a reference 4. **Handoff early** — after repeated failures, `@skill.human_handoff` Demo PIN and booking refs in `make show-demo-data` are chosen to be speakable. ## Verbatim vs rephrase | Mechanism | Use when | | --------------------------------------- | ------------------------------------------------- | | `utter:` + `responses.yml` | Exact compliance / recording / legal wording | | `metadata.rephrase: true` on a response | Greeting can vary slightly while staying on-brand | | Free LLM prose | Normal conversation | For voice, prefer verbatim for anything a regulator or brand team would quote. ## Inspector workflow ```bash make inspect ``` 1. Enable the microphone 2. Watch partial transcripts while speaking 3. Compare text and voice behaviour — skills should work in both modes 4. If ASR splits a booking ref, try speaking slower or use the typed channel to isolate Skills logic from speech ## Talking point Latency and turn-taking failures are voice-specific. Progressive control and short spoken utterances are the harness — not a bigger prompt. ## Next [Chapter 10 — Coding agents and flywheel](/library/tutorials/voice-ai-agent/10-flywheel/) --- # Chapter 10 — Coding agents and the flywheel Source: https://rasa.community/library/tutorials/voice-ai-agent/10-flywheel/ Author: Rasa team Published: 2026-08-14 ## Goal Leave with a practice loop: simulate, tighten a control lever, re-inspect — assisted by a coding agent that understands Skills files. ## Project context for coding agents Skills are files. A coding agent that can edit a repo can iterate on them. Point it at the companion’s `AGENTS.md` and `.cursor/rules/rasa-skills.mdc`. Minimal instructions (also in the companion): ```text This project is a Rasa Skills agent. Skills live under skills// as skill.md files with optional tools.py, references/, memory.yml, responses.yml. Before editing a skill, understand the control levers: tool_constraints, requires, requires_confirmation, if: markers, ordered_block, utter. Conditions are expressions with fully namespaced memory, for example session.flight_status.booking_ref or session.project.authenticated. Put skill-owned tools in skills//tools.py. Share via tools/ + import_tools only when two or more skills need the same function. After changing a skill, run make validate and make inspect. ``` Useful prompts: - “Gate cancel_booking so it only appears once a booking is selected.” - “This skill should ask before submitting — add requires_confirmation.” - “Split change_booking so find_booking is a reusable @skill.” ## Keep a human in the loop Coding agents apply known levers well. They are not a substitute for talking to the agent afterward. Especially for constraints and prerequisites: those change what the LLM is **allowed** to do, so failures look different from a wording mistake. ## The conversation flywheel 1. Write a skill from your best guess 2. Simulate conversations in Inspector (happy path, then break it) 3. Deploy when ready — real traffic surfaces phrasing you did not invent 4. Review transcripts; add a constraint or tighten instructions 5. Re-simulate with cases informed by production Each turn makes the next simulation more realistic. ## Closing exercise 1. Deliberately break a happy path (skip confirmation wording, change bag details mid-flow) 2. Add one tighter constraint or confirmation utterance 3. `make train && make inspect` 4. Confirm the failure mode is gone ## What you built | Capability | Skill | | --------------- | ------------------------------------------------ | | Orientation | `intro` | | FAQ | `trip_faq` | | Itinerary | `check_itinerary` | | Flight status | `flight_status` | | Baggage report | `report_baggage` | | Auth + change | `authenticate`, `find_booking`, `change_booking` | | Handoff / close | `human_handoff`, `goodbye` | | Voice | Deepgram ASR + TTS via Inspector | ## Where to go next - Companion repo: [rasa-voice-agent-tutorial](https://github.com/RasaHQ/rasa-community-resources/tree/52986f71071e419533f5f00335440a7ff1aef500/tutorials/rasa-voice-agent-tutorial) - Showcase on this site: [/showcase/](/showcase/) - Voice failure modes course: [Why Voice AI Agents Fail](https://github.com/RasaHQ/why-voice-ai-agents-fail) - Production channels beyond Inspector: see Rasa speech integration docs for your telephony stack You now have the Skills progressive-control spectrum and a voice harness you can reuse in any domain. --- # A phone appointment is not a branch visit Source: https://rasa.community/library/casebook/banking-advisor-appointment--conversation-designer/ Author: Rod Rivera Published: 2026-09-18T09:00:00.000Z A customer has chosen a phone appointment. Just before confirming, they ask to come into a branch instead. The assistant replies, “Of course,” and books the phone slot it already had. The reply accepted the correction. The booking did not. This is a fictional scheduling exercise at Northgate Bank. It starts with a small design problem: how should the conversation change when the customer's choice invalidates the appointment the system is about to reserve? You will produce a revised confirmation, a state table and three replay tests that an engineer can use to check the implementation. ## Start with the detail that changed In this example, the customer wants help with online banking. The proposed appointment is with a suitable support advisor by phone, on Tuesday at 10:00. That is the offer on screen when the customer speaks: > **Customer:** Actually, I need to come into a branch. It needs to be step-free. > > **Assistant:** I'll look for a step-free branch with someone who can help > with online banking. The Tuesday phone appointment isn't booked. Do you want > me to give up that phone slot while I look? > > **Customer:** Yes, I'd rather come in. The second sentence is permissible only if the system knows that no booking has been made. If a booking request has already been sent and its result is unknown, the assistant needs the different response in the timeout section below. Do not turn “step-free” into a request for the customer's medical history. It is an appointment requirement. Nor should the assistant treat the nearest branch as suitable merely because it has a free slot: advisor capability, access and the customer's choice must all match. For the exercise, assume the scheduling service can report whether a held slot has been released. If your service cannot do that, the designer and scheduling owner need to resolve that gap before promising the customer a replacement. ## Rewrite the confirmation from the new choice A confirmation such as “Shall I book that?” leaves the customer guessing what “that” refers to. Write it from the appointment record the booking tool will receive. Here is the revised offer after the old hold has been released and a suitable branch slot found: > **Assistant:** The Market Street branch has step-free access and an online > banking support advisor available on Wednesday at 11:30. Would you like that > appointment? > > **Customer:** Yes, book Wednesday. > > **Assistant:** Your branch appointment is booked for Wednesday at 11:30. > Your booking reference is NG-204. The Tuesday phone slot has been released. Market Street and NG-204 are illustrative values, not results from a live service. The last turn needs two separate observations: a receipt for the new booking and confirmation that the earlier hold was released. If only one is known, show only that part as complete. The user has confirmed a particular purpose, channel, location and time. Changing one of those details invalidates that confirmation. A previous “yes” to Tuesday by phone cannot authorise Wednesday in a branch. :::checkpoint{id="appointment-correction-confirmation" question="The customer agreed to Tuesday by phone, then asked for a branch visit. Can the assistant book the nearest available branch slot?" options="Use the earlier confirmation,Show the new appointment and ask for confirmation,Book if the advisor is qualified" answer="1"} The changed offer needs its own confirmation. Advisor capability is necessary, but it does not substitute for the customer's choice of place and time. ::: ## Give each waiting state its own message Use this table as a handoff to engineering. It describes the information the interface needs; it is not a claim that every scheduling API provides it. | State | What the customer sees or can do | Evidence required to move on | | ------------------------ | ---------------------------------------------------- | -------------------------------------------------------------- | | Phone slot offered | Accept, change the appointment, or decline | The exact purpose, channel and time shown in the offer | | Change requested | Confirm giving up the phone slot, or keep it | The customer's choice; no branch booking yet | | Releasing the old hold | Wait, ask for help, or stop searching | The scheduling service reports the hold released | | Branch slot offered | Check access, place and time; accept or change again | A matching advisor and a current slot hold | | Booking submitted | See that confirmation is pending; ask for help | A booking receipt or an authoritative failure for this request | | Branch booking confirmed | Read the details and reference | The receipt matches the newly confirmed appointment | There are two tempting shortcuts to reject during review. Keeping the phone slot while announcing that the branch visit is booked can leave an unwanted hold. Releasing it before asking the customer can take away an option they wanted to keep. The dialogue above asks which tradeoff the customer wants, then makes the chosen action visible. If there is no suitable branch slot, offer a specific next step: keep searching other dates, arrange a callback from the scheduling team, or stop. Do not quietly replace the specialist with a general appointment. Ask the scheduling owner to confirm who receives callback requests and what response time can be promised; do not invent one in the script. ## When the booking response never arrives A timeout after submission is a different state from a rejected booking. It does not establish that the service did nothing. A useful message at this point is: > **Assistant:** I haven't received confirmation of the Wednesday booking yet. > I'll check the request already sent before trying again. You can wait here or > ask the scheduling team to follow it up. Engineering needs to retain the original request identifier and look up that request. Sending a new identifier for the same booking can create a duplicate. A customer's request to change the appointment again must not be squeezed into that earlier identifier either: first establish what happened to the submitted booking, then handle the change as a separate decision. Agree one escalation rule for this exercise: if the first status lookup cannot establish the outcome, stop automated booking retries and send the scheduling team the request reference, the requested appointment and the unresolved status. Tell the customer that follow-up has been requested only after the receiving service acknowledges it. Avoid copying the entire conversation into that record. ## Use the small lab for the claim it actually tests The [downloadable lab](/casebook/casebook.zip) contains a synthetic appointment fixture. Extract it into a new folder and run this command from that folder with Python 3.11 or later: ```bash python3 casebook.py --case banking-advisor-appointment --prove ``` The result includes ten fixture observations and three removed-predicate tests. For each of `purpose_matched`, `slot_held` and `channel_confirmed`, false, missing and the string `"true"` must block the synthetic action. These checks help explain why a conversational “yes” cannot fill in a missing service fact. They do **not** test branch accessibility, slot release or the dialogue above. You can also inspect the lost-response behaviour yourself. Use a fresh database for these two commands so an earlier run does not become a replay: ```bash python3 casebook.py --case banking-advisor-appointment --database appointment-demo.sqlite --request-id branch-1 --lose-ack python3 casebook.py --case banking-advisor-appointment --database appointment-demo.sqlite --request-id branch-1 --lookup ``` The first result is `pending`, without a reference. The lookup returns `succeeded`: the synthetic action had been stored even though its acknowledgement was lost. This is the reason the waiting message cannot say “nothing was booked”. There is no real bank, Rasa conversation or external scheduling integration in this proof. Test the actual adapter separately before using the design in a customer service. ## Rehearse the correction before approving the flow Give a support colleague the dialogue prototype and these three tasks. Let them use the available choices; do not explain the state table first. 1. Choose Tuesday by phone, then ask for a step-free branch visit. After the revised offer, ask which appointment is currently booked. The expected answer is neither: the branch visit still needs confirmation. 2. Confirm the branch visit, then withhold the booking response. Ask whether trying again is safe. The prototype must show that the original outcome is still being checked, and provide the scheduling-team route. 3. Ask for a branch visit when no suitable advisor is available. The colleague should be able to choose other dates, a callback or stopping, without being steered into a general appointment they did not request. Record their answer at each transition alongside the prototype's displayed state. Count each misunderstanding once per person per task; retain the number of people who tried that task. These sessions reveal wording problems, not a population-level success rate. Do not approve the flow if a participant believes a pending appointment is confirmed or cannot find the change and stop options. Revise that transition and replay it. Then ask engineering to demonstrate the same sequence with the real service: release the old hold, reject the old confirmation and inspect the new booking receipt. The state table and those observations are the design handoff; a friendly transcript alone is not enough. --- # Preventing SIM swap fraud in a voice agent Source: https://rasa.community/library/casebook/telco-sim-swap/ Author: Rod Rivera Published: 2026-09-21T21:00:00.000Z The caller sounds calm. They have lost their phone, they say, and they have a new SIM in their hand. Could the number move across tonight? Here is an illustrative script of an agent without the rule this piece builds. The numbers and the code are made up. > **Caller:** I lost my phone this afternoon. I've got a new SIM from the shop > and I need my number on it. > > **Assistant:** I can help with that. For security, I've texted a six-digit > code to your number ending 0142. Can you read it to me? > > **Caller:** Four, one, nine, zero, two, seven. > > **Assistant:** Thank you, that's verified. What's the number on the new SIM? > > **Caller:** It ends 3381. > > **Assistant:** Done. Your number is queued to move to the new SIM, reference > SWP-7C21E0A4. Every step looks careful. The assistant verified before acting, used a one-time code rather than a date of birth, and asked for the new SIM number. It may also have handed the number to someone who was already receiving its texts. Consider who can read that code. If the phone really is lost, a code texted to it reaches the owner only if their texts also arrive on another device of theirs. So a caller who reads it back is one of three people: the owner on another of their devices, an owner whose phone is not lost after all, or someone else who receives texts for 0142, for example through a forwarding rule. Whichever it is, the code shows only that someone can read texts sent to the line that is about to move. **The line being replaced cannot be the only witness to its own replacement.** This is a fictional build study. Telecom of Rasa is an invented carrier, and the policy is a teaching contract rather than any real carrier's procedure. The code is in `examples/mantle-voice-telco-skills` in `RasaHQ/rasa-community-resources` at revision `69e27b6`, where the assistant is called Telano. You will see why the obvious check fails, write the rule that holds, put it inside the tool that queues the swap, and prove it with tests that need no model and no key. ## Why a code texted to the line proves nothing here The step-up pattern in the same repository, `patterns/voice-auth-stepup`, ranks factors: a spoken passphrase earns medium trust and a one-time code earns high trust (`authpolicy/challenges.py`, lines 103–121). The [guarded card reissue tutorial](/library/tutorials/guarding-irreversible-actions/) builds on the same idea. The pattern is candid about the ladder's weak rung. Its own docstring says a code read aloud on the same call "is closer to a second knowledge factor than to real possession, because a caller who has been socially engineered will read it out to the attacker." A SIM swap adds a second, worse problem. The code still reached a real phone, and no one had to be talked into anything. What makes it worthless is **where it was delivered**: to the thing the action changes. A ranking of factors cannot see that, because it never looks at the action's target. :::checkpoint{id="sim-swap-what-the-code-proves" question="A caller reads back, correctly, a one-time code that was texted to 555-0142, then asks to move 555-0142 to a new SIM. What has the correct code proved?" options="That the caller owns the account,That someone can read texts sent to 555-0142,That the caller holds the new SIM,Nothing at all" answer="1"} It proves possession of whatever receives texts for 555-0142, which is the one thing an attacker who has taken over the line also has. It says nothing about who owns the account, so it cannot approve moving that line. ::: ## Prevent the swap with an allowlist, not a ban list The tempting fix is a second rule: refuse a code sent to the target line. The Telano policy has that rule, but it is not what protects the swap. What protects it is an allowlist: two approved channels, and a refusal for everything else. The first is a push to the carrier's app on a device that was registered to the account before the call started; the tool compares the device's registration time with the start of the call (`test_device_registered_during_the_call_does_not_count`). The push goes to a device, not to the phone number being moved. The second is a store visit with photo ID, which a voice agent cannot perform but a store's system can record. The push is still confirmed by a spoken code. In a real deployment `send_swap_verification` would deliver a code to the app; the companion simulates that with a fixture code. The caller reads it back, and `confirm_swap_verification` compares it in constant time, then drops it. Two wrong codes lock the call and hand off to a person, and asking for another push does not reset the count (`skills/sim_swap/tools.py`, lines 75, 260–261 and 589–592). So the pattern's warning still applies in one form: an owner who is talked into reading the in-app code to an attacker defeats this check. Independence from the line closes the circular hole; it does not close that one. The decisions run in a fixed order, and anything that matches no allowed channel falls through to a refusal: :::diagram{title="The order of checks in evaluate_swap, ending in a default refusal"} ```dot rankdir=TB; start [label="Request to move a line"]; k [label="PIN, date of birth or\nsecurity answer?", shape=diamond]; t [label="Verification issued\nfor a different line?", shape=diamond]; c [label="Code sent to the\nline being moved?", shape=diamond]; p [label="Challenge passed?", shape=diamond]; a [label="App push to a device registered\nbefore the call, or store ID check?", shape=diamond]; r1 [label="knowledge_only", class="blocked"]; r2 [label="target_changed", class="blocked"]; r3 [label="circular_verification", class="blocked"]; r4 [label="no_verification", class="blocked"]; r5 [label="no_verification\n(default refusal)", class="blocked"]; ok [label="allowed", class="ok"]; m [label="Record missing or malformed,\nor no target line?", shape=diamond]; r0 [label="no_verification", class="blocked"]; start -> m; m -> r0 [label="yes"]; m -> k [label="no"]; k -> r1 [label="yes"]; k -> t [label="no"]; t -> r2 [label="yes"]; t -> c [label="no"]; c -> r3 [label="yes"]; c -> p [label="no"]; p -> r4 [label="no"]; p -> a [label="yes"]; a -> ok [label="yes"]; a -> r5 [label="no"]; ``` ::: The same order in code (`lib/sim_swap.py`, lines 250–272): ```python if record.channel in KNOWLEDGE_CHANNELS: return refuse(KNOWLEDGE_ONLY) if line_digits(record.issued_for_line) != line_digits(target): return refuse(TARGET_CHANGED) if record.channel in LINE_CHANNELS and line_digits(record.destination) == line_digits( target ): return refuse(CIRCULAR_VERIFICATION) if not record.passed: return refuse(NO_VERIFICATION) if record.channel == APP_PUSH and record.registered_before_call: return SwapDecision(True, ALLOWED, target, record.channel) if record.channel == STORE_ID_CHECK: return SwapDecision(True, ALLOWED, target, record.channel) # A code to some other number, or a push to a device enrolled during this # call. Not circular, not knowledge, and not independent either. return refuse(NO_VERIFICATION) ``` The allowlist lives in those last lines: only two branches return `True`, and everything else falls through to the final refusal (lines 264–272). An SMS to some other number ends there. The companion test module records what happened when the author deleted two of the named checks on 21 September 2026. With the circular check gone, only the SMS test failed, and only because the reason code changed: **the swap was still refused, as `no_verification`.** Deleting the knowledge check gave the same result. Those two rules make a refusal explain itself; the allowlist does the refusing. The module also declares a set of the same two channels (lines 102–106): ```python #: The only channels that can authorise a swap. This set is a declaration, not #: a router: adding a name here does NOT create an approval path, because no #: code allows a channel merely for being in it. A new channel needs its own #: explicit branch in `evaluate_swap` and in the send tool. INDEPENDENT_CHANNELS = frozenset({APP_PUSH, STORE_ID_CHECK}) ``` `evaluate_swap` never reads it; the send tool uses it as a gate before it issues anything. That split was a real trap. An earlier version of the send tool gated on the set alone and issued every accepted channel other than the store check as an app push. With `"email"` patched into the set, that version accepted it and wrote a push record, which is exactly what the policy allows. The send tool now handles each channel by name and refuses anything else, and `test_widening_the_channel_set_does_not_open_a_new_path` holds it there (`tools.py`, lines 414–420). If you copy this pattern, derive both the send step and the decision from one declaration, and test what happens when someone widens it. ## Bind the verification to one line The wrong-line check is different, and it is not a label. The verification record stores the line it was issued for (`issued_for_line`) alongside the channel, the destination and whether the challenge passed. Without that binding, a caller who verifies one of their lines properly could use the same verification to move another. The harm is concrete: an attacker who talks the owner into approving a push "for the tablet line" could then move the main number on that approval. `test_changing_the_target_line_discards_the_verification` checks the mirror case. A caller verifies the main line, 555-0142, with an app push, then asks to move the tablet's data line, 555-0187. The policy returns `target_changed`, the tool discards the verification, and switching back to 555-0142 does not bring it back. The deletion record shows what the check is worth: with it removed, the policy returned `allowed` in that test, and in `test_refusal_returns_no_swap_reference` the tool queued a swap of 555-0142 on a push issued for 555-0187. The record never stores the code, which would otherwise travel into memory, logs and tool results that a voice agent reads aloud. It is also read strictly (`lib/sim_swap.py`, lines 222–225): ```python # `type(...) is bool` rather than isinstance: True is also an int, and 1 is # not a verification outcome. if type(passed) is not bool or type(registered) is not bool: return None ``` A `passed` value of the string `"true"`, an unknown channel or a missing field all become `no_verification` (`test_malformed_verification_fails_closed`). In this project the model cannot write the verification slot at all: `memory.yml` does not mark it `llm_settable`, and Rasa Pro 3.20.0rc1 defaults that setting to false. So the strict read is defence in depth: it keeps the guard honest if another writer, or a later configuration change, puts a friendly-looking string there. ## Put the guard inside the tool The skill's frontmatter says `request_sim_swap` needs a target line and a spoken confirmation. That keeps the conversation coherent, but the step-up pattern's `authpolicy/guard.py` explains the limit: frontmatter is a routing control. Its `requires` condition shapes which tool the model is offered, and the confirmation is a pause that the conversation resolves; neither decides what the process does when the function runs. So the decision that binds is the first statement of the tool (`skills/sim_swap/tools.py`, lines 624–638, docstring omitted): ```python async def request_sim_swap( line: str = "", new_iccid: str = "", context: ToolContext = None, ) -> ToolResult: # (docstring omitted) try: require_independent_verification(line, context) except SwapRefused as exc: return _refuse(exc, context) ``` If the skill prose is edited, the model swapped or the YAML mistyped, the function still refuses. A refusal returns `ok: false`, a reason code and a hint for the agent. It writes no swap row and returns no reference, status or SIM number that could be read out as if the swap had happened. ## Run the proof without a model The tests call the policy and the tools directly, with a fake context and a throwaway database, so they show what the tool does whatever the model chooses. The example pins Rasa Pro 3.20.0rc1 and uses `uv`. Clone `RasaHQ/rasa-community-resources`, check out `69e27b6`, change into `examples/mantle-voice-telco-skills`, then run: ::run{mac="make test" linux="make test"} The last lines of one run on 21 September 2026 (the timing varies between machines): ```text ---------------------------------------------------------------------- Ran 11 tests in 0.028s OK ``` The module docstring records every deletion check, and you can repeat them on your own copy. With the four guard lines removed from `request_sim_swap`, eight of the eleven tests failed, each by getting a swap reference or a queued row for a caller who had not verified independently. `test_resending_a_push_does_not_reset_the_attempt_budget` covers the lockout: one wrong code, a fresh push, a second wrong code, and the call is locked. These are offline checks of the policy and tools. They say nothing about how a model behaves in a live call or about any real carrier's systems. ## Say only what the receipt proves An accepted request returns `status: "queued"` and `active: false`, and `check_swap_status` reads the real state from the store rather than assuming it. `test_a_queued_request_is_not_reported_as_active` checks those tool fields. What the agent says is a separate layer: the skill tells it to read the reference, say the request is queued, and call the new SIM active only when the status tool reports it. :::callout{type="pitfall" title="The tool proves the fields, not the sentence"} A correct `active: false` does not stop a model from saying "your new SIM is working". Test the spoken confirmation in your own conversation tests; the companion's offline suite does not cover it. ::: ## Write the refusals the rule allows The skill tells the agent never to offer a text or call to the line being moved, never to fall back to a knowledge question, and to hand off to the identity team when no independent path exists. An illustrative script of that behaviour: > **Caller:** I lost my phone this afternoon. I've got a new SIM and I need my > number on it. > > **Assistant:** I can start that for the number ending 0142. A code sent to > that number can't approve moving it, so I can send one to the Telecom of > Rasa app on a device already registered to your account, or you can visit a > store with photo ID. > > **Caller:** Can't you just text me? > > **Assistant:** A text would go to the number we're moving, so it can't > approve moving it. If neither the app nor a store works for you, I'll pass > you to our identity team. The send tool returns only that a push went to a registered device, never which one, so the agent has nothing to reveal to a caller who has not verified. That is checked in `test_unverified_caller_learns_nothing_about_the_device`. Whether a model follows the rest of the script in a live call is something to test with your own conversations; the tool refuses either way. ## Questions you will hit while building it :::solution{title="What if the registered device is the lost phone?"} It may be. If the owner lost it, the push never reaches them and the store is their path. If a thief holds it unlocked with the app signed in, the push reaches the thief. This rule proves the approval came from a device registered before the call, independent of the line; it does not prove the device is still in the owner's hands. One mitigation, which the companion does not implement or test, is to require the app's own unlock before it shows the code. Treat a stolen handset as a separate risk in your review. ::: :::solution{title="Isn't an SMS one-time code the standard step-up for risky actions?"} For many actions it can be reasonable, with the caveat the step-up pattern itself gives. For a SIM swap the Telano rule does not use SMS at all, even to another number: only the app push and the store check count. A short allowlist is easier to review than a list of every way a code can be circular. ::: :::solution{title="What if no device was registered before the call and the customer cannot reach a store?"} Then the swap is not requested on this call. `send_swap_verification` reports that no independent path exists and asks for a handoff; the skill tells the agent to do the same when the caller cannot use the app. Some legitimate customers will be inconvenienced; that cost belongs in the product decision and in the [permission review](/library/guides/review-agent-permissions/), not in a fallback the model improvises. ::: :::cta{href="/library/tutorials/guarding-irreversible-actions/" label="Build the execution guard step by step"} The card reissue tutorial works through declaring risk, the execution guard, address provenance, retries and refusal paths. :::