Skip to content
RasaGet a free licence
Guides for AI teams

Guide · Evaluation / data specialist

Test passes but the real code path fails: derive the fixture

A hand-typed session made a handoff test pass while the agent never collected account_id. Build fixtures from the key list the running code reads.

by Rod Rivera

About 7 minutes

Key takeaways (5)
  • A fixture you typed by hand is evidence about what you assumed the code collects. It is not evidence about what the code collects.
  • Judge a fixture by where its keys came from. In the handoff pattern, a plausible hand-typed session kept the hand-built test green on a copy where the handoff tool no longer collected account_id.
  • Take the fixture’s keys from the list the running code reads, intersected with the keys your code writes. The test then fails when a key it names drops out of either list.
  • Assert that named keys are present in the derived list. A loop over a list the parser failed to read makes no assertions and passes.
  • Deriving the keys does not derive the values or the conversion between memory and your function. The companion test still types attempts by hand, in the shape the tool produces after parsing attempts_log.

In a scratch copy of the handoff pattern from RasaHQ/rasa-community-resources, delete one line, "account_id",, from the tuple of memory keys the handoff tool reads. Then run two tests that assert the same claim: after a transfer, the human desk asks none of the five questions in its opening script (DESK_OPENING_SCRIPT, handoffpkg/desk.py lines 42–48) again.

$ python3 -m unittest -v tests.test_handoff_context.TestPackageSurvivesHandoff.test_caller_is_never_asked_a_question_they_answered tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question
test_caller_is_never_asked_a_question_they_answered (tests.test_handoff_context.TestPackageSurvivesHandoff.test_caller_is_never_asked_a_question_they_answered)
The teaching claim, as a number rather than a sentence. ... ok
test_the_agent_path_retires_every_desk_question (tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question)
THE HEADLINE CLAIM, ON THE PATH THE AGENT ACTUALLY RUNS. ...
  test_the_agent_path_retires_every_desk_question (tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question) (key='account_id')
THE HEADLINE CLAIM, ON THE PATH THE AGENT ACTUALLY RUNS. ... FAIL
test_the_agent_path_retires_every_desk_question (tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question)
THE HEADLINE CLAIM, ON THE PATH THE AGENT ACTUALLY RUNS. ... FAIL

That is an offline run: Python 3.14.3, a scratch copy of the pattern at commit 69e27b6, network access denied and no licence set. The excerpt keeps the verbose lines and cuts the two tracebacks and the summary lines, Ran 2 tests and FAILED (failures=2). Both assertion messages appear in the section on dropping a key.

The first test checks a package built from a session typed into the test file. It stays green although the tool can no longer carry the account. The second builds its session from the lists the agent’s own code reads and writes, and fails twice. The broken copy recreates a defect the pattern really shipped, which the second test’s docstring records (lines 344–349 of patterns/voice-handoff-context/tests/test_handoff_context.py):

        This test exists because the claim once passed while being false. The
        eval suite built its own session dict containing `account_id`, so
        `intent.details` was populated and "which account is this about?" was
        retired — but the agent's own `_MEMORY_KEYS` did not collect
        `account_id`, so a real handoff shipped an empty `details` and the desk
        asked anyway. Green tests, broken claim.

No model misbehaved. The handoff tool did what its code said, the package builder did what its code said, and the assertion was correct about the input it was given. The input was the problem. The person writing the test had typed it, and it held a field the running agent never gathered. Rerunning the suite could not expose that, because every run built the same session.

A fixture typed by hand is evidence about what its author assumed the code collects. Take the fixture’s keys from the list the running code reads, and that assumption becomes something a test checks. The change is to how you build one fixture; the suite does not grow. It costs you coupling between your tests and your source layout, and the last two sections show where that coupling breaks.

The caller, accounts and charges in the pattern are fixture values. Every run in this guide is offline, with no model and no licence, and each saved receipt records its conditions.

Why the test passes but the real code path fails

The hand-built test is short (lines 185–187):

    def test_caller_is_never_asked_a_question_they_answered(self):
        """The teaching claim, as a number rather than a sentence."""
        self.assertEqual(unanswered_questions(self.package), ())

Its class builds self.package in setUp by passing demo_session() to build_package_from_session (lines 133–134). demo_session() is a session typed into the test file. The start of it, lines 77–94:

def demo_session() -> dict:
    """Session state at the moment the caller asks for a human."""
    session = {
        # --- allowlisted: the state that SHOULD cross -----------------------
        "customer_id": "cust_00417",
        "display_name": "Jordan Rivera",
        "verified_tier": "medium",
        "verified_factors": ["knowledge_passphrase"],
        "channel": "voice:+1-555-0100",
        "goal": "dispute_transaction",
        "goal_label": "Dispute a card transaction",
        "goal_stage": "blocked",
        "account_id": "acc_checking",
        "account_label": "Everyday Checking",
        "card_last_four": "4821",
        "dispute_amount": "$248.00",
        "dispute_merchant": "Northgate Fuel",
        "dispute_date": "2026-08-29",

By the usual standard it is a good fixture. The names match the memory schema and the values are plausible.

On a real call the handoff tool, transfer_to_human in skills/human_handoff/tools.py (line 148), builds its session by walking one named tuple (lines 155–160):

    # Step 1 — collect indiscriminately. This tool is not the safety boundary.
    session: dict = {}
    for key in _MEMORY_KEYS:
        value = context.memory.get(key)
        if value is not None:
            session[key] = value

A key that is not in _MEMORY_KEYS never reaches the builder, whatever memory holds. The tuple runs from line 68 to line 96. Its first lines, 68–84, carry a comment left when the defect was fixed:

_MEMORY_KEYS = (
    "customer_id",
    "display_name",
    "verified_tier",
    "verified_factors",
    "channel",
    "goal",
    "goal_label",
    "goal_stage",
    # The goal's parameters -> intent.details. Omitting these was a real defect:
    # the eval suite passed on a hand-built session while the live agent shipped
    # a package with an empty `details`, so the desk still asked "which account
    # is this about?". test_the_agent_path_retires_every_desk_question now
    # exercises THIS list rather than a hand-built one.
    "account_id",
    "account_label",
    "card_last_four",

The dispute skill had the same gap on the writing side. The loop in load_caller_context that writes the caller profile, skills/dispute_transaction/tools.py lines 83–91, carries a matching comment:

    for key in (
        "customer_id", "display_name", "verified_tier", "verified_factors", "channel",
        # The account and card the caller is asking about. These are what make
        # `intent.details` non-empty, which is what retires "which account is
        # this about?" on the desk. Omitting them was a real defect: the eval
        # suite passed on a hand-built session while the live agent produced a
        # package that still made the desk ask.
        "account_id", "account_label", "card_last_four",
    ):

So before the fix, the fixture held account_id, as the docstring says, and both lists the running agent uses lacked the three detail keys. The skill’s comment names the three keys, and the tool’s comment sits directly above them. Set today’s demo_session() beside what the two comments say:

KeyIn demo_session() todayWritten by the dispute skill before the fixIn _MEMORY_KEYS before the fixReaches intent.details on a call
account_idyesnonono
account_labelyesnonono
card_last_fouryesnonono

In measurement terms the hand-built test had a construct validity problem. Cronbach and Meehl set out the idea in 1955 (“Construct validity in psychological tests”, Psychological Bulletin 52(4), 281–302): does a test measure the thing it is named for? This one measured something real, the builder retiring the account question when account_id is present. It was named for a handoff on the agent’s path retiring it. A more realistic fixture would have left that gap open. Realism describes values, and the defect was in which keys existed.

Build the session from the list the tool reads

Three lists of keys are in play: the one typed into demo_session(), the keys the dispute skill’s tools write to memory, and _MEMORY_KEYS. Each test takes a different route to the same builder.

demo_session() typed by hand build_package_from_session hand-built test, called directly Keys the dispute skill writes transfer_to_human (real call) memory Derived session: collected and written, plus typed values parsed _MEMORY_KEYS which keys it reads parsed after converting some fields derived test, called directly
FigureThree key lists, three routes into one builder

The hand-built test calls build_package_from_session directly with demo_session(), and neither agent list touches its input. A real call reaches the builder through transfer_to_human. That route reads the keys in _MEMORY_KEYS from memory the skill’s tools wrote, converts some fields (lines 162–175) and then calls the builder (line 180). The derived test also calls the builder directly, with a session whose keys it parsed from the two agent lists.

The fix lives in the TestTheAgentPathSpecifically class. Two helper methods read the agent’s own lists. The first reads _MEMORY_KEYS out of the tool module (lines 306–323):

    def _agent_memory_keys(self) -> tuple[str, ...]:
        """Read _MEMORY_KEYS out of the tool module without importing rasa.

        The tool module imports `rasa.mantle`, which the eval suite deliberately
        does not depend on — `make test` runs with no license and no install.
        Parsing the literal keeps this test honest about the real list while
        staying offline.
        """
        source = (
            Path(__file__).resolve().parents[1]
            / "skills" / "human_handoff" / "tools.py"
        ).read_text()
        block = source.split("_MEMORY_KEYS = (", 1)[1].split(")", 1)[0]
        return tuple(
            line.strip().strip(",").strip('"')
            for line in block.splitlines()
            if line.strip().startswith('"')
        )

The second, _skill_memory_writes() (lines 383–398), finds the keys the dispute skill’s tools write. It searches skills/dispute_transaction/tools.py for memory.set(...) calls, _append(...) calls and the for key in (...) loop in load_caller_context. Its body, lines 393–398:

        keys = set(re.findall(r'memory\.set\(\s*"([^"]+)"', source))
        keys |= set(re.findall(r'_append\(\s*context,\s*"([^"]+)"', source))
        # The `for key in (...)` loop that writes the caller profile.
        for block in re.findall(r"for key in \(([^)]*)\):", source):
            keys |= set(re.findall(r'"([^"]+)"', block))
        return keys

Neither helper imports the module it reads. Both tool modules import from rasa.mantle, and the suite is meant to run with no Rasa install, so the tests parse the source text instead. The last section weighs that choice for your own project.

The test then uses both lists (lines 355–366):

        written = self._skill_memory_writes()
        collected = self._agent_memory_keys()

        # Every detail key the allowlist expects must be both written by a skill
        # and collected by the handoff tool, or it never reaches the package.
        for key in ("account_id", "account_label", "card_last_four"):
            with self.subTest(key=key):
                self.assertIn(key, written, f"no skill writes {key} to memory")
                self.assertIn(key, collected, f"the handoff tool does not collect {key}")

        # Now build a session from ONLY what the agent genuinely produces.
        session = {key: f"value_for_{key}" for key in collected if key in written}

The session is the intersection of what the skill writes and what the tool collects, because a key reaches the package only if both happen. The values are placeholders: value_for_account_id looks like nobody’s account. Apart from attempts, which the test adds by hand in its next lines, every key in the session comes from those two parsed lists.

A second test, test_every_key_the_agent_collects_is_either_allowlisted_or_withheld, builds its session from every key _agent_memory_keys() returns, with no intersection (lines 400–404). It asserts that each collected key is either on the allowlist or named in the package’s withheld_fields. A hand-picked fixture covers the keys its author thought of. This test covers every key the tool collects.

Here is the whole suite at the pinned revision, 69e27b6, run offline with Python 3.14.3 from patterns/voice-handoff-context. The excerpt keeps the three agent-path tests and the last lines of the run, and cuts the other 38 tests:

$ python3 -m unittest tests.test_handoff_context -v
test_every_key_the_agent_collects_is_either_allowlisted_or_withheld (tests.test_handoff_context.TestTheAgentPathSpecifically.test_every_key_the_agent_collects_is_either_allowlisted_or_withheld)
No third category. A key is a decision, and both outcomes are visible. ... ok
test_the_agent_collects_credentials_and_the_allowlist_stops_them (tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_collects_credentials_and_the_allowlist_stops_them)
End to end on the real key list: collected, then withheld. ... ok
test_the_agent_path_retires_every_desk_question (tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question)
THE HEADLINE CLAIM, ON THE PATH THE AGENT ACTUALLY RUNS. ... ok

----------------------------------------------------------------------
Ran 41 tests in 0.004s

OK

That shows the technique runs today. It is an offline predicate test on parsed source lists. It is not a recording of any call.

Drop a key and see which test notices

The opening run applied a mutant by hand: one small, deliberate change to production code, made to see whether the tests notice. That is mutation testing, which DeMillo, Lipton and Sayward set out in 1978 (“Hints on Test Data Selection: Help for the Practicing Programmer”, IEEE Computer 11(4), 34–41). The mutation testing guide covers choosing mutants and counting what they kill. Here one mutant is enough, because it is the original defect. In the scratch copy, delete account_id from _MEMORY_KEYS:

@@ -80,5 +80,4 @@
     # is this about?". test_the_agent_path_retires_every_desk_question now
     # exercises THIS list rather than a hand-built one.
-    "account_id",
     "account_label",
     "card_last_four",

Then run the whole suite and keep the failure lines:

$ python3 -m unittest tests.test_handoff_context 2>&1 | grep -E '^(FAIL|AssertionError|Ran|FAILED)'
FAIL: test_the_agent_path_retires_every_desk_question (tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question) (key='account_id')
AssertionError: 'account_id' not found in ('customer_id', 'display_name', 'verified_tier', 'verified_factors', 'channel', 'goal', 'goal_label', 'goal_stage', 'account_label', 'card_last_four', 'dispute_amount', 'dispute_merchant', 'dispute_date', 'attempts_log', 'questions_answered', 'factors_verified', 'confirmed_facts', 'handoff_reason', 'pin_attempt', 'otp_code') : the handoff tool does not collect account_id
FAIL: test_the_agent_path_retires_every_desk_question (tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question)
AssertionError: None is not true : intent.details is empty on the live path — the desk will re-ask which account this is about
Ran 41 tests in 0.009s
FAILED (failures=2)

This run was offline too, on the same scratch copy, with Python 3.14.3. Both failures come from the derived test: once at the named-key check, and once at the package itself, where intent.details has no account_id. That is one failing test with two failures; the other 40 tests pass, including the hand-built one from the opening run. The same claim, asserted against two fixtures, gives two answers on the same broken code, and only the fixture built from the agent’s lists gives the right one.

The derived test names three keys, so this check guards those three. A mutant that deletes a key no assertion names can pass, as the next section shows.

What the derived fixture still does not prove

Deriving the key set closes one gap. It leaves others, and a report on this suite should name them.

Some values are still typed. After building the session from the intersection, the test sets four entries itself (lines 367–372):

        session["verified_tier"] = "medium"
        session["goal"] = "dispute_transaction"
        session["questions_answered"] = "\n".join(DESK_OPENING_SCRIPT)
        session["attempts"] = [
            {"action": "send_otp_sms", "outcome": "failed", "code": "delivery_failed"}
        ]

Three of them overwrite keys the lists supplied. The fourth, attempts, is in neither list. The skill writes attempts_log, a newline-separated text field (_append, lines 63–69 of the dispute tools), in the action|outcome|code|detail form that _parse_attempts reads (handoff tool, lines 117–118). On a real call the tool pops that field and converts it before calling the builder (line 175):

    session["attempts"] = _parse_attempts(session.pop("attempts_log", None))

The test calls the builder directly. Its session keeps a placeholder attempts_log next to a hand-typed attempts in the converted shape, and the conversion never runs on the path it checks. The handoff measurement guide follows what that means for the desk’s question about earlier attempts.

The parsers read source text. _skill_memory_writes() treats any double-quoted string inside a for key in (...) block as a key, comments included. At 69e27b6 its set contains one entry that is not a key at all: 'which account is\n # this about?', lifted from the comment in the for key in (...) loop in load_caller_context. It does no harm here, but a key quoted in a comment inside that loop would count as written. The regex also counts memory.set("otp_code", ...), although the skill sets otp_code only on the branch where the tier check fails (dispute tools, lines 173–195). The stray entry comes from an offline run of both parsers.

The intersection also hides keys that are collected and never written. The same offline parser run prints the difference between the two sets. One of its four output lines:

collected, not written: ['dispute_amount', 'dispute_date', 'dispute_merchant', 'factors_verified', 'handoff_reason']

No tool in skills/dispute_transaction/tools.py writes dispute_amount, dispute_merchant or dispute_date, so the intersection drops them and the derived test passes without them. No assertion names them. The measurement guide shows what the desk loses as a result.

Finally, none of this tests the model. Every run here is an offline test on parsed lists and a builder function, or an import of a module. It does not show what a model says on a call. It does not catch other kinds of fixture drift either, such as a policy string copied into a test and later changed in one place only.

Apply it to your own project: parse the source or import the module

Start by finding the list your code under test reads or writes at run time: a tuple of memory keys, the fields a tool returns, the slots a flow fills. That list is the source for your fixture’s keys. The handoff pattern has two such lists, one on each side, and parses both from source. The other obvious option is to import them.

ApproachCollected keys come fromWritten keys come fromWhat it costs
Parse the source text, as the companion doesThe literal as written in the handoff toolA regex over the skill’s sourceNo Rasa install in the test job. The parsers break on reformatting, read comments, and cannot see keys built at run time.
Import the handoff tool moduleThe tuple Python evaluates, whatever its formattingStill a regex, or a run of the skill’s tools against a stand-in contextThe test job must install rasa-pro and everything the module imports.
Move the list into a module with no Rasa importOne object both the tool and the test importStill a regex, or a run of the skill’s toolsA small refactor of production code. The companion does not do this.

Importing covers only the collected side. The skill’s writes are calls inside tool functions, and no single module attribute lists them all. Running those tools against a stand-in context would record them; neither this guide nor the companion tests that option.

A licence is not what stands in the way of importing. The two variables rasa/utils/licensing.py names for the licence are RASA_LICENSE and RASA_PRO_LICENSE (lines 24–25 in the pinned wheel). With both unset, rasa-pro 3.20.0rc1 installed and network access denied, the handoff tool module itself imports. It pulls in ToolContext and tool from rasa.mantle.tools.decorator and ToolResult from rasa.mantle.tools.result (lines 27–28). This run used Python 3.12.13. <venv> stands for the virtual environment’s path, as in the saved receipt, and the output is unedited:

$ env -u RASA_LICENSE -u RASA_PRO_LICENSE sandbox-exec -p '(version 1)(allow default)(deny network*)' <venv>/bin/python -c 'import importlib.util, os, sys; print("RASA_LICENSE set:", "RASA_LICENSE" in os.environ, "| RASA_PRO_LICENSE set:", "RASA_PRO_LICENSE" in os.environ); spec = importlib.util.spec_from_file_location("handoff_tools", "skills/human_handoff/tools.py"); mod = importlib.util.module_from_spec(spec); spec.loader.exec_module(mod); import rasa; print("rasa", rasa.__version__); print("imported skills/human_handoff/tools.py; transfer_to_human:", type(mod.transfer_to_human).__name__); print("_MEMORY_KEYS:", mod._MEMORY_KEYS); print("licence modules loaded:", sorted(m for m in sys.modules if "licens" in m))'
21:30:33 - LiteLLM:WARNING: get_model_cost_map.py:454 - LiteLLM: Failed to fetch remote model cost map from https://raw.githubusercontent.com/BerriAI/litellm/main/model_prices_and_context_window.json: [Errno 8] nodename nor servname provided, or not known. Falling back to local backup.
RASA_LICENSE set: False | RASA_PRO_LICENSE set: False
rasa 3.20.0rc1
imported skills/human_handoff/tools.py; transfer_to_human: function
_MEMORY_KEYS: ('customer_id', 'display_name', 'verified_tier', 'verified_factors', 'channel', 'goal', 'goal_label', 'goal_stage', 'account_id', 'account_label', 'card_last_four', 'dispute_amount', 'dispute_merchant', 'dispute_date', 'attempts_log', 'questions_answered', 'factors_verified', 'confirmed_facts', 'handoff_reason', 'pin_attempt', 'otp_code')
licence modules loaded: ['rasa.utils.licensing']

The warning is LiteLLM failing to reach the network the sandbox denied, then falling back to a local copy. The import loaded rasa.utils.licensing and did not fail. In this run, importing the module needed no licence. The run shows nothing about running an agent. The cost of importing is the install.

Import the module when your test environment can afford the install. The collected-key parser goes, and its limits go with it; the written-key side and the other limits remain. Parse the source when the suite has to run with no Rasa install and no licence, as the handoff pattern’s suite deliberately does. Then assert named keys so an empty parse fails, and keep the parsers as narrow as the file allows.

Whichever you choose, finish the way the dropped-key section did. Delete a key your derived test names from the production list in a scratch copy, run the suite, and check that the derived test fails and says which key. If nothing fails, the fixture is not derived from the list you think it is. A key the test does not name can go unnoticed. With dispute_amount deleted from _MEMORY_KEYS, all 41 tests still pass (offline run on a scratch copy).

For the wider design of an evaluation set, including slices and denominators, see the evaluation set guide.

Should I delete the hand-built fixtures?

No. A hand-built session is the right input for testing a function on data you choose, including keys the running code never collects. demo_session() also carries values under keys that are not in _MEMORY_KEYS, such as card_number and ssn (test file, lines 66–74 and 126). A derived session could never contain them. The handoff summary guide covers how the package names withheld fields. Keep hand-built fixtures for that job. Do not use them as evidence that the agent’s path produces those inputs.

My code does not read a named list. It reads whatever memory holds.

Then the declarations are the nearest list, for example the keys in your memory.yml files. Deriving from declarations tells you a field exists, not that anything writes it, which is why the handoff test intersects collected keys with written keys. If neither list exists in code, fix that before the fixture: a transfer that reads everything has no list a test can hold it to.