Skip to content
RasaGet a free licence
Guides for AI teams

Guide · Evaluation / data specialist

186 mutations killed, and the guard still broke

A guard suite reported 186 of 186 mutations killed and still passed a one-word break. Test for the mutation nobody authored, and count the result honestly.

by Rod Rivera

About 9 minutes

Key takeaways (5)
  • A suite that kills every mutation its author wrote has measured only the mutations its author thought of. The count is true and says nothing about any other kind of edit.
  • Before trusting a mutation count, read the loop that produces it: what it changes, which fixture it checks against, and what it calls killed.
  • Deleting one word from the tuple that contract_hash hashes left all seven companion tests passing, including the one that reports 186 mutations killed, while coverage.py still reported the line as executed.
  • Aim your own mutants at whatever decides whether an earlier result still stands (identity hashes, fingerprints, replay checks), because a test that changes only one field protects only that field.
  • Report mutation results per population, with the denominator and the survivors named, not as one number beside a green tick.

Clone RasaHQ/rasa-community-resources, check out 69e27b6, change into tutorials/rasa-ai-team-casebook and run the lab’s offline suite. Seven tests pass. One of them, test_all_authored_oracles_and_mutations, loads 62 scenario fixtures and asserts three killed mutations for each, 186 in all.

Now delete one word. Line 97 of casebook.py is the body of contract_hash, which fixes a case’s identity once a request has been committed. Take 'rules' out of its tuple:

 def contract_hash(spec: dict) -> str:
-    return digest({k: spec.get(k) for k in ('slug', 'action', 'rules', 'receipt', 'provenance', 'intakeKind')})
+    return digest({k: spec.get(k) for k in ('slug', 'action', 'receipt', 'provenance', 'intakeKind')})

Run the same seven tests again:

test_all_authored_oracles_and_mutations (test_casebook.CasebookTests.test_all_authored_oracles_and_mutations) ... ok
test_arbitrary_keys_and_truthy_values_do_not_grant_authority (test_casebook.CasebookTests.test_arbitrary_keys_and_truthy_values_do_not_grant_authority) ... ok
test_changed_case_or_revision_cannot_reuse_committed_identity (test_casebook.CasebookTests.test_changed_case_or_revision_cannot_reuse_committed_identity) ... ok
test_concurrent_retries_persist_exactly_one_synthetic_effect (test_casebook.CasebookTests.test_concurrent_retries_persist_exactly_one_synthetic_effect) ... ok
test_missing_is_unknown_and_reconciliation_cannot_create_an_action (test_casebook.CasebookTests.test_missing_is_unknown_and_reconciliation_cannot_create_an_action) ... ok
test_path_and_identifier_validation (test_casebook.CasebookTests.test_path_and_identifier_validation) ... ok
test_receipt_failure_never_means_no_action (test_casebook.CasebookTests.test_receipt_failure_never_means_no_action) ... ok

----------------------------------------------------------------------
Ran 7 tests in 0.701s

OK

That is the author’s offline run, with Python 3.14.3, in a copy of the lab exported with git archive, so the author’s runs never touched the companion checkout. The same day, an earlier pair of runs edited the checkout in place and reverted it afterwards. They recorded the same result: OK on the unmodified tree, then OK again after the edit. Timings vary between runs and machines.

The suite is well covered, too. On the unmodified tree, coverage.py 7.16.1 reported casebook.py at 82% of 194 statements, with line 97 among the executed lines.

With that word gone, a pending request can be reconciled after any of its case’s rules has been removed, and a replayed request no longer notices its receipt-phase rules or its rules’ reasons changing. Neither the suite that reports 186 mutations killed nor the coverage report notices that.

This guide is for the evaluation engineer who owns a green, well-covered suite around an agent guard and wants to know whether it would catch someone breaking that guard. The stance: a suite that reports every one of its own mutations killed is not evidence against a mutation nobody wrote. The mutation worth running is one aimed at what decides the outcome, and the survivors you find should be counted in the open, even when that makes your headline number smaller.

Everything here is an offline Python check against synthetic fixtures. No language model and no external service is involved.

Mutation testing Python unit tests: what the 186 counts

Mutation testing is an old idea. DeMillo, Lipton and Sayward set it out in 1978 (“Hints on Test Data Selection”, IEEE Computer): make small deliberate changes to the code, called mutants, and see whether the tests fail. A mutant the tests catch is killed. One they miss survives, and a survivor marks a behaviour no test pins down. Tools such as mutmut and cosmic-ray apply the idea to Python source; this lab applies it by hand, to data.

Two other guides on this site already use single deletions as evidence. The fixture guide deletes one key and watches which test notices, and the action register guide asks for a test that fails when the guard line is deleted. This guide deals with what comes after that habit becomes a count: which population of edits the count covers, and which survivors make no difference at all.

The lab’s loop is in prove() (casebook.py, lines 212–223):

    # Exercise each broken predicate against the independently stored oracle.
    killed = []
    for rule in spec['rules']:
        mutant = json.loads(json.dumps(spec))
        mutant['rules'] = [r for r in mutant['rules'] if r['field'] != rule['field']]
        fixture = next(f for f in spec['variants'] if f['name'] == rule['field'] + '-false')
        # Mutant no longer considers the deleted predicate part of the contract.
        facts = {k: v for k, v in fixture['facts'].items() if k != rule['field']}
        with tempfile.TemporaryDirectory() as tmp:
            outcome = execute(mutant, facts, 'mutant', Path(tmp) / 'ledger.sqlite')
        assert outcome['status'] != fixture['expected']['status'], 'surviving predicate deletion'
        killed.append(rule['field'])

Read it for three things: what it changes, what it compares against, and what it counts as killed.

  • What it changes. One whole rule, removed from spec['rules']. Every case has exactly three rules (load_case rejects any other count, line 28), so every case yields exactly three mutants.
  • What it compares against. The stored fixture named <field>-false, the one written to fail on that rule. The mutant runs once, under the request ID 'mutant', on a fresh ledger in a temporary directory (lines 220–221).
  • What counts as killed. The mutant’s status differs from the status the fixture expects. For the handoff case’s redacted_summary rule, the fixture expects blocked, and in the author’s offline run of the same steps the mutant returned succeeded, so it counts as killed.

The test then asserts the count (tests/test_casebook.py, lines 18–25):

    def test_all_authored_oracles_and_mutations(self):
        cases = sorted((ROOT / 'examples').glob('*.json'))
        self.assertEqual(len(cases), 62)
        for path in cases:
            with self.subTest(case=path.stem):
                outcome = prove(load_case(path.stem), Path(self.tmp.name) / f'{path.stem}.sqlite')
                self.assertEqual(len(outcome['observations']), 10)
                self.assertEqual(len(outcome['mutationsKilled']), 3)

Sixty-two cases times three rules is 186. The project’s verification record states the count in the same terms: the method “checks 62 cases, 620 expected input outcomes and 186 predicate-deletion mutations”, then lists the per-case replay, lost-acknowledgement and reconciliation checks (VERIFICATION.md, lines 3–5). The compatibility receipt names each check rather than folding them into a single pass (COMPATIBILITY.json, lines 116–131):

    {
      "path": "tutorials/rasa-ai-team-casebook",
      "sourceHash": "0f945cd467cdbeab0c48d647b97915fe89e4c9639edbf2f9a4c7efbcf3627e93",
      "checks": [
        "locked-install",
        "installed-version",
        "mantle-import",
        "project-validation",
        "licensed-training",
        "project-unittests",
        "62-scenario-oracles",
        "186-mutations-killed",
        "62-sdk-tool-dispatches"
      ],
      "liveConversations": "not tested"
    },

Nothing in those records is overstated. 186-mutations-killed is a true, reproducible number, and it names its population: mutations shaped like “delete this rule’s predicate”. The mistake comes later, when that number is read as “the guard would survive being tampered with”. The suite generates its own mutants, chooses the fixture each is judged against, and declares the verdict. It cannot, by construction, report on an edit shaped any other way.

Kill it by breaking the code, not the data

An authored mutant edits the case. The edit that matters in review is one someone makes to the code the case runs through. So the second question is: if I change casebook.py, does any of the seven tests fail?

The kill criterion changes with the question. For a code mutant, “killed” means the suite that passes on the clean tree goes red. Work in a throwaway copy, never the checkout you report from. From the root of your clone, one way to get one is a Git worktree (we ran these commands on macOS):

git worktree add ../casebook-scratch 69e27b6
cd ../casebook-scratch/tutorials/rasa-ai-team-casebook

Run the suite once on the clean tree before you change anything. This is the lab’s own make proof target:

macOS
python3 -m unittest discover -s tests -v
Linux
python3 -m unittest discover -s tests -v
Windows
python -m unittest discover -s tests -v

On the clean tree, the author’s run printed seven ok lines, then Ran 7 tests in 0.732s and OK. In a separate run of the two worktree commands in a local clone, the suite passed in the new worktree, git status showed nothing changed, and git worktree remove ../casebook-scratch cleaned it up.

Applying the mutant is the step that differs by platform, because in-place editing is spelt differently in BSD sed, GNU sed and PowerShell. Each version below keeps a copy of the original, makes the one-word edit to line 97, shows the changed line, runs the suite and puts the original back. The author ran the three edits on macOS with the system sed, GNU sed 4.9 and PowerShell 7.4.6, and all three produced the same line 97. None of them was run on Windows itself, where Set-Content writes Windows line endings, so what is known to match there is the text of line 97, not the bytes of the file.

macOS
cp casebook.py casebook.py.orig
sed -i '' "97s/'rules', //" casebook.py
diff casebook.py.orig casebook.py
python3 -m unittest discover -s tests -v
cp casebook.py.orig casebook.py
Linux (GNU sed)
cp casebook.py casebook.py.orig
sed -i "97s/'rules', //" casebook.py
diff casebook.py.orig casebook.py
python3 -m unittest discover -s tests -v
cp casebook.py.orig casebook.py
Windows (PowerShell)
Copy-Item casebook.py casebook.py.orig
$lines = Get-Content casebook.py
$lines[96] = $lines[96].Replace("'rules', ", "")
Set-Content casebook.py $lines
Compare-Object (Get-Content casebook.py.orig) (Get-Content casebook.py)
python -m unittest discover -s tests -v
Copy-Item casebook.py.orig casebook.py -Force

The diff or Compare-Object line is there so you can see the edit landed on line 97 before you trust a green run. If it prints nothing, the edit missed and the suite is testing the unmodified file. PowerShell arrays count from zero, which is why line 97 is $lines[96]. On Windows the interpreter may be py rather than python, depending on how Python was installed.

A first attempt shows the suite doing its job. In evaluate(), make a missing fact count as true by changing facts.get(rule['field']) on line 86 to facts.get(rule['field'], True). In the author’s offline run, the suite went red at once with FAILED (failures=62), one failure per case. The first, in the banking-advisor-appointment case, named the fixture that caught it:

AssertionError: ('purpose_matched-missing', {'status': 'succeeded', 'reason': 'verified_fixture_receipt', 'effects': 1}, {'status': 'blocked', 'reason': 'wrong_advisor_capability', 'effects': 0})

All 62 fixture files carry the same set of variants: accepted, plus a -false, a -missing and a -string variant for each of the three rules. That is where the ten observations per case come from, and why a mutation nobody listed in prove() was killed by the 62-scenario oracle anyway. The suite is not weak. It is strong exactly where its author looked, and the fixtures show where that was: the rule evaluator.

contract_hash is somewhere else.

case fixture spec['rules'] evaluate() lines 84–88 contract_hash() line 97 stored request row contract_hash column replay and reconcile lines 117 and 151 186 rule deletions prove(), lines 212–223 killed by prove() missing fact treated as true line 86 killed by fixtures 'rules' dropped from the tuple line 97 survives all 7 tests
FigureWhere each mutant lands, and which tests can see it

The first two mutants land in the evaluator, which the fixtures exercise from ten directions per case. The third lands in the identity hash, whose value only matters when a stored request is compared with a case that has changed since.

What contract_hash decides

Here is the function in full (casebook.py, lines 96–97):

def contract_hash(spec: dict) -> str:
    return digest({k: spec.get(k) for k in ('slug', 'action', 'rules', 'receipt', 'provenance', 'intakeKind')})

execute() stores that hash beside every committed request. When the same request ID comes back, the stored fingerprint and hash both have to match the case as it is now; if they do, the stored result is replayed (lines 116–119):

        if previous:
            if previous['fingerprint'] != fingerprint or previous['contract_hash'] != contract_hash(spec):
                return result('conflict', 'request_identity_changed', effects=1, replay=True)
            return result(previous['status'], previous['reason'], previous['reference'], 1, True)

reconcile(), which re-checks the receipt facts and turns a pending row into succeeded when they pass (lines 153–157), first compares the stored case and hash with the current ones (lines 151–152). It does not compare the fingerprint:

        if row['case_id'] != spec['slug'] or row['contract_hash'] != contract_hash(spec):
            return result('conflict', 'receipt_contract_changed', effects=1)

The fingerprint is the other half of that identity (lines 109–111):

    bound_facts = {r['field']: facts.get(r['field']) for r in spec['rules']
                   if r['phase'] == 'request'}
    fingerprint = digest({'case': spec['slug'], 'facts': bound_facts})

Read what it binds. It hashes the case slug plus the names and values of the request-phase facts, so adding or removing a request-phase rule changes it. It never sees receipt-phase rules, and it never sees a rule’s other attributes, such as its reason. And only execute() compares it: reconcile() does not check the fingerprint at all. So on replay, receipt rules and rule reasons are bound to a stored request only by the 'rules' entry in contract_hash, and at reconciliation every rule is, request-phase rules included. Take that entry out, and a replayed request no longer notices its receipt rules being removed or its rules’ reasons being changed, while a pending request can be reconciled after any of its rules has been removed.

Why does no test notice? In tests/test_casebook.py, only one test changes a case and then comes back with a request ID that is already committed (lines 38–46):

    def test_changed_case_or_revision_cannot_reuse_committed_identity(self):
        execute(self.spec, self.spec['facts'], 'request-1', self.database)
        other = load_case('banking-transfer')
        self.assertEqual(execute(other, other['facts'], 'request-1', self.database)['status'], 'conflict')
        changed = json.loads(json.dumps(self.spec))
        changed['provenance']['revision'] = 'changed'
        self.assertEqual(execute(changed, changed['facts'], 'request-1', self.database)['status'], 'conflict')
        self.assertEqual(reconcile(changed, changed['facts'], 'request-1', self.database)['status'], 'conflict')
        self.assertEqual(count_effects(self.database), 1)

It swaps in a different case, which the slug in the fingerprint already catches, then changes provenance.revision and edits nothing else. The hash lists six fields; the test exercises one of them directly.

This matters beyond the lab’s own tests because the Rasa tool reuses request IDs by design. The Mantle adapter in tools/cases.py builds the ID from an operator-set run ID and the case, not from anything per call (lines 9–13 and 20–24):

def identity(case_id):
    # The lab operator assigns a stable run ID. A real backend uses the
    # authenticated subject + action revision, resolved from trusted context.
    run_id = os.environ.get('CASEBOOK_RUN_ID', 'rehearsal-1')
    return f'{run_id}-{case_id}'
@tool(description='Rehearse one supported synthetic case. Returns a lab outcome, never a real customer action.')
async def rehearse_case(case_id: str, context: ToolContext = None) -> ToolResult:
    try:
        spec = load_case(case_id)
        outcome = execute(spec, spec['facts'], identity(case_id), database())

Read together with lines 116–119, the code says that a second rehearsal of the same case against the same ledger replays whatever the first one stored, and that contract_hash is the check meant to refuse that replay when the case’s rules have changed in between. That is a reading of the code only. This guide did not run rehearse_case, with or without the mutant.

Try it first: write the test that kills this mutant

Commit a request that stops at pending, remove the receipt rule it is waiting on, then try to reconcile under the changed case. The unmodified guard should refuse with conflict. This is the author’s scratch test, not part of the companion project:

import json
from pathlib import Path
import tempfile
import unittest
from casebook import execute, load_case, reconcile


class ContractBindingTests(unittest.TestCase):
    def test_changed_rules_cannot_reconcile_a_committed_request(self):
        spec = load_case('contextual-handoff')
        with tempfile.TemporaryDirectory() as tmp:
            database = Path(tmp) / 'ledger.sqlite'
            facts = dict(spec['facts'], desk_acknowledged=False)
            self.assertEqual(execute(spec, facts, 'r1', database)['status'], 'pending')
            changed = json.loads(json.dumps(spec))
            changed['rules'] = [r for r in changed['rules'] if r['field'] != 'desk_acknowledged']
            later = {k: v for k, v in spec['facts'].items() if k != 'desk_acknowledged'}
            self.assertEqual(reconcile(changed, later, 'r1', database)['status'], 'conflict')

In the author’s offline runs it passed against the unmodified casebook.py. With 'rules' removed from the tuple, the last lines of the run were:

AssertionError: 'succeeded' != 'conflict'
- succeeded
+ conflict


----------------------------------------------------------------------
Ran 1 test in 0.003s

FAILED (failures=1)

Read what the mutant did. A handoff was waiting for the desk to acknowledge it. Someone then deleted the acknowledgement rule from the case, and reconciliation marked the old request succeeded against a check it never passed. The request did not change. The rules it was judged by did.

Count what you tried, not only what you killed

Once one field in the tuple has survived, the fair next step is to try the others. In offline runs, the author deleted each of the six fields from line 97 in turn, one per run, and reran the seven tests:

Field removed from contract_hashSeven-test suiteReading
provenance1 failureKilled by test_changed_case_or_revision_cannot_reuse_committed_identity
slugOKEquivalent: the fingerprint and case_id still bind it
rulesOKSurvivor: no rule edit is bound at reconcile; on replay, receipt-phase rules and rule reasons are not
actionOKSurvivor
receiptOKSurvivor
intakeKindOKSurvivor

The slug row is the part mutation testing makes you do by hand. The case slug is already in the fingerprint on line 111 and is checked against case_id on line 151, so removing it from the hash changes nothing a caller could observe. That is an equivalent mutant, and no test should be written to kill it. Equivalent mutants are a standing cost of the technique: some survivors are real gaps, some are changes that make no difference, and only reading the code tells them apart.

The other three survivors are descriptive strings. action and receipt are strings in all 62 fixtures; intakeKind is a string in the 4 cases that have an independent intake and absent from the other 58. Whether a change to one of them should invalidate a committed request is a decision for the lab’s owner. The tuple on line 97 records that its author thought it should.

An honest report of this suite reads more like this:

  • predicate deletion, generated by prove(): 186 of 186 killed;
  • missing fact treated as true, in evaluate(): killed by the 62-scenario oracles;
  • single-field deletion from contract_hash: 6 tried, 1 killed, 1 equivalent, 4 survived.

That is less satisfying than “186 killed”, and more useful. It names each population and its denominator, and it tells the next reviewer where to aim. The evaluation set guide asks the same of agent evaluations: report passed, failed and not-run counts for every slice. The project’s COMPATIBILITY.json already names checks instead of reporting one pass, so the form is there to extend.

What this costs, and what it does not show

The cost is time and judgement. Six hand-written mutants took one line each to make, and one of the six, slug, needed code reading to rule out. Every extra mutant, whether you write it or a tool generates it, is another possible survivor that someone has to read and classify as a gap or an equivalent. You are choosing to spend reviewer time on fewer, targeted mutants instead of a larger number that is easier to put in a report.

The survivor is also narrower than it looks. On replay through execute(), 155 of the 186 rules across the 62 fixtures are request-phase, so removing one changes the fingerprint and replay refuses anyway; there the survivor bites only through the other 31 and through rule attributes such as reason. Reconciliation has no such backstop. Every fixture carries a provenance.revision, and the one test that edits a committed case changes exactly that field. If everyone who edits a case’s rules also bumps its revision, the provenance check catches the change and the survivor never fires. The mutant shows that the rules binding is untested, not that the lab is unsafe as used.

These checks do not show how a model behaves, how the Mantle tool behaves in a live conversation, or anything about a real service. The lab is a synthetic teaching contract, and the counts above are finite offline results, not rates.