Guide · Evaluation / data specialist
186 mutations killed, and the guard still broke
A guard suite reported 186 of 186 mutations killed and still passed a one-word break. Test for the mutation nobody authored, and count the result honestly.
Key takeaways (5)
- A suite that kills every mutation its author wrote has measured only the mutations its author thought of. The count is true and says nothing about any other kind of edit.
- Before trusting a mutation count, read the loop that produces it: what it changes, which fixture it checks against, and what it calls killed.
- Deleting one word from the tuple that contract_hash hashes left all seven companion tests passing, including the one that reports 186 mutations killed, while coverage.py still reported the line as executed.
- Aim your own mutants at whatever decides whether an earlier result still stands (identity hashes, fingerprints, replay checks), because a test that changes only one field protects only that field.
- Report mutation results per population, with the denominator and the survivors named, not as one number beside a green tick.
Clone RasaHQ/rasa-community-resources, check out 69e27b6, change into
tutorials/rasa-ai-team-casebook and run the lab’s offline suite. Seven tests
pass. One of them, test_all_authored_oracles_and_mutations, loads 62 scenario
fixtures and asserts three killed mutations for each, 186 in all.
Now delete one word. Line 97 of casebook.py is the body of contract_hash,
which fixes a case’s identity once a request has been committed. Take
'rules' out of its tuple:
def contract_hash(spec: dict) -> str:
- return digest({k: spec.get(k) for k in ('slug', 'action', 'rules', 'receipt', 'provenance', 'intakeKind')})
+ return digest({k: spec.get(k) for k in ('slug', 'action', 'receipt', 'provenance', 'intakeKind')})
Run the same seven tests again:
test_all_authored_oracles_and_mutations (test_casebook.CasebookTests.test_all_authored_oracles_and_mutations) ... ok
test_arbitrary_keys_and_truthy_values_do_not_grant_authority (test_casebook.CasebookTests.test_arbitrary_keys_and_truthy_values_do_not_grant_authority) ... ok
test_changed_case_or_revision_cannot_reuse_committed_identity (test_casebook.CasebookTests.test_changed_case_or_revision_cannot_reuse_committed_identity) ... ok
test_concurrent_retries_persist_exactly_one_synthetic_effect (test_casebook.CasebookTests.test_concurrent_retries_persist_exactly_one_synthetic_effect) ... ok
test_missing_is_unknown_and_reconciliation_cannot_create_an_action (test_casebook.CasebookTests.test_missing_is_unknown_and_reconciliation_cannot_create_an_action) ... ok
test_path_and_identifier_validation (test_casebook.CasebookTests.test_path_and_identifier_validation) ... ok
test_receipt_failure_never_means_no_action (test_casebook.CasebookTests.test_receipt_failure_never_means_no_action) ... ok
----------------------------------------------------------------------
Ran 7 tests in 0.701s
OK
That is the author’s offline run, with Python 3.14.3, in
a copy of the lab exported with git archive, so the author’s runs never
touched the companion checkout. The same day, an earlier pair of runs edited
the checkout in place and reverted it afterwards. They recorded the same
result: OK on the unmodified tree, then OK again after the edit. Timings
vary between runs and machines.
The suite is well covered, too. On the unmodified tree, coverage.py
7.16.1 reported casebook.py at 82% of 194 statements, with line 97 among
the executed lines.
With that word gone, a pending request can be reconciled after any of its case’s rules has been removed, and a replayed request no longer notices its receipt-phase rules or its rules’ reasons changing. Neither the suite that reports 186 mutations killed nor the coverage report notices that.
This guide is for the evaluation engineer who owns a green, well-covered suite around an agent guard and wants to know whether it would catch someone breaking that guard. The stance: a suite that reports every one of its own mutations killed is not evidence against a mutation nobody wrote. The mutation worth running is one aimed at what decides the outcome, and the survivors you find should be counted in the open, even when that makes your headline number smaller.
Everything here is an offline Python check against synthetic fixtures. No language model and no external service is involved.
Mutation testing Python unit tests: what the 186 counts
Mutation testing is an old idea. DeMillo, Lipton and Sayward set it out in 1978 (“Hints on Test Data Selection”, IEEE Computer): make small deliberate changes to the code, called mutants, and see whether the tests fail. A mutant the tests catch is killed. One they miss survives, and a survivor marks a behaviour no test pins down. Tools such as mutmut and cosmic-ray apply the idea to Python source; this lab applies it by hand, to data.
Two other guides on this site already use single deletions as evidence. The fixture guide deletes one key and watches which test notices, and the action register guide asks for a test that fails when the guard line is deleted. This guide deals with what comes after that habit becomes a count: which population of edits the count covers, and which survivors make no difference at all.
The lab’s loop is in prove() (casebook.py, lines 212–223):
# Exercise each broken predicate against the independently stored oracle.
killed = []
for rule in spec['rules']:
mutant = json.loads(json.dumps(spec))
mutant['rules'] = [r for r in mutant['rules'] if r['field'] != rule['field']]
fixture = next(f for f in spec['variants'] if f['name'] == rule['field'] + '-false')
# Mutant no longer considers the deleted predicate part of the contract.
facts = {k: v for k, v in fixture['facts'].items() if k != rule['field']}
with tempfile.TemporaryDirectory() as tmp:
outcome = execute(mutant, facts, 'mutant', Path(tmp) / 'ledger.sqlite')
assert outcome['status'] != fixture['expected']['status'], 'surviving predicate deletion'
killed.append(rule['field'])
Read it for three things: what it changes, what it compares against, and what it counts as killed.
- What it changes. One whole rule, removed from
spec['rules']. Every case has exactly three rules (load_caserejects any other count, line 28), so every case yields exactly three mutants. - What it compares against. The stored fixture named
<field>-false, the one written to fail on that rule. The mutant runs once, under the request ID'mutant', on a fresh ledger in a temporary directory (lines 220–221). - What counts as killed. The mutant’s status differs from the status the
fixture expects. For the handoff case’s
redacted_summaryrule, the fixture expectsblocked, and in the author’s offline run of the same steps the mutant returnedsucceeded, so it counts as killed.
The test then asserts the count (tests/test_casebook.py, lines 18–25):
def test_all_authored_oracles_and_mutations(self):
cases = sorted((ROOT / 'examples').glob('*.json'))
self.assertEqual(len(cases), 62)
for path in cases:
with self.subTest(case=path.stem):
outcome = prove(load_case(path.stem), Path(self.tmp.name) / f'{path.stem}.sqlite')
self.assertEqual(len(outcome['observations']), 10)
self.assertEqual(len(outcome['mutationsKilled']), 3)
Sixty-two cases times three rules is 186. The project’s verification record
states the count in the same terms: the method “checks 62 cases, 620 expected
input outcomes and 186 predicate-deletion mutations”, then lists the
per-case replay, lost-acknowledgement and reconciliation checks
(VERIFICATION.md, lines 3–5). The compatibility receipt names each check
rather than folding them into a single pass (COMPATIBILITY.json, lines
116–131):
{
"path": "tutorials/rasa-ai-team-casebook",
"sourceHash": "0f945cd467cdbeab0c48d647b97915fe89e4c9639edbf2f9a4c7efbcf3627e93",
"checks": [
"locked-install",
"installed-version",
"mantle-import",
"project-validation",
"licensed-training",
"project-unittests",
"62-scenario-oracles",
"186-mutations-killed",
"62-sdk-tool-dispatches"
],
"liveConversations": "not tested"
},
Nothing in those records is overstated. 186-mutations-killed is a true,
reproducible number, and it names its population: mutations shaped like
“delete this rule’s predicate”. The mistake comes later, when that number
is read as “the guard would survive being tampered with”. The suite
generates its own mutants, chooses the fixture each is judged against, and
declares the verdict. It cannot, by construction, report on an edit shaped any
other way.
Kill it by breaking the code, not the data
An authored mutant edits the case. The edit that matters in review is one
someone makes to the code the case runs through. So the second question is:
if I change casebook.py, does any of the seven tests fail?
The kill criterion changes with the question. For a code mutant, “killed” means the suite that passes on the clean tree goes red. Work in a throwaway copy, never the checkout you report from. From the root of your clone, one way to get one is a Git worktree (we ran these commands on macOS):
git worktree add ../casebook-scratch 69e27b6
cd ../casebook-scratch/tutorials/rasa-ai-team-casebook
Run the suite once on the clean tree before you change anything. This is the
lab’s own make proof target:
- macOS
python3 -m unittest discover -s tests -v- Linux
python3 -m unittest discover -s tests -v- Windows
python -m unittest discover -s tests -v
On the clean tree, the author’s run printed seven ok lines, then
Ran 7 tests in 0.732s and OK. In a separate run of the two worktree
commands in a local clone, the suite passed in the new worktree,
git status showed nothing changed, and
git worktree remove ../casebook-scratch cleaned it up.
Applying the mutant is the step that differs by platform, because in-place
editing is spelt differently in BSD sed, GNU sed and PowerShell. Each version
below keeps a copy of the original, makes the one-word edit to line 97, shows
the changed line, runs the suite and puts the original back. The author ran
the three edits on macOS with the system sed, GNU sed 4.9 and PowerShell 7.4.6,
and all three produced the same line 97. None of them was run on Windows
itself, where Set-Content writes Windows line endings, so what is known to
match there is the text of line 97, not the bytes of the file.
macOS
cp casebook.py casebook.py.orig
sed -i '' "97s/'rules', //" casebook.py
diff casebook.py.orig casebook.py
python3 -m unittest discover -s tests -v
cp casebook.py.orig casebook.pyLinux (GNU sed)
cp casebook.py casebook.py.orig
sed -i "97s/'rules', //" casebook.py
diff casebook.py.orig casebook.py
python3 -m unittest discover -s tests -v
cp casebook.py.orig casebook.pyWindows (PowerShell)
Copy-Item casebook.py casebook.py.orig
$lines = Get-Content casebook.py
$lines[96] = $lines[96].Replace("'rules', ", "")
Set-Content casebook.py $lines
Compare-Object (Get-Content casebook.py.orig) (Get-Content casebook.py)
python -m unittest discover -s tests -v
Copy-Item casebook.py.orig casebook.py -ForceThe diff or Compare-Object line is there so you can see the edit landed on
line 97 before you trust a green run. If it prints nothing, the edit missed and
the suite is testing the unmodified file. PowerShell arrays count from zero,
which is why line 97 is $lines[96]. On Windows the interpreter may be py
rather than python, depending on how Python was installed.
A first attempt shows the suite doing its job. In evaluate(), make a missing
fact count as true by changing facts.get(rule['field']) on line 86 to
facts.get(rule['field'], True). In the author’s offline run, the suite went red at once with FAILED (failures=62), one failure per
case. The first, in the banking-advisor-appointment case, named the fixture
that caught it:
AssertionError: ('purpose_matched-missing', {'status': 'succeeded', 'reason': 'verified_fixture_receipt', 'effects': 1}, {'status': 'blocked', 'reason': 'wrong_advisor_capability', 'effects': 0})
All 62 fixture files carry the same set of variants: accepted, plus a
-false, a -missing and a -string variant for each of the three rules.
That is where the ten observations per case come from, and why a mutation
nobody listed in prove() was killed by the 62-scenario oracle anyway. The
suite is not weak. It is strong exactly where its author looked, and the
fixtures show where that was: the rule evaluator.
contract_hash is somewhere else.
The first two mutants land in the evaluator, which the fixtures exercise from ten directions per case. The third lands in the identity hash, whose value only matters when a stored request is compared with a case that has changed since.
What contract_hash decides
Here is the function in full (casebook.py, lines 96–97):
def contract_hash(spec: dict) -> str:
return digest({k: spec.get(k) for k in ('slug', 'action', 'rules', 'receipt', 'provenance', 'intakeKind')})
execute() stores that hash beside every committed request. When the same
request ID comes back, the stored fingerprint and hash both have to match the
case as it is now; if they do, the stored result is replayed
(lines 116–119):
if previous:
if previous['fingerprint'] != fingerprint or previous['contract_hash'] != contract_hash(spec):
return result('conflict', 'request_identity_changed', effects=1, replay=True)
return result(previous['status'], previous['reason'], previous['reference'], 1, True)
reconcile(), which re-checks the receipt facts and turns a pending row
into succeeded when they pass (lines 153–157), first compares the stored
case and hash with the current ones (lines 151–152). It does not compare the
fingerprint:
if row['case_id'] != spec['slug'] or row['contract_hash'] != contract_hash(spec):
return result('conflict', 'receipt_contract_changed', effects=1)
The fingerprint is the other half of that identity (lines 109–111):
bound_facts = {r['field']: facts.get(r['field']) for r in spec['rules']
if r['phase'] == 'request'}
fingerprint = digest({'case': spec['slug'], 'facts': bound_facts})
Read what it binds. It hashes the case slug plus the names and values of the
request-phase facts, so adding or removing a request-phase rule changes it.
It never sees receipt-phase rules, and it never sees a rule’s other
attributes, such as its reason. And only execute() compares it:
reconcile() does not check the fingerprint at all. So on replay, receipt
rules and rule reasons are bound to a stored request only by the 'rules'
entry in contract_hash, and at reconciliation every rule is, request-phase
rules included. Take that entry out, and a replayed request no longer notices
its receipt rules being removed or its rules’ reasons being changed, while a
pending request can be reconciled after any of its rules has been removed.
Why does no test notice? In tests/test_casebook.py, only one test changes a
case and then comes back with a request ID that is already committed
(lines 38–46):
def test_changed_case_or_revision_cannot_reuse_committed_identity(self):
execute(self.spec, self.spec['facts'], 'request-1', self.database)
other = load_case('banking-transfer')
self.assertEqual(execute(other, other['facts'], 'request-1', self.database)['status'], 'conflict')
changed = json.loads(json.dumps(self.spec))
changed['provenance']['revision'] = 'changed'
self.assertEqual(execute(changed, changed['facts'], 'request-1', self.database)['status'], 'conflict')
self.assertEqual(reconcile(changed, changed['facts'], 'request-1', self.database)['status'], 'conflict')
self.assertEqual(count_effects(self.database), 1)
It swaps in a different case, which the slug in the fingerprint already
catches, then changes provenance.revision and edits nothing else. The hash
lists six fields; the test exercises one of them directly.
This matters beyond the lab’s own tests because the Rasa tool reuses request
IDs by design. The Mantle adapter in tools/cases.py builds the ID from an
operator-set run ID and the case, not from anything per call (lines 9–13 and
20–24):
def identity(case_id):
# The lab operator assigns a stable run ID. A real backend uses the
# authenticated subject + action revision, resolved from trusted context.
run_id = os.environ.get('CASEBOOK_RUN_ID', 'rehearsal-1')
return f'{run_id}-{case_id}'
@tool(description='Rehearse one supported synthetic case. Returns a lab outcome, never a real customer action.')
async def rehearse_case(case_id: str, context: ToolContext = None) -> ToolResult:
try:
spec = load_case(case_id)
outcome = execute(spec, spec['facts'], identity(case_id), database())
Read together with lines 116–119, the code says that a second rehearsal of
the same case against the same ledger replays whatever the first one stored,
and that contract_hash is the check meant to refuse that replay when the
case’s rules have changed in between. That is a reading of the code only. This
guide did not run rehearse_case, with or without the mutant.
Try it first: write the test that kills this mutant
Commit a request that stops at pending, remove the receipt rule it is
waiting on, then try to reconcile under the changed case. The unmodified guard
should refuse with conflict. This is the author’s scratch test, not part of
the companion project:
import json
from pathlib import Path
import tempfile
import unittest
from casebook import execute, load_case, reconcile
class ContractBindingTests(unittest.TestCase):
def test_changed_rules_cannot_reconcile_a_committed_request(self):
spec = load_case('contextual-handoff')
with tempfile.TemporaryDirectory() as tmp:
database = Path(tmp) / 'ledger.sqlite'
facts = dict(spec['facts'], desk_acknowledged=False)
self.assertEqual(execute(spec, facts, 'r1', database)['status'], 'pending')
changed = json.loads(json.dumps(spec))
changed['rules'] = [r for r in changed['rules'] if r['field'] != 'desk_acknowledged']
later = {k: v for k, v in spec['facts'].items() if k != 'desk_acknowledged'}
self.assertEqual(reconcile(changed, later, 'r1', database)['status'], 'conflict')In the author’s offline runs it passed against the
unmodified casebook.py. With 'rules' removed from the tuple, the last lines
of the run were:
AssertionError: 'succeeded' != 'conflict'
- succeeded
+ conflict
----------------------------------------------------------------------
Ran 1 test in 0.003s
FAILED (failures=1)Read what the mutant did. A handoff was waiting for the desk to acknowledge
it. Someone then deleted the acknowledgement rule from the case, and
reconciliation marked the old request succeeded against a check it never
passed. The request did not change. The rules it was judged by did.
Count what you tried, not only what you killed
Once one field in the tuple has survived, the fair next step is to try the others. In offline runs, the author deleted each of the six fields from line 97 in turn, one per run, and reran the seven tests:
Field removed from contract_hash | Seven-test suite | Reading |
|---|---|---|
provenance | 1 failure | Killed by test_changed_case_or_revision_cannot_reuse_committed_identity |
slug | OK | Equivalent: the fingerprint and case_id still bind it |
rules | OK | Survivor: no rule edit is bound at reconcile; on replay, receipt-phase rules and rule reasons are not |
action | OK | Survivor |
receipt | OK | Survivor |
intakeKind | OK | Survivor |
The slug row is the part mutation testing makes you do by hand. The case slug
is already in the fingerprint on line 111 and is checked against case_id on
line 151, so removing it from the hash changes nothing a caller could observe.
That is an equivalent mutant, and no test should be written to kill it.
Equivalent mutants are a standing cost of the technique: some survivors are
real gaps, some are changes that make no difference, and only reading the code
tells them apart.
The other three survivors are descriptive strings. action and receipt are
strings in all 62 fixtures; intakeKind is a string in the 4 cases that have
an independent intake and absent from the other 58. Whether a change to one of
them should invalidate a committed request is a decision for the lab’s owner.
The tuple on line 97 records that its author thought it should.
An honest report of this suite reads more like this:
- predicate deletion, generated by
prove(): 186 of 186 killed; - missing fact treated as true, in
evaluate(): killed by the 62-scenario oracles; - single-field deletion from
contract_hash: 6 tried, 1 killed, 1 equivalent, 4 survived.
That is less satisfying than “186 killed”, and more useful. It names each
population and its denominator, and it tells the next reviewer where to aim.
The evaluation set guide asks the
same of agent evaluations: report passed, failed and not-run counts for every
slice. The project’s COMPATIBILITY.json already names checks instead of
reporting one pass, so the form is there to extend.
What this costs, and what it does not show
The cost is time and judgement. Six hand-written mutants took one line each to
make, and one of the six, slug, needed code reading to rule out. Every extra
mutant, whether you write it or a tool generates it, is another possible
survivor that someone has to read and classify as a gap or an equivalent. You
are choosing to spend reviewer time on fewer, targeted mutants instead of a
larger number that is easier to put in a report.
The survivor is also narrower than it looks. On replay through execute(),
155 of the 186 rules across the 62 fixtures are request-phase, so removing one
changes the fingerprint and replay refuses anyway; there the survivor bites
only through the other 31 and through rule attributes such as reason.
Reconciliation has no such backstop. Every fixture carries a
provenance.revision, and the one test that edits a committed case changes
exactly that field. If everyone who edits a case’s rules also bumps its
revision, the provenance check catches the change and the survivor never
fires. The mutant shows that the rules binding is untested, not that the lab
is unsafe as used.
These checks do not show how a model behaves, how the Mantle tool behaves in a live conversation, or anything about a real service. The lab is a synthetic teaching contract, and the counts above are finite offline results, not rates.