Guide · Risk reviewer / domain owner
How to audit a document an AI agent generates
We audited an agent that derives documents from records. Its tests pass; the engine call path, free-text reasons and a replaceable risk warning do not.
Key takeaways (5)
- A careful read of an agent’s document checks one copy on one day. It cannot show where a figure came from, and it does not repeat for the next thousand documents.
- Ask for evidence from the path the agent runs. Every companion tool is a plain function; the two we ran through Rasa Pro 3.20.0rc1’s tool invoker did their work and then reported only a generic error, so the tested refusal never reached the agent.
- Ask which test fails when the protection is removed. In the companion the guarantee is the tool signatures, pinned by TestTheModelCannotWriteTheArtifact; the test on the llm_settable flag of sourced fields checks a condition that never decides the outcome, because the kind check refuses first.
- List every way the model can write to the document. In the companion, a free-text reason on every edit is printed verbatim into the record, and the model can blank or replace the regulated risk warning while the document still renders.
- A refusal stops the next rendering, not the copy already written. Ask which file leaves the building and what re-checks it on the way out.
The record below is illustrative, not a recorded model run: it is the example the deriving-a-document tutorial opens on, and the client, Marged Ellis, and her portfolio are invented fixture data. An adviser asks an agent for a client suitability record. Here is the first sentence of its body; the example goes on from there:
Your portfolio is currently valued at approximately £486,000, held across
a balanced mix of global equities (around 58%), sterling corporate bonds
(roughly 30%), and property funds (about 12%).
Every figure sits close to the invented custodian extract behind it: £486,210.44 in total, equities at 57.7%, bonds at 30.4% and a property row at 11.9%. A reviewer holding the extract would tick all four. What the page does not say is which figures were read from a record and which were rounded, remembered from an older valuation or composed so the sentence would feel complete.
The property figure is the one to look at. The extract has a property row, but the record the companion project assembles from the same fixture never cites it. The rendered record’s Completeness section says so:
## Completeness
3 declared field(s) have no source and render blank:
- `property_value`
- `property_weight`
- `include_property_breakdown`
The prose version states a property weighting that the assembled record has no citation for, in the same confident voice as the figures that do.
This guide is for the person asked to approve an agent that writes documents a client or regulator will keep: summaries, letters, case records. Its position is that a fluent document proves nothing, and neither does a green test suite unless the tests run the path the agent actually uses. Sign off on a list of every way the model can write to the document, each tested through the engine you will deploy, and on a refusal that reaches both the agent and any copy already written.
We tested that position on the best case available: the companion project
behind the tutorial, tutorials/rasa-document-artifact-tutorial in
RasaHQ/rasa-community-resources at commit 69e27b6, which was designed so
that a model cannot write a figure. Its 41 tests pass. We still found five gaps
a sign-off should catch, and this guide is built on them. The scenario and its
field list are synthetic: the fixture’s README says every person, holding and
valuation in it is “invented for this tutorial” (data/source/README.md,
lines 1–5), and the property row is POS-PF4402-003 in
data/source/holdings.json (lines 43–47). Nothing here claims the design satisfies any
regulator’s suitability, record-keeping or disclosure rule, and which fields
your own documents need is a decision for your team.
Why reading the output is not the control
The natural control is a careful review of what the agent produced. It fails in three ways, and none of them is about how careful the reviewer is.
- It checks agreement, not origin. “About 12%” agrees with 11.9% within rounding. The reviewer cannot see whether the model read that row, recalled a figure from an earlier conversation or produced a number that fitted.
- It checks a copy, not a process. A clean review of this record says nothing about the next one.
- It happens once. A custodian restates a valuation after the review and before the letter goes out, and the reviewed figure is now wrong under the firm’s name.
The tutorial’s full example goes on to say the allocation “remains well suited
to your Balanced risk profile”. That is a suitability judgement, and none of
the fixture’s three sources (holdings.json, factfind.json and
references/disclosures.json) contains it or the word “suited”. A reviewer may agree with
it; that agreement is the reviewer’s judgement, not evidence about where the
sentence came from.
What the model can touch in this design
The companion project takes the pen away. The conversation edits a structured
state, never the document, and the renderer rebuilds the whole document from
that state each time it runs (docpkg/render.py, render_markdown). Most
fields are sourced: the model names a record and the record supplies the
value. Two are negotiated: the model picks a value from a short fixed list,
such as who the record is addressed to (docpkg/state.py, lines 114–119).
The model can only fill in the parameters the engine offers it. We asked Rasa Pro
3.20.0rc1 to build the schema for each document tool, offline, with the
engine’s own agent_tool_schema_from_custom_tool
(rasa/mantle/llm/tool_schemas.py, lines 577–610 in the wheel). The output of
that run:
point_field_at_record: properties=['field_key', 'source_id', 'record_id', 'record_field', 'reason'] required=['field_key', 'source_id', 'record_id', 'record_field', 'reason'] additionalProperties=False
choose_option: properties=['field_key', 'option', 'reason'] required=['field_key', 'option', 'reason'] additionalProperties=False
clear_document_field: properties=['field_key', 'reason'] required=['field_key', 'reason'] additionalProperties=False
render_document: properties=[] required=[] additionalProperties=False
Before reading on: from those four schemas, list every way the model can change what the document says
There are four, and a sign-off should name all of them.
- Which record a sourced field cites, through
point_field_at_record. It cannot type a figure, but it can cite the wrong record. - A value from a fixed list, through
choose_option, for the two negotiated fields. - Removing any field, through
clear_document_field, which takes anyfield_key. - Free text, through
reason, which every write tool requires and which the renderer prints into the record’s revision history.
The fourth is easy to miss, because nothing in the schema says where a reason ends up.
The list names the inputs. What it does not show is where each one ends up, including the file the render tool leaves behind:
Auditing a document an AI agent generates: five findings
We ran each check offline, in a fresh export of the
companion at 69e27b6. The engine checks used an installed Rasa Pro
3.20.0rc1 under Python 3.12.13; the suite was also run under Python 3.14.3
with no engine. The scripts are kept with this guide’s brief as receipts,
each run’s output saved verbatim with the command on its first line.
| Finding | What we ran | Result |
|---|---|---|
| 1. The refusal does not reach the agent | Tools called through the engine’s call_tool_func | State changed; the agent was told only “error” |
| 2. The boundary test checks a flag that never decides | llm_settable set to true on total_value | The test’s condition breaks; the tool still refuses |
| 3. The model writes prose into the record | A reason containing a suitability sentence and a pipe | Printed verbatim, unescaped, in the history |
| 4. The model can blank or replace the risk warning | clear_field, then a re-citation, on risk_warning | Both render: blank, then “Property funds” |
| 5. A refusal leaves the earlier copy | Render, restate the extract, render again | Refused; the first file still holds £486,210.44 |
None of these needs a model, a licence or a network, and none is a claim about what a model would choose to do in a conversation. Each is a path the code leaves open.
Finding 1: the refusal the tests prove is not what the agent receives
Engineering will show you the suite. It passes with or without the engine
installed; the last lines of our make test run, the command engineering
would use, under Python 3.14.3:
Ran 41 tests in 0.053s
OK
Among those tests is one written for exactly the question a reviewer asks: does
a refusal reach the agent? Its docstring reads “The path the AGENT runs, not
just the function underneath it” (tests/test_derived_document.py, line 296),
and another test in the same file warns against “asserting on a convenient internal path
while the path that actually runs was different” (lines 115–120). But the
refusal test calls td.render_document() directly, as a plain function (lines
295–305).
The agent does not call it that way. In Rasa Pro 3.20.0rc1 the engine runs a
local tool with return await func(context=ctx, **tool_args), None
(rasa/mantle/orchestration/tool_execution/invoker.py, line 234). The
companion’s tools are declared with plain def, so the function runs to the
end, side effects included, and only then does awaiting its result raise a
TypeError. The invoker’s catch-all (lines 161–172) turns that into a generic
error for the model.
This is a defect in the companion, not an engine quirk. The engine’s decorator
documents the contract: tool() is described as “Mark an async function as an
LLM-callable tool”, and a tool function’s type is
Callable[..., Awaitable[ToolResult]] (rasa/mantle/tools/decorator.py,
lines 80 and 270).
We called the tools through ToolInvoker.call_tool_func, the engine method
that runs a builder tool (call_builder_tool hands every call to it, line
106), with a stand-in for the session object. The
only session method it uses on this path is recording_write, which in the
engine sets and restores a marker around the call and lets any exception
through (rasa/mantle/state/tracker_handle.py, lines 525–545). Ours does
nothing, so it cannot change what the tool returns. The first call re-points
the total at the cash line (log timestamps removed):
== 1. point_field_at_record through the engine
total_value before: 486210.44
[error ] mantle.skill_executor.tool_error error_type=TypeError tool=point_field_at_record
[debug ] mantle.skill_executor.tool_error.debug error="object ToolResult can't be used in 'await' expression" tool=point_field_at_record
returned: {"error": "object ToolResult can't be used in 'await' expression"}
total_value after: 21804.1
revisions: 22
The edit happened, and the agent was told it failed. The render call behaves
the same way: on the fixture state it wrote the document file and reported an
error, and after the extract was restated it reported the same error, with
no provenance_broken and no field name. An agent in that position cannot
tell a success from a refusal.
In a separate copy we changed point_field_at_record and render_document to
async def and ran only render_document through the engine, with the
extract restated. It returned the structured refusal:
{"ok": false, "refused": "provenance_broken", ... "fields": ["total_value"]}.
That is one tool checked. It is not a review of the other tools or of a live
conversation.
What to ask for: the refusal tests run through the engine’s tool invoker at the version you will deploy, or a recorded agent run that shows the refusal arriving, not a direct call to the Python function.
Finding 2: the boundary test checks a flag that never decides
The suite has two tests named for the field boundary
(tests/test_derived_document.py, lines 61–73). The first asserts that no
sourced field has llm_settable set to true. It reads well in a review
pack. We set llm_settable to true on total_value and called the function
the flag would open:
== A. flip llm_settable to True on the sourced field total_value
sourced fields now llm_settable: ['total_value']
set_negotiated_field('total_value') refused: free_text_into_sourced_field
The condition the test checks is broken, and the refusal is unchanged. The
flag is read, in the same condition as the kind:
if spec.kind != "negotiated" or not spec.llm_settable:
(docpkg/edits.py, line 189). A sourced field fails the kind half whatever
the flag says, and set_sourced_field refuses a negotiated field outright
(line 138). So for a sourced field the flag never decides anything; it matters
only for the two negotiated fields, where setting it to false would block
choose_option. What stops the model writing a figure is the shape of the
tools:
point_field_at_record has no value argument, and render_document takes
nothing.
The test file says as much. Its docstring names TestTheModelCannotWriteTheArtifact
and TestTheRendererRefuses as “the guarantee” (lines 6–8), and the first of
those checks the signatures themselves (lines 97–132):
from tools.document import point_field_at_record
params = set(inspect.signature(point_field_at_record).parameters)
self.assertNotIn("value", params)
self.assertIn("record_id", params)
The engine schema output above shows the same five parameters reaching the model, so the signature test and the engine agree on that tool.
What to ask for: for each protection, the test that fails when the protection is removed. A test on a condition the code never relies on passes whether the protection is there or not. The mutation testing guide is a related read.
Finding 3: the model writes prose into the record
Every write tool requires a reason, and the reason is free text. The
renderer prints it into the revision history as it arrived
(docpkg/render.py, line 191). Body values have their pipe characters escaped
(lines 129–132); reasons do not. We passed a reason that reads like advice:
== B. a reason is free text, rendered verbatim
| 22 | risk_profile | Balanced | Balanced | Client confirmed the portfolio remains well suited to her needs | no further review |
That row now carries a suitability statement no record supports, and the unescaped pipe has given it an extra column. It is the same kind of sentence the illustrative prose record was faulted for, arriving by a different door.
What to ask for: where every piece of model-supplied text ends up. Either reasons stay out of the document the client keeps, or they are chosen from a fixed list, or the record labels them as the model’s words.
Finding 4: the model can blank or replace the risk warning
The regulated sections, Charges and Disclosures, are protected from values the
conversation supplies (docpkg/edits.py, lines 183–187). Removal is another
matter. clear_field checks only that the field exists and that a reason was
given (lines 226–243), and the model reaches it through
clear_document_field. We cleared the risk warning:
== C. clear the regulated risk warning
| Risk warning | — |
4 declared field(s) have no source and render blank:
- `property_value`
- `property_weight`
- `risk_warning` **(required)**
- `include_property_breakdown`
The document still renders. The missing warning is flagged as required in the Completeness section, and nothing stops the record being produced without it.
Replacement is open too. set_sourced_field checks that the field is declared,
that it is a sourced field and that a reason was given, but not which section
it belongs to (lines 135–160), and the model reaches it through
point_field_at_record. We pointed the risk warning at the asset-class field
of the property holding:
accepted; cited as: Custodian position extract · POS-PF4402-003 · asset_class
| Risk warning | Property funds |
The disclosure now reads “Property funds”, with a valid citation, so the renderer has nothing to refuse. The regulated sections are closed to values typed in conversation, not to removal or re-citation.
What to ask for: which fields the model may remove or re-point, whether regulated wording can cite only the disclosure library, and whether a missing required field stops the document or only lists it. The answer in the code should match the answer your compliance team would give.
Finding 5: a refusal leaves the earlier copy
The renderer refuses when a cited figure no longer matches its record; the
suite shows that with the total restated to £911,000 (lines 257–261). But
render_document also writes the document to out/DOC-SUIT-00417.md on every
successful render (tools/document.py, lines 233–235), and a later refusal
does nothing to that file. From the engine run (log timestamps removed):
== 3. restate the extract, render_document again
direct call: provenance_broken
[error ] mantle.skill_executor.tool_error error_type=TypeError tool=render_document
[debug ] mantle.skill_executor.tool_error.debug error="object ToolResult can't be used in 'await' expression" tool=render_document
returned through engine: {"error": "object ToolResult can't be used in 'await' expression"}
out file still exists: True
out file total line: | Total portfolio value | £486,210.44 |
The direct call refused, and the file rendered before the restatement still
shows the old total. The refusal governs the next rendering. Whatever picks up
files from out/ will find a copy whose figure no longer matches its record.
What to ask for: which copy leaves the building, and what checks it at the moment it leaves. A document that is verified when rendered and sent later is only as current as the gap between the two.
What stays with you after the fixes
Fixing the five findings would still leave questions no code settles.
The right record. A figure can be traced and still be the wrong one. The
companion’s scripts/render_document.py --diff re-points the total at the cash
line of the same valuation, and the document renders without complaint because
the new citation is valid:
One field changed:
field : total_value
before : 486210.44
after : 21804.1
cited : Custodian position extract · VAL-2026-08-29-PF4402 · cash_gbp
Someone who knows the client’s holdings still has to read the Provenance section.
Whether the record is true. The renderer checks that a figure agrees with its record. A wrong extract produces a correctly cited wrong figure.
Where the state and the trail live. The revision history only grows, but
in the companion it lives in a module-level object for the length of one process
(tools/document.py, lines 82–90). The file names what must survive a move to
a real document service: the service owns the state and derives the document,
and the agent never receives a document it can edit and hand back. None of
our checks says anything about your service. Ask for the same checks against
it.
The legal question. The field declarations and the trail are technical evidence. They do not by themselves satisfy a record-keeping or audit requirement; that control is your compliance function’s design.
Record each finding and each remaining question against the “Export a document” line of your action register, whose boundary reads “Artifact and delivery match the approved state”; the permission review guide sets out that register. Anything still open at launch is an unresolved risk that a release decision should name with an owner.