Skip to content
RasaGet a free licence
Guides for AI teams

Guide · Risk reviewer / domain owner

How to audit a document an AI agent generates

We audited an agent that derives documents from records. Its tests pass; the engine call path, free-text reasons and a replaceable risk warning do not.

by Rod Rivera

About 4 minutes

Key takeaways (5)
  • A careful read of an agent’s document checks one copy on one day. It cannot show where a figure came from, and it does not repeat for the next thousand documents.
  • Ask for evidence from the path the agent runs. Every companion tool is a plain function; the two we ran through Rasa Pro 3.20.0rc1’s tool invoker did their work and then reported only a generic error, so the tested refusal never reached the agent.
  • Ask which test fails when the protection is removed. In the companion the guarantee is the tool signatures, pinned by TestTheModelCannotWriteTheArtifact; the test on the llm_settable flag of sourced fields checks a condition that never decides the outcome, because the kind check refuses first.
  • List every way the model can write to the document. In the companion, a free-text reason on every edit is printed verbatim into the record, and the model can blank or replace the regulated risk warning while the document still renders.
  • A refusal stops the next rendering, not the copy already written. Ask which file leaves the building and what re-checks it on the way out.

The record below is illustrative, not a recorded model run: it is the example the deriving-a-document tutorial opens on, and the client, Marged Ellis, and her portfolio are invented fixture data. An adviser asks an agent for a client suitability record. Here is the first sentence of its body; the example goes on from there:

Your portfolio is currently valued at approximately £486,000, held across
a balanced mix of global equities (around 58%), sterling corporate bonds
(roughly 30%), and property funds (about 12%).

Every figure sits close to the invented custodian extract behind it: £486,210.44 in total, equities at 57.7%, bonds at 30.4% and a property row at 11.9%. A reviewer holding the extract would tick all four. What the page does not say is which figures were read from a record and which were rounded, remembered from an older valuation or composed so the sentence would feel complete.

The property figure is the one to look at. The extract has a property row, but the record the companion project assembles from the same fixture never cites it. The rendered record’s Completeness section says so:

## Completeness

3 declared field(s) have no source and render blank:

- `property_value`
- `property_weight`
- `include_property_breakdown`

The prose version states a property weighting that the assembled record has no citation for, in the same confident voice as the figures that do.

This guide is for the person asked to approve an agent that writes documents a client or regulator will keep: summaries, letters, case records. Its position is that a fluent document proves nothing, and neither does a green test suite unless the tests run the path the agent actually uses. Sign off on a list of every way the model can write to the document, each tested through the engine you will deploy, and on a refusal that reaches both the agent and any copy already written.

We tested that position on the best case available: the companion project behind the tutorial, tutorials/rasa-document-artifact-tutorial in RasaHQ/rasa-community-resources at commit 69e27b6, which was designed so that a model cannot write a figure. Its 41 tests pass. We still found five gaps a sign-off should catch, and this guide is built on them. The scenario and its field list are synthetic: the fixture’s README says every person, holding and valuation in it is “invented for this tutorial” (data/source/README.md, lines 1–5), and the property row is POS-PF4402-003 in data/source/holdings.json (lines 43–47). Nothing here claims the design satisfies any regulator’s suitability, record-keeping or disclosure rule, and which fields your own documents need is a decision for your team.

Why reading the output is not the control

The natural control is a careful review of what the agent produced. It fails in three ways, and none of them is about how careful the reviewer is.

  • It checks agreement, not origin. “About 12%” agrees with 11.9% within rounding. The reviewer cannot see whether the model read that row, recalled a figure from an earlier conversation or produced a number that fitted.
  • It checks a copy, not a process. A clean review of this record says nothing about the next one.
  • It happens once. A custodian restates a valuation after the review and before the letter goes out, and the reviewed figure is now wrong under the firm’s name.

The tutorial’s full example goes on to say the allocation “remains well suited to your Balanced risk profile”. That is a suitability judgement, and none of the fixture’s three sources (holdings.json, factfind.json and references/disclosures.json) contains it or the word “suited”. A reviewer may agree with it; that agreement is the reviewer’s judgement, not evidence about where the sentence came from.

What the model can touch in this design

The companion project takes the pen away. The conversation edits a structured state, never the document, and the renderer rebuilds the whole document from that state each time it runs (docpkg/render.py, render_markdown). Most fields are sourced: the model names a record and the record supplies the value. Two are negotiated: the model picks a value from a short fixed list, such as who the record is addressed to (docpkg/state.py, lines 114–119).

The model can only fill in the parameters the engine offers it. We asked Rasa Pro 3.20.0rc1 to build the schema for each document tool, offline, with the engine’s own agent_tool_schema_from_custom_tool (rasa/mantle/llm/tool_schemas.py, lines 577–610 in the wheel). The output of that run:

point_field_at_record: properties=['field_key', 'source_id', 'record_id', 'record_field', 'reason'] required=['field_key', 'source_id', 'record_id', 'record_field', 'reason'] additionalProperties=False
choose_option: properties=['field_key', 'option', 'reason'] required=['field_key', 'option', 'reason'] additionalProperties=False
clear_document_field: properties=['field_key', 'reason'] required=['field_key', 'reason'] additionalProperties=False
render_document: properties=[] required=[] additionalProperties=False
Before reading on: from those four schemas, list every way the model can change what the document says

There are four, and a sign-off should name all of them.

  1. Which record a sourced field cites, through point_field_at_record. It cannot type a figure, but it can cite the wrong record.
  2. A value from a fixed list, through choose_option, for the two negotiated fields.
  3. Removing any field, through clear_document_field, which takes any field_key.
  4. Free text, through reason, which every write tool requires and which the renderer prints into the record’s revision history.

The fourth is easy to miss, because nothing in the schema says where a reason ends up.

The list names the inputs. What it does not show is where each one ends up, including the file the render tool leaves behind:

Record choices and fixed-list values Document state Removed fields Reason text Record body (values, blanks, Provenance) values and blanks Revision history reasons, verbatim out/DOC-SUIT-00417.md render_document writes A later refusal leaves it in place
FigureWhere each model input lands

Auditing a document an AI agent generates: five findings

We ran each check offline, in a fresh export of the companion at 69e27b6. The engine checks used an installed Rasa Pro 3.20.0rc1 under Python 3.12.13; the suite was also run under Python 3.14.3 with no engine. The scripts are kept with this guide’s brief as receipts, each run’s output saved verbatim with the command on its first line.

FindingWhat we ranResult
1. The refusal does not reach the agentTools called through the engine’s call_tool_funcState changed; the agent was told only “error”
2. The boundary test checks a flag that never decidesllm_settable set to true on total_valueThe test’s condition breaks; the tool still refuses
3. The model writes prose into the recordA reason containing a suitability sentence and a pipePrinted verbatim, unescaped, in the history
4. The model can blank or replace the risk warningclear_field, then a re-citation, on risk_warningBoth render: blank, then “Property funds”
5. A refusal leaves the earlier copyRender, restate the extract, render againRefused; the first file still holds £486,210.44

None of these needs a model, a licence or a network, and none is a claim about what a model would choose to do in a conversation. Each is a path the code leaves open.

Finding 1: the refusal the tests prove is not what the agent receives

Engineering will show you the suite. It passes with or without the engine installed; the last lines of our make test run, the command engineering would use, under Python 3.14.3:

Ran 41 tests in 0.053s

OK

Among those tests is one written for exactly the question a reviewer asks: does a refusal reach the agent? Its docstring reads “The path the AGENT runs, not just the function underneath it” (tests/test_derived_document.py, line 296), and another test in the same file warns against “asserting on a convenient internal path while the path that actually runs was different” (lines 115–120). But the refusal test calls td.render_document() directly, as a plain function (lines 295–305).

The agent does not call it that way. In Rasa Pro 3.20.0rc1 the engine runs a local tool with return await func(context=ctx, **tool_args), None (rasa/mantle/orchestration/tool_execution/invoker.py, line 234). The companion’s tools are declared with plain def, so the function runs to the end, side effects included, and only then does awaiting its result raise a TypeError. The invoker’s catch-all (lines 161–172) turns that into a generic error for the model.

This is a defect in the companion, not an engine quirk. The engine’s decorator documents the contract: tool() is described as “Mark an async function as an LLM-callable tool”, and a tool function’s type is Callable[..., Awaitable[ToolResult]] (rasa/mantle/tools/decorator.py, lines 80 and 270).

We called the tools through ToolInvoker.call_tool_func, the engine method that runs a builder tool (call_builder_tool hands every call to it, line 106), with a stand-in for the session object. The only session method it uses on this path is recording_write, which in the engine sets and restores a marker around the call and lets any exception through (rasa/mantle/state/tracker_handle.py, lines 525–545). Ours does nothing, so it cannot change what the tool returns. The first call re-points the total at the cash line (log timestamps removed):

== 1. point_field_at_record through the engine
total_value before: 486210.44
[error    ] mantle.skill_executor.tool_error error_type=TypeError tool=point_field_at_record
[debug    ] mantle.skill_executor.tool_error.debug error="object ToolResult can't be used in 'await' expression" tool=point_field_at_record
returned: {"error": "object ToolResult can't be used in 'await' expression"}
total_value after: 21804.1
revisions: 22

The edit happened, and the agent was told it failed. The render call behaves the same way: on the fixture state it wrote the document file and reported an error, and after the extract was restated it reported the same error, with no provenance_broken and no field name. An agent in that position cannot tell a success from a refusal.

In a separate copy we changed point_field_at_record and render_document to async def and ran only render_document through the engine, with the extract restated. It returned the structured refusal: {"ok": false, "refused": "provenance_broken", ... "fields": ["total_value"]}. That is one tool checked. It is not a review of the other tools or of a live conversation.

What to ask for: the refusal tests run through the engine’s tool invoker at the version you will deploy, or a recorded agent run that shows the refusal arriving, not a direct call to the Python function.

Finding 2: the boundary test checks a flag that never decides

The suite has two tests named for the field boundary (tests/test_derived_document.py, lines 61–73). The first asserts that no sourced field has llm_settable set to true. It reads well in a review pack. We set llm_settable to true on total_value and called the function the flag would open:

== A. flip llm_settable to True on the sourced field total_value
sourced fields now llm_settable: ['total_value']
set_negotiated_field('total_value') refused: free_text_into_sourced_field

The condition the test checks is broken, and the refusal is unchanged. The flag is read, in the same condition as the kind: if spec.kind != "negotiated" or not spec.llm_settable: (docpkg/edits.py, line 189). A sourced field fails the kind half whatever the flag says, and set_sourced_field refuses a negotiated field outright (line 138). So for a sourced field the flag never decides anything; it matters only for the two negotiated fields, where setting it to false would block choose_option. What stops the model writing a figure is the shape of the tools: point_field_at_record has no value argument, and render_document takes nothing.

The test file says as much. Its docstring names TestTheModelCannotWriteTheArtifact and TestTheRendererRefuses as “the guarantee” (lines 6–8), and the first of those checks the signatures themselves (lines 97–132):

        from tools.document import point_field_at_record

        params = set(inspect.signature(point_field_at_record).parameters)
        self.assertNotIn("value", params)
        self.assertIn("record_id", params)

The engine schema output above shows the same five parameters reaching the model, so the signature test and the engine agree on that tool.

What to ask for: for each protection, the test that fails when the protection is removed. A test on a condition the code never relies on passes whether the protection is there or not. The mutation testing guide is a related read.

Finding 3: the model writes prose into the record

Every write tool requires a reason, and the reason is free text. The renderer prints it into the revision history as it arrived (docpkg/render.py, line 191). Body values have their pipe characters escaped (lines 129–132); reasons do not. We passed a reason that reads like advice:

== B. a reason is free text, rendered verbatim
| 22 | risk_profile | Balanced | Balanced | Client confirmed the portfolio remains well suited to her needs | no further review |

That row now carries a suitability statement no record supports, and the unescaped pipe has given it an extra column. It is the same kind of sentence the illustrative prose record was faulted for, arriving by a different door.

What to ask for: where every piece of model-supplied text ends up. Either reasons stay out of the document the client keeps, or they are chosen from a fixed list, or the record labels them as the model’s words.

Finding 4: the model can blank or replace the risk warning

The regulated sections, Charges and Disclosures, are protected from values the conversation supplies (docpkg/edits.py, lines 183–187). Removal is another matter. clear_field checks only that the field exists and that a reason was given (lines 226–243), and the model reaches it through clear_document_field. We cleared the risk warning:

== C. clear the regulated risk warning
| Risk warning | — |
4 declared field(s) have no source and render blank:

- `property_value`
- `property_weight`
- `risk_warning` **(required)**
- `include_property_breakdown`

The document still renders. The missing warning is flagged as required in the Completeness section, and nothing stops the record being produced without it.

Replacement is open too. set_sourced_field checks that the field is declared, that it is a sourced field and that a reason was given, but not which section it belongs to (lines 135–160), and the model reaches it through point_field_at_record. We pointed the risk warning at the asset-class field of the property holding:

accepted; cited as: Custodian position extract · POS-PF4402-003 · asset_class
| Risk warning | Property funds |

The disclosure now reads “Property funds”, with a valid citation, so the renderer has nothing to refuse. The regulated sections are closed to values typed in conversation, not to removal or re-citation.

What to ask for: which fields the model may remove or re-point, whether regulated wording can cite only the disclosure library, and whether a missing required field stops the document or only lists it. The answer in the code should match the answer your compliance team would give.

Finding 5: a refusal leaves the earlier copy

The renderer refuses when a cited figure no longer matches its record; the suite shows that with the total restated to £911,000 (lines 257–261). But render_document also writes the document to out/DOC-SUIT-00417.md on every successful render (tools/document.py, lines 233–235), and a later refusal does nothing to that file. From the engine run (log timestamps removed):

== 3. restate the extract, render_document again
direct call: provenance_broken
[error    ] mantle.skill_executor.tool_error error_type=TypeError tool=render_document
[debug    ] mantle.skill_executor.tool_error.debug error="object ToolResult can't be used in 'await' expression" tool=render_document
returned through engine: {"error": "object ToolResult can't be used in 'await' expression"}
out file still exists: True
out file total line: | Total portfolio value | £486,210.44 |

The direct call refused, and the file rendered before the restatement still shows the old total. The refusal governs the next rendering. Whatever picks up files from out/ will find a copy whose figure no longer matches its record.

What to ask for: which copy leaves the building, and what checks it at the moment it leaves. A document that is verified when rendered and sent later is only as current as the gap between the two.

What stays with you after the fixes

Fixing the five findings would still leave questions no code settles.

The right record. A figure can be traced and still be the wrong one. The companion’s scripts/render_document.py --diff re-points the total at the cash line of the same valuation, and the document renders without complaint because the new citation is valid:

One field changed:
  field  : total_value
  before : 486210.44
  after  : 21804.1
  cited  : Custodian position extract · VAL-2026-08-29-PF4402 · cash_gbp

Someone who knows the client’s holdings still has to read the Provenance section.

Whether the record is true. The renderer checks that a figure agrees with its record. A wrong extract produces a correctly cited wrong figure.

Where the state and the trail live. The revision history only grows, but in the companion it lives in a module-level object for the length of one process (tools/document.py, lines 82–90). The file names what must survive a move to a real document service: the service owns the state and derives the document, and the agent never receives a document it can edit and hand back. None of our checks says anything about your service. Ask for the same checks against it.

The legal question. The field declarations and the trail are technical evidence. They do not by themselves satisfy a record-keeping or audit requirement; that control is your compliance function’s design.

Record each finding and each remaining question against the “Export a document” line of your action register, whose boundary reads “Artifact and delivery match the approved state”; the permission review guide sets out that register. Anything still open at launch is an unresolved risk that a release decision should name with an owner.