Skip to content
RasaGet a free licence
Guides for AI teams

Guide · Conversation designer / UX specialist

Stop customers repeating themselves after a handoff

To stop customers repeating themselves after a handoff, build the human agent’s first screen from typed fields and test the promise on the agent’s own path.

by Rod Rivera

About 8 minutes

Key takeaways (5)
  • Compute the screen a human agent reads first from typed fields on every read. The companion pattern is built on the premise that a desk reads prose first, so it leaves no place to store a summary paragraph that could go stale.
  • Specify the first screen as the desk questions it retires and the field that retires each one. A question with no field behind it will be asked again, whatever the summary says.
  • Show withheld fields by name and never by value, so the desk knows a one-time code was held back by policy and does not ask the caller to read it out.
  • The transfer promise is spoken before the transfer tool runs and may be reworded by a model. Check it before release against the screen the agent’s own tools produce.
  • In the companion, the agent’s own path drops which charge the caller disputed while the “no repeat” check still passes. Write the promise no wider than the questions you test.

Assistant: I’ll bring in a specialist and pass them everything we’ve covered, so you won’t need to repeat any of it. Shall I go ahead?

Caller: Yes, go ahead.

The desk agent’s screen:

### TODAY — one free-text handoff_reason (every example in the catalog)

========================================================================
INBOUND HANDOFF ho_today — Dispute needs a second factor the caller cannot complete on this line.
========================================================================
CALLER      UNIDENTIFIED
TRUST       tier=unverified — Identity NOT established. Treat every identifying detail as unconfirmed.
ASKING FOR  unknown (stage: stated)

ALREADY TRIED (do not retry these):
  · (nothing)

DO NOT ASK — the caller already answered these:
  · (nothing)

YOU MAY:
  · Answer general questions
  · Explain what the agent already did
  · DO NOT discuss account-specific details — identity not established
  · DO NOT action irreversible changes — step the caller up first
========================================================================

The caller will now be asked 5 question(s) they already answered:
  ✗ Can I take your name?
  ✗ Can you confirm your date of birth?
  ✗ Which account is this about?
  ✗ What are you calling about today?
  ✗ Have you tried anything already?

Desk agent: Can I take your name? … Can you confirm your date of birth? … Which account is this about? … What are you calling about today? … Have you tried anything already?

The caller answers all five again, seconds after being told they would not have to.

The dialogue is an illustrative script, not a recorded call; the assistant’s line is the confirmation text in the companion project. The screen is real output from make compare in patterns/voice-handoff-context (RasaHQ/rasa-community-resources at 69e27b6), run offline. It shows what a desk receives when a handoff carries nothing but a handoff_reason string. This companion’s own transfer_to_human sends much more; the script builds the “TODAY” screen from the reason alone (scripts/agent_desk.py, lines 107–113) to stand in for the rest of the catalogue. By “the catalogue” the pattern’s authors mean the other example projects in rasa-community-resources: its README says every human_handoff skill there “captures one free-text handoff_reason string and opens a ticket” (README.md, lines 14–17), and the baseline test’s docstring says the same (tests/test_handoff_context.py, lines 190–194). The companion’s agent.yml calls the assistant “a concise, competent banking assistant” and names no bank. The caller, account and merchant are fixture values, and the desk is a terminal fixture standing in for a contact-centre screen.

Make the summary a view over typed fields

test_the_catalogs_current_handoff_answers_nothing pins the baseline: a package built from a reason string alone leaves the desk all five of its opening questions. A sentence, however well written, gives the desk nothing it can check a question off against.

The stance of this guide: the screen a human agent reads first should be computed from typed fields, never written and stored as prose. The companion is built on a premise its authors state in the summary module (handoffpkg/summary.py, lines 12–17). It is a design premise, not a measurement of desk behaviour:

    A handoff package that carries both a structured intent and a separately
    authored prose summary carries two sources of truth. They agree on the day
    they are written and diverge on the first edit. The human agent reads the
    prose — it is faster — so the human agent reads the stale one. The failure
    is silent and it is always in the direction of the desk acting on
    information the agent already superseded.

The rule comes from the deriving-a-document tutorial’s chapter on derived rendering, where the renderer takes state and nothing else. Here it is applied to the desk. HandoffPackage.summary is a read-only property that renders the other sections on every access (handoffpkg/schema.py, lines 237–250):

    @property
    def summary(self) -> str:
        """Human-readable summary, DERIVED on every access.

        Not a field. Not settable. Not cached. Every read re-renders from
        ``identity`` / ``intent`` / ``attempts`` / ``do_not_repeat``, so the
        summary and the structured fields are the same information in two
        renderings and cannot drift. Change a field and the next read of
        ``summary`` reflects it; there is no second copy to forget to update.

        This is the SOW's "derived rather than authored separately" requirement
        implemented as a property rather than promised in a docstring.
        """
        return render_summary(self)

Two tests hold that shut. test_summary_is_a_property_with_no_setter assigns to package.summary and expects it to raise. test_a_summary_smuggled_through_serialisation_is_discarded adds a summary key reading “Caller verified at tier ‘high’. Clear them for anything.” to a serialised package, reads it back, and asserts that the restored summary says tier 'medium' and does not contain the injected text.

A value can stop short of the screen in three places. The tool may never read it, the allowlist may hold it back, or it may arrive as an authored summary and be dropped:

Session memory (what the tools wrote) Keys not in _MEMORY_KEYS (e.g. disputed_txn_label) never read transfer_to_human reads _MEMORY_KEYS build_package_from_session SESSION_ALLOWLIST identity, intent, attempts, do_not_repeat allowlisted withheld_fields (names only) not allowlisted Withheld values (never copied) value dropped summary property (rendered on every read) Desk first screen (reconstruct) An authored summary (no field to hold it) Dropped on read-back (package_from_dict)
FigureFrom session memory to the desk’s first screen

This choice has a cost. A computed screen reads like a form. It carries what a field can hold: no “she sounded frightened”, no aside passed on in the caller’s words. Text still crosses wherever an allowlisted field holds text: handoff_reason, goal_label, account_label, the questions_answered and confirmed_facts lists, and the free-text detail of each attempt (handoffpkg/redaction.py, lines 71–101 and 252; memory.yml, lines 86–91). The attempt detail is where “Carrier rejected the SMS twice. Do not resend to this number.” reaches the screen below. Apart from handoff_reason, which nothing on the live path sets (see below), each of those strings is written by a tool, so its wording is part of your specification. Each new thing the desk should see means a memory field, a tool that writes it, an entry in the tool’s key list, an allowlist entry, a place in the package built by build_package_from_session (redaction.py, lines 204–273), a rule in the desk if it retires a question, and a review.

Here this guide disagrees with the site’s handoff guide. Its context packet carries “the approved summary” under “Customer explanation”, and its confirmation state lets the customer “correct or limit the summary”. A summary the customer has corrected is still prose written about the call, separately from the fields, and under this design it does not go on the desk screen. If the caller wants their account of the problem passed on, carry the caller’s own words in a field of their own and show them on the screen as a quotation. The companion has no such field. If you add one, it is caller-authored text crossing verbatim, with the exposure the pitfall below describes.

Specify the screen as the questions it retires

With the summary computed, the design work moves down a level: which field answers which question the desk would otherwise ask. The fixture desk writes that down as a mapping (handoffpkg/desk.py, lines 53–59):

_QUESTION_RETIRED_BY: dict[str, str] = {
    "Can I take your name?": "identity.display_name",
    "Can you confirm your date of birth?": "identity.verified_tier",
    "Which account is this about?": "intent.details.account_id",
    "What are you calling about today?": "intent.goal",
    "Have you tried anything already?": "attempts",
}

unanswered_questions() (desk.py, lines 178–191) walks that table against a package and asks _package_answers (from line 194) whether each field is populated. Those rules (lines 208–228) are where the design decisions sit:

Desk questionRetired byStill asked when
Can I take your name?identity.display_namethe name is empty
Can you confirm your date of birth?identity.verified_tierthe tier is unverified
Which account is this about?intent.details.account_idno account id crossed
What are you calling about today?intent.goalthe goal is empty or unknown
Have you tried anything already?attemptsthe list is empty, even if the true answer is “no”

The date-of-birth question is retired by the verification tier, not by a date of birth. The desk needs to know that identity was established and how strongly; it never needs the value.

The screen the agent’s own tools produce

The second half of make compare shows a full screen, but it is rendered from DEMO_SESSION, a dictionary typed by hand in scripts/agent_desk.py (lines 35–74). To see what the running agent would send, we called the dispute skill’s tools and transfer_to_human directly, with a stand-in for the memory context, on an unmodified copy of the pattern. No model, network or licence is involved. Two values are typed in. The first is the charge the caller picks, txn_c101; the model can set disputed_txn_id on a call, because the dispute skill declares it llm_settable (skills/dispute_transaction/memory.yml, lines 3–6). The second is the one-line handoff_reason, and nothing on the live path can set that one (see below); the script types it only so the transfer can run. The script is editorial/receipts/design-the-handoff-summary/agent-path-desk-screen.py in the site repository, and the first line of its saved transcript records the exact command. The handoff id is random on each run (tools.py, line 177). Its unedited output, Python 3.12.13 and rasa-pro 3.20.0rc1:

attempts_log as written by the dispute tools:
verify_passphrase|succeeded||Caller answered the knowledge factor on the first try.
send_otp_sms|failed|delivery_failed|Carrier rejected the SMS twice. Do not resend to this number.
raise_dispute|blocked|insufficient_tier|Dispute needs tier 'high'; caller is at 'medium'.
disputed_txn_label in memory: Northgate Fuel $248.00
disputed_txn_label in _MEMORY_KEYS: False
dispute_date in memory: False
withheld_fields: ['otp_code', 'pin_attempt']
desk_still_needs_to_ask: []
hint: Context package delivered. Tell the caller the handoff id and that they will not need to repeat themselves. Do not read back any withheld field.

========================================================================
INBOUND HANDOFF ho_ed47b54e — Dispute needs a second factor the caller cannot complete on this line.
========================================================================
CALLER      Jordan Rivera (cust_00417) via voice
TRUST       tier=medium — Verified with a knowledge factor. Account-specific information is in scope.
ASKING FOR  Dispute a card transaction [account_id=acc_checking, account_label=Everyday Checking, card_last_four=4821] (stage: blocked)

ALREADY TRIED (do not retry these):
  · verify_passphrase → succeeded (Caller answered the knowledge factor on the first try.)
  · send_otp_sms → failed (Carrier rejected the SMS twice. Do not resend to this number.)
  · raise_dispute → blocked (Dispute needs tier 'high'; caller is at 'medium'.)

DO NOT ASK — the caller already answered these:
  · Can I take your name?
  · Can you confirm your date of birth?
  · Which account is this about?
  · What are you calling about today?
  · verified: knowledge_passphrase
  · confirmed: Caller consented to the call being recorded.

YOU MAY:
  · Answer general questions
  · Explain what the agent already did
  · Discuss account-specific details
  · DO NOT action irreversible changes — step the caller up first

WITHHELD BY POLICY (present in the session, not transferred):
  · otp_code
  · pin_attempt
  Do not ask the caller to read these out to you.
========================================================================

The attempts_log lines at the top are what the dispute tools write. transfer_to_human turns them into the package’s attempts (tools.py, line 175), which is why the last desk question counts as retired even though it is missing from the DO NOT ASK list. That list shows the questions the tools recorded as text; retirement is computed from fields. Specify one of them as the source for both, or the screen will say one thing while the check says another.

Look at the ALREADY TRIED block next. The line telling the desk not to resend the SMS matters more than the name, because without it the desk agent’s first helpful move is to repeat the step that already failed twice.

Then look at ASKING FOR. The account and the card are there. The charge the caller disputed is not. The dispute skill writes it: raise_dispute reads disputed_txn_id (skills/dispute_transaction/tools.py, line 156) and stores the merchant and amount as disputed_txn_label (line 170). Both keys are declared in the skill’s own memory.yml (lines 3 and 9), and the output above shows Northgate Fuel $248.00 sitting in memory. But the handoff tool’s key list (skills/human_handoff/tools.py, lines 68–96) names neither key. It names dispute_amount, dispute_merchant and dispute_date instead, which no tool writes, and no tool writes a date at all. So the charge never leaves memory. desk_still_needs_to_ask is still empty, because “Which charge is it?” is not one of the five questions.

Name what was withheld

The last block on that screen is the easiest one to leave out of a spec. A credential that is absent and a credential that was held back look the same on a screen that only shows what arrived. The pattern sends the names of withheld keys (handoffpkg/redaction.py, lines 104–108):

# Keys whose *names* are still useful to the desk even though their values must
# not cross. Naming them lets the summary say "withheld by policy" rather than
# leaving the desk to wonder, and stops an agent asking the caller to repeat a
# credential the system already has.
ANNOUNCE_WITHHELD = True

That comment states the authors’ intent: a named gap is meant to stop a helpful desk agent asking the caller to fill it. On the agent’s path above, the two names are otp_code and pin_attempt, the credentials the dispute tools write. test_withheld_names_cross_but_values_do_not plants seven sensitive values, checks that every name reaches withheld_fields, and checks that the planted PIN, 4242, appears nowhere in the serialised package. test_every_key_the_agent_collects_is_either_allowlisted_or_withheld checks there is no silent third outcome for a key the tool reads.

For the screen, specify the wording as well as the list. “Do not ask the caller to read these out to you” is an instruction to a person. The skill carries the same rule on the assistant’s side (skills/human_handoff/skill.md, lines 21–28):

Do NOT ask the caller to summarise the call, restate who they are, confirm their
account again, or repeat anything they already told you. All of it is already in
session state and `transfer_to_human` transfers it. Asking is the exact failure
this skill exists to prevent.

Never put a PIN, one-time code, passphrase, full card number or token into
`handoff_reason`. Those are withheld from the transfer by policy, and writing one
into the reason line would carry it across anyway.

Check the promise against the agent’s own path

Now the sentence the caller heard (skills/human_handoff/responses.yml, lines 1–10):

# Confirmation wording for the handoff. The promise made here — "you will not
# need to repeat yourself" — is one the context package has to actually keep,
# and tests/test_handoff_context.py is what keeps it honest.
responses:
  utter_confirm_handoff:
    - text: >-
        I'll bring in a specialist and pass them everything we've covered, so you
        won't need to repeat any of it. Shall I go ahead?
      metadata:
        rephrase: true

Two facts about this sentence change how you test it. Both come from reading the Rasa Pro 3.20.0rc1 wheel, not from a live run.

  • It is spoken before the package exists: it is the confirmation that requires_confirmation attaches to transfer_to_human (skill.md, lines 7–13). The engine’s confirmation gate says “The tool itself is never invoked here”, only “after the LLM confirms” (rasa/mantle/orchestration/tool_execution/constraints.py, lines 301–303).
  • It may not be the sentence spoken: rephrase: true is read from the response metadata into the response definition (rasa/mantle/content/responses.py, lines 353 and 362), which documents it as “the orchestrator rewords text via a dedicated LLM call” (line 122). The confirmation ask carries that flag (constraints.py, line 503).

The approved wording is where the promise starts, not a guarantee of what the caller hears. Nothing in the companion checks the spoken sentence, so the place to hold it is a check that runs before release, against a package the agent’s own tools produced.

The same reading turns up a gap before the sentence is ever spoken. The skill tells the model to “set handoff_reason” (skill.md, line 18), and transfer_to_human is offered only when that field has a value (requires: session.project.handoff_reason, line 9). But handoff_reason is a project memory field (memory.yml, lines 94–96). At 3.20.0rc1 the wheel rejects llm_settable on project fields: “Project fields cannot be written by the LLM” (rasa/mantle/memory/validation.py, lines 81–95). The set_fields tool offers only llm_settable fields and collect targets (rasa/mantle/llm/tool_schemas.py, lines 26–29 and 525–526), no tool in the pattern writes the field, and a tool whose requires is falsy is left out of the model’s tool list (tool_schemas.py, lines 974–990). So at this revision nothing on the live path can set the reason, and the transfer tool would not be offered at all. The pattern has no hooks file and no seed: true field, the other two ways a project field is written (a grep of the pattern is saved with this guide’s receipts). This comes from reading the source, not from a live call. Put a writer for the reason line in your specification: a tool that composes it from fields, so its wording is yours and never the caller’s. We have not built or tested that tool. This companion has already shipped a green test over a hand-built session while the running agent did not collect account_id; deriving test fixtures from production code covers that defect and the method that fixes it.

What to demand in your specification

The agent’s real path loses the disputed charge, the full demo screen is hand-built, and the promise is repeated regardless of what the desk still needs. So demand three things before you sign off the wording:

  1. Review only desk screens rendered from a package the agent’s own tools produced.
  2. Map each part of the promise to a desk question, and each question to its field, the tool that writes it and its entry in _MEMORY_KEYS.
  3. Make the line spoken after the transfer depend on desk_still_needs_to_ask.

The tool already reports that list. Here is the part of its result the assistant sees (skills/human_handoff/tools.py, lines 207–214):

            "withheld_fields": list(package.withheld_fields),
            "desk_still_needs_to_ask": list(unanswered_questions(package)),
            "freetext_risk": freetext_risk,
            "hint": (
                "Context package delivered. Tell the caller the handoff id and "
                "that they will not need to repeat themselves. Do not read back "
                "any withheld field."
            ),

The hint repeats the promise whatever desk_still_needs_to_ask contains, and so does the skill’s closing instruction (skill.md, lines 30–33). The companion does not branch on it; the branch belongs in your specification.

Try it: what should the assistant say when desk_still_needs_to_ask is not empty?

Write the line before you open this. One proposal, not built or tested in the companion, for a list holding “Which account is this about?”:

Assistant: You’re through to a specialist, reference ho_1a2b3c4d. They have your details and what we’ve tried. They may need to check which account this is about.

The reference is illustrative. The sentence names the one gap instead of denying it, so the caller is not surprised by the question and the promise stays true.

Make the promise no wider than the test

“Everything we’ve covered” is a promise about the whole conversation. The check covers five questions. Close that gap from the design side, in one of two ways:

  1. Widen the list: add “Which charge is it?” to the opening script and its entry in _QUESTION_RETIRED_BY in the same change. A question without a mapping makes unanswered_questions raise KeyError (desk.py, line 188), and transfer_to_human calls it on every transfer (tools.py, line 208). With both in place the check fails at this revision, and that failure is the point. Making it pass takes four more changes, none of which we have tried: collect disputed_txn_label in _MEMORY_KEYS; add it to SESSION_ALLOWLIST, or it is only named as withheld (redaction.py, lines 150–154); add it to the fixed detail_keys that build intent.details (lines 218–230); and add a branch to _package_answers, which returns False for any field it does not know (desk.py, line 229).
  2. Narrow the sentence: promise what the fields carry, and say what is still to come. A proposed rewording, not tested in the companion: “I’ll pass the specialist your details, what you’re calling about and what we’ve already tried, so you won’t have to go through it again. They’ll need to complete one extra security check before they can raise the dispute. Shall I go ahead?”

What these checks prove, and what they do not

The six tests named in this guide pass when run together, and so does the full 41-test suite, make test (both run offline; the output is saved with this guide’s receipts). Every run here is offline, on the package, the tools and the fixture desk. Four claims rest on reading source instead of a run: when the confirmation is spoken, that rephrase rewords it, that nothing can set handoff_reason, and that the model can set disputed_txn_id. None of it shows what a model says on a live call, such as whether it asks the caller to summarise despite the skill’s instruction or how it rewords the confirmation. Test those with your own conversations.

The allowlist is a design boundary, not a compliance control. The pattern’s docstring says PCI, HIPAA and similar regimes impose obligations “that a dataclass does not discharge” (redaction.py, lines 28–30), and that it does not redact the call recording, the transcript or logs written before the handoff (lines 31–34).

The fixture desk is a local directory: transfer_to_human writes the package as a JSON file into fixtures/desk_queue (tools.py, lines 188–189). Whether a real contact-centre system accepted and queued the case is a separate question, and nothing here answers it.

Questions you will hit while specifying it

Why not have a model write the summary at handoff time?

Because it would be another thing that can disagree with the fields. The rendering function’s docstring (handoffpkg/summary.py, lines 102–106) puts it this way: “A summary produced by an LLM at handoff time would be a sixth piece of state”, one that is “unversioned, unreproducible and free to contradict the other five.” A rendered view is the same every time for the same package, so you can test it. If the desk wants friendlier wording, change the renderer, not the data.

Should the assistant read the desk preview back to the caller?

No. The tool’s result includes the rendered screen so the assistant knows what was passed on, and the comment on it says it is returned “so the agent can honestly tell the caller what was passed on”, and “never so it can read it back to them” (tools.py, lines 198–199). Reading a screen aloud also risks speaking account details to whoever is on the line. Tell the caller the handoff id and what happens next, in a sentence you have checked against desk_still_needs_to_ask.