Skip to content
RasaGet a free licence
Guides for AI teams

Guide · AI product manager

Measure a handoff before you claim it worked

How to measure AI agent human handoff success: count the desk questions a context package retires, and keep that apart from handle time and CSAT.

by Rod Rivera

About 7 minutes

Key takeaways (5)
  • At launch you can defend one number: how many questions from a fixed desk opening script the handoff package answers. Handle time, CSAT and real repeat rates need production data.
  • Fix the denominator before you count. The handoff tool also returns questions_retired, which has no fixed denominator and does not belong on a slide.
  • The count checks that fields are populated, not that they are right, and it cannot see facts off the script. On the companion agent it reads five out of five while the package omits which charge was disputed.
  • The agent is instructed to promise callers they will not need to repeat themselves, whatever the package holds. Change the wording or record the cost in the review.
  • Split the slide into what is proven offline, what is not yet measured (with an owner, a data source and a date) and the changes you owe.

Aurora Home Insurance, a fictional insurer, has shipped a redesigned agent-to-human handoff. Its voice agent now passes a structured context package to the human desk instead of a one-line reason. The launch-review slide says:

Redesigned handoff packages cut average handle time by 30%.

The only evidence behind it is a passing offline test suite. Here is what that evidence supports, and what it does not:

What you can claimWhat you cannot claim yetWhat to measure instead
On the tested path the package answers all five desk opening questions; the old handoff answers noneCallers stopped repeating themselvesScripted questions the desk still asked, from call review
Every handoff is scored against the same five questionsThe package carries everything the case needsDesk questions that are not on the script, per handoff
All 41 offline tests passHandle time fell by 30%, or callers are happier with transfersTimed transferred calls before and after; a satisfaction measure scoped to transfers

The artefact linked from the slide is the tail of a make test run in the companion pattern:

----------------------------------------------------------------------
Ran 41 tests in 0.004s

OK

No call was recorded and no desk agent was timed. A search of all 25 files in the pattern finds no “30%” and no mention of handle time. The team had a real result and reported a different one.

Aurora is invented; the code is real. It is patterns/voice-handoff-context in RasaHQ/rasa-community-resources at 69e27b6, the revision pinned for Rasa Pro 3.20.0rc1. Its fixture call is a disputed card payment on a bank account, which the fictional team kept for its pilot.

The position this guide takes: at launch, report the count of fixed desk questions the package retires, proven offline, and nothing larger. The cost is a smaller slide, a measurement you still owe, and a script that sees only the questions written into it.

What does the passing test prove?

This section gives you the denominator the test counts against, and the before-and-after it proves.

“Callers repeat themselves less” has no denominator. The pattern fixes one before any caller exists: five questions a desk agent asks when nothing was transferred, each mapped to the package field that makes it unnecessary (handoffpkg/desk.py, lines 53–59):

_QUESTION_RETIRED_BY: dict[str, str] = {
    "Can I take your name?": "identity.display_name",
    "Can you confirm your date of birth?": "identity.verified_tier",
    "Which account is this about?": "intent.details.account_id",
    "What are you calling about today?": "intent.goal",
    "Have you tried anything already?": "attempts",
}

The denominator is the same five for every handoff, whatever the caller or case, so one handoff compares with the next and the old design with the new. The counting function, unanswered_questions, calls itself “The measurable form of the teaching claim” (lines 178–185).

The baseline test builds a package from one free-text handoff_reason and gets all five questions back; the agent-path test gets none back. One offline run on 21 September 2026 passed both, along with the permissions test this guide uses later.

Show the command and output for the three tests

From patterns/voice-handoff-context at 69e27b6, on Python 3.14.3 with no network and no model:

python3 -m unittest -v \
  tests.test_handoff_context.TestPackageSurvivesHandoff.test_the_catalogs_current_handoff_answers_nothing \
  tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question \
  tests.test_handoff_context.TestPackageSurvivesHandoff.test_desk_permissions_follow_the_tier_not_the_package
test_the_catalogs_current_handoff_answers_nothing (tests.test_handoff_context.TestPackageSurvivesHandoff.test_the_catalogs_current_handoff_answers_nothing)
Baseline: one free-text reason retires none of the desk's questions. ... ok
test_the_agent_path_retires_every_desk_question (tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question)
THE HEADLINE CLAIM, ON THE PATH THE AGENT ACTUALLY RUNS. ... ok
test_desk_permissions_follow_the_tier_not_the_package (tests.test_handoff_context.TestPackageSurvivesHandoff.test_desk_permissions_follow_the_tier_not_the_package)
A medium-tier caller must not be actioned for an irreversible change. ... ok

----------------------------------------------------------------------
Ran 3 tests in 0.001s

OK

The five questions are the pattern author’s, written for a fixture desk. Take yours from your desk’s call guide or QA scorecard, version the list, and date every figure against a version. If the script changes, the number changes, though no caller behaved differently.

Our desk's script has eight questions. Do we just extend the list?

Extend both structures together: unanswered_questions looks each question up with _QUESTION_RETIRED_BY[question] (line 188), so a question without a mapping raises a KeyError on every transfer. A question mapped to a field _package_answers does not recognise falls through to return False (line 229) and counts as unanswered on every handoff. Add a test for each new question, and restate the baseline before comparing figures.

Which count belongs on the slide?

The handoff tool returns two counts in the same response, and only one of them has a fixed denominator.

questions_retireddesk_still_needs_to_ask
What it countsLines the skills appended to questions_answeredScript questions whose package field is empty
DenominatorNone fixedThe five-question opening script
On the dispute path4[], so nothing left to ask
Sourceskills/human_handoff/tools.py, line 205The same file, line 208
Use it forNothing on a slidePackage coverage, tracked as a regression check

They disagree because the dispute skill appends four script questions (skills/dispute_transaction/tools.py, lines 98–103 and 135) and never “Have you tried anything already?”. The package still carries three attempts, so the scripted check counts that question as answered. Pick one definition and write it into the review.

An offline run of the real dispute and handoff tools printed both counts side by side.

Show the output from the real tools on the dispute path

editorial/receipts/measure-a-handoff-that-loses-context/agent-path-repro.py in this site’s repository calls load_caller_context, list_recent_charges, raise_dispute for charge txn_c101, then transfer_to_human, with a memory-only stand-in context. It ran on an unmodified copy of the pattern with Rasa Pro 3.20.0rc1 installed, because the tools import rasa.mantle. The run on 21 September 2026, on Python 3.12.13, printed, in part:

dispute_* keys in memory: ['disputed_txn_id', 'disputed_txn_label']
transferred: {"attempts_carried": 3, "goal": "dispute_transaction", "identity": true, "questions_retired": 4, "verified_tier": "medium"}
desk_still_needs_to_ask: []
hint: Context package delivered. Tell the caller the handoff id and that they will not need to repeat themselves. Do not read back any withheld field.
intent.details: {"account_id": "acc_checking", "account_label": "Everyday Checking", "card_last_four": "4821"}
ASKING FOR line: ASKING FOR  Dispute a card transaction [account_id=acc_checking, account_label=Everyday Checking, card_last_four=4821] (stage: blocked)

Where does the handoff lose the disputed charge?

This section shows the one fact the companion agent drops, and why the count cannot notice.

Caller picks the charge Dispute skill writes disputed_txn_label Handoff tool collects dispute_amount dispute_merchant dispute_date no key matches intent.details: account_id, account_label, card_last_four nothing to carry desk_still_needs_to_ask: [] five of five script check Desk screen: which charge? not stated what the desk sees
FigureOn the dispute path, the charge the caller picked never reaches the package

The handoff tool collects the disputed amount, merchant and date, but nothing on the dispute path writes them. The regression test checks three other keys, so the package names a card dispute on Everyday Checking but not which charge, and the count says the desk needs to ask nothing.

Show the source for this finding
  • skills/human_handoff/tools.py, lines 85–87: the handoff tool collects dispute_amount, dispute_merchant and dispute_date.
  • handoffpkg/redaction.py, lines 92–94 and 222–224: the allowlist passes those three keys into intent.details.
  • skills/dispute_transaction/tools.py, line 170: raise_dispute writes disputed_txn_label. No dispute tool writes the three keys above.
  • skills/human_handoff/tools.py, lines 68–96: the handoff tool’s key list names neither disputed_txn_label nor disputed_txn_id.
  • tests/test_handoff_context.py, lines 360–363: the regression test checks only account_id, account_label and card_last_four.

Our inference, which no test checks, is that the desk will have to ask which charge it is. Stop customers repeating themselves after a handoff reaches the same finding from the design side.

Three more limits decide what the count is worth:

LimitWhat the code doesWhat it means for the review
Presence, not correctnessThe check asks whether a field is populated. In the test the caller’s name is a placeholder string, and it counts.A plumbing check, never a team target: the cheapest way to five is to fill every field.
Partly on the agent’s pathThe test was added after an earlier suite missed the live agent’s empty details, and it still sets four values itself.Only the name and account questions retire from keys the agent writes; attempts_log parsing is untested.
Emptiness is absenceAn empty attempts list does not retire “Have you tried anything already?”The count errs towards reporting a repeat, the safe direction.
Show the source for the three limits
  • handoffpkg/desk.py, lines 208–229: _package_answers checks that a field is populated. The agent-path test’s display name is value_for_display_name.
  • skills/human_handoff/tools.py, lines 77–81: the comment recording that an earlier suite passed while the live agent shipped empty details.
  • handoffpkg/desk.py, line 206: emptiness counts as absence, so an empty attempts list retires nothing.

tests/test_handoff_context.py, lines 366–372, indentation removed, shows the four values the test sets itself. The session already holds the keys the handoff tool collects and the dispute skill writes.

session["verified_tier"] = "medium"
session["goal"] = "dispute_transaction"
session["questions_answered"] = "\n".join(DESK_OPENING_SCRIPT)
session["attempts"] = [
    {"action": "send_otp_sms", "outcome": "failed", "code": "delivery_failed"}
]

Can the agent promise callers they won’t repeat themselves?

Not unconditionally, and this section shows the two instructions that contradict each other.

What the desk may do is worked out from the transferred tier when the package is read. The demo caller is hard-coded at medium, and the permissions test asserts that account details may be discussed but irreversible changes may not.

Raising a dispute needs high (_DISPUTE_MIN_TIER in skills/dispute_transaction/tools.py, lines 58–60), and the comment there says the medium caller “is what makes the handoff happen”. So the desk screen tells the human to step the caller up before the dispute can go through. That is correct design, and it hurts a handle-time slide.

Show the source for the promise and the tier rule
  • skills/human_handoff/tools.py, lines 210–214: the hint that tells the model to promise no repeats.
  • skills/human_handoff/skill.md, lines 30–33: the skill’s closing step.
  • skills/dispute_transaction/tools.py, line 44: the demo caller’s tier, hard-coded to medium.
  • test_desk_permissions_follow_the_tier_not_the_package: the permissions test.

handoffpkg/desk.py, lines 112–121, with the leading indentation removed, is the tier rule:

tier = package.identity.verified_tier
actions = ["Answer general questions", "Explain what the agent already did"]
if tier_at_least(tier, "medium"):
    actions.append("Discuss account-specific details")
else:
    actions.append("DO NOT discuss account-specific details — identity not established")
if tier_at_least(tier, "high"):
    actions.append("Action irreversible changes (transfers, card reissue, SIM swap)")
else:
    actions.append("DO NOT action irreversible changes — step the caller up first")

How do you measure AI agent human handoff success after launch?

This section gives you the order of work for the numbers only real calls can produce.

The whole suite runs from patterns/voice-handoff-context:

Any system
make test

In the saved run on 21 September 2026 (Python 3.14.3), all 41 tests passed over synthetic sessions, with no network, model or licence. None places a call, times a human or reads a production session.

Version your desk’s opening script

Fix the questions and date the list before counting anything. Figures taken against different versions cannot be compared.

Recompute package coverage from stored packages

The tool computes the count on every transfer but does not store it. It does save each package to a file named after its handoff, so recompute the count from those files and call it package coverage. Coverage is not a repeat rate: a desk agent who asks for the date of birth out of habit, from a package that retired it, shows up only in call review.

Show the source

handoffpkg/desk.py, lines 247–250: deliver writes the whole package to JSON. skills/human_handoff/tools.py, line 189: the file is named after the handoff_id. The tool returns desk_still_needs_to_ask without storing it.

Sample transferred calls and count what the desk asked

Join each call to its package by handoff_id. Count scripted questions the desk asked although the package answered them, and separately list every question not on the script with what it asked for.

Report the three figures side by side

Put coverage, the repeat rate and the off-script list together with the sample size. Without the off-script list, the fixed script hides the loss the dispute path shows.

Take handle time and satisfaction from their owners

Handle time comes from contact-centre reporting on a comparable queue mix; satisfaction from your CSAT or complaints process, scoped to transfers.

The companion tests none of these steps. Step three depends on what your contact centre records:

Calls are recorded, not transcribed

A QA reviewer lists the populated script questions from each sampled package, listens to the desk side, and marks each as asked or not, noting every other question too. The repeat rate is script questions asked again over populated script questions in the sample, not a rate for all calls.

Calls are transcribed

Search the desk agent’s turns for each populated script question. Agents paraphrase, so a text match will miss “and your name is?” for “Can I take your name?”. Questions that match nothing on the script go into a separate list for a reviewer to label. Have QA listen to a sample the search cleared, and report how many it missed.

What goes on a launch-review slide you can defend?

Here is Aurora’s slide rewritten so that someone in the room can check every line.

Part of the slideWhat it saysEvidence or owner
Proven offlineOn the fixture dispute path the package answers all five opening-script questions; the one-line handoff answers noneThe baseline and agent-path tests at 69e27b6, output above
Limits of the proofPopulated, not correct; tier, goal and attempts set by the test; no disputed amount, merchant or date, yet five of fiveThe agent-path test and the dispute-path run above
Not yet measuredHandle time, the desk repeat rate, off-script desk questions, satisfaction with transfersAn owner, a data source and a date for each
Changes owedDispute tools write the amount, merchant and date under test, or the gap is recorded; the promise is reworded or its cost accepted; the script is versionedEngineering, conversation design and desk operations
Not on this slideRedaction. The README says the allowlist “governs session-state KEYS. It cannot police free text” and “It is not a compliance control.”A separate slide

The same suite shows the package withholding seven planted sensitive values and a field nobody anticipated. Once the owed changes are decided, that is reasonable to ship on offline evidence. Announcing a business effect before anyone has counted a call is not.

Show the source

SENSITIVE_VALUES (tests/test_handoff_context.py, lines 66–74) plants the seven values, and test_no_sensitive_value_appears_anywhere_in_the_package (lines 224–233) searches the serialised package for each. test_a_field_nobody_anticipated_is_withheld_by_default covers the new field. Both passed in the full make test run.

Design a handoff a customer can trust covers consent, the context packet and recovery when the receiving team is unavailable. The tool and memory scope tutorial builds a bank agent that makes the caller verify exactly once.