Guide · AI product manager
Measure a handoff before you claim it worked
How to measure AI agent human handoff success: count the desk questions a context package retires, and keep that apart from handle time and CSAT.
Key takeaways (5)
- At launch you can defend one number: how many questions from a fixed desk opening script the handoff package answers. Handle time, CSAT and real repeat rates need production data.
- Fix the denominator before you count. The handoff tool also returns questions_retired, which has no fixed denominator and does not belong on a slide.
- The count checks that fields are populated, not that they are right, and it cannot see facts off the script. On the companion agent it reads five out of five while the package omits which charge was disputed.
- The agent is instructed to promise callers they will not need to repeat themselves, whatever the package holds. Change the wording or record the cost in the review.
- Split the slide into what is proven offline, what is not yet measured (with an owner, a data source and a date) and the changes you owe.
Aurora Home Insurance, a fictional insurer, has shipped a redesigned agent-to-human handoff. Its voice agent now passes a structured context package to the human desk instead of a one-line reason. The launch-review slide says:
Redesigned handoff packages cut average handle time by 30%.
The only evidence behind it is a passing offline test suite. Here is what that evidence supports, and what it does not:
| What you can claim | What you cannot claim yet | What to measure instead |
|---|---|---|
| On the tested path the package answers all five desk opening questions; the old handoff answers none | Callers stopped repeating themselves | Scripted questions the desk still asked, from call review |
| Every handoff is scored against the same five questions | The package carries everything the case needs | Desk questions that are not on the script, per handoff |
| All 41 offline tests pass | Handle time fell by 30%, or callers are happier with transfers | Timed transferred calls before and after; a satisfaction measure scoped to transfers |
The artefact linked from the slide is the tail of a make test run in the
companion pattern:
----------------------------------------------------------------------
Ran 41 tests in 0.004s
OK
No call was recorded and no desk agent was timed. A search of all 25 files in the pattern finds no “30%” and no mention of handle time. The team had a real result and reported a different one.
Aurora is invented; the code is real. It is patterns/voice-handoff-context in
RasaHQ/rasa-community-resources at 69e27b6, the revision pinned for Rasa
Pro 3.20.0rc1. Its fixture call is a disputed card payment on a bank account,
which the fictional team kept for its pilot.
The position this guide takes: at launch, report the count of fixed desk questions the package retires, proven offline, and nothing larger. The cost is a smaller slide, a measurement you still owe, and a script that sees only the questions written into it.
What does the passing test prove?
This section gives you the denominator the test counts against, and the before-and-after it proves.
“Callers repeat themselves less” has no denominator. The pattern fixes one
before any caller exists: five questions a desk agent asks when nothing was
transferred, each mapped to the package field that makes it unnecessary
(handoffpkg/desk.py, lines 53–59):
_QUESTION_RETIRED_BY: dict[str, str] = {
"Can I take your name?": "identity.display_name",
"Can you confirm your date of birth?": "identity.verified_tier",
"Which account is this about?": "intent.details.account_id",
"What are you calling about today?": "intent.goal",
"Have you tried anything already?": "attempts",
}
The denominator is the same five for every handoff, whatever the caller or
case, so one handoff compares with the next and the old design with the new.
The counting function, unanswered_questions, calls itself “The measurable
form of the teaching claim” (lines 178–185).
The baseline test builds a package from one free-text handoff_reason and gets
all five questions back; the agent-path test gets none back. One offline run on 21 September 2026 passed both, along with the permissions
test this guide uses later.
Show the command and output for the three tests
From patterns/voice-handoff-context at 69e27b6, on Python 3.14.3 with no
network and no model:
python3 -m unittest -v \
tests.test_handoff_context.TestPackageSurvivesHandoff.test_the_catalogs_current_handoff_answers_nothing \
tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question \
tests.test_handoff_context.TestPackageSurvivesHandoff.test_desk_permissions_follow_the_tier_not_the_packagetest_the_catalogs_current_handoff_answers_nothing (tests.test_handoff_context.TestPackageSurvivesHandoff.test_the_catalogs_current_handoff_answers_nothing)
Baseline: one free-text reason retires none of the desk's questions. ... ok
test_the_agent_path_retires_every_desk_question (tests.test_handoff_context.TestTheAgentPathSpecifically.test_the_agent_path_retires_every_desk_question)
THE HEADLINE CLAIM, ON THE PATH THE AGENT ACTUALLY RUNS. ... ok
test_desk_permissions_follow_the_tier_not_the_package (tests.test_handoff_context.TestPackageSurvivesHandoff.test_desk_permissions_follow_the_tier_not_the_package)
A medium-tier caller must not be actioned for an irreversible change. ... ok
----------------------------------------------------------------------
Ran 3 tests in 0.001s
OKThe five questions are the pattern author’s, written for a fixture desk. Take yours from your desk’s call guide or QA scorecard, version the list, and date every figure against a version. If the script changes, the number changes, though no caller behaved differently.
Our desk's script has eight questions. Do we just extend the list?
Extend both structures together: unanswered_questions looks each question up
with _QUESTION_RETIRED_BY[question] (line 188), so a question without a
mapping raises a KeyError on every transfer. A question mapped to a field
_package_answers does not recognise falls through to return False (line 229) and counts as unanswered on every handoff. Add a test for each new
question, and restate the baseline before comparing figures.
Which count belongs on the slide?
The handoff tool returns two counts in the same response, and only one of them has a fixed denominator.
questions_retired | desk_still_needs_to_ask | |
|---|---|---|
| What it counts | Lines the skills appended to questions_answered | Script questions whose package field is empty |
| Denominator | None fixed | The five-question opening script |
| On the dispute path | 4 | [], so nothing left to ask |
| Source | skills/human_handoff/tools.py, line 205 | The same file, line 208 |
| Use it for | Nothing on a slide | Package coverage, tracked as a regression check |
They disagree because the dispute skill appends four script questions
(skills/dispute_transaction/tools.py, lines 98–103 and 135) and never “Have
you tried anything already?”. The package still carries three attempts, so the
scripted check counts that question as answered. Pick one definition and write
it into the review.
An offline run of the real dispute and handoff tools printed both counts side by side.
Show the output from the real tools on the dispute path
editorial/receipts/measure-a-handoff-that-loses-context/agent-path-repro.py
in this site’s repository calls load_caller_context, list_recent_charges,
raise_dispute for charge txn_c101, then transfer_to_human, with a
memory-only stand-in context. It ran on an unmodified copy of the pattern with
Rasa Pro 3.20.0rc1 installed, because the tools import rasa.mantle. The run
on 21 September 2026, on Python 3.12.13, printed, in part:
dispute_* keys in memory: ['disputed_txn_id', 'disputed_txn_label']
transferred: {"attempts_carried": 3, "goal": "dispute_transaction", "identity": true, "questions_retired": 4, "verified_tier": "medium"}
desk_still_needs_to_ask: []
hint: Context package delivered. Tell the caller the handoff id and that they will not need to repeat themselves. Do not read back any withheld field.
intent.details: {"account_id": "acc_checking", "account_label": "Everyday Checking", "card_last_four": "4821"}
ASKING FOR line: ASKING FOR Dispute a card transaction [account_id=acc_checking, account_label=Everyday Checking, card_last_four=4821] (stage: blocked)Where does the handoff lose the disputed charge?
This section shows the one fact the companion agent drops, and why the count cannot notice.
The handoff tool collects the disputed amount, merchant and date, but nothing on the dispute path writes them. The regression test checks three other keys, so the package names a card dispute on Everyday Checking but not which charge, and the count says the desk needs to ask nothing.
Show the source for this finding
skills/human_handoff/tools.py, lines 85–87: the handoff tool collectsdispute_amount,dispute_merchantanddispute_date.handoffpkg/redaction.py, lines 92–94 and 222–224: the allowlist passes those three keys intointent.details.skills/dispute_transaction/tools.py, line 170:raise_disputewritesdisputed_txn_label. No dispute tool writes the three keys above.skills/human_handoff/tools.py, lines 68–96: the handoff tool’s key list names neitherdisputed_txn_labelnordisputed_txn_id.tests/test_handoff_context.py, lines 360–363: the regression test checks onlyaccount_id,account_labelandcard_last_four.
Our inference, which no test checks, is that the desk will have to ask which charge it is. Stop customers repeating themselves after a handoff reaches the same finding from the design side.
Three more limits decide what the count is worth:
| Limit | What the code does | What it means for the review |
|---|---|---|
| Presence, not correctness | The check asks whether a field is populated. In the test the caller’s name is a placeholder string, and it counts. | A plumbing check, never a team target: the cheapest way to five is to fill every field. |
| Partly on the agent’s path | The test was added after an earlier suite missed the live agent’s empty details, and it still sets four values itself. | Only the name and account questions retire from keys the agent writes; attempts_log parsing is untested. |
| Emptiness is absence | An empty attempts list does not retire “Have you tried anything already?” | The count errs towards reporting a repeat, the safe direction. |
Show the source for the three limits
handoffpkg/desk.py, lines 208–229:_package_answerschecks that a field is populated. The agent-path test’s display name isvalue_for_display_name.skills/human_handoff/tools.py, lines 77–81: the comment recording that an earlier suite passed while the live agent shipped empty details.handoffpkg/desk.py, line 206: emptiness counts as absence, so an emptyattemptslist retires nothing.
tests/test_handoff_context.py, lines 366–372, indentation removed, shows the
four values the test sets itself. The session already holds the keys the
handoff tool collects and the dispute skill writes.
session["verified_tier"] = "medium"
session["goal"] = "dispute_transaction"
session["questions_answered"] = "\n".join(DESK_OPENING_SCRIPT)
session["attempts"] = [
{"action": "send_otp_sms", "outcome": "failed", "code": "delivery_failed"}
]Can the agent promise callers they won’t repeat themselves?
Not unconditionally, and this section shows the two instructions that contradict each other.
What the desk may do is worked out from the transferred tier when the package
is read. The demo caller is hard-coded at medium, and the permissions test
asserts that account details may be discussed but irreversible changes may
not.
Raising a dispute needs high (_DISPUTE_MIN_TIER in
skills/dispute_transaction/tools.py, lines 58–60), and the
comment there says the medium caller “is what makes the handoff happen”. So the
desk screen tells the human to step the caller up before the dispute can go
through. That is correct design, and it hurts a handle-time slide.
Show the source for the promise and the tier rule
skills/human_handoff/tools.py, lines 210–214: the hint that tells the model to promise no repeats.skills/human_handoff/skill.md, lines 30–33: the skill’s closing step.skills/dispute_transaction/tools.py, line 44: the demo caller’s tier, hard-coded tomedium.test_desk_permissions_follow_the_tier_not_the_package: the permissions test.
handoffpkg/desk.py, lines 112–121, with the leading indentation removed, is
the tier rule:
tier = package.identity.verified_tier
actions = ["Answer general questions", "Explain what the agent already did"]
if tier_at_least(tier, "medium"):
actions.append("Discuss account-specific details")
else:
actions.append("DO NOT discuss account-specific details — identity not established")
if tier_at_least(tier, "high"):
actions.append("Action irreversible changes (transfers, card reissue, SIM swap)")
else:
actions.append("DO NOT action irreversible changes — step the caller up first")How do you measure AI agent human handoff success after launch?
This section gives you the order of work for the numbers only real calls can produce.
The whole suite runs from patterns/voice-handoff-context:
- Any system
make test
In the saved run on 21 September 2026 (Python 3.14.3), all 41 tests passed over synthetic sessions, with no network, model or licence. None places a call, times a human or reads a production session.
Version your desk’s opening script
Fix the questions and date the list before counting anything. Figures taken against different versions cannot be compared.
Recompute package coverage from stored packages
The tool computes the count on every transfer but does not store it. It does save each package to a file named after its handoff, so recompute the count from those files and call it package coverage. Coverage is not a repeat rate: a desk agent who asks for the date of birth out of habit, from a package that retired it, shows up only in call review.
Show the source
handoffpkg/desk.py, lines 247–250: deliver writes the whole package to
JSON. skills/human_handoff/tools.py, line 189: the file is named after the
handoff_id. The tool returns desk_still_needs_to_ask without storing it.
Sample transferred calls and count what the desk asked
Join each call to its package by handoff_id. Count scripted questions the desk
asked although the package answered them, and separately list every question
not on the script with what it asked for.
Report the three figures side by side
Put coverage, the repeat rate and the off-script list together with the sample size. Without the off-script list, the fixed script hides the loss the dispute path shows.
Take handle time and satisfaction from their owners
Handle time comes from contact-centre reporting on a comparable queue mix; satisfaction from your CSAT or complaints process, scoped to transfers.
The companion tests none of these steps. Step three depends on what your contact centre records:
Calls are recorded, not transcribed
A QA reviewer lists the populated script questions from each sampled package, listens to the desk side, and marks each as asked or not, noting every other question too. The repeat rate is script questions asked again over populated script questions in the sample, not a rate for all calls.
Calls are transcribed
Search the desk agent’s turns for each populated script question. Agents paraphrase, so a text match will miss “and your name is?” for “Can I take your name?”. Questions that match nothing on the script go into a separate list for a reviewer to label. Have QA listen to a sample the search cleared, and report how many it missed.
What goes on a launch-review slide you can defend?
Here is Aurora’s slide rewritten so that someone in the room can check every line.
| Part of the slide | What it says | Evidence or owner |
|---|---|---|
| Proven offline | On the fixture dispute path the package answers all five opening-script questions; the one-line handoff answers none | The baseline and agent-path tests at 69e27b6, output above |
| Limits of the proof | Populated, not correct; tier, goal and attempts set by the test; no disputed amount, merchant or date, yet five of five | The agent-path test and the dispute-path run above |
| Not yet measured | Handle time, the desk repeat rate, off-script desk questions, satisfaction with transfers | An owner, a data source and a date for each |
| Changes owed | Dispute tools write the amount, merchant and date under test, or the gap is recorded; the promise is reworded or its cost accepted; the script is versioned | Engineering, conversation design and desk operations |
| Not on this slide | Redaction. The README says the allowlist “governs session-state KEYS. It cannot police free text” and “It is not a compliance control.” | A separate slide |
The same suite shows the package withholding seven planted sensitive values and a field nobody anticipated. Once the owed changes are decided, that is reasonable to ship on offline evidence. Announcing a business effect before anyone has counted a call is not.
Show the source
SENSITIVE_VALUES (tests/test_handoff_context.py, lines 66–74) plants the
seven values, and test_no_sensitive_value_appears_anywhere_in_the_package
(lines 224–233) searches the serialised package for each.
test_a_field_nobody_anticipated_is_withheld_by_default covers the new field.
Both passed in the full make test run.
Design a handoff a customer can trust covers consent, the context packet and recovery when the receiving team is unavailable. The tool and memory scope tutorial builds a bank agent that makes the caller verify exactly once.