Skip to content
RasaGet a free licence
Guides for AI teams

Guide · Platform / operations engineer

How to measure voice agent latency to the answer

If your voice agent’s latency looks good but callers still wait, time each turn to the message that answers them, not to the first sound.

by Rod Rivera

About 8 minutes

  • 1.98 s

    median time from the end of the caller’s speech to the first sound

  • 8.4 s

    median time from the caller’s message to the answer

  • 2.80 s

    median answer time on 8 repeated calls with lower reasoning, down from 7.92 s

Source: Scripted test calls to a sample Rasa voice agent, 30 September 2026
Key takeaways (3)
  • If your agent speaks a holding line first, time to first audio measures that line, not the answer.
  • Time each turn from the caller’s message to the first agent message that is not a holding line.
  • Test the model’s reasoning setting against that clock on the same scripted calls, and check the calls still pass.

Your voice agent’s latency dashboard looks healthy. Time to first audio sits under two seconds. Yet callers say the agent is slow.

Here is one turn from a scripted test call to a sample order-status agent. The agent sent a holding line after a second and a half. The answer came seven and a half seconds after that.

Voice callRecorded test call normal-delivered-10482, 30 September 2026, trimmed: the skill-start call is left out. Times are seconds after the caller’s message reached Rasa.
  1. Caller0.0 s

    Hi. Can you tell me where my order is? It’s one zero four eight two

  2. Agent1.5 s

    Alright, let me pull up that order for you.

    Annotation:

    A holding line, recorded at 1.5 s. Its audio started 1.66 s after the caller stopped speaking. Both usual latency figures stop here.

  3. Tool3.7 s

    track_order(order_number=“WS-10482”) → answered

  4. Agent9.0 s

    Order one zero four eight two was delivered on Monday at 2:12 PM, observed by a Larkspur Parcel carrier scan. Your status reference is WS-ST-20260930-B75EF3B9.

    Annotation:

    The answer. Neither latency figure stops here.

The first message is a holding line: a short “I’m on it” message the agent sends while a tool runs. Your first-audio figure stops at that line, not at the answer.

The fix is to time each turn from the caller’s message to the first agent message that is not a holding line. Rasa’s agent runtime already records both timestamps, so you need no new instrumentation.

This matters because the two numbers can be far apart. Across the sample’s 19 test calls, the median time to first audio was 1.98 s. The median time to the answer was 8.4 s.

Once you time the answer, you can tune for it. On the same 8 scripted calls, one line of model configuration cut the median wait for the answer from 7.92 s to 2.80 s.

What you need:

  • The sample is the companion project examples/mantle-voice-retail-order-status-gemini at commit 4aa0c44. It is a Rasa voice agent for Willow Shop, a fictional retailer. It pins rasa-pro 3.21.0.dev5 and runs on gemini-3.8-flash. Deepgram Flux turns the caller’s speech into text, and Rime Mist v3 speaks the replies.
  • The Rasa behaviour described here is from Rasa Pro 3.21.0.dev5, the version the sample pins. That covers the holding-line tag, the order of text and tools, and where the model settings live. Check it on your version.
  • Reading the stored results needs only Python 3. The sample commits the trackers from its test runs.
  • Rerunning the calls is billed and needs uv, which the test harness uses to start Rasa. In the project folder, make install installs the project and make env copies .env.example to .env. Fill in RASA_LICENSE, GEMINI_API_KEY, DEEPGRAM_API_KEY and RIME_API_KEY there.

Why first audio stops at the holding line

In Rasa’s agent runtime, the model can return some text and a tool call in the same response. The runtime sends that text first, then runs the tool. This is one comment line from Rasa’s orchestrator code (rasa/mantle/orchestration/orchestrator.py), quoted as written:

# Text + tools (no hangup): emit text as a filler before the tools run.

On a call this is useful. The caller hears that the agent is working while the tool runs. But it also makes the first sound of the turn something the model said before it had the data. In the sample’s main test run, 24 of 30 caller messages got more than one agent message, and the first was a holding line every time.

So look at what each of the usual figures measures:

  • End of speech to first audio: the test harness’s own clock. It runs from the end of the caller’s speech to the first agent audio with sound in it.
  • user_perceived_latency_ms: the figure Rasa writes at the end of each turn. Its code documents it as the time from the turn’s input to its first output. On voice, the first spoken audio counts as that output.
  • Caller’s message to the answer: read from the conversation record. It stops at the answer.

The first two did what their definitions say. They timed the first output, and the first output was the holding line. The drawing below shows one turn to scale:

caller's message 0.0 s 0 s holding line sent 1.5 s track_order answered 3.7 s answer message 9.0 s 9 s 3 s 6 s
  1. The answer clock starts when the caller’s message reaches Rasa.
  2. First audio and user_perceived_latency_ms stop at the holding line.
  3. The answer clock stops at the answer message.
FigureOne turn of call normal-delivered-10482, drawn to scale

Where the rest of the wait goes

A turn that needs a tool is a chain of model calls, one after another. Each call reasons before it writes anything. The run’s usage log records every model call, so you can see the chain for the turn above:

Model callWhat it producedTimeReasoning tokens of total
1the holding line and the call to start work1.46 s176 of 205
2the track_order call2.12 s282 of 306
3the answer5.23 s1,119 of 1,171

The three model calls take almost the whole wait, and the order lookup barely shows. The call that writes the answer took 5.23 s of the 9.0 s, so that is the call to make faster.

Most of that time went on reasoning. Across the main run, about 95% of the model’s output tokens were reasoning. That is why the reasoning setting is the first thing to test.

How to time the answer from the tracker

Rasa records each conversation as a tracker. A tracker is the list of events in the conversation, in order, each with a timestamp. It includes every caller message, as a user event, and every agent message, as a bot event.

The sample’s make run starts Rasa with rasa run --enable-api on port 5005. With the API on, you can fetch any tracker over HTTP, as the sample’s test harness does. The path takes the conversation ID, which Rasa calls the sender ID. This one is from the call above. Replace it with one of your own:

CONVERSATION_ID=retail-order-status-gemini-voice-normal-delivered-10482-20260930T033956

Then save the tracker to a file:

Any system
curl -o tracker.json http://localhost:5005/conversations/$CONVERSATION_ID/tracker

Without the error rule, a fast apology counts as a fast answer. The sample’s run with a setting the model rejected shows this. Every caller turn there got the apology within 0.3 s, so the answer clock would have looked better than on any working run.

With the rule, all three of its caller turns count as errors. Chart the error count next to the answer clock.

The sample never transfers the call to a person. Its delivery-help tool opens a request with a reference, and the agent keeps talking. If yours does, find what your transfer message carries in its metadata, and set those turns aside in the same way.

The reusable part

These steps are one function. Save it as answer_clock.py in your own project. It is ours, not part of the companion repository:

ERRORS = {"utter_model_call_error", "utter_model_request_hook_error", "utter_tool_call_hook_error"}


def answer_ms(events):
    """Time from each caller message to its answer, in ms.

    The answer is the first bot message after the caller's message that is
    not a holding line (filler) and not an error apology. A turn whose first
    such message is an error apology is counted in `errors`, not timed.
    """
    waits, errors = [], 0
    for i, e in enumerate(events):
        if e["event"] != "user" or (e.get("text") or "").startswith("/"):
            continue
        for later in events[i + 1:]:
            if later["event"] == "user":
                break
            if later["event"] != "bot":
                continue
            meta = later.get("metadata") or {}
            if meta.get("utter_action") in ERRORS:
                errors += 1
                break
            if meta.get("mantle_response_source") != "filler":
                waits.append((later["timestamp"] - e["timestamp"]) * 1000)
                break
    return waits, errors

The three names in ERRORS are the default responses in Rasa’s agent runtime that speak the “something went wrong” apology.

Then feed it any saved tracker, such as the tracker.json from the curl command above:

import json

from answer_clock import answer_ms

events = json.load(open("tracker.json"))["events"]
waits, errors = answer_ms(events)
print([round(w / 1000, 2) for w in waits], "errors:", errors)

Put this in whatever job already reads your trackers: a nightly test run, a monitoring script or your analytics export. You do not need to change the agent.

Without Rasa

The method does not depend on Rasa. If your agent sends a holding line, mark it as one on the message when you send it. Mark error apologies too. Then keep two timestamps per turn: when the caller’s message arrived, and when the first message that is neither went out.

Show a script that prints all three clocks for a test run

Save this as three_clocks.py, next to answer_clock.py. It is ours, not part of the companion repository. It reads the harness’s results.json for the first two clocks and the stored trackers for the answer clock. Run it from examples/mantle-voice-retail-order-status-gemini/case-build/results/ at commit 4aa0c44, with the folder that holds both files on PYTHONPATH.

import json, sys
from pathlib import Path

from answer_clock import answer_ms


def p(values, pct):  # nearest rank, as the harness computes it
    s = sorted(values)
    return s[max(1, -(-pct * len(s) // 100)) - 1] / 1000


def calls(run):
    return json.loads((run / "results.json").read_text())["conversations"]

# Optional second run: keep only the calls that run placed.
run = Path(sys.argv[1])
only = {c["id"] for c in calls(Path(sys.argv[2]))} if len(sys.argv) > 2 else None
first_audio, perceived, answer, errors = [], [], [], 0
for conv in calls(run):
    if only and conv["id"] not in only:
        continue
    for turn in conv["turns"]:
        first_audio.append(turn["extra"]["eos_to_first_audible_ms"])
        first = [b for b in (turn["extra"]["latency_breakdown"] or [])[:1] if b]
        perceived += [b["user_perceived_latency_ms"] for b in first]
    tracker = json.loads((run / "trackers" / f"{conv['id']}.json").read_text())
    waits, errs = answer_ms(tracker["events"])
    answer += waits
    errors += errs

for name, v in [("end of speech -> first audio", first_audio),
                ("user_perceived_latency_ms", perceived),
                ("user event -> answer message", answer)]:
    print(f"{name:30} p50 {p(v, 50):5.2f} s  p95 {p(v, 95):5.2f} s  n={len(v)}")
print(f"turns that ended in an error apology: {errors}")

Its output over the main run:

$ python3 three_clocks.py 2026-09-30-gemini-3.8-flash-default
end of speech -> first audio   p50  1.98 s  p95  7.74 s  n=27
user_perceived_latency_ms      p50  1.98 s  p95  3.60 s  n=25
user event -> answer message   p50  8.29 s  p95 15.15 s  n=30
turns that ended in an error apology: 0

With 30 values, the middle falls between the 15th (8.29 s) and the 16th (8.43 s). This script reports the 15th and the companion README reports the 16th. Both round to 8.4 s.

How to test the reasoning setting against the answer clock

Change one setting, replay the same scripted calls, and compare the answer clock and the outcomes. In the sample, the setting to change was reasoning_effort. It tells the model how much to reason before it replies.

The sample’s model group sets no reasoning option, so the model runs at its provider’s default. A variant of the build adds one line. In your own Rasa project, the model group is in integrations.yml:

integrations.yml, model group: main run and the variant

Main run

model_groups:
  - id: orchestrator
    models:
      - provider: gemini
        model: gemini-3.8-flash
        api_key: ${GEMINI_API_KEY}

No reasoning setting is sent.

Variant with low reasoning

model_groups:
  - id: orchestrator
    models:
      - provider: gemini
        model: gemini-3.8-flash
        api_key: ${GEMINI_API_KEY}
        reasoning_effort: low

The only change is the last line.

The variant placed 8 of the 19 test calls again. These are the results, from the companion README:

Same 8 callsProvider defaultreasoning_effort: low
Calls that passed the checks8 of 88 of 8
Caller’s message to the answer, median7.92 s2.80 s
Caller’s message to the answer, 95th pct15.79 s7.68 s
End of speech to first audio, median1.98 s1.35 s
Reasoning tokens / all output tokens29,023 / 30,7515,915 / 7,638
Model cost for the 8 calls0.25 USD0.14 USD

The first-audio median moved by less than a second. On the same 8 calls, the answer median moved by about five seconds. A dashboard that charted only first audio would have hidden most of the gain.

The script in the earlier “Show a script” panel prints the same clocks for both columns. Pass a second run folder to keep only the calls that run placed. Its output, from the case-build/results/ folder:

Provider default, same 8 calls
$ python3 three_clocks.py 2026-09-30-gemini-3.8-flash-default 2026-09-30-gemini-3.8-flash-thinking-low
end of speech -> first audio   p50  1.98 s  p95  9.66 s  n=13
user_perceived_latency_ms      p50  1.98 s  p95  7.74 s  n=12
user event -> answer message   p50  7.92 s  p95 15.79 s  n=15
turns that ended in an error apology: 0
reasoning_effort: low
$ python3 three_clocks.py 2026-09-30-gemini-3.8-flash-thinking-low
end of speech -> first audio   p50  1.35 s  p95  2.26 s  n=13
user_perceived_latency_ms      p50  1.20 s  p95  2.23 s  n=13
user event -> answer message   p50  2.80 s  p95  7.68 s  n=15
turns that ended in an error apology: 0

The checks read which tools ran and what they returned, not the wording. At low, the agent gave the same kinds of answers as before. One reply was worded differently: the agent once read a status reference character by character.

The same line can slow a different model

When your model group sets no reasoning_effort, Rasa checks whether the model accepts none. If it does, Rasa sends none. For the Gemini model in this sample it sent nothing, so the provider’s default applied.

A text build in the companion, examples/mantle-text-retail-guided-selling-gpt, runs on GPT-5.5. There, Rasa’s default was already none, so adding reasoning_effort: low raised the reasoning. Turn time here is from sending the request to receiving the full reply:

Same 22 conversationslownone (Rasa’s default)
Passed21 of 2220 of 22
Turn time, median11.05 s7.48 s

Lower reasoning was faster on both builds. What differed is where the default sat.

The GPT build kept low anyway. At none, one conversation ended on a promise to check stock with no tool call after it. Another got a generic “can’t help” reply to an in-scope request.

How to check the fix works

After any change to the model or its settings:

  1. Place one live call and fetch its tracker. Confirm the agent answered, with no error apology.
  2. Replay the same scripted calls you used before the change. Compare the answer clock’s median and 95th percentile, not only first audio.
  3. Confirm the same calls still pass your checks.
Show how to rerun the sample's comparison

These are live, billed calls to Gemini, Deepgram and Rime. The harness starts Rasa through uv, so install uv first.

Then, in examples/mantle-voice-retail-order-status-gemini/, run make caller-audio. Most of the recorded caller audio files are not in git. This target renders the missing ones with Gemini’s text-to-speech, which is also billed. The new files say the same words, but their bytes differ.

Then, from the root of the companion repository, with the keys in the project’s .env:

python3 scripts/case_builds/run_build.py examples/mantle-voice-retail-order-status-gemini \
  --budget-usd 4 --variant thinking-low \
  --only normal-item-not-number,normal-split-order-both-parcels,normal-label-then-delay-help,adversarial-label-means-delivered,adversarial-stale-best-guess,recovery-stale-to-delivery-help,correction-second-parcel,correction-drop-stale-switch-parcel

--budget-usd caps the build’s total recorded spend. --variant picks the configuration change, and --only lists the calls to place. Point three_clocks.py at the new results folder.

Trade-offs

The answer clock is longer and noisier: it includes tool time, which you may not control. Expect a wider spread than first audio shows.

It reads server timestamps, not the caller’s ear: it stops when the answer message is recorded. Text-to-speech starts after that, so the caller hears the answer a little later.

Keep the holding line if it helps: it may make the wait feel shorter than silence would. The sample did not test how callers felt. Chart both clocks: first audio for the holding line, the answer clock for the answer.

Limits

  • All figures come from one sample agent with scripted synthetic callers, on one day, from one laptop. They are not production latencies.
  • Each column of the reasoning comparison is one run of 8 calls. The default column reuses those calls from the main run.
  • The answer clock starts at a caller message. If speech-to-text never turns a caller’s speech into a message, the clock cannot see that turn. The Deepgram Flux guide shows how to check for that.