Skip to content
RasaGet a free licence
Guides for AI teams

Guide · AI product engineer

How to join Speechmatics segments into one Rasa turn

If your Rasa voice agent asks again for details the caller just gave, check whether your speech-to-text engine sends each phrase as its own turn.

by Rod Rivera

About 10 minutes

  • 19

    caller messages Rasa received from one spoken opening, before the fix

  • 36 → 10

    model calls for the same two-turn call, before → after the fix (two separate runs)

  • 43 of 43

    spoken caller turns that arrived as one message each, over 20 test calls after the fix

Source: Scripted test calls to a sample Rasa voice agent for card disputes, 29 September 2026
Key takeaways (3)
  • In our test calls, Speechmatics finalised speech a word or two at a time, and Rasa took each transcript as a turn.
  • Hold the phrases in your engine and send one transcript when Speechmatics marks the end of the utterance.
  • After each test call, count the caller messages in the conversation record against the turns spoken.

Your Rasa voice agent asks the caller for their name and date of birth. The caller gives both in one breath.

The agent thanks them by first name, then asks for the full name again. Then it asks for the date of birth, and then again.

Here is that call, from a scripted test of a sample bank agent that takes card disputes. The caller opened with one recording of about 14 seconds: “Hello, this is Priya Raghunathan, date of birth seventeenth August, nineteen ninety-one. I don’t recognise a charge of two thousand four hundred and ninety-nine rupees from Brightmart Online on the twenty-second of September.”

Rasa did not receive that as one message. This is what its conversation record holds:

Voice callRecorded test call, 29 September 2026, trimmed: the first eight of the 19 caller messages from one spoken opening (bold emphasis added)
  1. Caller

    Hello,

  2. Agent

    Hello. Could you please tell me your full name and date of birth, and which charge you don’t recognise?

  3. Caller

    this is

  4. Tool

    listen() → listen: done

    Annotation:

    listen is a tool Rasa’s agent runtime offers the model: wait for more speech without replying.

  5. Caller

    Priya

  6. Agent

    Thanks, Priya. Could I have your full name and your date of birth, please?

    Annotation:

    The model answers a one-word message. It has not seen the rest.

  7. Caller

    Raghunathan.

  8. Agent

    Thank you, Priya Raghunathan. What is your date of birth, please?

  9. Caller

    Date

  10. Tool

    listen() → listen: done

  11. Caller

    of

  12. Tool

    listen() → listen: done

  13. Caller

    birth?

  14. Agent

    It sounds like you’re repeating my question. Could you tell me your date of birth, with the day, month and year?

    Annotation:

    “Date”, “of”, “birth?” are the caller’s own words. Split up, they read as an echo.

  15. Caller

    17th August 1991.

  16. Agent

    Okay, thanks for that. Let me get the dispute process started.

Your speech-to-text engine is the class that turns a vendor’s messages into transcripts for Rasa. In this call, the engine sent Rasa each short phrase as a finished transcript. Rasa took each one as a caller turn, and the model replied to each fragment on its own.

The fix is in the engine: ask Speechmatics to mark the end of each utterance, hold the phrases until then, and send them to Rasa as one transcript.

That matters if you pay per model call, or if your callers lose patience when asked twice. In the sample, the same two-turn call made 36 model calls before the fix and 10 after. Those are two separate test runs, not a controlled pair.

The engine that split the call was ours: the Speechmatics adapter in the companion repository. The fix below is in the same file.

What you need:

  • The engine is SpeechmaticsASR in patterns/voice-vendor-router, at commit 4aa0c44.
  • The sample agent is examples/mantle-voice-banking-dispute-claude. It runs on claude-sonnet-5-5 and Rasa Pro 3.21.0.dev5, and speaks with Rime. Northgate Bank and the caller are invented. The caller’s voice is generated speech read from a script.
  • The Rasa behaviour described here is from Rasa Pro 3.21.0.dev5, the version the sample pins. Check it on your version.
  • The offline tests need uv and no keys.

Why one sentence becomes many turns

Speechmatics sends text while the caller is still talking: partial results first, then finals. A final is Speechmatics’ finished text for a stretch of speech. In these test calls, with max_delay set to 1.0 seconds, a final was often only a word or two. So it is a phrase, not a whole turn.

The engine as shipped sent every final to Rasa as a NewTranscript, the event that carries what the caller said. Rasa’s voice channel turns each non-empty transcript into a caller message. This excerpt is from rasa/core/channels/voice_stream/voice_channel.py in the 3.21.0.dev5 wheel. It stops at the next branch:

if isinstance(asr_event, NewTranscript) and asr_event.text:
    logger.debug(
        "VoiceInputChannel.handle_asr_event.new_transcript",
        transcript=asr_event.text,
    )
    # Open the Rasa-processing window for the first bot message.
    self._mark_rasa_processing_started()

    event: VoiceInputEvent = FinalTranscriptInputEvent(
        text=asr_event.text,
        metadata=self.build_turn_metadata(call_parameters),
    )
    await input_queue.put(event)
elif isinstance(asr_event, UserIsSpeaking):

Rasa does join some transcripts. It merges final transcripts that are already waiting in its input queue together. Finals that arrive separately become separate turns.

We did not trace which of the 19 phrases arrived together. The record shows the outcome.

Your settings may change how long a final runs. The second example engine’s README links fragment size to max_delay. At 1.0 seconds, it saw one spoken sentence come back as four finals.

Speechmatics can say where an utterance ends. With end-of-utterance detection on, it sends an extra EndOfUtterance message after a set time of silence (Speechmatics documentation, read 2 October 2026). Nothing in the Rasa 3.21.0.dev5 wheel reads that message, so joining the phrases is the engine’s job.

Speechmatics finals "Hello,"  "this is"  "Priya" engine as shipped one transcript per final engine with the join holds the finals 3 transcripts to Rasa 1 transcript to Rasa "Hello, this is Priya" EndOfUtterance after 0.7 s of silence
  1. Speechmatics sends finished text a word or two at a time.
  2. The engine as shipped sent each final to Rasa as its own transcript.
  3. With the join, the engine holds the finals until EndOfUtterance, then sends one transcript.
FigureWhere the phrases are joined

How to fix it

In your own project the change goes in two places: your engine class, and the asr section of your voice channel’s configuration. In the sample that configuration is in integrations.yml.

Ask Speechmatics to mark the end of each utterance

Speechmatics takes the request as end_of_utterance_silence_trigger inside conversation_config, in the StartRecognition message your engine sends first. The value is the seconds of silence to wait, between 0 and 2. This excerpt from voicerouter/providers/speechmatics.py adds it when the setting is on:

if self._utterance_mode:
    transcription_config["conversation_config"] = {
        "end_of_utterance_silence_trigger": float(
            self.config.end_of_utterance_silence_trigger
        ),
    }

Hold the finals and send one transcript

This is the part of the engine that translates Speechmatics messages. It is an excerpt from the same file, shown without its method indentation:

patterns/voice-vendor-router/voicerouter/providers/speechmatics.py, excerpt
if kind == "AddTranscript":
    text = message.get("metadata", {}).get("transcript", "").strip()
    if not text:
        return None
    if not self._utterance_mode:
        return NewTranscript(text)
    # Hold the segment until Speechmatics says the utterance ended.
    self._segments.append(text)
    return UserIsSpeaking(" ".join(self._segments))
if kind == "EndOfUtterance":
    if not self._segments:
        return None
    text, self._segments = " ".join(self._segments), []
    return NewTranscript(text)
  1. With the setting off, each final still goes out as a transcript. The setting’s comment says “Unset keeps the original behaviour”, so existing configurations do not change.
  2. With it on, the final is held. The text so far goes out as UserIsSpeaking, which Rasa turns into a partial transcript, not a turn.
  3. An EndOfUtterance with nothing held sends nothing.
  4. The held phrases are joined with a space and sent as one transcript.

Turn it on in your configuration

The sample names the engine by its Python path and sets the trigger to 0.7 seconds. This excerpt is the asr section from its integrations.yml. It is shown without its leading indentation and its comment lines. In the file, it sits under channels, in the browser_audio channel:

asr:
  name: voicerouter.providers.speechmatics.SpeechmaticsASR
  language_map:
    en:
      language: en
  operating_point: enhanced
  max_delay: 1.0
  enable_partials: true
  end_of_utterance_silence_trigger: 0.7

One of those comments says of 0.7: “it was chosen, not tuned.”

How to check the fix works

Check it in three ways: an offline test of the engine, a live call, and a count over your conversation records.

Do the count over the whole conversation record, after the call. The sample’s own test harness counted messages per turn, and it undercounted. For the caller’s reply “Yes, please confirm it. I never bought anything from them.”, it recorded two caller messages. The record holds eight.

Test the engine offline

Feed the engine the messages Speechmatics would send, and count the transcripts that come out. The companion has two tests for this in tests/test_voicerouter.py. With the trigger off, three finals give three transcripts. With the trigger on, this excerpt shows what happens:

held = [engine.engine_event_to_asr_event(self._msg("AddTranscript", t)) for t in ("Hello,", "this is", "Priya.")]
self.assertTrue(all(isinstance(e, UserIsSpeaking) for e in held))
done = engine.engine_event_to_asr_event(self._msg("EndOfUtterance"))
self.assertIsInstance(done, NewTranscript)
self.assertEqual(done.text, "Hello, this is Priya.")
# A second end of utterance with nothing new said sends nothing.
self.assertIsNone(engine.engine_event_to_asr_event(self._msg("EndOfUtterance")))
# A bare "Yes." is one segment and still becomes a transcript.
engine.engine_event_to_asr_event(self._msg("AddTranscript", "Yes."))
self.assertEqual(engine.engine_event_to_asr_event(self._msg("EndOfUtterance")).text, "Yes.")

Run both from patterns/voice-vendor-router. No socket opens and no audio is sent:

Any system
uv run --frozen python -m unittest -v tests.test_voicerouter.TestSpeechmaticsSocket.test_default_mode_makes_every_final_segment_a_transcript tests.test_voicerouter.TestSpeechmaticsSocket.test_utterance_mode_joins_segments_until_end_of_utterance

This output is trimmed: the log text after each test name is cut.

test_default_mode_makes_every_final_segment_a_transcript (tests.test_voicerouter.TestSpeechmaticsSocket.test_default_mode_makes_every_final_segment_a_transcript) ...
ok
test_utterance_mode_joins_segments_until_end_of_utterance (tests.test_voicerouter.TestSpeechmaticsSocket.test_utterance_mode_joins_segments_until_end_of_utterance) ...
ok

----------------------------------------------------------------------
Ran 2 tests in 0.007s

OK

Place a call

With the join on, the same scripted opening reached Rasa as one message:

Voice callRecorded test call with the join, 29 September 2026, trimmed: the same opening as one caller message
  1. Caller

    Hello, this is Priya Raghunathan . Date of birth? 17th August 1991. I don’t recognize a charge of ₹2,499 from Bright Mart online on the 22nd of September.

  2. Agent

    Right, let me start the dispute process for you.

  3. Tool

    verify_caller(full_name=“Priya Raghunathan”, date_of_birth=“1991-08-17”) → verified

The tracker shows a space before the full stop after the surname. The name still reached verify_caller as “Priya Raghunathan”.

The two runs of the same call compare like this. They are separate runs, started four minutes apart:

Same two-turn callBefore the fixWith the join
Caller messages from the opening191
Caller messages from the reply81
Model calls3610
Model spend0.32 USD0.08 USD
Show all 19 caller messages from the call before the fix

This is the whole opening from the before-fix record, with the agent’s replies and tool calls between the messages. Run it from case-build/results/estimate-adapter-as-shipped/ in the sample. The agent tried to look up the charge after “online”, before the caller had given the date.

$ python3 - <<'EOF'
import json
events = json.load(open("trackers/normal-brightmart-unrecognised.json"))["events"]
n = 0
for e in events:
    if e["event"] == "user" and e["text"] == "Yes.":
        break  # the caller's second turn starts here
    if e["event"] == "user" and not e["text"].startswith("/"):
        n += 1
        print(f'user {n:>2}  {e["text"]}')
    elif e["event"] == "bot":
        print(f'bot      {e["text"].strip()}')
    elif e["event"] == "tool_executed":
        print(f'tool     {e["tool_name"]}')
EOF
bot      Northgate Bank disputes. To help with a charge, please tell me your full name and date of birth.
user  1  Hello,
bot      Hello. Could you please tell me your full name and date of birth, and which charge you don't recognise?
user  2  this is
tool     listen
user  3  Priya
bot      Thanks, Priya. Could I have your full name and your date of birth, please?
user  4  Raghunathan.
bot      Thank you, Priya Raghunathan. What is your date of birth, please?
user  5  Date
tool     listen
user  6  of
tool     listen
user  7  birth?
bot      It sounds like you're repeating my question. Could you tell me your date of birth, with the day, month and year?
user  8  17th August 1991.
bot      Okay, thanks for that. Let me get the dispute process started.
tool     activate
tool     verify_caller
bot      Thank you, Priya, you're verified. Which charge don't you recognise? Please tell me the merchant, amount, date and card ending if you know them.
user  9  I
tool     listen
user 10  don't
tool     listen
user 11  recognize a
tool     listen
user 12  charge
tool     listen
user 13  of
tool     listen
user 14  ₹2,499 from
tool     listen
user 15  Bright
tool     listen
user 16  Mart
tool     listen
user 17  online
bot      Okay, let me look up that Bright Mart charge.
tool     select_transaction
tool     select_transaction
bot      I couldn't find a match for Bright Mart at 2,499 rupees. Could you check the merchant name, amount, and date, and tell me the card ending if you have it?
user 18  on
tool     listen
user 19  the 22nd of September.
tool     select_transaction
bot      I still can't find a match for Bright Mart, 2,499 rupees, on the 22nd of September. Could you check the merchant name and amount, and tell me the last four digits of the card?

Count caller messages against turns spoken

A test call can look fine while it splits. So after each test run, count the caller messages in each conversation record. Compare them with the turns your caller spoke.

Rasa calls that record the tracker, and each caller message in it is a user event. The Deepgram Flux guide uses the same count to catch a turn that produced no message. A split is the other direction.

This script does the count. It is ours, written for the sample’s test harness, which saves a results.json and a trackers folder per run. Save it as recount.py in the sample’s case-build/results/ folder. Commands such as /session_start are left out of the count.

import json, pathlib, sys

run = pathlib.Path(sys.argv[1])
spoken = {c["id"]: len(c["turns"]) for c in json.loads((run / "results.json").read_text())["conversations"]}
total = 0
for path in sorted((run / "trackers").glob("*.json")):
    events = [e["text"] for e in json.loads(path.read_text())["events"]
              if e["event"] == "user" and not (e.get("text") or "").startswith("/")]
    total += len(events)
    flag = "" if len(events) == spoken[path.stem] else "  <- mismatch"
    print(f"{path.stem:40} spoken {spoken[path.stem]}  user events {len(events)}{flag}")
print(f"total: spoken {sum(spoken.values())}, user events {total}")

If your harness is different, keep the tracker loop. Take the number of turns spoken from your own call script.

Here it is on the run before the fix, and on a 20-call run with the join:

Before the fix
$ python3 recount.py estimate-adapter-as-shipped
normal-brightmart-unrecognised           spoken 2  user events 27  <- mismatch
total: spoken 2, user events 27
With the join, 20 calls
$ python3 recount.py 2026-09-29-claude-sonnet-5.5
adversarial-ambiguous-file-both          spoken 3  user events 3
adversarial-refund-now                   spoken 2  user events 2
adversarial-skip-confirmation            spoken 2  user events 2
adversarial-spouse-transaction           spoken 2  user events 2
adversarial-unknown-charge               spoken 2  user events 2
adversarial-unsure-at-confirmation       spoken 1  user events 1
adversarial-wrong-birth-date             spoken 3  user events 3
correction-other-charge-at-confirmation  spoken 1  user events 1
correction-recognises-gym                spoken 2  user events 2
correction-self-corrected-amount         spoken 2  user events 2
normal-brightmart-unrecognised           spoken 2  user events 2
normal-date-said-numerically             spoken 2  user events 2
normal-dispute-and-block-card            spoken 3  user events 3
normal-identity-first-then-charge        spoken 3  user events 3
normal-second-customer-arjun             spoken 2  user events 2
recovery-ack-lost-found-by-key           spoken 2  user events 2
recovery-ambiguous-then-amount           spoken 3  user events 3
recovery-case-service-down               spoken 2  user events 2
short-reply-no                           spoken 2  user events 2
short-reply-yes                          spoken 2  user events 2
total: spoken 43, user events 43

Every call in the 20-call run matched. The longest of those turns had about 19 seconds of speech. The short replies “Yes.” and “No.” were heard too.

Trade-offs

Every turn now waits for silence. With the join, Rasa gets no transcript until the caller has been quiet for 0.7 seconds. The runs recorded the time from the end of the caller’s speech to the first agent audio:

RunMedian, end of speech to first agent audioTurns
Before the fix, same call230 ms2
With the join, same call2,400 ms2
With the join, 20-call run3,093 ms42

The sources are each run’s summary.md, line 18, and the sample’s README.

The first two rows do not compare like with like. Before the fix, the agent was already replying to fragments while the caller was still speaking. In the second turn, its first reply after “Yes.” is recorded before “Please”. So the 230 ms is a fast reply to part of a sentence, not to the whole turn.

How much of the wait the 0.7 seconds adds on its own was not measured. In the 20-call run, the median time from the end of speech to Rasa starting to process the turn was 1,587 ms. That includes the 0.7 seconds, Speechmatics’ own processing and the network. No run changed the trigger with the join in place, so no run splits that time apart.

A long pause ends the utterance. If a caller stops for longer than the trigger in the middle of a sentence, Speechmatics ends the utterance there. None of the 43 scripted turns split, but the callers read from a script. These runs cannot show how often it happens with real callers.

Why not raise max_delay instead?

The second example engine’s README suggests it: “Raise max_delay for whole utterances at the cost of latency.” No run here tested that. Speechmatics documents max_delay as “the delay in seconds between the end of a spoken word and returning the Final transcript results” (Speechmatics realtime API reference, read 2 October 2026). That setting is about each word. The message that marks the end of an utterance is EndOfUtterance, so that is the one to join on.

Limits

  • The calls are scripted, with generated caller voices, on one sample agent and one day. They are observations, not rates for real callers.
  • The before and after model-call counts come from two separate runs, started four minutes apart.
  • The Rasa behaviour is read from the 3.21.0.dev5 code. Check it on your version.
  • The trigger was set to 0.7 seconds and never varied, so this page gives no advice on tuning it.

Starting from scratch? The voice AI agent tutorial builds an agent on Deepgram for speech in and out instead.