Guide · AI product engineer
How to join Speechmatics segments into one Rasa turn
If your Rasa voice agent asks again for details the caller just gave, check whether your speech-to-text engine sends each phrase as its own turn.
19
caller messages Rasa received from one spoken opening, before the fix
36 → 10
model calls for the same two-turn call, before → after the fix (two separate runs)
43 of 43
spoken caller turns that arrived as one message each, over 20 test calls after the fix
Key takeaways (3)
- In our test calls, Speechmatics finalised speech a word or two at a time, and Rasa took each transcript as a turn.
- Hold the phrases in your engine and send one transcript when Speechmatics marks the end of the utterance.
- After each test call, count the caller messages in the conversation record against the turns spoken.
Your Rasa voice agent asks the caller for their name and date of birth. The caller gives both in one breath.
The agent thanks them by first name, then asks for the full name again. Then it asks for the date of birth, and then again.
Here is that call, from a scripted test of a sample bank agent that takes card disputes. The caller opened with one recording of about 14 seconds: “Hello, this is Priya Raghunathan, date of birth seventeenth August, nineteen ninety-one. I don’t recognise a charge of two thousand four hundred and ninety-nine rupees from Brightmart Online on the twenty-second of September.”
Rasa did not receive that as one message. This is what its conversation record holds:
Caller
Hello,
Agent
Hello. Could you please tell me your full name and date of birth, and which charge you don’t recognise?
Caller
this is
Tool
listen() → listen: done
Annotation:listenis a tool Rasa’s agent runtime offers the model: wait for more speech without replying.Caller
Priya
Agent
Thanks, Priya. Could I have your full name and your date of birth, please?
Annotation:The model answers a one-word message. It has not seen the rest.
Caller
Raghunathan.
Agent
Thank you, Priya Raghunathan. What is your date of birth, please?
Caller
Date
Tool
listen() → listen: done
Caller
of
Tool
listen() → listen: done
Caller
birth?
Agent
It sounds like you’re repeating my question. Could you tell me your date of birth, with the day, month and year?
Annotation:“Date”, “of”, “birth?” are the caller’s own words. Split up, they read as an echo.
Caller
17th August 1991.
Agent
Okay, thanks for that. Let me get the dispute process started.
Your speech-to-text engine is the class that turns a vendor’s messages into transcripts for Rasa. In this call, the engine sent Rasa each short phrase as a finished transcript. Rasa took each one as a caller turn, and the model replied to each fragment on its own.
The fix is in the engine: ask Speechmatics to mark the end of each utterance, hold the phrases until then, and send them to Rasa as one transcript.
That matters if you pay per model call, or if your callers lose patience when asked twice. In the sample, the same two-turn call made 36 model calls before the fix and 10 after. Those are two separate test runs, not a controlled pair.
The engine that split the call was ours: the Speechmatics adapter in the companion repository. The fix below is in the same file.
What you need:
- The engine is
SpeechmaticsASRinpatterns/voice-vendor-router, at commit4aa0c44. - The sample agent is
examples/mantle-voice-banking-dispute-claude. It runs onclaude-sonnet-5-5and Rasa Pro 3.21.0.dev5, and speaks with Rime. Northgate Bank and the caller are invented. The caller’s voice is generated speech read from a script. - The Rasa behaviour described here is from Rasa Pro 3.21.0.dev5, the version the sample pins. Check it on your version.
- The offline tests need uv and no keys.
Why one sentence becomes many turns
Speechmatics sends text while the caller is still talking: partial results
first, then finals. A final is Speechmatics’ finished text for a stretch of
speech. In these test calls, with max_delay set to 1.0 seconds, a final was
often only a word or two. So it is a phrase, not a whole turn.
The engine as shipped sent every final to Rasa as a NewTranscript, the
event that carries what the caller said. Rasa’s voice channel turns each
non-empty transcript into a caller message. This excerpt is from
rasa/core/channels/voice_stream/voice_channel.py in the 3.21.0.dev5 wheel.
It stops at the next branch:
if isinstance(asr_event, NewTranscript) and asr_event.text:
logger.debug(
"VoiceInputChannel.handle_asr_event.new_transcript",
transcript=asr_event.text,
)
# Open the Rasa-processing window for the first bot message.
self._mark_rasa_processing_started()
event: VoiceInputEvent = FinalTranscriptInputEvent(
text=asr_event.text,
metadata=self.build_turn_metadata(call_parameters),
)
await input_queue.put(event)
elif isinstance(asr_event, UserIsSpeaking):
Rasa does join some transcripts. It merges final transcripts that are already waiting in its input queue together. Finals that arrive separately become separate turns.
We did not trace which of the 19 phrases arrived together. The record shows the outcome.
Your settings may change how long a final runs. The second example engine’s
README links fragment size to max_delay. At 1.0 seconds, it saw one
spoken sentence come back as four finals.
Speechmatics can say where an utterance ends. With end-of-utterance
detection on, it sends an extra EndOfUtterance message after a set time of
silence
(Speechmatics documentation,
read 2 October 2026). Nothing in the Rasa 3.21.0.dev5 wheel reads that
message, so joining the phrases is the engine’s job.
- Speechmatics sends finished text a word or two at a time.
- The engine as shipped sent each final to Rasa as its own transcript.
- With the join, the engine holds the finals until
EndOfUtterance, then sends one transcript.
How to fix it
In your own project the change goes in two places: your engine class, and
the asr section of your voice channel’s configuration. In the sample that
configuration is in integrations.yml.
Ask Speechmatics to mark the end of each utterance
Speechmatics takes the request as end_of_utterance_silence_trigger inside
conversation_config, in the StartRecognition message your engine sends
first. The value is the seconds of silence to wait, between 0 and 2. This
excerpt from
voicerouter/providers/speechmatics.py
adds it when the setting is on:
if self._utterance_mode:
transcription_config["conversation_config"] = {
"end_of_utterance_silence_trigger": float(
self.config.end_of_utterance_silence_trigger
),
}Hold the finals and send one transcript
This is the part of the engine that translates Speechmatics messages. It is an excerpt from the same file, shown without its method indentation:
if kind == "AddTranscript":
text = message.get("metadata", {}).get("transcript", "").strip()
if not text:
return None
if not self._utterance_mode:
return NewTranscript(text)
# Hold the segment until Speechmatics says the utterance ended.
self._segments.append(text)
return UserIsSpeaking(" ".join(self._segments))
if kind == "EndOfUtterance":
if not self._segments:
return None
text, self._segments = " ".join(self._segments), []
return NewTranscript(text)- With the setting off, each final still goes out as a transcript. The setting’s comment says “Unset keeps the original behaviour”, so existing configurations do not change.
- With it on, the final is held. The text so far goes out as
UserIsSpeaking, which Rasa turns into a partial transcript, not a turn. - An
EndOfUtterancewith nothing held sends nothing. - The held phrases are joined with a space and sent as one transcript.
Turn it on in your configuration
The sample names the engine by its Python path and sets the trigger to
0.7 seconds. This excerpt is the asr section from its
integrations.yml.
It is shown without its leading indentation and its comment lines. In the
file, it sits under channels, in the browser_audio channel:
asr:
name: voicerouter.providers.speechmatics.SpeechmaticsASR
language_map:
en:
language: en
operating_point: enhanced
max_delay: 1.0
enable_partials: true
end_of_utterance_silence_trigger: 0.7One of those comments says of 0.7: “it was chosen, not tuned.”
How to check the fix works
Check it in three ways: an offline test of the engine, a live call, and a count over your conversation records.
Do the count over the whole conversation record, after the call. The sample’s own test harness counted messages per turn, and it undercounted. For the caller’s reply “Yes, please confirm it. I never bought anything from them.”, it recorded two caller messages. The record holds eight.
Test the engine offline
Feed the engine the messages Speechmatics would send, and count the
transcripts that come out. The companion has two tests for this in
tests/test_voicerouter.py.
With the trigger off, three finals give three transcripts. With the trigger
on, this excerpt shows what happens:
held = [engine.engine_event_to_asr_event(self._msg("AddTranscript", t)) for t in ("Hello,", "this is", "Priya.")]
self.assertTrue(all(isinstance(e, UserIsSpeaking) for e in held))
done = engine.engine_event_to_asr_event(self._msg("EndOfUtterance"))
self.assertIsInstance(done, NewTranscript)
self.assertEqual(done.text, "Hello, this is Priya.")
# A second end of utterance with nothing new said sends nothing.
self.assertIsNone(engine.engine_event_to_asr_event(self._msg("EndOfUtterance")))
# A bare "Yes." is one segment and still becomes a transcript.
engine.engine_event_to_asr_event(self._msg("AddTranscript", "Yes."))
self.assertEqual(engine.engine_event_to_asr_event(self._msg("EndOfUtterance")).text, "Yes.")
Run both from patterns/voice-vendor-router. No socket opens and no audio
is sent:
- Any system
uv run --frozen python -m unittest -v tests.test_voicerouter.TestSpeechmaticsSocket.test_default_mode_makes_every_final_segment_a_transcript tests.test_voicerouter.TestSpeechmaticsSocket.test_utterance_mode_joins_segments_until_end_of_utterance
This output is trimmed: the log text after each test name is cut.
test_default_mode_makes_every_final_segment_a_transcript (tests.test_voicerouter.TestSpeechmaticsSocket.test_default_mode_makes_every_final_segment_a_transcript) ...
ok
test_utterance_mode_joins_segments_until_end_of_utterance (tests.test_voicerouter.TestSpeechmaticsSocket.test_utterance_mode_joins_segments_until_end_of_utterance) ...
ok
----------------------------------------------------------------------
Ran 2 tests in 0.007s
OK
Place a call
With the join on, the same scripted opening reached Rasa as one message:
Caller
Hello, this is Priya Raghunathan . Date of birth? 17th August 1991. I don’t recognize a charge of ₹2,499 from Bright Mart online on the 22nd of September.
Agent
Right, let me start the dispute process for you.
Tool
verify_caller(full_name=“Priya Raghunathan”, date_of_birth=“1991-08-17”) → verified
The tracker shows a space before the full stop after the surname. The name
still reached verify_caller as “Priya Raghunathan”.
The two runs of the same call compare like this. They are separate runs, started four minutes apart:
| Same two-turn call | Before the fix | With the join |
|---|---|---|
| Caller messages from the opening | 19 | 1 |
| Caller messages from the reply | 8 | 1 |
| Model calls | 36 | 10 |
| Model spend | 0.32 USD | 0.08 USD |
Show all 19 caller messages from the call before the fix
This is the whole opening from the before-fix record, with the agent’s
replies and tool calls between the messages. Run it from
case-build/results/estimate-adapter-as-shipped/ in the sample. The
agent tried to look up the charge after “online”, before the caller had
given the date.
$ python3 - <<'EOF'
import json
events = json.load(open("trackers/normal-brightmart-unrecognised.json"))["events"]
n = 0
for e in events:
if e["event"] == "user" and e["text"] == "Yes.":
break # the caller's second turn starts here
if e["event"] == "user" and not e["text"].startswith("/"):
n += 1
print(f'user {n:>2} {e["text"]}')
elif e["event"] == "bot":
print(f'bot {e["text"].strip()}')
elif e["event"] == "tool_executed":
print(f'tool {e["tool_name"]}')
EOF
bot Northgate Bank disputes. To help with a charge, please tell me your full name and date of birth.
user 1 Hello,
bot Hello. Could you please tell me your full name and date of birth, and which charge you don't recognise?
user 2 this is
tool listen
user 3 Priya
bot Thanks, Priya. Could I have your full name and your date of birth, please?
user 4 Raghunathan.
bot Thank you, Priya Raghunathan. What is your date of birth, please?
user 5 Date
tool listen
user 6 of
tool listen
user 7 birth?
bot It sounds like you're repeating my question. Could you tell me your date of birth, with the day, month and year?
user 8 17th August 1991.
bot Okay, thanks for that. Let me get the dispute process started.
tool activate
tool verify_caller
bot Thank you, Priya, you're verified. Which charge don't you recognise? Please tell me the merchant, amount, date and card ending if you know them.
user 9 I
tool listen
user 10 don't
tool listen
user 11 recognize a
tool listen
user 12 charge
tool listen
user 13 of
tool listen
user 14 ₹2,499 from
tool listen
user 15 Bright
tool listen
user 16 Mart
tool listen
user 17 online
bot Okay, let me look up that Bright Mart charge.
tool select_transaction
tool select_transaction
bot I couldn't find a match for Bright Mart at 2,499 rupees. Could you check the merchant name, amount, and date, and tell me the card ending if you have it?
user 18 on
tool listen
user 19 the 22nd of September.
tool select_transaction
bot I still can't find a match for Bright Mart, 2,499 rupees, on the 22nd of September. Could you check the merchant name and amount, and tell me the last four digits of the card?Count caller messages against turns spoken
A test call can look fine while it splits. So after each test run, count the caller messages in each conversation record. Compare them with the turns your caller spoke.
Rasa calls that record the tracker, and each caller message in it is a
user event. The
Deepgram Flux guide
uses the same count to catch a turn that produced no message. A split is the
other direction.
This script does the count. It is ours, written for the sample’s test
harness, which saves a results.json and a trackers folder per run. Save
it as recount.py in the sample’s case-build/results/ folder. Commands
such as /session_start are left out of the count.
import json, pathlib, sys
run = pathlib.Path(sys.argv[1])
spoken = {c["id"]: len(c["turns"]) for c in json.loads((run / "results.json").read_text())["conversations"]}
total = 0
for path in sorted((run / "trackers").glob("*.json")):
events = [e["text"] for e in json.loads(path.read_text())["events"]
if e["event"] == "user" and not (e.get("text") or "").startswith("/")]
total += len(events)
flag = "" if len(events) == spoken[path.stem] else " <- mismatch"
print(f"{path.stem:40} spoken {spoken[path.stem]} user events {len(events)}{flag}")
print(f"total: spoken {sum(spoken.values())}, user events {total}")
If your harness is different, keep the tracker loop. Take the number of turns spoken from your own call script.
Here it is on the run before the fix, and on a 20-call run with the join:
Before the fix
$ python3 recount.py estimate-adapter-as-shipped
normal-brightmart-unrecognised spoken 2 user events 27 <- mismatch
total: spoken 2, user events 27With the join, 20 calls
$ python3 recount.py 2026-09-29-claude-sonnet-5.5
adversarial-ambiguous-file-both spoken 3 user events 3
adversarial-refund-now spoken 2 user events 2
adversarial-skip-confirmation spoken 2 user events 2
adversarial-spouse-transaction spoken 2 user events 2
adversarial-unknown-charge spoken 2 user events 2
adversarial-unsure-at-confirmation spoken 1 user events 1
adversarial-wrong-birth-date spoken 3 user events 3
correction-other-charge-at-confirmation spoken 1 user events 1
correction-recognises-gym spoken 2 user events 2
correction-self-corrected-amount spoken 2 user events 2
normal-brightmart-unrecognised spoken 2 user events 2
normal-date-said-numerically spoken 2 user events 2
normal-dispute-and-block-card spoken 3 user events 3
normal-identity-first-then-charge spoken 3 user events 3
normal-second-customer-arjun spoken 2 user events 2
recovery-ack-lost-found-by-key spoken 2 user events 2
recovery-ambiguous-then-amount spoken 3 user events 3
recovery-case-service-down spoken 2 user events 2
short-reply-no spoken 2 user events 2
short-reply-yes spoken 2 user events 2
total: spoken 43, user events 43Every call in the 20-call run matched. The longest of those turns had about 19 seconds of speech. The short replies “Yes.” and “No.” were heard too.
Trade-offs
Every turn now waits for silence. With the join, Rasa gets no transcript until the caller has been quiet for 0.7 seconds. The runs recorded the time from the end of the caller’s speech to the first agent audio:
| Run | Median, end of speech to first agent audio | Turns |
|---|---|---|
| Before the fix, same call | 230 ms | 2 |
| With the join, same call | 2,400 ms | 2 |
| With the join, 20-call run | 3,093 ms | 42 |
The sources are each run’s summary.md, line 18, and the sample’s
README.
The first two rows do not compare like with like. Before the fix, the agent was already replying to fragments while the caller was still speaking. In the second turn, its first reply after “Yes.” is recorded before “Please”. So the 230 ms is a fast reply to part of a sentence, not to the whole turn.
How much of the wait the 0.7 seconds adds on its own was not measured. In the 20-call run, the median time from the end of speech to Rasa starting to process the turn was 1,587 ms. That includes the 0.7 seconds, Speechmatics’ own processing and the network. No run changed the trigger with the join in place, so no run splits that time apart.
A long pause ends the utterance. If a caller stops for longer than the trigger in the middle of a sentence, Speechmatics ends the utterance there. None of the 43 scripted turns split, but the callers read from a script. These runs cannot show how often it happens with real callers.
Why not raise max_delay instead?
The second example engine’s README suggests it: “Raise max_delay for
whole utterances at the cost of latency.” No run here tested that.
Speechmatics documents max_delay as “the delay in seconds between the end
of a spoken word and returning the Final transcript results”
(Speechmatics realtime API reference,
read 2 October 2026). That setting is about each word. The message that marks
the end of an utterance is EndOfUtterance, so that is the one to join on.
Limits
- The calls are scripted, with generated caller voices, on one sample agent and one day. They are observations, not rates for real callers.
- The before and after model-call counts come from two separate runs, started four minutes apart.
- The Rasa behaviour is read from the 3.21.0.dev5 code. Check it on your version.
- The trigger was set to 0.7 seconds and never varied, so this page gives no advice on tuning it.
Starting from scratch? The voice AI agent tutorial builds an agent on Deepgram for speech in and out instead.