Skip to content
RasaGet a free licence
Guides for AI teams

Guide · AI product engineer

Why a Rasa voice agent on Deepgram Flux can miss a short yes

When Deepgram Flux opens and closes a turn with no interim update, the Rasa handler commits nothing, so the agent never hears the reply.

by Rod Rivera

About 5 minutes

  • 4 of 10

    times one recorded "Yes." produced no user event in test calls

  • 16 of 16

    longer "Yes, ..." replies in the same test runs that the agent heard

Source: Scripted voice test calls of a sample card-blocking agent, 29 September 2026
Key takeaways (3)
  • A one-word reply can arrive with no interim update, and the Flux handler then commits nothing and logs nothing.
  • The fix is small: when the turn ends with nothing buffered, use the final message's own transcript.
  • Test voice calls by counting user events against the turns your caller spoke.

Your voice agent asks the caller to confirm something. The caller says “yes”. The agent says nothing. About 30 seconds after the caller stopped speaking, it asks whether they are still there. Longer answers work. Only the one-word ones go missing, and only some of the time.

Here is one of those calls, from a sample bank agent that blocks lost cards. The caller’s audio was the same recorded “Yes.” every time it was used:

Voice callLive run, 29 September 2026, trimmed: the tracker after the caller says “Yes.”
  1. Agent00:35

    Do you mean the debit card ending 4 4 1 7 on your everyday current account? Blocking that card will not replace it or change your other cards.

  2. Caller

    Yes.

    Annotation:

    0.33 seconds of speech. The tracker has no user event for it. The caller’s line is the audio the test played, so it is not in the tracker.

  3. Agent01:15

    Are you still on the line? I can continue whenever you are ready.

    Annotation:

    The silence prompt. The channel’s silence_timeout is 30 seconds.

  4. Agent01:19

    Please say yes to block the debit card ending four four one seven, or no if that is not the right card.

The same recording was lost in 4 of 10 attempts.

A direct test against Flux reproduces a cause that fits these calls. A very short word can arrive as “turn started” and then “turn ended”, with no update carrying the words in between. The Rasa handler builds the reply only from updates, so it commits nothing. The fix is to commit the transcript that the “turn ended” message already carries.

The calls ran on Rasa Pro 3.21.0.dev5, a prerelease. The released 3.20.1 has the same end-of-turn logic. Replayed offline over the same saved Flux messages, both handlers lost 26 of 54 attempts. We did not run a live call on 3.20.1.

This matters because the failure hides. Nothing in the logs says a turn was lost. The call then fails its block_card check with “0 matching call(s), need >= 1”, which reads like the model failed to act. It never heard the answer.

Why the short reply goes missing

Deepgram Flux is a speech-to-text model that also decides when the caller has finished a turn. It sends three kinds of message:

  • StartOfTurn when it hears the caller start.
  • Update, an interim message with the words so far. Flux sends one about every 0.24 seconds of audio for as long as the connection is open, and most of them are empty.
  • EndOfTurn when it decides the caller has finished.

Rasa turns these into its own events. UserIsSpeaking means the caller is talking. NewTranscript is the final reply, which becomes a user event in the tracker. Here is the part of the Rasa handler that matters, as an excerpt:

rasa/core/channels/voice_stream/asr/deepgram/engine.py, _DeepgramV2.parse_event (excerpt, rasa-pro 3.21.0.dev5)
        event_type = data.get("event")
        transcript = data.get("transcript", "")

        if event_type == "StartOfTurn":
            self._accumulated_transcript = ""
            return None

        if event_type == "Update":
            # v2 sends the full current-turn transcript on every Update; replace.
            if not transcript:
                return None
            self._accumulated_transcript = transcript
            return UserIsSpeaking(text=transcript)

        if event_type == "EndOfTurn":
            if self._accumulated_transcript:
                committed = self._accumulated_transcript
                self._accumulated_transcript = ""
                return NewTranscript(text=committed)
            return None
  1. Every message’s transcript is read here, including the one on EndOfTurn.
  2. StartOfTurn empties the buffer, the handler’s own copy of the reply.
  3. This is the only line that fills the buffer, and only from an Update with words.
  4. If no such Update arrived, the buffer is empty and the turn is dropped. Nothing is logged on this path.

The same pattern can appear in any speech-to-text adapter. The adapter rebuilds the final transcript from interim updates and ignores the final message’s own text. If you use another vendor behind your own adapter, check for it there too.

To see this outside a live call, the sample’s recordings were streamed straight to Flux, through Rasa’s own engine with the sample’s settings. There were 54 attempts: 44 of the same “Yes.” and 10 of “Yes, that’s the one.”. Each attempt had some silence before the word. Every message Flux sent was saved, then fed through the handler. Here are two “Yes.” attempts whose only difference is 20 milliseconds of extra silence:

Silence before the wordTurn messages with words (empty Updates omitted)What the handler returned
2.00 secondsStartOfTurn “Yes.”, then EndOfTurn “Yes.”Nothing: the turn is lost
2.02 secondsStartOfTurn “Yeah.”, Update “Yes.”, EndOfTurn “Yes.”NewTranscript("Yes.")

In the second attempt, Flux opened the turn a step earlier, so the word also went out as an Update.

Across the 54 attempts, 26 had no Update with words. The handler returned nothing on exactly those 26, and all of them were the short “Yes.”. All 10 longer replies had an Update with words and were heard.

The test also moved the word in 20 ms steps across one 240 ms window of Flux’s audio, twice over. At 3 of the 12 positions the word was lost in both passes, and at the other 9 it was heard in both:

2.00 lost 2.02 heard 2.04 lost 2.06 lost 2.08 heard 2.10 heard 2.12 heard 2.14 heard 2.16 heard 2.18 heard 2.20 heard 2.22 heard
  1. Lost: Flux sent EndOfTurn straight after StartOfTurn, so no Update carried the word.
  2. Heard: Flux opened the turn one step early, on “Yeah.”, so the word went out as an Update.
  3. Heard from here on: an Update with the word came before EndOfTurn.
FigureWhere the word started, in seconds of silence before it, and whether the handler kept it

Our inference from the saved messages, not something we measured in a live call: there, the word’s position would depend on everything streamed before the reply. That is the likely explanation for why the loss looks random.

How to fix it

The change is one branch. When EndOfTurn arrives and the buffer is empty, commit the transcript that EndOfTurn carries. Here it is as a change to the EndOfTurn branch of engine.py in the installed rasa-pro package:

         if event_type == "EndOfTurn":
             if self._accumulated_transcript:
                 committed = self._accumulated_transcript
                 self._accumulated_transcript = ""
                 return NewTranscript(text=committed)
+            if transcript:
+                return NewTranscript(text=transcript)
             return None

It was tested only by offline replay, not in a live call. The companion’s patched_handler_replay.py applies the same logic to the 54 saved attempts: lost replies went from 26 to none, and none was committed twice.

Check whether your release has the same code

The two releases read are in the versions paragraph near the top. The one difference: in 3.20.1, StartOfTurn returns an empty UserIsSpeaking instead of nothing.

Find the installed handler

Run this with the Python from your agent’s environment. It prints the path of engine.py:

python -c "import rasa.core.channels.voice_stream.asr.deepgram.engine as m; print(m.__file__)"

Find the end-of-turn branch

Open that file. Find _DeepgramV2.parse_event and the if event_type == "EndOfTurn": branch.

Look at what it returns with an empty buffer

If it returns None without reading EndOfTurn’s own transcript, your agent can lose short replies.

What the change costs

With the change, Rasa can send a final reply with no UserIsSpeaking before it. In 3.21.0.dev5 we read the code that consumes these events and found nothing that needs one. We did not read that code in 3.20.1.

How to test for lost turns

A lost turn leaves no transcript to compare, so start from what the caller said. The sample’s test driver does this. Before each caller turn, it counts the user events in the tracker. It plays the audio, waits for the bot, and counts again. A turn with no new user event was not heard.

Counting user events also shows the opposite failure. In the guide to a speech adapter that split one opening into 19 turns, one spoken opening became 19 user events. Here a spoken turn produced none. For splits, that guide recounts user events over the whole tracker after the call.

This is the check from the companion folder, unheard_turns.py, as an excerpt. It reads the sample’s own results file. It flags a turn with no user event, or a wait as long as the silence timeout:

SILENCE_TIMEOUT_MS = 30_000  # integrations.yml: silence_timeout: 30


def unheard_turns(path):
    run = json.load(open(path))
    for call in run["conversations"]:
        for i, turn in enumerate(call["turns"]):
            extra = turn.get("extra") or {}
            if extra.get("mode") != "audio":
                continue
            events = len(extra.get("heard") or [])  # new user events in the tracker
            wait = turn.get("latency_ms") or 0
            if events == 0 or wait >= SILENCE_TIMEOUT_MS:
                yield call["id"], i, turn["user"], events, wait

Map the two checks to your own harness’s fields:

  • heard is the list of new user events after the caller’s turn. Fail the run if it is empty.
  • latency_ms is the time from the end of the caller’s speech to the bot’s next audio. Set SILENCE_TIMEOUT_MS to the silence_timeout in your integrations.yml, in milliseconds.

Run over the sample’s four recorded test runs, 75 spoken caller turns in all, it flags the four lost replies and nothing else. The recorded command and output, from the sample’s folder:

$ cd case-build/results && python3 ../flux-short-reply/unheard_turns.py 2026-09-29-gpt-5.5-default/results.json 2026-09-29-gpt-5.5-reasoning-low/results.json 2026-09-29-gpt-5.5-default-verify-fix/results.json 2026-09-29-gpt-5.5-default-fix-rerun/results.json; echo "exit $?"
recovery-verification-retry turn 2: said 'Yes.', user events 0, wait 30,223 ms
normal-block-then-asked-replacement turn 1: said 'Yes.', user events 0, wait 30,045 ms
adversarial-lost-debit-two-debits turn 2: said 'Yes.', user events 0, wait 30,274 ms
correction-other-card-at-confirmation turn 3: said 'Yes.', user events 0, wait 30,203 ms
4 unheard turn(s)
exit 1

It exits with 1 when it finds a lost turn, so it can fail a CI job.

On live traffic, with no script, look for the silence prompt straight after a caller spoke. In the sample’s server logs, kept locally, it was the only sign of a lost turn.

Reproduce it yourself

The saved Flux messages and the replay scripts are in the companion repository. The replays run offline and need no Deepgram key. make install runs uv sync to create the sample’s .venv with the pinned rasa-pro, so it needs uv and a network connection:

git clone https://github.com/RasaHQ/rasa-community-resources.git
cd rasa-community-resources
git checkout cc1aff712c15d9cfed9c6e56928e4650754d6614
cd examples/mantle-voice-banking-block-card-gpt
make install

Then replay the saved messages through the installed handler, unchanged. This writes its report to /tmp/replay.txt; change that path if you like:

.venv/bin/python case-build/flux-short-reply/flux-short-reply-probe.py replay \
  --frames case-build/flux-short-reply/flux-short-reply-frames.jsonl --out /tmp/replay.txt

Look for the line that starts all attempts. In the recorded replay it reads:

all attempts, both WAVs and sets: 54; no-Update sequences 26; handler None 26; None set == no-Update set: True

So the handler returned nothing on exactly the 26 attempts with no Update. Then replay them with the fix, using the companion’s patched_handler_replay.py:

Any system
cd case-build/flux-short-reply && ../../.venv/bin/python patched_handler_replay.py

It prints one row per attempt, then these totals (trimmed to the totals):

attempts: 54
unchanged handler returned None: 26
patched handler returned None:   0
attempts whose committed text changed: 26
attempts with more than one NewTranscript under the patch: 0

The probe’s probe and sweep subcommands capture new messages from Flux. They need a DEEPGRAM_API_KEY. The full capture commands are on the first line of flux-short-reply-frames.jsonl.

Limits

  • The live calls come from a sample agent for an invented bank, with synthetic caller voices over 16 kHz browser audio. There were no real callers, and 8 kHz phone audio was not tested.
  • Ten attempts of one recording show that the problem happens, not how often it will happen to your callers.
  • Flux’s messages were not captured inside the four lost live calls. The direct test reproduces a cause that fits them.
  • The fix was replayed through the handler only, not in a live call.
  • With interruptions turned on, a final reply that arrives while the bot is speaking may be handled differently. The sample ran with interruptions off, so this case is not settled.
  • Only two releases were read: 3.21.0.dev5 and 3.20.1.
  • The sample ran only at an eot_threshold of 0.7. A longer confirmation question was not tested either.
Is this a bug in Deepgram Flux?

Flux returned “Yes.” on both the StartOfTurn and the EndOfTurn in every one of the 26 attempts with no Update, so the words reached Rasa every time. The part that drops them is the handler’s buffer logic. We did not check Flux’s documentation, so this page makes no claim about what Flux promises.