Skip to content
RasaGet a free licence
Guides for AI teams

Guide · Platform / operations engineer

Why TTS failover returns after a Rasa voice agent restart

A redeploy clears the voice router’s memory of a failing vendor; it never fixes the vendor. Read the failure class and pick wait, reset, re-route or escalate.

by Rod Rivera

About 9 minutes

Key takeaways (6)
  • A failover alert says a vendor failed. The failure class in the same log line says whether waiting, a fix or a person is needed, so route alerts on the class.
  • In the companion voice router, a rejected key or a broken request disables the vendor until the process restarts. A later success never re-enables it.
  • A restart empties the router’s process-wide health registry. Unless the vendor was fixed first, the next call finds the same fault and fails over again.
  • The log text "skipping briefly" is not the parking time. A quota failure with no Retry-After header parks the vendor for 900 seconds.
  • In the companion, reset_shared_registries() is called only by the tests and the drill. No endpoint exposes it, so in practice the reset is a restart after the fix.
  • What Rasa 3.20.0rc1 does with a TTS error raised mid-stream by these engines was not established. Do not build a runbook on the caller hearing silence.

The alert shows two log lines. They come from an offline drill in the companion project, with the vendor calls stubbed, in which the primary text-to-speech (TTS) vendor starts returning HTTP 402 (out of credits) on the third sentence of a call:

2026-09-21 21:19:57 [warning  ] voicerouter.tts.failing_over   attempt=1 from_provider=rime verdict='quota (HTTP 402) — skipping briefly' voice_changes=True
2026-09-21 21:19:57 [info     ] voicerouter.tts.failover       from_provider=rime reason='served after failover' to_provider=deepgram

Picture the on-call engineer who sees those lines. The agent has switched to a backup voice, and something looks stuck, so they redeploy the voice service. The alerts stop. On the next real call, the router tries Rime first again. Rime returns 402 again, the same two lines fire and the pager goes off again.

Nothing about Rime’s account changed. The redeploy wiped the router’s in-process record that Rime was out of credits, a record that had been keeping later calls off Rime for 900 seconds, so the next conversation paid for the discovery again. The scene is illustrative; the log lines and the 900 seconds come from the drill run described below.

The signal worth routing is the failure’s class, not the fact that a vendor failed. Page a person on a failure that only a person can fix. Let the router handle a rate limit. Redeploy only after the fix, because a redeploy on its own resets the router’s memory and does nothing to the vendor. The cost of this stance is work up front: your alerting has to read a field from the log or metric rather than count failovers, and someone has to own each class before the first incident.

This guide uses the voice router in RasaHQ/rasa-community-resources at revision 69e27b6: the pattern patterns/voice-vendor-router and the example agent examples/mantle-voice-routed-skills, whose assistant Vela answers for the companion’s seeded demo bank (its README, lines 207 and 212). Both pin Rasa Pro 3.20.0rc1 (rasa-pro==3.20.0rc1 on line 8 of each pyproject.toml). The router applies the circuit-breaker pattern to speech vendors, and it honours the HTTP Retry-After header when a vendor sends one. Rime, Deepgram and OpenAI appear because they are the example’s configured TTS chain. Nothing here measures any vendor’s reliability.

Why a Rasa voice agent needs a router to fail over

A Rasa voice channel holds one ASR (speech-to-text) engine and one TTS engine. The example’s integrations.yml names voicerouter.RoutedTTS as its TTS engine, and that router holds a chain of real engines (lines 63–90): rime, then deepgram, then openai. Rasa itself does not walk a chain. Mantle speaks a response through TurnContext.send, which calls the channel’s send_text_message. On that path, the handler that catches a vendor error is this (rasa/core/channels/voice_stream/voice_channel.py, lines 619–625, in the 3.20.0rc1 wheel). We did not examine the separate streaming-response path (send_response_chunk and its background audio sender), which has its own error handling:

        try:
            audio_stream = self.tts_engine.synthesize(text)
        except TTSError as e:
            logger.error("voice_channel.tts_synthesis_error", error=str(e))
            voice_tracing.get_current_synthesize_speech_recorder().record_error(e)
            # TODO: add message that works without tts, e.g. loading from disc
            audio_stream = self.chunk_audio(generate_silence(self.audio_format))

That code logs the error and plays silence, and it tries no other vendor. Its reach is narrower than it looks. It catches only an error raised by calling synthesize(). In the base engine, in Rasa’s RimeTTS and DeepgramTTS (tts/deepgram.py, lines 184–195) and in RoutedTTS, synthesize is an async generator, and calling an async generator runs none of its body. Their errors surface later, at line 677, where _stream_audio_to_channel reads the stream, outside that try. We checked this with a short offline probe, run with the router pattern’s virtual environment. It is not part of the companion. <companion> stands for your checkout path, and the vendor is a stub that always raises:

import asyncio, inspect, sys
sys.path.insert(0, "<companion>/patterns/voice-vendor-router")
from rasa.core.channels.voice_stream.tts.tts_engine import TTSError
from rasa.core.channels.voice_stream.tts.rime import RimeTTS
from voicerouter.routed_tts import RoutedTTS
from voicerouter.base import BuiltProvider, ProviderSpec, RouterPolicy
from voicerouter.health import reset_shared_registries
print("RimeTTS.synthesize is async generator:", inspect.isasyncgenfunction(RimeTTS.synthesize))
print("RoutedTTS.synthesize is async generator:", inspect.isasyncgenfunction(RoutedTTS.synthesize))
class Dead:
    async def connect(self): pass
    async def synthesize(self, text, config=None):
        raise TTSError("401 Unauthorized"); yield b""
reset_shared_registries()
r = RoutedTTS([BuiltProvider(ProviderSpec(name="rime", label="rime", config={}), Dead())], RouterPolicy())
try:
    stream = r.synthesize("hello")   # the call Rasa wraps in try/except TTSError
    print("calling synthesize(): no exception raised")
except TTSError as e:
    print("calling synthesize(): raised", e)
async def consume():
    try:
        async for _ in stream: pass
    except TTSError as e:
        print("iterating the stream: raised TTSError:", str(e)[:90])
asyncio.run(consume())

Its output, with the router’s two log lines filtered out:

RimeTTS.synthesize is async generator: True
RoutedTTS.synthesize is async generator: True
calling synthesize(): no exception raised
iterating the stream: raised TTSError: voicerouter: no TTS provider available — rime: disabled (auth) — 1 attempted for this utte

We traced that error up through send_text_message (lines 1038–1122) and Mantle’s TurnContext.send (rasa/mantle/orchestration/turn_context.py, lines 290–299). Neither catches it, and we did not establish what the turn or the call does after that. Treat “every TTS vendor is out” as an unknown outcome for the caller, not as silence you can plan around.

synthesize(text) is called (line 620) logged as tts_synthesis_error, silence played (lines 621-625) engine raises on the call itself the stream is read (line 677, no try) async generator: nothing runs yet send_text_message, TurnContext.send: no handler; outcome not established TTSError
FigureWhere a TTSError surfaces in the Rasa 3.20.0rc1 voice channel

For ASR there are two layers, and the first one matters more. Rasa’s base ASR engine reads the vendor socket inside its own try (rasa/core/channels/voice_stream/asr/asr_engine.py, lines 353–364):

    async def stream_asr_events(self) -> AsyncIterator[ASREvent]:
        """Stream the events returned by the ASR system as it is fed audio bytes."""
        if self.asr_socket is None:
            raise ConnectionException("Websocket not connected.")

        try:
            async for message in self.asr_socket:
                asr_event = self.engine_event_to_asr_event(message)
                if asr_event:
                    yield asr_event
        except Exception as e:
            logger.warning(f"Error while streaming ASR events: {e}")

A socket error during a read is logged at warning level, and the stream then ends as if the call were over. DeepgramASR does not override this method. Only an error that escapes the engine, such as the ConnectionException above, reaches the voice channel’s own handler, which logs it at error level and re-raises (voice_channel.py, lines 1372–1380):

        except Exception as e:
            logger.error(
                "voice_channel.receive_asr_events.error",
                call_id=call_parameters.call_id,
                error=repr(e),
            )
            # Close the active transcription span as failed before re-raising.
            voice_tracing.get_current_transcribe_audio_recorder().record_error(e).end()
            raise

Of the engine’s own errors, the router sees only the ones that escape the engine, just as that channel handler does. RoutedASR treats a stream that ends without an error as the end of the call (routed_asr.py, lines 247–250), so a read error the engine has already swallowed never reaches it. An offline probe (a fake Deepgram socket that raises ConnectionError on the first read, no network) shows both cases. Case A is a read that raises; case B is a socket that was never connected. We cut four start-up lines from the output: Rasa’s asr.engine.initialized line, the engine class: DeepgramASR line and the two voicerouter.asr.ready lines. The rest is unedited:

DeepgramASR overrides stream_asr_events: False
2026-09-21 21:32:18 [warning  ] Error while streaming ASR events: socket dropped mid-call (probe)
A: socket drops during a read -> RoutedASR.stream_asr_events ended without raising
A: socket drops during a read -> deepgram health: failures = 0 | active provider: deepgram
2026-09-21 21:32:18 [warning  ] voicerouter.asr.stream_failed  error='Websocket not connected.' provider=deepgram verdict='unknown (ConnectionException) — skipping briefly'
2026-09-21 21:32:18 [info     ] voicerouter.asr.failover       from_provider=deepgram reason='connected after failover' to_provider=backup
2026-09-21 21:32:18 [info     ] voicerouter.asr.connected      provider=backup
2026-09-21 21:32:18 [info     ] voicerouter.asr.resumed        note='audio sent during the failure was not transcribed' provider=backup
B: socket was never connected -> RoutedASR.stream_asr_events ended without raising
B: socket was never connected -> deepgram health: failures = 1 | active provider: backup

In case A the router records nothing and does not fail over, so voicerouter.asr.stream_failed and voicerouter.asr.exhausted never fire when a read raises inside the engine. The probe does not cover a real socket that closes cleanly. What the call does after the stream ends was not established.

Read the failure class, not the failover

Every vendor error the router sees goes through classify() in voicerouter/failures.py (lines 206–263), which turns an HTTP status, an AWS error code or the message text into a class. The class decides two things: how long the vendor is skipped (lines 70–76) and whether the caller may hear a different voice (lines 88–93):

DEFAULT_COOLDOWNS: dict[FailureKind, float] = {
    FailureKind.AUTH: 0.0,         # permanent; cooldown unused
    FailureKind.CONFIG: 0.0,       # permanent; cooldown unused
    FailureKind.QUOTA: 900.0,      # 15 minutes
    FailureKind.RATE_LIMIT: 20.0,  # overridden by Retry-After when present
    FailureKind.TRANSIENT: 15.0,
}
_JUSTIFIES_VOICE_CHANGE = {
    FailureKind.AUTH,
    FailureKind.CONFIG,
    FailureKind.QUOTA,
    FailureKind.UNAVAILABLE,
}

classify() supplies the triggers in the table below. The “waiting” column paraphrases the module docstring (lines 6–10). When a verdict carries the vendor’s Retry-After, that value replaces the default, with a floor of one second (lines 266–274). Quota, rate-limit and transient verdicts can carry it; unavailable verdicts never do (lines 248 and 261). Classes with no default use the policy’s cooldown_seconds, which the example sets to 60.

ClassTypical triggerWaiting fixes it?Router skips the vendor forVoice changes?
authHTTP 401 or 403NeverUntil restart or resetYes
configHTTP 400, 404, 405, 415 or 422NeverUntil restart or resetYes
quotaHTTP 402, or a 429 with quota textAfter a top-up900 s, or Retry-AfterYes
unavailableConnection refused, timeoutOftencooldown_seconds (60 here)Yes
rate_limitHTTP 429Yes, in seconds20 s, or Retry-AfterAfter one same-vendor retry
transientHTTP 5xxYes, in seconds15 s, or Retry-AfterAfter one same-vendor retry

The last column comes from routed_tts.py, lines 356–377. A retryable failure gets one more attempt on the same vendor (same_provider_retries: 1, after retry_backoff_ms: 250 in the example) as long as the wait is at most one second. If that retry fails too, or the vendor asks for seven seconds, the router moves on. Anything classify() cannot place is unknown: it is treated like transient, with the policy’s cooldown_seconds as its park.

Where the router keeps circuit state

Each vendor has a ProviderHealth record. The branch that matters for on-call work is in record_failure (voicerouter/health.py, lines 75–85):

        if is_permanent(verdict):
            # No threshold for these. One 401 is enough to know.
            self.disabled = True
            self.opened_at = time.monotonic()
            self.current_cooldown = float("inf")
            return verdict

        if self.consecutive_failures >= self.failure_threshold:
            self.opened_at = time.monotonic()
            self.current_cooldown = cooldown_for(verdict, self.cooldown_seconds)
        return verdict

record_success() resets the failure count and closes the circuit, but it leaves disabled alone. A disabled vendor is dropped from the candidate list. A vendor that is only cooling down stays on the list, below the healthy ones, because a vendor rate-limited ten seconds ago still beats nothing (routed_tts.py, lines 194–236).

closed (tried in order) open (cooling down, lower priority) quota, rate limit, transient, unavailable disabled (dropped from the list) auth or config process restart half-open (next try is a probe) cooldown elapses probe succeeds probe fails process restart or reset_shared_registries()
FigureOne vendor's circuit in the router, and what a restart does to it

Solid arrows are the router’s own decisions. The dashed arrows are the only ways out of disabled, and a restart takes them whether or not anyone fixed the vendor. While Rime is out of credits, the first sentence after each 900-second park is a probe. If the account is still empty, the probe fails, Rime is parked again and the failover alert fires once more. While the backups keep serving, the alert repeats once per park, and less often when calls are sparse. That is the router probing on schedule, not a stuck router. If the backups also fail on a sentence, the router tries Rime again before its park ends, because a parked vendor stays on the candidate list (routed_tts.py, lines 197–199).

The offline tests pin this behaviour. From patterns/voice-vendor-router, using the pattern’s own virtual environment (Rasa Pro 3.20.0rc1), run on 21 September 2026:

$ PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -m unittest -v \
    tests.test_voicerouter.TestHealth.test_permanent_failure_disables_and_success_does_not_revive_it \
    tests.test_voicerouter.TestHealth.test_quota_parks_far_longer_than_a_transient_error \
    tests.test_voicerouter.TestHealth.test_registry_is_shared_per_kind_so_it_outlives_a_call \
    tests.test_voicerouter.TestHealth.test_reset_clears_it \
    tests.test_voicerouter.TestCandidateSelection.test_disabled_providers_are_dropped_entirely \
    tests.test_voicerouter.TestCandidateSelection.test_cooling_providers_drop_below_healthy_ones_but_stay
test_permanent_failure_disables_and_success_does_not_revive_it (tests.test_voicerouter.TestHealth.test_permanent_failure_disables_and_success_does_not_revive_it) ... ok
test_quota_parks_far_longer_than_a_transient_error (tests.test_voicerouter.TestHealth.test_quota_parks_far_longer_than_a_transient_error) ... ok
test_registry_is_shared_per_kind_so_it_outlives_a_call (tests.test_voicerouter.TestHealth.test_registry_is_shared_per_kind_so_it_outlives_a_call) ... ok
test_reset_clears_it (tests.test_voicerouter.TestHealth.test_reset_clears_it) ... ok
test_disabled_providers_are_dropped_entirely (tests.test_voicerouter.TestCandidateSelection.test_disabled_providers_are_dropped_entirely) ... 2026-09-21 21:20:12 [info     ] voicerouter.tts.ready          providers=['a', 'b'] skipped=[]
ok
test_cooling_providers_drop_below_healthy_ones_but_stay (tests.test_voicerouter.TestCandidateSelection.test_cooling_providers_drop_below_healthy_ones_but_stay) ... 2026-09-21 21:20:12 [info     ] voicerouter.tts.ready          providers=['a', 'b'] skipped=[]
ok

----------------------------------------------------------------------
Ran 6 tests in 0.001s

OK

These are offline checks of the router’s decisions with stubbed errors. They say nothing about how a real vendor fails or what a caller hears. make test in the same directory runs the whole suite (Makefile, lines 116–117).

What to alert on

The router gives you three places to read the class:

  • Log lines: voicerouter.tts.failing_over carries the verdict, including the class. voicerouter.tts.failover names the vendor that took over. There are also voicerouter.tts.retrying_same_provider and voicerouter.tts.failed_mid_stream. On the ASR side, add Rasa’s own warning Error while streaming ASR events to the router’s voicerouter.asr.stream_failed. When a read raises inside the engine, the engine logs that warning and the router logs nothing.
  • Metrics: voicerouter/metrics.py (lines 135–163) creates an OpenTelemetry counter voicerouter.tts.failure with the attributes provider and failure_kind, plus a voicerouter.tts.failover counter. Whether they are exported depends on the deployment’s tracing configuration (lines 116–123). Key alerts on failure_kind, not on the failover count.
  • health_snapshot(): this method on the router lists each vendor’s state, failure count, last class and reopens_in. At this revision, a search of the companion finds only one caller, drill_failover.py, and nothing serves it over HTTP, so the drill is where you will see it.

What a restart resets, and what it cannot

Rasa builds a fresh engine pair for every call (voice_channel.py, lines 1493–1496):

        """Run streaming tasks and teardown for one call."""
        # Media engines (ASR/TTS)
        asr_engine, tts_engine = self._get_asr_and_tts_engines(model_metadata)
        async with asr_engine, tts_engine:

If the router kept health on its own object, every call would rediscover a dead vendor. So by default it keeps health in a module-level dictionary that lives as long as the process (health.py, lines 156–161):

# Deliberately in-process only. A shared store (Redis) would extend this across
# workers and is the obvious next step, but it introduces a dependency and a
# failure mode of its own, and process scope already removes the large majority
# of the waste.

_SHARED: dict[str, "HealthRegistry"] = {}

A restart empties that dictionary, and a redeploy starts a new process. Either way, every vendor goes back to closed, the configured order applies again and the first sentence of the next call goes to Rime. The ASR and TTS registries are separate (test_registry_is_shared_per_kind_so_it_outlives_a_call), so a TTS fault on Rime does not mark anything on the ASR side. One comment near the top of health.py (lines 13–15) still says the state “lives for the life of the call”. The code at lines 151–177 and the test say otherwise; trust those. All of this covers one process. The registry is not shared across workers or replicas, so this guide makes no claim about how several of them behave together.

The reset function exists (health.py, lines 180–187):

def reset_shared_registries() -> None:
    """Forget everything. For tests, and for an operator escape hatch.

    A provider disabled by a rejected key stays disabled for the life of the
    process, which is right until someone fixes the key — at which point there
    has to be a way to say so without a restart.
    """
    _SHARED.clear()

test_reset_clears_it shows that it re-enables a disabled vendor in the running process. It clears the whole dictionary, so the TTS and ASR registries go together, not just the vendor you fixed. Nothing calls it at runtime, though. At this revision, a search of the companion finds only the definition, the tests and drill_failover.py. There is no endpoint or command for it. And for Rime, a new key can only arrive by restarting: the engine requires RIME_API_KEY and reads it from os.environ when it builds its headers (rasa/core/channels/voice_stream/tts/rime.py, lines 69 and 145). So in this project, the working reset for a fixed key is a restart that comes after the fix.

The runbook: wait, reset, re-route or escalate

Rehearse the credits case before you need it. In examples/mantle-voice-routed-skills, run make drill SCENARIO=credits. It reads the live integrations.yml, builds the real router over stub engines and fails the first vendor from turn 3. We ran the script that target calls, PYTHONDONTWRITEBYTECODE=1 NO_COLOR=1 .venv/bin/python scripts/drill_failover.py --scenario credits, on 21 September 2026. Two edits below: the terminal’s bold codes are removed from the heading line, and the script’s closing three-line summary is cut:

Routing rules read from examples/mantle-voice-routed-skills/integrations.yml
Vendor calls are stubbed — the routing decisions below are the real ones.
2026-09-21 21:19:57 [info     ] voicerouter.tts.ready          providers=['rime', 'deepgram', 'openai'] skipped=[]

credits — Primary runs out of credits (HTTP 402) from turn 3
  chain: rime > deepgram > openai

  1. [rime            ] Hi, you're through to Northwind. This is Vela.
  2. [rime            ] One moment.
2026-09-21 21:19:57 [warning  ] voicerouter.tts.failing_over   attempt=1 from_provider=rime verdict='quota (HTTP 402) — skipping briefly' voice_changes=True
2026-09-21 21:19:57 [info     ] voicerouter.tts.failover       from_provider=rime reason='served after failover' to_provider=deepgram
  3. [deepgram        ] Your current balance is two thousand four hundred fifty dollars.   <- rime -> deepgram, failover
  4. [deepgram        ] Transferring four hundred pounds to Sam Rivera - shall I go ahead?
  5. [deepgram        ] Okay, got it.
  6. [deepgram        ] That's done. Is there anything else?

  after the call, rime: open, 1 failure(s), last was quota, retried in 900.0s

The drill starts each scenario with an empty registry for two reasons. It runs as a new process, just as a redeploy does, and drill_failover.py calls reset_shared_registries() before each scenario (line 175). So every run shows the discovery a redeployed service makes on its next call.

When a failover alert fires, work through it in order:

  1. Read the class from failure_kind or the verdict (quota, auth, rate_limit and so on). Ignore “briefly”.
  2. rate_limit or transient: wait. The router retries the same vendor or parks it for seconds. Do nothing unless the rate of these failures keeps rising, which is a capacity conversation with the vendor, not an incident.
  3. unavailable: wait and watch. The vendor is parked for the policy’s cooldown_seconds, then probed. If the backup is serving, no one needs to act tonight.
  4. quota: escalate to whoever owns the vendor account. After the top-up, the next probe closes the circuit on its own. The probe is the first sentence after the 900-second park ends, or after the vendor’s Retry-After if it sent one. A restart afterwards only brings that probe forward. A restart before the top-up only buys one more failover, with the backup serving that call.
  5. auth or config: page, fix, then restart. The vendor stays disabled until the process restarts or something calls reset_shared_registries(). Fix the key or the configuration first. For Rime, the fixed key reaches the process through its environment, so the restart delivers the fix.
  6. Re-route if the vendor will be out for a long time. Take it out of the chain so the router stops probing it. See the constraints below.
  7. Page on a TTS error that reads no TTS provider available, and on repeated Error while streaming ASR events warnings. What the caller experiences in either case was not established. voicerouter.asr.exhausted fires only for ASR errors that escape the engine, so do not rely on it for a Deepgram socket drop.

Questions from the first incident

Why not set health_scope: call, so a restart changes nothing?

policy.health_scope: call makes the router keep health on the per-call engine, which is Rasa’s own engine lifetime (voicerouter/base.py, lines 67–69). A restart then changes nothing because nothing is remembered: every call tries the failed vendor first and pays for the failover on its opening sentence. The process-scoped default exists to stop exactly that.

Will the voice switch back to Rime on its own, mid-call?

Yes, if Rime recovers. Once Rime’s park ends, is_available() returns true (health.py, lines 46–58), and _candidates() puts healthy vendors first in configured order (routed_tts.py, lines 194–236). With the default selection: order, the next sentence goes to Rime. If Rime answers, it serves from then on, so a caller can hear Vela’s own voice come back partway through a call. If Rime fails again, it is parked again and the backup carries on.

What if Rime fails halfway through a sentence?

The router does not fail over within that sentence. Once audio has started, a failure logs voicerouter.tts.failed_mid_stream with the note “sentence truncated; provider marked unhealthy” and ends the sentence (routed_tts.py, lines 348–354). The next sentence goes to the next candidate. The module docstring gives the reason: failing over mid-sentence would replay the first half in a second voice (lines 21–25). Treat failed_mid_stream as a cut-off sentence the caller may notice, not as a failover.

Who owns each alert

A failover alert that pages the same person whatever the class trains that person to silence it. Split it by class before the first incident:

  • Platform on-call owns auth and config pages, the no TTS provider available page, repeated ASR stream warnings and the restart that follows a fix. They need access to the secret store, or a named person who has it.
  • The vendor account owner owns quota. It usually needs a purchase decision, not a restart, so route it to someone who can make one.
  • Nobody is paged for rate_limit, transient or a single unavailable. Put their counts on a dashboard and review them in working hours.
  • The team that owns the voice decides whether a long outage justifies re-routing. A backup voice is a product decision: the stacks README warns that Vela’s voice changes when Rime is out.

Record who can restart the service and under what condition. For how to prove that a stop or rollback works before you rely on it, see roll out an agent with a working stop button.

Everything here is an offline test, a stubbed drill or a reading of the pinned wheel and companion source. None of it is a live Mantle run, so it shows the router’s decisions, not what a caller hears or how any vendor fails in production. The 402 is the drill’s synthetic scenario. Your paging tool and escalation policy are yours to set; this guide only says which signal each one should key on.