Guide · Platform / operations engineer
Why TTS failover returns after a Rasa voice agent restart
A redeploy clears the voice router’s memory of a failing vendor; it never fixes the vendor. Read the failure class and pick wait, reset, re-route or escalate.
Key takeaways (6)
- A failover alert says a vendor failed. The failure class in the same log line says whether waiting, a fix or a person is needed, so route alerts on the class.
- In the companion voice router, a rejected key or a broken request disables the vendor until the process restarts. A later success never re-enables it.
- A restart empties the router’s process-wide health registry. Unless the vendor was fixed first, the next call finds the same fault and fails over again.
- The log text "skipping briefly" is not the parking time. A quota failure with no Retry-After header parks the vendor for 900 seconds.
- In the companion, reset_shared_registries() is called only by the tests and the drill. No endpoint exposes it, so in practice the reset is a restart after the fix.
- What Rasa 3.20.0rc1 does with a TTS error raised mid-stream by these engines was not established. Do not build a runbook on the caller hearing silence.
The alert shows two log lines. They come from an offline drill in the companion project, with the vendor calls stubbed, in which the primary text-to-speech (TTS) vendor starts returning HTTP 402 (out of credits) on the third sentence of a call:
2026-09-21 21:19:57 [warning ] voicerouter.tts.failing_over attempt=1 from_provider=rime verdict='quota (HTTP 402) — skipping briefly' voice_changes=True
2026-09-21 21:19:57 [info ] voicerouter.tts.failover from_provider=rime reason='served after failover' to_provider=deepgram
Picture the on-call engineer who sees those lines. The agent has switched to a backup voice, and something looks stuck, so they redeploy the voice service. The alerts stop. On the next real call, the router tries Rime first again. Rime returns 402 again, the same two lines fire and the pager goes off again.
Nothing about Rime’s account changed. The redeploy wiped the router’s in-process record that Rime was out of credits, a record that had been keeping later calls off Rime for 900 seconds, so the next conversation paid for the discovery again. The scene is illustrative; the log lines and the 900 seconds come from the drill run described below.
The signal worth routing is the failure’s class, not the fact that a vendor failed. Page a person on a failure that only a person can fix. Let the router handle a rate limit. Redeploy only after the fix, because a redeploy on its own resets the router’s memory and does nothing to the vendor. The cost of this stance is work up front: your alerting has to read a field from the log or metric rather than count failovers, and someone has to own each class before the first incident.
This guide uses the voice router in RasaHQ/rasa-community-resources at
revision 69e27b6: the pattern patterns/voice-vendor-router and the example
agent examples/mantle-voice-routed-skills, whose assistant Vela answers for
the companion’s seeded demo bank (its README, lines 207 and 212). Both pin Rasa
Pro 3.20.0rc1 (rasa-pro==3.20.0rc1 on line 8 of each pyproject.toml). The
router applies the circuit-breaker pattern to speech vendors, and it honours the
HTTP Retry-After header when a vendor sends one. Rime, Deepgram and OpenAI
appear because they are the example’s configured TTS chain. Nothing here
measures any vendor’s reliability.
Why a Rasa voice agent needs a router to fail over
A Rasa voice channel holds one ASR (speech-to-text) engine and one TTS engine.
The example’s integrations.yml names voicerouter.RoutedTTS as its TTS
engine, and that router holds a chain of real engines (lines 63–90):
rime, then deepgram, then openai. Rasa itself does not walk a chain.
Mantle speaks a response through TurnContext.send, which calls the channel’s
send_text_message. On that path, the handler that catches a vendor error is
this (rasa/core/channels/voice_stream/voice_channel.py, lines 619–625, in the
3.20.0rc1 wheel). We did not examine the separate streaming-response path
(send_response_chunk and its background audio sender), which has its own
error handling:
try:
audio_stream = self.tts_engine.synthesize(text)
except TTSError as e:
logger.error("voice_channel.tts_synthesis_error", error=str(e))
voice_tracing.get_current_synthesize_speech_recorder().record_error(e)
# TODO: add message that works without tts, e.g. loading from disc
audio_stream = self.chunk_audio(generate_silence(self.audio_format))
That code logs the error and plays silence, and it tries no other vendor. Its
reach is narrower than it looks. It catches only an error raised by calling
synthesize(). In the base engine, in Rasa’s RimeTTS and DeepgramTTS
(tts/deepgram.py, lines 184–195) and in RoutedTTS, synthesize is an async
generator, and calling an async generator runs none of
its body. Their errors surface later, at line 677, where
_stream_audio_to_channel reads the stream, outside that try. We checked this
with a short offline probe, run with the router pattern’s virtual environment.
It is not part of the companion. <companion> stands for your checkout path,
and the vendor is a stub that always raises:
import asyncio, inspect, sys
sys.path.insert(0, "<companion>/patterns/voice-vendor-router")
from rasa.core.channels.voice_stream.tts.tts_engine import TTSError
from rasa.core.channels.voice_stream.tts.rime import RimeTTS
from voicerouter.routed_tts import RoutedTTS
from voicerouter.base import BuiltProvider, ProviderSpec, RouterPolicy
from voicerouter.health import reset_shared_registries
print("RimeTTS.synthesize is async generator:", inspect.isasyncgenfunction(RimeTTS.synthesize))
print("RoutedTTS.synthesize is async generator:", inspect.isasyncgenfunction(RoutedTTS.synthesize))
class Dead:
async def connect(self): pass
async def synthesize(self, text, config=None):
raise TTSError("401 Unauthorized"); yield b""
reset_shared_registries()
r = RoutedTTS([BuiltProvider(ProviderSpec(name="rime", label="rime", config={}), Dead())], RouterPolicy())
try:
stream = r.synthesize("hello") # the call Rasa wraps in try/except TTSError
print("calling synthesize(): no exception raised")
except TTSError as e:
print("calling synthesize(): raised", e)
async def consume():
try:
async for _ in stream: pass
except TTSError as e:
print("iterating the stream: raised TTSError:", str(e)[:90])
asyncio.run(consume())
Its output, with the router’s two log lines filtered out:
RimeTTS.synthesize is async generator: True
RoutedTTS.synthesize is async generator: True
calling synthesize(): no exception raised
iterating the stream: raised TTSError: voicerouter: no TTS provider available — rime: disabled (auth) — 1 attempted for this utte
We traced that error up through send_text_message (lines 1038–1122) and
Mantle’s TurnContext.send (rasa/mantle/orchestration/turn_context.py, lines
290–299). Neither catches it, and we did not establish what the turn or the call
does after that. Treat “every TTS vendor is out” as an unknown outcome for the
caller, not as silence you can plan around.
For ASR there are two layers, and the first one matters more. Rasa’s base ASR
engine reads the vendor socket inside its own try
(rasa/core/channels/voice_stream/asr/asr_engine.py, lines 353–364):
async def stream_asr_events(self) -> AsyncIterator[ASREvent]:
"""Stream the events returned by the ASR system as it is fed audio bytes."""
if self.asr_socket is None:
raise ConnectionException("Websocket not connected.")
try:
async for message in self.asr_socket:
asr_event = self.engine_event_to_asr_event(message)
if asr_event:
yield asr_event
except Exception as e:
logger.warning(f"Error while streaming ASR events: {e}")
A socket error during a read is logged at warning level, and the stream then
ends as if the call were over. DeepgramASR does not override this method. Only
an error that escapes the engine, such as the ConnectionException above,
reaches the voice channel’s own handler, which logs it at error level and
re-raises (voice_channel.py, lines 1372–1380):
except Exception as e:
logger.error(
"voice_channel.receive_asr_events.error",
call_id=call_parameters.call_id,
error=repr(e),
)
# Close the active transcription span as failed before re-raising.
voice_tracing.get_current_transcribe_audio_recorder().record_error(e).end()
raise
Of the engine’s own errors, the router sees only the ones that escape the
engine, just as that channel handler does. RoutedASR treats a stream that ends without
an error as the end of the call (routed_asr.py, lines 247–250), so a read
error the engine has already swallowed never reaches it. An offline probe (a
fake Deepgram socket that raises ConnectionError on the first read, no
network) shows both cases. Case A is a read that raises; case B is a socket
that was never connected. We cut four start-up lines from the output: Rasa’s
asr.engine.initialized line, the engine class: DeepgramASR line and the two
voicerouter.asr.ready lines. The rest is unedited:
DeepgramASR overrides stream_asr_events: False
2026-09-21 21:32:18 [warning ] Error while streaming ASR events: socket dropped mid-call (probe)
A: socket drops during a read -> RoutedASR.stream_asr_events ended without raising
A: socket drops during a read -> deepgram health: failures = 0 | active provider: deepgram
2026-09-21 21:32:18 [warning ] voicerouter.asr.stream_failed error='Websocket not connected.' provider=deepgram verdict='unknown (ConnectionException) — skipping briefly'
2026-09-21 21:32:18 [info ] voicerouter.asr.failover from_provider=deepgram reason='connected after failover' to_provider=backup
2026-09-21 21:32:18 [info ] voicerouter.asr.connected provider=backup
2026-09-21 21:32:18 [info ] voicerouter.asr.resumed note='audio sent during the failure was not transcribed' provider=backup
B: socket was never connected -> RoutedASR.stream_asr_events ended without raising
B: socket was never connected -> deepgram health: failures = 1 | active provider: backup
In case A the router records nothing and does not fail over, so
voicerouter.asr.stream_failed and voicerouter.asr.exhausted never fire when
a read raises inside the engine. The probe does not cover a real socket that
closes cleanly. What the call does after the stream ends was not
established.
Read the failure class, not the failover
Every vendor error the router sees goes through classify() in
voicerouter/failures.py (lines 206–263), which turns an HTTP status, an AWS
error code or the message text into a class. The class decides two things: how long the vendor
is skipped (lines 70–76) and whether the caller may hear a different voice
(lines 88–93):
DEFAULT_COOLDOWNS: dict[FailureKind, float] = {
FailureKind.AUTH: 0.0, # permanent; cooldown unused
FailureKind.CONFIG: 0.0, # permanent; cooldown unused
FailureKind.QUOTA: 900.0, # 15 minutes
FailureKind.RATE_LIMIT: 20.0, # overridden by Retry-After when present
FailureKind.TRANSIENT: 15.0,
}
_JUSTIFIES_VOICE_CHANGE = {
FailureKind.AUTH,
FailureKind.CONFIG,
FailureKind.QUOTA,
FailureKind.UNAVAILABLE,
}
classify() supplies the triggers in the table below. The “waiting” column
paraphrases the module docstring (lines 6–10). When a verdict carries the
vendor’s Retry-After, that value replaces the default, with a floor of one
second (lines 266–274). Quota, rate-limit and transient verdicts can carry it;
unavailable verdicts never do (lines 248 and 261). Classes with no default use
the policy’s cooldown_seconds, which the example sets to 60.
| Class | Typical trigger | Waiting fixes it? | Router skips the vendor for | Voice changes? |
|---|---|---|---|---|
auth | HTTP 401 or 403 | Never | Until restart or reset | Yes |
config | HTTP 400, 404, 405, 415 or 422 | Never | Until restart or reset | Yes |
quota | HTTP 402, or a 429 with quota text | After a top-up | 900 s, or Retry-After | Yes |
unavailable | Connection refused, timeout | Often | cooldown_seconds (60 here) | Yes |
rate_limit | HTTP 429 | Yes, in seconds | 20 s, or Retry-After | After one same-vendor retry |
transient | HTTP 5xx | Yes, in seconds | 15 s, or Retry-After | After one same-vendor retry |
The last column comes from routed_tts.py, lines 356–377. A retryable failure
gets one more attempt on the same vendor (same_provider_retries: 1, after
retry_backoff_ms: 250 in the example) as long as the wait is at most one
second. If that retry fails too, or the vendor asks for seven seconds, the
router moves on. Anything classify() cannot place is unknown: it is treated
like transient, with the policy’s cooldown_seconds as its park.
Where the router keeps circuit state
Each vendor has a ProviderHealth record. The branch that matters for on-call
work is in record_failure (voicerouter/health.py, lines 75–85):
if is_permanent(verdict):
# No threshold for these. One 401 is enough to know.
self.disabled = True
self.opened_at = time.monotonic()
self.current_cooldown = float("inf")
return verdict
if self.consecutive_failures >= self.failure_threshold:
self.opened_at = time.monotonic()
self.current_cooldown = cooldown_for(verdict, self.cooldown_seconds)
return verdict
record_success() resets the failure count and closes the circuit, but it
leaves disabled alone. A disabled vendor is dropped from the candidate list.
A vendor that is only cooling down stays on the list, below the healthy ones,
because a vendor rate-limited ten seconds ago still beats nothing
(routed_tts.py, lines 194–236).
Solid arrows are the router’s own decisions. The dashed arrows are the only
ways out of disabled, and a restart takes them whether or not anyone fixed the
vendor. While Rime is out of credits, the first sentence after each 900-second
park is a probe. If the account is still empty, the probe fails, Rime is parked
again and the failover alert fires once more. While the backups keep serving,
the alert repeats once per park, and less often when calls are sparse. That is
the router probing on schedule, not a stuck router. If the backups also fail on
a sentence, the router tries Rime again before its park ends, because a parked
vendor stays on the candidate list (routed_tts.py, lines 197–199).
The offline tests pin this behaviour. From patterns/voice-vendor-router, using
the pattern’s own virtual environment (Rasa Pro 3.20.0rc1), run on 21 September
2026:
$ PYTHONDONTWRITEBYTECODE=1 .venv/bin/python -m unittest -v \
tests.test_voicerouter.TestHealth.test_permanent_failure_disables_and_success_does_not_revive_it \
tests.test_voicerouter.TestHealth.test_quota_parks_far_longer_than_a_transient_error \
tests.test_voicerouter.TestHealth.test_registry_is_shared_per_kind_so_it_outlives_a_call \
tests.test_voicerouter.TestHealth.test_reset_clears_it \
tests.test_voicerouter.TestCandidateSelection.test_disabled_providers_are_dropped_entirely \
tests.test_voicerouter.TestCandidateSelection.test_cooling_providers_drop_below_healthy_ones_but_stay
test_permanent_failure_disables_and_success_does_not_revive_it (tests.test_voicerouter.TestHealth.test_permanent_failure_disables_and_success_does_not_revive_it) ... ok
test_quota_parks_far_longer_than_a_transient_error (tests.test_voicerouter.TestHealth.test_quota_parks_far_longer_than_a_transient_error) ... ok
test_registry_is_shared_per_kind_so_it_outlives_a_call (tests.test_voicerouter.TestHealth.test_registry_is_shared_per_kind_so_it_outlives_a_call) ... ok
test_reset_clears_it (tests.test_voicerouter.TestHealth.test_reset_clears_it) ... ok
test_disabled_providers_are_dropped_entirely (tests.test_voicerouter.TestCandidateSelection.test_disabled_providers_are_dropped_entirely) ... 2026-09-21 21:20:12 [info ] voicerouter.tts.ready providers=['a', 'b'] skipped=[]
ok
test_cooling_providers_drop_below_healthy_ones_but_stay (tests.test_voicerouter.TestCandidateSelection.test_cooling_providers_drop_below_healthy_ones_but_stay) ... 2026-09-21 21:20:12 [info ] voicerouter.tts.ready providers=['a', 'b'] skipped=[]
ok
----------------------------------------------------------------------
Ran 6 tests in 0.001s
OK
These are offline checks of the router’s decisions with stubbed errors. They say
nothing about how a real vendor fails or what a caller hears. make test in the
same directory runs the whole suite (Makefile, lines 116–117).
What to alert on
The router gives you three places to read the class:
- Log lines:
voicerouter.tts.failing_overcarries the verdict, including the class.voicerouter.tts.failovernames the vendor that took over. There are alsovoicerouter.tts.retrying_same_providerandvoicerouter.tts.failed_mid_stream. On the ASR side, add Rasa’s own warningError while streaming ASR eventsto the router’svoicerouter.asr.stream_failed. When a read raises inside the engine, the engine logs that warning and the router logs nothing. - Metrics:
voicerouter/metrics.py(lines 135–163) creates an OpenTelemetry countervoicerouter.tts.failurewith the attributesproviderandfailure_kind, plus avoicerouter.tts.failovercounter. Whether they are exported depends on the deployment’s tracing configuration (lines 116–123). Key alerts onfailure_kind, not on the failover count. health_snapshot(): this method on the router lists each vendor’s state, failure count, last class andreopens_in. At this revision, a search of the companion finds only one caller,drill_failover.py, and nothing serves it over HTTP, so the drill is where you will see it.
What a restart resets, and what it cannot
Rasa builds a fresh engine pair for every call
(voice_channel.py, lines 1493–1496):
"""Run streaming tasks and teardown for one call."""
# Media engines (ASR/TTS)
asr_engine, tts_engine = self._get_asr_and_tts_engines(model_metadata)
async with asr_engine, tts_engine:
If the router kept health on its own object, every call would rediscover a dead
vendor. So by default it keeps health in a module-level dictionary that lives
as long as the process (health.py, lines 156–161):
# Deliberately in-process only. A shared store (Redis) would extend this across
# workers and is the obvious next step, but it introduces a dependency and a
# failure mode of its own, and process scope already removes the large majority
# of the waste.
_SHARED: dict[str, "HealthRegistry"] = {}
A restart empties that dictionary, and a redeploy starts a new process. Either
way, every vendor goes back to closed, the configured order applies again and
the first sentence of the next call goes to Rime. The ASR and TTS registries are
separate (test_registry_is_shared_per_kind_so_it_outlives_a_call), so a TTS
fault on Rime does not mark anything on the ASR side. One comment near the top
of health.py (lines 13–15) still says the state “lives for the life of the
call”. The code at lines 151–177 and the test say otherwise; trust those. All
of this covers one process. The registry is not shared across workers or
replicas, so this guide makes no claim about how several of them behave
together.
The reset function exists (health.py, lines 180–187):
def reset_shared_registries() -> None:
"""Forget everything. For tests, and for an operator escape hatch.
A provider disabled by a rejected key stays disabled for the life of the
process, which is right until someone fixes the key — at which point there
has to be a way to say so without a restart.
"""
_SHARED.clear()
test_reset_clears_it shows that it re-enables a disabled vendor in the running
process. It clears the whole dictionary, so the TTS and ASR registries go
together, not just the vendor you fixed. Nothing calls it at runtime, though. At
this revision, a search of the companion finds only the definition, the tests
and drill_failover.py. There is no endpoint or command for it. And for Rime, a new key can only arrive by restarting: the
engine requires RIME_API_KEY and reads it from os.environ when it builds
its headers (rasa/core/channels/voice_stream/tts/rime.py, lines 69 and 145).
So in this project, the working reset for a fixed key is a restart that comes
after the fix.
The runbook: wait, reset, re-route or escalate
Rehearse the credits case before you need it. In
examples/mantle-voice-routed-skills, run make drill SCENARIO=credits. It
reads the live integrations.yml, builds the real router over stub engines and
fails the first vendor from turn 3. We ran the script that target calls,
PYTHONDONTWRITEBYTECODE=1 NO_COLOR=1 .venv/bin/python scripts/drill_failover.py --scenario credits,
on 21 September 2026. Two edits below: the terminal’s bold codes are removed
from the heading line, and the script’s closing three-line summary is cut:
Routing rules read from examples/mantle-voice-routed-skills/integrations.yml
Vendor calls are stubbed — the routing decisions below are the real ones.
2026-09-21 21:19:57 [info ] voicerouter.tts.ready providers=['rime', 'deepgram', 'openai'] skipped=[]
credits — Primary runs out of credits (HTTP 402) from turn 3
chain: rime > deepgram > openai
1. [rime ] Hi, you're through to Northwind. This is Vela.
2. [rime ] One moment.
2026-09-21 21:19:57 [warning ] voicerouter.tts.failing_over attempt=1 from_provider=rime verdict='quota (HTTP 402) — skipping briefly' voice_changes=True
2026-09-21 21:19:57 [info ] voicerouter.tts.failover from_provider=rime reason='served after failover' to_provider=deepgram
3. [deepgram ] Your current balance is two thousand four hundred fifty dollars. <- rime -> deepgram, failover
4. [deepgram ] Transferring four hundred pounds to Sam Rivera - shall I go ahead?
5. [deepgram ] Okay, got it.
6. [deepgram ] That's done. Is there anything else?
after the call, rime: open, 1 failure(s), last was quota, retried in 900.0s
The drill starts each scenario with an empty registry for two reasons. It runs
as a new process, just as a redeploy does, and drill_failover.py calls
reset_shared_registries() before each scenario (line 175). So every run shows
the discovery a redeployed service makes on its next call.
When a failover alert fires, work through it in order:
- Read the class from
failure_kindor the verdict (quota,auth,rate_limitand so on). Ignore “briefly”. rate_limitortransient: wait. The router retries the same vendor or parks it for seconds. Do nothing unless the rate of these failures keeps rising, which is a capacity conversation with the vendor, not an incident.unavailable: wait and watch. The vendor is parked for the policy’scooldown_seconds, then probed. If the backup is serving, no one needs to act tonight.quota: escalate to whoever owns the vendor account. After the top-up, the next probe closes the circuit on its own. The probe is the first sentence after the 900-second park ends, or after the vendor’sRetry-Afterif it sent one. A restart afterwards only brings that probe forward. A restart before the top-up only buys one more failover, with the backup serving that call.authorconfig: page, fix, then restart. The vendor stays disabled until the process restarts or something callsreset_shared_registries(). Fix the key or the configuration first. For Rime, the fixed key reaches the process through its environment, so the restart delivers the fix.- Re-route if the vendor will be out for a long time. Take it out of the chain so the router stops probing it. See the constraints below.
- Page on a TTS error that reads
no TTS provider available, and on repeatedError while streaming ASR eventswarnings. What the caller experiences in either case was not established.voicerouter.asr.exhaustedfires only for ASR errors that escape the engine, so do not rely on it for a Deepgram socket drop.
Questions from the first incident
Why not set health_scope: call, so a restart changes nothing?
policy.health_scope: call makes the router keep health on the per-call engine,
which is Rasa’s own engine lifetime (voicerouter/base.py, lines 67–69). A
restart then changes nothing because nothing is remembered: every call tries
the failed vendor first and pays for the failover on its opening sentence. The
process-scoped default exists to stop exactly that.
Will the voice switch back to Rime on its own, mid-call?
Yes, if Rime recovers. Once Rime’s park ends, is_available() returns true
(health.py, lines 46–58), and _candidates() puts healthy vendors first in
configured order (routed_tts.py, lines 194–236). With the default
selection: order, the next sentence goes to Rime. If Rime answers, it serves from then on,
so a caller can hear Vela’s own voice come back partway through a call. If Rime
fails again, it is parked again and the backup carries on.
What if Rime fails halfway through a sentence?
The router does not fail over within that sentence. Once audio has started, a
failure logs voicerouter.tts.failed_mid_stream with the note “sentence
truncated; provider marked unhealthy” and ends the sentence (routed_tts.py,
lines 348–354). The next sentence goes to the next candidate. The module
docstring gives the reason: failing over mid-sentence would replay the first
half in a second voice (lines 21–25). Treat failed_mid_stream as a cut-off
sentence the caller may notice, not as a failover.
Who owns each alert
A failover alert that pages the same person whatever the class trains that person to silence it. Split it by class before the first incident:
- Platform on-call owns
authandconfigpages, theno TTS provider availablepage, repeated ASR stream warnings and the restart that follows a fix. They need access to the secret store, or a named person who has it. - The vendor account owner owns
quota. It usually needs a purchase decision, not a restart, so route it to someone who can make one. - Nobody is paged for
rate_limit,transientor a singleunavailable. Put their counts on a dashboard and review them in working hours. - The team that owns the voice decides whether a long outage justifies re-routing. A backup voice is a product decision: the stacks README warns that Vela’s voice changes when Rime is out.
Record who can restart the service and under what condition. For how to prove that a stop or rollback works before you rely on it, see roll out an agent with a working stop button.
Everything here is an offline test, a stubbed drill or a reading of the pinned wheel and companion source. None of it is a live Mantle run, so it shows the router’s decisions, not what a caller hears or how any vendor fails in production. The 402 is the drill’s synthetic scenario. Your paging tool and escalation policy are yours to set; this guide only says which signal each one should key on.