Key takeaways (3)
- Choose Rasa if you want the runtime to run the call and enforce the read-back from configuration.
- Choose LangGraph or Strands if you want to write and own the voice loop yourself.
- Whichever you pick, test the guard against a copy without it, and against a yes that comes too early.
You have now seen the same Cedar Clinic refill agent built on Rasa, LangGraph and Strands Agents. This chapter helps you choose between them. It starts with the choice, then shows what each framework runs for you and what you write yourself.
How to choose
Choose by what you want to own. On the 17 main recorded calls, all three guards held, so the difference is in the work you do, not the outcome.
The guard is the safety check that stops a refill going out without a clear yes, or for a second patient.
| If you want | Choose | What you get |
|---|---|---|
| The runtime to run the call, and the guard as configuration | Rasa | Turn-taking, fillers, silence check-ins and streaming come with the runtime; the read-back is a rule |
| To write the loop, with a graph, checkpoints and a built-in pause | LangGraph | A pause that is part of the run, resumed by your loop with the caller’s answer |
| To write the loop, with typed checks around each tool call | Strands Agents | Deny and Confirm interventions around a model-driven loop |
Choose Rasa if you want to write as little voice code as possible. The voice loop is the code that listens, takes turns and speaks. In Rasa it is a channel block in YAML, and the runtime runs it. With a built-in speech vendor such as Deepgram, you write no speech adapter at all. The read-back guard is a rule in the skill, which the runtime enforces. A skill is a task the agent knows how to do, written as instructions.
The turn loop belongs to the runtime. So a rule about timing goes in a channel subclass and a tool, rather than in a loop you edit.
Choose LangGraph if you want to write the voice loop yourself and like
working with a graph. You get checkpoints, and a pause, interrupt(), that
holds the run until your loop resumes it. You write the loop in Python and
own every line. Barge-in, the caller talking over the agent, and a speech
cache are yours to add to the loop. This build has neither.
Choose Strands Agents if you want to write the loop yourself and prefer
typed checks around each tool call. The guard returns Deny to refuse a send
or Confirm to pause it for the caller. As with LangGraph, the loop is
yours, and so is every change to it.
You can also put an agent framework behind a voice framework that supplies the loop. LiveKit Agents, for example, documents a plugin that runs a LangGraph workflow as its agent’s model. This series uses LangGraph and Strands on their own, so you see the loop you would write.
Whichever you choose, do two things. Keep a copy of your agent with the guard removed, and check that your tests fail on it. And test the guard with a yes that arrives before the read-back has finished playing. If you choose Rasa, the voice agent tutorial builds a voice agent with Deepgram for speech in and out.
What the runtime gives you, and what you write
- In Rasa the voice loop is configuration. With the fix, a subclass of the browser audio channel records when the caller began speaking and when the read-back finished playing.
- The guard is a rule in the skill. Tools write the memory it reads.
- With the fix, the send tool refuses an answer that began before the read-back finished playing.
- In LangGraph and Strands the voice loop is Python you write. The fix’s timing check sits in it, before the paused run is resumed.
- The guard wraps the send tool, and your loop has to know when the run is paused for the read-back.
Job by job, from chapters 2 to 4:
| Job | Rasa | LangGraph | Strands Agents |
|---|---|---|---|
| Browser protocol | The built-in browser audio channel | A WebSocket server you write | A WebSocket server you write |
| End of turn | The speech engine holds Speechmatics’ pieces of text until its end-of-utterance event, then passes them on as one turn | One end-of-utterance event per caller turn, queued | One agent turn per end of utterance |
| Streaming the model’s reply | Into any speech engine that accepts streamed text | Cut at sentence ends in your loop | Cut at sentence ends in your loop |
| Fillers | Written by the model, spoken by the runtime | A fixed phrase per tool | Spoken after the tools finish, if nothing was said |
| Playback markers | The runtime | Written in your loop | Written in your loop |
| Silence check-in | A timeout setting, and a built-in skill that checks in three times, then hangs up | Written in your loop, after 30 seconds | Written in your loop, after 30 seconds |
| Barge-in | The runtime (off in this build; beta when on) | Not written | Not written |
| Pause for the read-back | Held by the runtime | interrupt() in middleware, resumed by your loop | A Confirm intervention, resumed by your loop |
| Who judges the yes | Rasa’s main model | A separate model call in the guard | A fixed rule in code |
A filler is a short message such as “One moment while I check your details.” A playback marker tells the server when the caller’s browser has finished playing a message. The main model is the model that runs each of the agent’s turns; Rasa calls it the orchestrator.
How much code each build needed
Each number below counts the non-blank lines that are not only comments, with Python docstrings left out, grouped by the job each file declares.
| Job | Rasa | LangGraph | Strands Agents |
|---|---|---|---|
| Voice loop | 58 lines of YAML (34 with Deepgram) | 315 lines of Python | 261 lines of Python |
| Speech adapter | None for Deepgram; 213 for Speechmatics | Shared clients: 184 for Deepgram, 251 for Speechmatics | The same shared clients |
| Read-back guard | 47, of which 26 are configuration: the rule, both memory files, the responses and the tool code | 112 | 92 |
The counts say how much was written, not how long it took. Rasa’s built-in Deepgram listener takes no extra vocabulary, so the medicine
names cannot be given to it. No build used vocabulary on Deepgram: the
shared Deepgram clients send none either. Run make count in the companion to print every count by job.
Pick a speech engine that accepts streamed text
Rasa streams the model’s reply into the text-to-speech engine when the engine accepts streamed text. Every built-in engine does. With Deepgram, Rasa streams the text straight into the speech socket.
A custom engine that takes whole utterances, such as the Speechmatics preview here, makes Rasa wait for each whole reply. LangGraph and Strands stream the model in their own loops and send the speech engine a sentence at a time.
So for Rasa the advice is short. Use a built-in engine, or give a custom one
streaming_input and send_text_chunk, as
chapter 2
shows.
Can you keep Speechmatics and still stream in Rasa?
If the vendor’s text-to-speech accepts streamed text, write the engine the
way the built-in ones are written. Set streaming_input to True and
implement send_text_chunk. The Speechmatics preview used here takes one
whole utterance per HTTP request, so this build’s engine could not.
The companion’s deepgram-tts variant keeps Speechmatics for listening and
uses Rasa’s built-in Deepgram engine for speaking. A test checks that
nothing else differs.
Where the timing fix goes in each build
Chapter 5 adds an opt-in fix to each build. It counts a yes only if the caller began it after the read-back finished playing.
| Build | Where the check runs |
|---|---|
| Rasa | In the send tool, using timings that a subclass of the browser audio channel records |
| LangGraph | In your voice loop, before it resumes the paused run |
| Strands Agents | In your voice loop, before it answers the pending Confirm |
All three fixed builds stopped every early yes in the replays. They also let every normal yes through on the live calls made with them. The fix is tested on replays and 11 live calls, not proven.
Details
- Who wrote the builds: an AI coding agent wrote all three. The LangGraph and Strands builds were committed after the Rasa one. The LangGraph build was written against the finished voice protocol and test runner.
- Speech adapters: Rasa’s Speechmatics engine is a copy of adapters the companion had already written and tested live. The shared Deepgram clients were written for this comparison.
- Prompt text: Rasa restates the shared instructions in files its runtime reads, and those lines count as Rasa’s. LangGraph and Strands import the text, so it counts for neither.
- Barge-in: off in the Rasa build, as it is by default in this release, and not written in the other two.
- The model endpoint: LangGraph and Strands both needed OpenAI’s
Responses API for this model with tools at low reasoning effort. In LangChain that is one documented
argument,
use_responses_api=True. In Strands it is a different model class.