Key takeaways (3)
- LangGraph gives you the agent loop, state for each call, and a pause that waits for the caller's answer.
- You write the voice loop yourself: audio in and out, turn order, fillers and the silence check-in.
- Count the caller's yes only after the read-back question has finished playing.
You build the Cedar Clinic prescription-refill voice agent on LangGraph. The caller gives their name and date of birth, then names a medicine. The agent reads the medicine back and sends a refill request only after a clear yes.
LangGraph is LangChain’s low-level orchestration runtime for long-running,
stateful agents. LangGraph’s docs recommend starting one level up, with
LangChain’s create_agent, which runs on LangGraph. This build does that.
You get a working agent loop, saved state for each call and a built-in way to pause for the caller. You write the voice loop: the code that listens, takes turns and speaks. The chapter walks through both.
What you need. Python 3.11 or 3.12 with uv, an OpenAI API key and a Speechmatics API
key. LangGraph and LangChain are MIT-licensed, so there is no licence key.
Live calls are billed by OpenAI and Speechmatics. The offline tests are free.
Get the code at the commit this series uses:
git clone https://github.com/RasaHQ/rasa-community-resources
cd rasa-community-resources
git checkout 41dd184955425f1d1686cdb39c91a0fcc1829442
cd tutorials/voice-agent-three-frameworks/langgraph
The steps below read the companion’s finished build, one part at a time. To
start your own project, copy the four Python files in the langgraph folder:
agent.py, guard.py, voice_loop.py and server.py. They import two shared
packages from the companion. Replace cedar_clinic, the clinic’s business
rules, with your own. Keep cedar_speech, the Speechmatics clients, or swap
in your speech vendor.
What LangGraph gives you, and what you write
| Part of the agent | What LangGraph and LangChain give you | What you write |
|---|---|---|
| The agent loop | create_agent: the model calls tools until it is done | The prompt and the list of tools |
| State for each call | A checkpointer, and state fields the model cannot set | Which fields to keep, and the tools that set them |
| The confirmation step | Middleware hooks, interrupt() and resume | When to ask, and how to judge the reply |
| The voice loop | Streaming of model text, tool calls and custom events | Audio, turn order, speech, fillers and silence |
| Barge-in (the caller talking over the agent) | No audio parts, so nothing built in | Not written in this build |
The business rules live in a shared Python package, cedar_clinic. It looks
up records, applies the clinic’s rules and writes the audit log: the clinic’s
own record of every tool call and outcome. All three builds call the same
package, so this chapter only covers the LangGraph code around it.
Step 1: Define the tools and the state they keep
The agent has five tools. Each one calls cedar_clinic and returns its result
to the model. This excerpt from
langgraph/agent.py
builds the agent:
def build_agent(model: Any = None, *, classify: Any = None, checkpointer: Any = None):
"""The compiled graph. One checkpointer per process; thread_id is the conversation id."""
model = model if model is not None else make_model(streaming=True)
return create_agent(
model=model,
tools=TOOLS,
system_prompt=SYSTEM_PROMPT,
# concern-begin: refill-guard
middleware=[RefillGuard(classify or llm_classifier(make_model()))],
# concern-end
checkpointer=checkpointer or InMemorySaver(),
)
The checkpointer keeps each call’s history under a thread id. The voice loop
sets that id to the conversation id, so every call has its own history. The
prompt comes from cedar_clinic, so it is the same text in all three builds.
Each tool gets a ToolRuntime, which gives it the call’s state. It returns a
Command that updates that state. This excerpt is the tool that checks the
caller’s name and date of birth:
@tool("verify_patient", **_spec("verify_patient"))
async def verify_patient(full_name: str, date_of_birth: str, runtime: ToolRuntime) -> Command:
result = clinic.verify_patient(_conversation_id(runtime), full_name, date_of_birth)
update: dict = {}
# concern-begin: refill-guard
# The patient id goes to state, never to the model. The first verified
# patient stays: a call verified as one patient cannot become another.
if result["status"] == "verified":
current = runtime.state.get("patient_id")
if not current:
update = {"patient_id": result["patient_id"], "patient_first_name": result["first_name"]}
elif current != result["patient_id"]:
result = {"status": "not_verified", "reason": "already_verified_as_another_patient",
"next_step": "This call is verified for a different patient. Do not act for this one."}
# concern-end
return _reply(runtime, clinic.for_model(result), **update)
This is part of the guard: the safety check that stops a refill going out without a clear yes, or for a second patient. The first verified patient stays for the whole call. A second patient on the same call is refused.
The select_medication tool works the same way. It writes the selected record
and its label to state, and a new selection always replaces the old one.
Keep these fields out of the model’s reach. This excerpt from
langgraph/guard.py
declares them:
class RefillState(AgentState):
"""Agent state plus what only tools write. Private: not settable from the graph's input."""
patient_id: NotRequired[Annotated[str, PrivateStateAttr]]
patient_first_name: NotRequired[Annotated[str, PrivateStateAttr]]
selected_record_id: NotRequired[Annotated[str, PrivateStateAttr]]
selected_label: NotRequired[Annotated[str, PrivateStateAttr]]
PrivateStateAttr leaves a field out of the graph’s input and output. So
nothing outside the tools can set the patient or the selected medicine. No
tool takes a patient id from the model either.
Step 2: Add the confirmation step
The agent must read the medicine back and wait for a yes before it sends. In
this build, that is middleware. The middleware hooks belong to LangChain’s
create_agent. The guard uses two of them.
The first hook hides the send tool until a medicine is selected. This is an
excerpt from guard.py:
async def awrap_model_call(self, request: ModelRequest, handler: Callable) -> Any:
"""Hide send_refill_request until a record entry is selected."""
if not request.state.get("selected_record_id"):
request = request.override(tools=[t for t in request.tools if getattr(t, "name", None) != SEND])
return await handler(request)
The second hook runs around every tool call. For the send tool, it checks the
record, then pauses the run with interrupt(). This excerpt is the rest of
the guard class
(guard.py):
async def awrap_tool_call(self, request: ToolCallRequest, handler: Callable) -> Any:
call = request.tool_call
if call["name"] != SEND:
return await handler(request)
state = request.state
selected = state.get("selected_record_id") or ""
record_id = str(call["args"].get("record_id") or "").strip().upper()
if not selected or record_id != selected:
return _blocked(call, "medication_not_resolved",
"Call select_medication first and send only the record_id it returned.")
question = clinic.confirmation_question(state.get("selected_label") or "")
# Everything above runs again on resume and has no side effects.
resumed = interrupt({"kind": "confirm_refill", "question": question, "record_id": record_id})
answer = str(resumed.get("text") if isinstance(resumed, dict) else resumed or "").strip()
confirmed = bool(answer) and await self.classify(question, answer)
conversation_id = request.runtime.config["configurable"]["thread_id"]
clinic.record_confirmation(conversation_id, record_id, confirmed, mechanism=MECHANISM,
question=question, answer=answer)
if not confirmed:
request.runtime.stream_writer({"say": clinic.DECLINED_TEXT})
return ToolMessage(json.dumps({
"status": "declined", "effects": 0, "caller_answer": answer,
"caller_was_told": clinic.DECLINED_TEXT,
"next_step": ("Nothing was sent and the caller has been told so; do not repeat it. If the caller "
"named a different medicine, call select_medication with it. Otherwise answer "
"what they said or ask what else they need."),
}), tool_call_id=call["id"], name=SEND)
# A progress event for the voice loop, which may say a filler while the request is sent.
request.runtime.stream_writer({"confirmed": record_id})
result = await handler(request)
return _with_answer(result, answer)
Here is what happens, in order:
- A send for any record other than the selected one is refused before the tool runs.
interrupt()stops the run with the clinic’s read-back question. The voice loop speaks it.- The caller’s next turn resumes the run. A separate model call judges whether the words are a yes.
- The answer goes to the clinic’s audit log. Only a yes runs the tool.
- A no speaks the clinic’s decline and hands the caller’s words back to the model. So “No, wait, not that one. I meant my budesonide inhaler.” gets the inhaler selected and read back on the same turn.
The yes judge uses structured output, so the model must return true or false.
This excerpt from guard.py builds it:
def llm_classifier(model: Any) -> Classifier:
"""Yes or no from the model, with structured output: the judgement Rasa's engine also leaves to the model."""
structured = model.with_structured_output(CallerAnswer, method="json_schema")
async def classify(question: str, answer: str) -> bool:
result = await structured.ainvoke(CLASSIFY_PROMPT.format(question=question, answer=answer))
return bool(result.confirmed)
return classify
Its prompt says that a no, a different medicine, a question without a yes, or anything unclear is not a yes.
Why not LangChain's HumanInTheLoopMiddleware?
It is built for a reviewer who approves, edits, rejects or responds to a
proposed tool call. A caller answers in words, so something still has to
turn the words into a decision. The question also has to read the medicine
back from state, and the answer has to reach the clinic’s audit log.
LangGraph’s interrupts page
says interrupts “can be placed anywhere in your code and can be conditional”.
So this guard calls interrupt() from its own middleware, and the tool body
stays simple.
Step 3: Write the voice loop
LangGraph and LangChain do not handle audio. So you write the voice loop, in
voice_loop.py and server.py. This build uses Starlette and uvicorn for
the server. It speaks the same browser audio protocol as Rasa, so one web page
works for all three builds.
The loop does these jobs:
- Audio in: It forwards the caller’s audio to Speechmatics speech-to-text.
- End of turn: Speechmatics sends one final transcript when the caller pauses for 0.7 seconds. The loop answers those turns one at a time, in order.
- Speaking while the model writes: It cuts the model’s text at sentence ends and sends each sentence to text-to-speech straight away.
- Fillers: If the model starts a tool call before saying anything, it speaks a short fixed phrase, such as “Let me look at your record.”
- Playback markers: It sends markers with the audio, and the browser sends them back once that audio has played.
- Silence check-in: After 30 seconds with no caller speech, it asks “Are you still there? I can help when you’re ready.”
- Events: It records what the caller and the agent said, for the test runner.
This excerpt from voice_loop.py streams the agent and speaks a filler when
a tool call starts:
async for mode, data in self.agent.astream(payload, self.config,
stream_mode=["messages", "custom", "updates"]):
if mode == "messages":
chunk, meta = data
if meta.get("langgraph_node") != "model":
continue
if chunk.id != current_id:
self._finish(current)
current, current_id = None, chunk.id
delta = chunk.text
if delta:
current = current or Message(clock)
self._feed(current, delta)
spoke = True
calls = getattr(chunk, "tool_call_chunks", None) or []
if calls and not spoke:
self._speak_whole(FILLERS.get(calls[0].get("name") or "", "One moment."), clock)
spoke = True
The loop also has to know about the guard’s pause. When the run is paused,
the caller’s next turn is the answer, so the loop sends it as a resume. This
excerpt is from
voice_loop.py:
# concern-begin: refill-guard
# The guard asked the confirmation question on an earlier caller turn and the run is paused
# in interrupt(). This caller turn is the answer: it is the resume value, and only a caller
# turn ever resumes it, so the answer always comes from a later turn than the question.
if state.interrupts:
payload = Command(resume={"text": text, "turn": self.turn_count})
# concern-end
Another part of the loop speaks the interrupt’s question, and the decline the guard writes to the custom stream.
This build does not do barge-in, and it has no cache for repeated speech.
Audio that arrives while the agent speaks is answered after the current
turn. To add barge-in, you would cancel the task running astream, stop the
audio, and decide what the saved state keeps of a half-spoken reply.
To use Deepgram for speech instead, the companion’s launcher starts the same server with Deepgram clients. No file in the LangGraph folder changes.
Count a yes only after the question has played
The loop answers transcripts in the order they arrive. That keeps the answer on a later turn than the question. But it does not prove the caller heard the question first.
A test replay showed the gap. It copies a live call on which speech-to-text
split the caller’s words. The replay sent “Yes, please.” the moment the caller
had finished saying it. In the replay, the yes arrived before the read-back
had played. On that live call, the transcript came after it. This is one
replay, from
results/langgraph/2026-10-01-late-transcript-replay/summary.md. SENT is the
replay’s transcript, and user is the build logging it:
16.07 bot_turn_ended
18.85 SENT Of my omeprazole.
18.85 user Of my omeprazole.
20.58 bot Let me look at your record.
20.70 SENT Yes, please.
20.70 user Yes, please.
26.81 bot I can send a request about this recorded medication, omeprazole twenty milligram capsules, one capsule before breakfast, to the prescribing team. Would you like me to do that?
26.81 bot_turn_ended
29.81 bot Right, I'll send that request for review.
32.54 bot Your request reference is R Q, seven six zero four. It is awaiting prescribing team review, and they will contact you with the outcome.
The early yes resumed the run, and the request went out. That happened in 3 of 3 replays. The same gap showed up in all three builds, so it is not a LangGraph problem. It is a voice loop problem, and the fix goes in the loop.
The companion’s opt-in
fix.diff
changes only voice_loop.py, plus a test for it. It records when the caller
began speaking, and when the browser confirmed the last turn had played. Then
it checks both before it resumes. This hunk from fix.diff is the core of the
change:
@@ -244,6 +263,15 @@
# in interrupt(). This caller turn is the answer: it is the resume value, and only a caller
# turn ever resumes it, so the answer always comes from a later turn than the question.
if state.interrupts:
+ # The fix: an answer that began before the question finished playing is not consent.
+ # Ask the question again instead of resuming.
+ if self.turn_played_at is None or (onset or clock.started) < self.turn_played_at:
+ log.info("%s: answer began before the question finished playing; asking again", self.id)
+ for item in state.interrupts:
+ value = getattr(item, "value", None)
+ if isinstance(value, dict) and value.get("question"):
+ self._speak_whole(value["question"], clock)
+ return
payload = Command(resume={"text": text, "turn": self.turn_count})
# concern-end
current: Optional[Message] = None
With the fix, the early yes no longer counted. The agent asked the question again, and nothing was sent in 3 of 3 replays. The same held for an “Okay.” said over the filler and a “Yeah.” said during the read-back. On live calls, the fixed copy passed 4 of 4, each with one read-back and then the send (3 calls and 1 call).
Step 4: Run it
From the langgraph folder:
make install # langgraph, langchain, langchain-openai, starlette, uvicorn, cedar_clinic, cedar_speech
make env # fill OPENAI_API_KEY and SPEECHMATICS_API_KEY
make run # ws://localhost:5006/webhooks/browser_audio/websocket
make web # in another shell: the voice page on http://127.0.0.1:8765/
Open the voice page and talk to the agent. The test patients are fictional. Try “I’m Maria Alvarez, born March 14th, 1968. Lisinopril, please.” The agent reads the medicine back and waits for your yes.
To try the timing fix, make a fixed copy from the tutorial folder:
make fix-copy FW=langgraph # langgraph-fix/ with fix.diff applied
The copy leaves out your .env and installed packages. So run make install
and make env again inside langgraph-fix. make env creates .env there
with both keys blank, so fill them in again before make run.
Step 5: Test it
Start with the offline tests. They use a scripted model and fake speech, so they need no keys and no network:
make test
The guard tests drive the real graph with a scripted model. Here is that group on its own, with its output:
uv run --locked python -m unittest discover -s tests -k GuardTests -v
test_a_no_sends_nothing_and_the_model_hears_the_answer (test_guard.GuardTests.test_a_no_sends_nothing_and_the_model_hears_the_answer) ... ok
test_a_send_for_another_record_than_the_selected_one_is_refused (test_guard.GuardTests.test_a_send_for_another_record_than_the_selected_one_is_refused) ... ok
test_a_send_without_a_selection_is_refused_before_the_tool_runs (test_guard.GuardTests.test_a_send_without_a_selection_is_refused_before_the_tool_runs) ... ok
test_a_tool_called_beside_the_paused_send_is_not_run_again_on_resume (test_guard.GuardTests.test_a_tool_called_beside_the_paused_send_is_not_run_again_on_resume) ... ok
test_guard_state_cannot_be_set_from_the_graph_input (test_guard.GuardTests.test_guard_state_cannot_be_set_from_the_graph_input) ... ok
test_send_is_not_offered_until_a_medicine_is_selected (test_guard.GuardTests.test_send_is_not_offered_until_a_medicine_is_selected) ... ok
test_the_model_never_sees_the_patient_id (test_guard.GuardTests.test_the_model_never_sees_the_patient_id) ... ok
test_the_question_pauses_the_run_and_only_a_later_yes_sends (test_guard.GuardTests.test_the_question_pauses_the_run_and_only_a_later_yes_sends) ... ok
----------------------------------------------------------------------
Ran 8 tests in 0.050s
OK
The last test is the main one. It runs one turn and checks that the run paused with the clinic’s question and sent nothing. Then it resumes with “Yes, please send it.” and checks the audit log. These tests prove the guard’s rules. They do not show what a real model does.
For that, play the shared recorded calls to the build:
make spec # the 17 recorded calls over browser audio (billed, capped at 4 USD a run)
The runner plays each call and judges it from the clinic’s audit log, with the same checks for every build. On the live run, this build passed 16 of the 17 calls, and the guard was never broken.
To test the timing gap, run the early-yes replay from the tutorial folder. Run it against the shipped build, then against the fixed copy:
make late-transcript-replay FW=langgraph LABEL=my-replay # billed
make late-transcript-replay FW=langgraph CWD=langgraph-fix LABEL=my-replay-fix # billed
The first should send the request, and the second should ask again. Chapter 5 shows how to prove your tests catch a missing guard.
Limits
- Cedar Clinic is fictional, and nothing here is clinical advice.
- The runs in this chapter used GPT-5.5 with Speechmatics. One extra run of the 17 calls used Deepgram.
- The early-yes replays simulate only speech-to-text. They show that the shipped loop accepts an early yes, not how often callers give one.
- The fix was tested on replays and a few live calls, not proven.