Skip to content
RasaGet a free licence
Guides for AI teams

Guide · AI product engineer

How to handle misheard names in a voice agent

When your voice agent gets numbers right but fails calls on names, measure names separately and list the known ones for speech to text.

by Rod Rivera

About 9 minutes

  • 5 of 7

    failed test calls in the card-dispute sample that failed on one misheard name

  • 6 → 11

    of 11 rerun calls passing: no list → seven vocabulary entries

  • 6 → 6

    of the same 11 calls passing: no list → a list of three other merchants

Source: Scripted test calls with synthetic callers to a sample Rasa card-dispute voice agent, 29 September and 2 October 2026
Key takeaways (3)
  • Count names separately from numbers in your test calls. One overall error rate cannot show which words were missed.
  • For each name a tool looks up, decide whether it must match exactly or can accept small spelling changes.
  • A speech-to-text vocabulary list helped only the names it held, so build it from your own merchants and products.

Your voice agent passes the amounts, dates and card numbers in your test calls. Then a call fails because speech to text misheard one name: a merchant, a branch or a surname.

Here is one call to a sample card-dispute agent. Every number arrives intact. The merchant does not.

Voice callLive run, 29 September 2026: normal-second-customer-arjun (claude-sonnet-5-5, Speechmatics enhanced, synthetic caller). Trimmed: the greeting, the activate call and the session end, with its cancel_skill call and closing question, are left out.
  1. Caller

    Good afternoon, this is Arjun Raghunathan, born on the fifth of February, nineteen eighty-nine. My debit card ending five five three zero shows five thousand nine hundred and ninety-nine rupees at Brightmart Online. I didn’t make that purchase.

    Annotation:

    The script line the synthetic caller’s audio was made from.

  2. Heard

    Good afternoon. This is Arjun Raghunathan. Born on the 5th of February 1989. My debit card ending 5530 shows ₹5,999 at Breitbart Online. I didn’t make that purchase.

    Annotation:

    The tracker’s user text. Name, date, card ending and amount are all right.

  3. Agent

    Right, let me get your dispute started now.

  4. Tool

    verify_caller(full_name=“Arjun Raghunathan”, date_of_birth=“1989-02-05”) → verified

  5. Tool

    select_transaction(merchant=“Breitbart Online”, amount_rupees=“5999”, card_last_four=“5530”) → blocked, transaction_ambiguous, matches 0

  6. Agent

    Thank you, Arjun, you’re verified. I couldn’t find one charge that matches those details, so could you check the merchant name, the amount and the date, and tell me again?

  7. Heard

    Yes, I confirm .

  8. Agent

    Okay, I don’t have a matching charge to confirm yet. Could you tell me the date of the charge, and check the merchant name and amount on your statement?

The tool looks charges up by merchant, so it selected nothing and filed nothing. The scripted caller had no answer to the agent’s question.

The fix: count names separately when you test, decide how each tool matches a name, and list the names you know for the recogniser.

That matters if your tools look records up by a name the caller says. In the sample’s test run, 5 of the 7 failed calls failed on one misheard name. Yet the run’s word error rate, the share of words the recogniser got wrong, was a healthy-looking 0.042. A vocabulary list of seven names took the calls with a name miss from 6 passing to 11 of 11.

What you need:

  • The samples are four companion projects in RasaHQ/rasa-community-resources at commit 4aa0c44. All four are verified with rasa-pro 3.21.0.dev5 and use fictional organisations. The lead sample is examples/mantle-voice-banking-dispute-claude, a card-dispute agent on Claude that uses Speechmatics for speech to text.
  • Speechmatics is not a Rasa built-in. The sample reaches it through an adapter in the companion’s patterns/voice-vendor-router package. The vocabulary setting below belongs to that adapter.
  • The offline checks here need only Python 3, with no keys or network. The voice test calls are live and billed, and you do not need to rerun them.

Why names fail when numbers do not

The sample’s test harness tags every word a tool acts on, by kind, and checks it in the heard text. The dispute build used Speechmatics. The other three used Deepgram Flux, one in its multilingual form for Spanish.

Kind of wordDisputeAdvisorRoadsideCollections
Amounts17/17––13/13
Dates and times45/4832/34––
Card, account and policy numbers4/4–15/1624/24
Exits, mile markers, street numbers––10/10–
Purpose and topic words–21/21–7/7
Names47/586/1112/167/9

A number counts as right when its value is right, however it was written. A name counts only when it is written the way the script spells it. A dash means the build had no word of that kind.

Numbers survive because the model and the tools absorb the formatting. Speechmatics wrote “₹5,999” for “five thousand nine hundred and ninety-nine rupees”, and Claude passed the right amount to the tool. A misheard name is different. In the dispute sample, Claude passed misheard merchant names on as heard, so the tool’s matching rule decided what happened next.

The word error rate hides this. It is averaged over the caller’s turns, and no turn’s rate says which words were wrong.

Step 1: Count each kind of word your tools use

In your own test scripts, tag every word a tool needs, with its kind. The sample stores the tags next to each caller line. This is an excerpt of the tags for Arjun’s first line, from the dispute sample’s case-build/conversations.json:

"asr_tokens": [
  {"token": "Arjun", "kind": "name"},
  {"token": "Raghunathan", "kind": "name"},
  {"token": "5", "kind": "date"},
  {"token": "1989", "kind": "date"},
  {"token": "5530", "kind": "card_ending"},
  {"token": "5999", "kind": "amount"},
  {"token": "Brightmart", "kind": "name"}
]

After each call, the harness looks for each tagged word in the text the recogniser returned. It matches whole words and ignores case. For numbers, it checks a second time after turning spoken numbers into digits. Names get no second check: a name is right as written or it is wrong.

Report each kind next to the word error rate. The next two steps act at different points between the caller and the tool:

caller says "Brightmart" speech to text Step 3: vocabulary list heard text Step 1: count it tool's match rule Step 2: exact or tolerant right record selected match no match, caller asked again no match
FigureWhere each fix acts on a spoken name

Step 2: Decide how each tool matches a name

In the dispute run, 11 of 58 names did not come back as written. Six cost nothing: “Bright Mart” and “Lake View” still matched their merchants. The other five each failed a call:

SaidHeardWhat the tool did
Brightmart“Breitbart Online”no merchant match, nothing selected
Brightmart“Breitbart Online”no merchant match, nothing selected
Brightmart“bright mark”no merchant match, nothing selected
Hollins“Haaland’s Fitness”no merchant match, nothing selected
Raghunathan“Priya Ragunathan”caller not verified

None of the failed calls filed a wrong dispute or promised a refund.

“Bright Mart” passed and “bright mark” failed because of the merchant rule. This is the sample’s merchant_matches function, from lib/disputes.py. The select_transaction tool uses it to search the verified caller’s own charges:

def merchant_matches(spoken: Any, merchant: str) -> bool:
    """Every word the caller gave appears in the merchant name, or the reverse.

    Word boundaries are also ignored, because speech-to-text splits and joins
    brand names ("Brightmart" came back as "Bright Mart" live): the spoken
    name with its spaces removed may appear inside the merchant's, or the
    reverse, if it is at least six characters long.
    """
    said, name = merchant_words(spoken), merchant_words(merchant)
    if not said:
        return False
    if said <= name or name <= said:
        return True
    said_c, name_c = _compact(spoken), _compact(merchant)
    return len(said_c) >= 6 and (said_c in name_c or name_c in said_c)

With the spaces removed, “brightmart” sits inside “brightmartonline”, so a split name matches. A misspelt name does not: “brightmark” is not inside it.

The caller’s own name gets a stricter rule. The verify_caller check lower-cases the name and drops stray punctuation and extra spaces. After that, the name must equal the customer record exactly. So “Priya Ragunathan”, one letter off, fails even with the right date of birth.

You can see both rules without a model or network. Run these from the dispute project folder at commit 4aa0c44:

$ python3 --version
Python 3.14.3
$ python3 -c "from lib import disputes as nd; s = nd.LedgerService(); print([nd.verify_caller(s, n, '1991-08-17')['status'] for n in ('priya  RAGHUNATHAN.', 'Priya Ragunathan', 'Priya Raghu Nathan')])"
['verified', 'not_verified', 'not_verified']
$ python3 -c "from lib import disputes as nd; print([nd.merchant_matches(m, 'Brightmart Online') for m in ('Bright Mart', 'bright mark online', 'Breitbart Online')])"
[True, False, False]
$ python3 -m unittest -v tests.test_guard.SelectionTests.test_verification_is_exact_on_name_and_date tests.test_guard.SelectionTests.test_brand_names_split_or_joined_by_speech_to_text_still_match
test_verification_is_exact_on_name_and_date (tests.test_guard.SelectionTests.test_verification_is_exact_on_name_and_date) ... ok
test_brand_names_split_or_joined_by_speech_to_text_still_match (tests.test_guard.SelectionTests.test_brand_names_split_or_joined_by_speech_to_text_still_match) ... ok

----------------------------------------------------------------------
Ran 2 tests in 0.000s

OK

The tests also check that “Lake View Fuel” finds its charge and that “Mart” alone finds none.

Three of the four builds look something up by a spoken name. The fourth looks accounts up by their last digits. The three use these rules:

NameSampleRuleA miss leads to
Caller’s full nameDisputeexact, after case, punctuation and spacesnot verified; the caller is asked again
MerchantDisputespaces ignored, among the verified caller’s chargesno match; the caller is asked to check
BranchAdvisorspaces ignored, close spellings acceptedbranch ignored; the search proposes other slots
Policyholder’s surnameRoadsideexact, letters onlynot found; the caller is asked for the digits

This guide suggests a simple test. Be tolerant where a loose match picks among records the caller has already proved are theirs. Be exact where the name is part of the proof. The dispute sample’s code fits it: the merchant search covers only the verified caller’s charges, and the name check is exact. The advisor’s branch matcher is tolerant across all branches. None of this is identity-verification advice.

Tolerance also has a ceiling. The branch matcher accepts a spelling that scores at least 0.8 on Python’s difflib similarity ratio. It took “Ashcom” for the Ashcombe branch, but not “Ashken” or “hashem”. For a name you must match exactly, the fix has to come before the text exists.

Step 3: Add the names you know to the vocabulary list

Speechmatics’ custom dictionary documentation gives one reason a word goes missing: “it’s not in the vocabulary for that language, for example a company or person’s name”. It adds: “Adding custom words can improve the likelihood they will be output.” So a custom vocabulary is advice the vendor already gives. What this guide adds is the count per kind of word and the held-out test below. You pass the words in the additional_vocab setting. Each entry can carry optional sounds_like spellings.

In your own project, edit the asr block in integrations.yml of each voice channel that uses the Speechmatics adapter. This excerpt is the sample’s block with the test’s seven entries added and comments removed. In the file it sits four spaces deep under browser_audio. It is shown here without that indentation:

asr:
  name: voicerouter.providers.speechmatics.SpeechmaticsASR
  language_map:
    en:
      language: en
  operating_point: enhanced
  max_delay: 1.0
  enable_partials: true
  end_of_utterance_silence_trigger: 0.7
  additional_vocab:
    - content: Brightmart
      sounds_like: [bright mart]
    - content: Lakeview
      sounds_like: [lake view]
    - content: Hollins Fitness
    - content: Cinnabar Streaming
    - content: Saffron Table
    - content: Metro Cabs
    - content: Raghunathan
      sounds_like: [raghu nathan]

Six entries are the sample ledger’s merchants. One is the scripted customers’ surname.

The adapter copies the list into the settings it sends when recognition starts. This excerpt is from voicerouter/providers/speechmatics.py:

        if self.config.additional_vocab:
            transcription_config["additional_vocab"] = [
                {"content": entry} if isinstance(entry, str) else dict(entry)
                for entry in self.config.additional_vocab
            ]

The adapter’s unit test checks what is sent, and it passes. Whether recognition improves needs test calls.

What the list changed

The sample reran the 11 calls that had a name miss, with the list added and nothing else changed. Both runs played the same recorded caller audio, a few minutes apart.

The same caller line, without and with the vocabulary

Avoid: Default, 29 September 2026

Excerpt of the heard line:

My debit card ending 5530 shows ₹5,999 at Breitbart Online.

select_transaction(merchant="Breitbart Online", amount_rupees="5999", card_last_four="5530") returned blocked, transaction_ambiguous, 0 matches. Nothing filed.

Prefer: With additional_vocab, minutes later

Excerpt of the heard line:

My debit card ending 5530 shows ₹5,999 at Brightmart online.

select_transaction(merchant="Brightmart", amount_rupees="5999", card_last_four="5530") returned selected, then file_dispute on NB-TXN-3201 succeeded.

  • 6 of 11

    calls passed without the list

  • 11 of 11

    passed with it

  • 22 of 33

    names heard as written without it

  • 33 of 33

    with it

Source: Dispute sample, the 11 calls with a name miss, without and with the seven-entry list; live test calls with synthetic callers, 29 September 2026

All five calls that had failed on a name passed.

The list helps only the names it holds

That list came from the names in the test script. So a second test held out Brightmart Online, Lakeview Fuel, Hollins Fitness and the surname. The list held only the ledger’s other three merchants. On 2 October the sample ran the same 11 calls with that short list, and again with no list as a same-day control:

Same 11 callsPassedHeld-out merchants as writtenSurname as written
29 Sept, no list6/111/1110/11
29 Sept, all seven entries11/1111/1111/11
2 Oct, three other merchants6/111/1110/11
2 Oct, no list (control)6/111/1110/11

The held-out names were misheard the same way in both runs, and every call ended the same way. None of the listed names turned up in any transcript.

How to check the fix works

Rerun the offline match tests

After you change a match rule, rerun the guard tests. In the dispute sample’s folder, this needs only Python 3:

Any system
python3 -m unittest discover -s tests -v

The suite passes. Tests that need Rasa installed are skipped.

Replay the same recordings

Record each scripted caller line once and replay that file in every run. The sample names each file from a hash of the voice and the words.

Compare runs from the same day

The same recording can come back differently on another day. Between the sample’s two test days, two calls changed outcome with no list at all.

Count names as written

Check each tagged name as written, and count the calls that pass.

Look for listed words nobody said

Search for listed words in turns where the caller did not say them. The sample found none.

When a name is still lost, make the tool refuse to guess

The three Deepgram Flux builds ran with no vocabulary list. In two cases there, a tool that refuses to guess stopped a lost word from becoming a wrong action. The first is the advisor’s purpose rule. A caller asks for an appointment about an ISA, which the sample books under investments. Flux hears “ASA”:

Voice callLive rerun, 30 September 2026: recovery-hold-lapsed (claude-sonnet-5-5, Deepgram Flux, synthetic British caller). Trimmed: the greeting, the load_session_customer and activate calls, and the closing turn are left out.
  1. Heard

    Can I book a video call about my ASA for Friday at a quarter to two?

    Annotation:

    The tracker’s user text. The script line, and the audio made from it, said “ISA”.

  2. Agent

    Right, let me get that booking started for you.

  3. Agent

    Just to check, did you mean an ISA?

  4. Event

    The caller’s scripted “Yes.” produced no user event.

  5. Event

    Two agent turns omitted: a silence check-in and an error reply.

  6. Heard

    Yes. Please try again.

    Annotation:

    The script line was “Yes, please try again.”

  7. Agent

    Okay, I’ll look for video slots for your ISA on Friday.

  8. Tool

    find_advisor_slots(purpose=“ISA”, channel=“video”, day=“Friday”, part_of_day=“afternoon”) → blocked, wrong_advisor_capability, purpose_not_stated

    Annotation:

    Claude’s guess was right. The guard does not take Claude’s word for it.

  9. Agent

    Sorry, I need to check this with you directly. In your own words, what would you like the appointment to be about?

Claude guessed “ISA”, and the search still refused. Its purpose rule reads the caller’s own messages, not what the model passed, and none of them named a purpose. You can check that from the advisor project folder, examples/mantle-voice-banking-advisor-appointment-claude:

$ python3 -c "from lib.appointments import stated_purpose as p; print(p(['Can I book a video call about my ASA for Friday at a quarter to two?', 'Yes.', 'Yes. Please try again.'])); print(p(['Can I book a video call about my ISA for Friday at a quarter to two?']))"
None
investments

The cost is one more question. In return, a guess about a lost word cannot become a search for the wrong purpose.

The second is the roadside policy lookup. It finds the policy by its six digits, then requires the spoken surname to match exactly. “Raman” came back as “Ramen” in half the lines that said it, and each lookup failed. Because the digits pick the policy, a misheard surname costs a repeat but cannot open someone else’s policy.

Trade-offs

Exact matching costs a repeat: the right customer, misheard, is asked again. One dispute call failed that way.

Tolerant matching lets more through: the sample did not test what a looser branch match would accept.

A vocabulary list needs the names in advance: it suits a known set, such as merchants in a ledger, branches and product names.

Limits

  • The callers, audio and run sizes are as the info callout describes.
  • Each vocabulary condition is one run on one day.
  • The Deepgram Flux builds were not tested with a vocabulary list.

Related: Flux dropping one-word replies traces a lost “Yes.” like the ISA call’s, and TTS failover after a restart covers the same adapter package.