Skip to content
RasaGet a free licence
Guides for AI teams

Guide · AI product engineer

Haiku returned nothing and Rasa logged facts_count 0

Rasa Mantle memory discovery on Anthropic: Haiku gave Rasa 3.20.0's fact extractor no tool call in 20 replays, even after Please continue.

by Rod Rivera

About 8 minutes

  • 0 of 20

    replays of Rasa's extractor request, unchanged, that got a tool call from Haiku

  • 20 of 20

    replays that got both facts once the final user turn asked, or the tool was forced

  • HTTP 400

    Opus answer to the extractor request in the acceptance run, logged as a WARNING

Source: One captured Rasa Pro 3.20.0 extractor request replayed to the Anthropic API, plus live quickstart runs
Key takeaways (7)
  • On Rasa Pro 3.20.0 with claude-haiku-4-5, an invented customer stated her name and her contact preference, and no discovered fact reached memory in either of two runs. The reply was correct, and the log recorded a completed discovery with facts_count 0.
  • The post-turn fact extractor puts its task in the system prompt and sends the conversation as messages, ending on the agent's reply. claude-haiku-4-5 answered that with empty content or with text instead of the tool call; in the acceptance run, claude-opus-5-5 refused it with HTTP 400, logged as one WARNING.
  • Replayed to the Anthropic API 20 times per variant, the captured request got the tool call 0 of 20 times unchanged and 0 of 20 with "Continue." appended, and 20 of 20 with "Record the facts now." appended, with the tool forced, or with the task text sent as the final user message. Anthropic's documented "Please continue" also got 0 of 20. A hand-built request with no Rasa got the tool call 20 of 20, so the loss belongs to Rasa's own request, not to message order in general.
  • Discovery runs after the reply, on each turn that switches into a new flow, and every failure is swallowed. Acceptance passed on both models and asserts nothing about memory, so it cannot tell you whether this feature works.
  • Test a background write by stating a fact and asserting the memory_set event for system.__discovered__. The reply proves nothing either way: in an Opus run where our proxy appended "Record the facts now." and the facts were recorded, the reply told the customer her preference had not been recorded.
  • On this request, three options got both facts 20 of 20 at the API level: forcing the tool, sending the task text as the final user message, and appending a user turn that asks for the task. The 3.20.0 extractor does none of these, and we did not patch Rasa.
  • A bisect found that removing both earlier tool-call pairs from the history brought the call back, 140 of 140 across four forms of the edit, while removing either pair alone did not. We do not know why, and a minimal request with one tool-call pair still got the call. Anthropic documents empty end_turn replies after tool results and when a completed response is sent back.

This is the second customer turn of a live run of the quickstart agent, Juniper (a fictional plant shop), on Rasa Pro 3.20.0 with its model group switched to Anthropic’s claude-haiku-4-5. Priya Nair is an invented customer. The bot lines are what the REST channel returned; the event lines summarise the capture proxy, the server log and the tracker.

Chat transcriptRun A2: claude-haiku-4-5 on Rasa Pro 3.20.0
  1. Customer

    Hi, my name is Priya Nair and I only want updates by text message, never email. How many monstera do you have?

  2. Bot

    Got it — text only, no email. Let me check the monstera stock for you.

  3. Tool

    check_stock → {“found”: true, “plant”: “monstera”, “quantity”: 7}

  4. Bot

    We have 7 monsteras in stock. Would you like to check anything else?

  5. Event

    record_discovered_facts request → HTTP 200, stop_reason end_turn, content []

    Annotation:

    The call that should write her name and preference to memory. The model sent back nothing.

  6. Event

    mantle.processor.discover_facts.completed, facts_count 0 (INFO)

  7. Event

    memory events in the tracker: []

The reply is right, and the log calls the discovery completed. The extractor’s request ended on the agent’s reply above, and Haiku sent back no content at all:

assistant: [{"type": "text", "text": "We have **7 monsteras** in stock. Would you like to check anything else?"}]
response: "content":[], "stop_reason":"end_turn", "output_tokens":3

Those are the last message of the request and three fields of the response, cut from the capture. Anthropic’s page Handling stop reasons describes this reply: “exactly 2–3 tokens with no content”, “particularly after tool results”, with “Sending Claude’s completed response back without adding anything” among the common causes. “If you still get empty responses after fixing the message structure”, it advises a continuation prompt in a new user message; its code example uses “Please continue”. On Rasa’s request that wording got the tool call 0 of 20 times. We sent the captured request straight to the Anthropic API, 20 calls per run and one change per arm; some arms ran more than once, and the table gives the totals:

Change to A2’s captured extractor requestTool calledFacts in the call
none: ends on the agent’s reply0 of 20none; empty content
user turn “Please continue” appended (the page’s wording)0 of 20none; text only
user turn “Continue.” appended0 of 20none; text only
user turn “Record the facts now.” appended20 of 20name and preference, 20 of 20
tool_choice forced to record_discovered_facts20 of 20name and preference, 20 of 20
the extractor’s task text appended as the final user turn20 of 20name and preference, 20 of 20
both earlier tool-call pairs removed, four forms, seven runs140 of 140preference 140; name 8
the activate tool-call pair removed alone0 of 20none; empty content
the check_stock tool-call pair removed alone, two runs0 of 40none; empty content
six other single changes, listed below0 of 20 eachnone; empty content

The counts are summaries computed from the per-call rows the scripts wrote. Of the arms that add a final user turn or force the tool, the continuation from the example got nothing, and an asking turn, a forced tool or the task as the final user turn got both facts. The only other change that brought the call back was removing both earlier tool-call pairs. The page is general advice, not a description of Rasa’s extractor, and why Rasa’s request behaves this way we did not find: our minimal request got the call with “Please continue” and with a tool-call pair in its history.

The cost lands on your code, not Rasa’s. In the 3.20.0 wheel’s Python engine, nothing but the helper that lists already-captured keys reads the discovered facts, so a lost fact shows up wherever your own code reads system.__discovered__. The reply will not warn you: in one Opus run where our proxy appended “Record the facts now.” and the facts were recorded, the agent told Priya “your preference hasn’t been recorded anywhere”.

Our position: on Rasa Pro 3.20.0 with Claude, do not build on discovered facts, and test any write an agent makes in the background by asserting the write itself. Acceptance passed in both runs we made of it, and in the Haiku probes the only log line about discovery reads like success. The cost is a slower, live check that needs a conversation written to trigger the feature. We have not reported this upstream.

Both acceptance runs passed while the extractor got nothing back

The quickstart at commit 8a0320c, its 3.20.0 version, used an OpenAI model group. Our working copies came from a later commit whose examples/quickstart is identical. We switched that one group in examples/quickstart/integrations.yml as shown; the Opus runs used the same lines with model: claude-opus-5-5 and no temperature:

 model_groups:
   - id: orchestrator
     models:
-      - provider: openai
-        model: gpt-4.1-mini
-        api_key_env: OPENAI_API_KEY
+      - provider: anthropic
+        model: claude-haiku-4-5
+        api_key_env: ANTHROPIC_API_KEY
         temperature: 0.0

The live acceptance check trains the project and runs four synthetic cases. Its assertions check that each case got a reply, the stock lookups, the refusal of a reservation, the tools called and that no tool errored; none of them reads memory. Both runs reported licensedTraining passed, acceptance passed. In each, request 4 to Anthropic was the fact extractor. The Haiku run’s line, then the Opus run’s, from the proxy summaries:

4: 200 roles=['user', 'assistant', 'user', 'assistant', 'user', 'assistant', 'user', 'assistant'] stop_reason=end_turn types=[] error=None
4: 400 roles=['user', 'assistant', 'user', 'assistant', 'user', 'assistant', 'user', 'assistant'] stop_reason=None types=[] error=This model does not support assistant message prefill. The conversation must end with a user message.

Neither acceptance conversation states a fact. For Haiku, an empty answer to such a conversation is a permitted one: the task text says the model may “not call it at all” when nothing new was stated. For Opus the call itself failed, and its server log holds one line about it, with the timestamp cut:

WARNING  rasa.mantle.memory.discovery.extractor  - {"error": "ProviderClientAPIException:\n\nOriginal error: litellm.BadRequestError: AnthropicException - {\"type\":\"error\",\"error\":{\"type\":\"invalid_request_error\",\"message\":\"This model does not support assistant message prefill. The conversation must end with a user message.\"},\"request_id\":\"req_011CfXaFueXZ1swLdR9ohDsw\"}\n", "event": "mantle.memory.discovery.extractor.discover_facts.llm_error", "conversation_id": "compatibility-stock-cf59550033d14c1a84d19b8979371ff3", "level": "warning"}

The one memory_set event in the Opus run has the key default_unsupported_request_handler.request_type, not system.__discovered__, so it is not a discovered fact. The probe runs then state the facts, with the server at LOG_LEVEL=INFO. Every live run, side by side:

RunModelLast message of the extractor requestAnthropic’s answerDiscovered facts in the trackerDiscovery in the server log
Acceptanceclaude-haiku-4-5the agent’s reply200, no contentnone (no fact stated)not captured: server at LOG_LEVEL=ERROR
Acceptanceclaude-opus-5-5the agent’s reply400, prefill refusednone (no fact stated)one WARNING, llm_error
A1claude-haiku-4-5the agent’s reply200, text and no tool call0INFO, facts_count 0
A2claude-haiku-4-5the agent’s reply200, no content0INFO, facts_count 0
B1, B2claude-haiku-4-5appended “Record the facts now.”200, record_discovered_facts call2: name, contact preferenceINFO, facts_count 2
C1, C2claude-opus-5-5appended “Record the facts now.”200, record_discovered_facts call3: name, contact preference, monsteraINFO, facts_count 3
D1, D2claude-haiku-4-5appended “Continue.”200, text and no tool call0INFO, facts_count 0

In A1, Haiku answered the extractor with text that began “I’ve noted your preference for text-only updates.” The extractor reads only tool calls, so that text went nowhere and the customer never saw it.

The extractor’s request ends on the agent’s own reply

ContextExtractor.discover_facts builds the request in rasa/mantle/memory/discovery/extractor.py. These are lines 164 to 189 of the Rasa Pro 3.20.0 wheel, sha256 a68fa7ef…e9b5:

        messages = [
            llm_message(
                ROLE_SYSTEM,
                render_template(
                    _DISCOVER_FACTS_TASK_TEMPLATE,
                    filled_scoped_entries=sorted(filled_scoped_entries),
                    known_discovered_keys=sorted(known_discovered_keys or ()),
                    memory_entries=memory_entries_for_prompt(defined_memory_fields),
                ),
            )
        ]
        history, _ = conversation_messages_from_tracker(
            tracker, max_utterances=max_utterances
        )
        messages.extend(history)

        try:
            response = await self._complete(
                self._llm_client, messages, tools=[_RECORD_DISCOVERED_FACTS_SCHEMA]
            )
        except Exception as exc:
            structlogger.warning(
                "mantle.memory.discovery.extractor.discover_facts.llm_error",
                error=str(exc),
            )
            return []

The task is the system message. A window of recent tracker history follows as it stands, and because discovery runs after the agent has answered, the last item is that answer. Nothing appends a user message, and the call passes the tool without forcing it. Any exception becomes a warning and an empty list. Lines 193 to 202 then read facts only from tool calls, so an answer with no tool call also yields an empty list.

The task text, lines 2 and 15 of discover_facts_task.jinja2, asks the model to review “the conversation above”, which in the request comes after it, and allows it not to call the tool:

Review the conversation above for facts the customer stated that are not already covered by one of the fields listed as "already filled" below. Only record facts that are clearly and specifically stated — skip vague or uncertain asides.
Call `record_discovered_facts` with any such facts now. If nothing new was stated, call it with an empty `facts` list or do not call it at all.

So “no tool call” is a legal answer to this prompt, and the extractor cannot tell it apart from a model that did not do the task.

Nothing between Rasa and the provider changes the order. LLMClient.acompletion (rasa/mantle/llm/client.py, lines 303 to 329) passes the message list to the provider client as it is, and the proxy recorded the same shape arriving at Anthropic. Anthropic calls a conversation that ends on an assistant message a prefill, as the Opus error shows. Opus refused it. Haiku accepted it and answered with no content or with text, never with the tool call.

What the calls themselves show, without an explanation: when Haiku called the tool on a request that still ended on the assistant turn, in the arms that removed both tool-call pairs (140 of 140 calls) and in our minimal request with a tool-call pair (20 of 20), its text began with two newlines and carried on from the agent’s reply, for example “In the meantime, let me record your communication preference:”. A1’s text carried on in the same way and called nothing. What makes the difference inside Rasa’s request, we did not find.

Discovery runs after the reply, on each turn that enters a new flow

The call site is rasa/mantle/processor.py. Lines 452 to 458 run discovery only after a turn that ended without an error and was not cancelled:

        if not turn_result.error and not turn_result.cancelled:
            # Skip post-turn discovery on barge-in: cancelled turns still
            # report error=None, but running another LLM call after interrupt
            # defeats the point of aborting promptly.
            await self._maybe_discover_facts(
                session, previous_flow_id, previous_skill_id
            )

Lines 482 to 484 then return early unless the turn switched into a new, non-null flow:

        new_flow_id = session.active_flow_id
        if new_flow_id is None or new_flow_id == previous_flow_id:
            return

So it runs on each turn that switches into a new flow, not on every turn. In our probes that happened once per conversation, on the second customer turn, and the log line names the switch: "previous_flow_id": null, "new_flow_id": "check_stock". We stated the facts on that same turn. The request carries a window of earlier history too, so a fact stated a turn or two before a flow switch may also reach the extractor; we did not test that.

The docstring of _maybe_discover_facts (lines 474 to 481) states the design: “Best-effort: any failure is logged and swallowed rather than surfaced, since a missed capture should never fail an otherwise-successful turn.” Lines 534 to 538 catch anything the extractor did not, as a warning. A missed capture is meant to be quiet, and it is.

agent reply already sent no error, not cancelled, new flow entered? processor.py 452-484 no discovery no extractor request system task + history last role: assistant yes claude-opus-5-5 HTTP 400 (acceptance run) claude-haiku-4-5 200, no tool call replay 20 of 20 + user turn "Continue." or "Please continue" + user turn asking for the task, or the tool forced empty list of facts caught, WARNING (INFO line: from the source, not observed) discover_facts.completed facts_count 0 (INFO) replay 20 of 20 each record_discovered_facts both facts replay 20 of 20 each
  1. Discovery waits for the reply and runs only on a clean turn that enters a new flow.
  2. The request’s last message is the reply the customer has already seen, and the task is only in the system prompt.
  3. A refused call, an empty answer and a text answer all end as the same INFO line.
  4. On this request, a final user turn that asks for the task or a forced tool got both facts; the only other change that got a call was removing both earlier tool-call pairs.
FigureThe post-turn discovery path on Rasa Pro 3.20.0, and where each variant ended

One captured request, 20 calls per change: what got the tool call

Every replay and bisect arm starts from A2’s extractor request body as the proxy recorded it (the whole body is saved with this article’s receipts and is not linked from this page). We re-sent it with our own headers to the Anthropic Messages API, with no Rasa, litellm or proxy in the path, at temperature 0.0, the value in the body. Twenty calls at that setting show that the answer was stable on this one request; they are not a rate across conversations.

The replay recorded one row per call: the status, the stop reason, the block types, the fact keys and the first 100 characters of any text. Call 1 of the two arms that separate a trailing user turn from what it asks:

The first replay of two variants of A2's request, as recorded

Avoid: “Continue.” appended

{"arm": "continue", "i": 1, "status": 200, "stop_reason": "end_turn", "block_types": ["text"], "tool_called": false, "facts": 0, "fact_keys": [], "text_head": "I don't have any additional information to share at the moment. What would you like to do next? I ca", "error": null}

A user turn at the end that asks nothing: text, no tool call. All 20 calls ended this way.

Prefer: “Record the facts now.” appended

{"arm": "ask", "i": 1, "status": 200, "stop_reason": "tool_use", "block_types": ["tool_use"], "tool_called": true, "facts": 2, "fact_keys": ["customer_name", "communication_preference"], "text_head": "", "error": null}

A user turn at the end that asks for the task: the tool call with both fact keys. All 20 calls ended this way.

Three variants got both facts 20 of 20 on this request: forcing the tool, sending the task text as the final user message, and a final user turn that asks for the task. Only the first two do not depend on our wording, which the “Continue.” and “Please continue” arms show matters. They worked at the API level, not in Rasa: the 3.20.0 extractor sends no final user turn and no tool_choice, and we did not patch it.

Then we bisected the unchanged request to find what makes it end silently. These single changes left it at 0 of 20, empty every time: the minimal three-sentence task text in place of Rasa’s; the task text without “or do not call it at all” (the tool’s own description still said “Omit this call entirely if nothing new was stated.”); a minimal tool schema; the opening greeting pair removed; the history’s own tools (activate, check_stock) declared; and the tool blocks kept but the plain-text messages sent as strings. Removing one tool-call pair left it at 0 as well: activate 0 of 20, and check_stock 0 of 40 over two runs (that arm also puts the two assistant texts into one turn; a second arm we meant as a separate change built the identical request, so we count it as the second run). Five of the totals, copied from the receipt:

  unchanged                          n=20 tool_called=0 empty=20 text_only=0 errors=0 fact keys (calls containing): {}
  drop_activate_pair                 n=20 tool_called=0 empty=20 text_only=0 errors=0 fact keys (calls containing): {}
  drop_check_stock_pair              n=20 tool_called=0 empty=20 text_only=0 errors=0 fact keys (calls containing): {}
  drop_both_pairs_blocks_kept        n=20 tool_called=20 empty=0 text_only=0 errors=0 fact keys (calls containing): {'communication_preference_text_only': 13, 'customer_name': 7, 'communication_preference': 7}
  drop_both_pairs_consecutive_assistant n=20 tool_called=20 empty=0 text_only=0 errors=0 fact keys (calls containing): {'communication_preference_text_only': 20}

Removing both tool-call pairs brought the call back in every form we tried: merged into plain strings (60 of 60 over three runs), merged into content blocks (40 of 40 over two runs), both assistant texts kept as separate blocks in one turn (20 of 20), and two assistant messages left in a row, sent unmerged (20 of 20). All 140 calls held the preference; the name appeared in 1, 0, 7 and 0 of them respectively.

Two arms in that list make a clean pair. They end on the same final assistant turn, the text the agent sent before the stock lookup followed by its final reply, and differ only in whether the activate pair is still in the history:

Arm (receipt name)activate paircheck_stock pairFinal assistant turnTool called
check_stock pair removed (drop_check_stock_pair)keptremovedboth texts in one turn0 of 40
both removed (drop_both_pairs_blocks_kept)removedremovedthe same turn20 of 20
activate pair removed (drop_activate_pair)removedkeptas captured0 of 20

Holding the final turn fixed, removing the activate pair as well took the result from 0 of 40 to 20 of 20; removing it alone got 0 of 20. We do not claim why, or that tool calls silence the model in general.

Our minimal Rasa-free request, with a three-sentence task in the system prompt, one tool offered and not forced and a history about the same customer ending on the assistant, got the call 20 of 20 as plain strings, 20 of 20 as content blocks, 20 of 20 with a check_stock tool-call pair added, and 20 of 20 with “Please continue” appended. So a tool call in the history does not by itself stop it; the effect is specific to this request. On Rasa’s request, “Please continue” got a text reply and no tool call in 20 of 20 (16 of its answers began “I’m ready to help! What would you like to do next?”), and “Continue.” got no tool call in 20 of 20 there and in 20 of 20 on the minimal request. What you can take to your own capture is the procedure: replay it, then remove one part of the history at a time, and see which arm gets the call.

The live runs in the table above agree: B1 and B2, whose extractor requests match A2’s apart from tool call ids and the appended turn, kept both facts. On Opus, which we did not replay, the same appended turn got the call in C1 and C2.

The replay script is 44 lines of Python that need an Anthropic key and python-dotenv, not Rasa; you run it against a request body you captured yourself. The bisect script is built the same way, with one arm per change.

Show the replay script as run (replay_a2.py)

It reads a captured request body and posts it to https://api.anthropic.com/v1/messages. This is the whole script as run, with one change: line 9 read our .env file by its absolute path, and here reads .env in the working directory. Nothing else differs, and the assert key on line 10 is kept. The recorded command was python replay_a2.py d320/runA2.jsonl replay-a2-haiku.jsonl 20, where d320 is our working directory for the probe captures. It imports dotenv, so install python-dotenv first.

"""Replay the extractor request Rasa 3.20.0 sent in probe run A2 (captured body, request 4 of
runA2.jsonl) straight to the Anthropic Messages API, N times per arm. The only differences
between arms are the ones named below. The API key is read from the project .env inside
this process and never printed or written."""
import json, sys, copy, urllib.request, urllib.error, time, pathlib
from dotenv import dotenv_values
CAP, OUT, N = pathlib.Path(sys.argv[1]), pathlib.Path(sys.argv[2]), int(sys.argv[3])
base = json.loads(CAP.read_text().splitlines()[4])['request']
key = dotenv_values('.env').get('ANTHROPIC_API_KEY')
assert key, 'ANTHROPIC_API_KEY missing'
system_text = base['system'] if isinstance(base['system'], str) else ' '.join(b.get('text', '') for b in base['system'])
def arm(name):
    b = copy.deepcopy(base)
    if name == 'unchanged': pass
    elif name == 'continue': b['messages'].append({'role': 'user', 'content': 'Continue.'})
    elif name == 'ask': b['messages'].append({'role': 'user', 'content': 'Record the facts now.'})
    elif name == 'tool_choice': b['tool_choice'] = {'type': 'tool', 'name': 'record_discovered_facts'}
    elif name == 'task_as_user': b['messages'].append({'role': 'user', 'content': system_text})
    return b
rows = []
for name in ('unchanged', 'continue', 'ask', 'tool_choice', 'task_as_user'):
    body = arm(name)
    for i in range(N):
        req = urllib.request.Request('https://api.anthropic.com/v1/messages', data=json.dumps(body).encode(),
              headers={'x-api-key': key, 'anthropic-version': '2023-06-01', 'content-type': 'application/json'}, method='POST')
        try:
            with urllib.request.urlopen(req, timeout=120) as r: status, resp = r.status, json.loads(r.read())
        except urllib.error.HTTPError as e:
            status, resp = e.code, json.loads(e.read() or b'{}')
            if e.code == 401: raise SystemExit('401: stop')
        blocks = resp.get('content') or []
        tu = [b for b in blocks if b.get('type') == 'tool_use' and b.get('name') == 'record_discovered_facts']
        facts = sum(len((b.get('input') or {}).get('facts') or []) for b in tu)
        rows.append({'arm': name, 'i': i + 1, 'status': status, 'stop_reason': resp.get('stop_reason'),
                     'block_types': [b.get('type') for b in blocks], 'tool_called': bool(tu), 'facts': facts,
                     'fact_keys': [f.get('key') for b in tu for f in ((b.get('input') or {}).get('facts') or [])],
                     'text_head': ' '.join(b.get('text', '') for b in blocks if b.get('type') == 'text')[:100],
                     'error': (resp.get('error') or {}).get('message')})
        time.sleep(0.3)
OUT.write_text('\n'.join(json.dumps(r) for r in rows) + '\n')
from collections import Counter
for name in ('unchanged', 'continue', 'ask', 'tool_choice', 'task_as_user'):
    rs = [r for r in rows if r['arm'] == name]
    print(f"{name:13s} n={len(rs)} tool_called={sum(r['tool_called'] for r in rs)} empty={sum(r['block_types']==[] for r in rs)} text_only={sum(r['block_types']==['text'] for r in rs)} errors={sum(r['status']!=200 for r in rs)} facts>=2={sum(r['facts']>=2 for r in rs)}")

Line 8 is specific to our capture: splitlines()[4] takes line 5 of our proxy log, the extractor request, and ['request'] takes the body the proxy recorded. Our A2 body is saved with this article’s receipts but not linked from this page, so to use the script, capture a request your own side call sends and replace that line with code that loads it as a dict with model, system, messages and tools. Then change one thing per arm.

Assert the memory write, not the reply

Discovery runs after the reply is sent, so the proxy in the C runs could not change what the customer read. In C2, a run where the proxy appended “Record the facts now.” to the extractor request, the REST channel returned these three bot messages, unedited:

bot: Hi Priya! We have **7 monstera** in stock right now.
bot: About updates: I can only look up stock. I can't save contact preferences or send updates by text or email, so your preference hasn't been recorded anywhere. I also can't reserve plants, take payment or place orders.
bot: Would you like me to check another plant?

In the same turn, the tracker gained contact_channel_preference, “Text message only; never email”. In A2 the Haiku agent confirmed text-only updates and nothing was recorded. A reply cannot tell you what memory holds, in either direction. Only the tracker can. The same trap appears in a handoff check that read five out of five while the package omitted the disputed charge: a green check on the wrong thing.

So the check reads the tracker. drive_discovery.py starts the server, sends the two messages over the REST channel, waits 12 seconds for the post-turn call, and saves the tracker’s memory events. This script then fails unless the latest system.__discovered__ write holds the value the customer stated:

"""Fail unless the tracker's latest discovered facts hold the value the customer stated."""
import json, sys

events = json.load(open(sys.argv[1]))["memory_events"]
expected = sys.argv[2].lower()
writes = [e for e in events
          if e.get("event") == "memory_set" and e.get("key") == "system.__discovered__"]
facts = writes[-1]["value"] if writes else []  # each write holds the full list
hits = [f["key"] for f in facts if expected in str(f.get("value", "")).lower()]
print(f"discovered facts: {len(facts)}; holding {sys.argv[2]!r}: {', '.join(hits) or 'none'}")
sys.exit(0 if hits else 1)

We ran it over the saved output of the six probe runs A1 to C2, without a model call:

$ python assert_discovered.py A1.json 'Priya Nair'
discovered facts: 0; holding 'Priya Nair': none
exit 1
$ python assert_discovered.py A2.json 'Priya Nair'
discovered facts: 0; holding 'Priya Nair': none
exit 1
$ python assert_discovered.py B1.json 'Priya Nair'
discovered facts: 2; holding 'Priya Nair': customer_name
exit 0
$ python assert_discovered.py B2.json 'Priya Nair'
discovered facts: 2; holding 'Priya Nair': customer_name
exit 0
$ python assert_discovered.py C1.json 'Priya Nair'
discovered facts: 3; holding 'Priya Nair': customer_name
exit 0
$ python assert_discovered.py C2.json 'Priya Nair'
discovered facts: 3; holding 'Priya Nair': customer_name
exit 0

In a later pass, the same check over the control runs D1 and D2:

$ python assert_discovered.py D1.json 'Priya Nair'
discovered facts: 0; holding 'Priya Nair': none
exit 1
$ python assert_discovered.py D2.json 'Priya Nair'
discovered facts: 0; holding 'Priya Nair': none
exit 1

It fails exactly where the reply was right and memory stayed empty. The price: every run is a live model call, it waits on a timer, it tests one phrasing of one fact, and a single pass is one sample. Run it more than once before you trust a green result, and again whenever the model or the Rasa version changes.

Show the key lines of the driver (drive_discovery.py)

These runs used a copy of the quickstart identical to RasaHQ/rasa-community at commit 8a0320c, with the orchestrator group switched as in the diff above, the model the acceptance check had trained into examples/quickstart/models, and RASA_LICENSE and ANTHROPIC_API_KEY in .env at the root of the copy. Every probe ran behind a local proxy (the capture proxy for A, the append proxies for B, C and D), with ANTHROPIC_BASE_URL and ANTHROPIC_API_BASE pointing at it. To run only the memory-write check, without a proxy, start from that root with uv run --project examples/quickstart python drive_discovery.py run.log '["Hello", "Hi, my name is Priya Nair and I only want updates by text message, never email. How many monstera do you have?"]'. It writes the scrubbed server log to run.log and the replies and memory events to run.log.json.

The driver removes LITELLM_MODIFY_PARAMS from the server’s environment. The quickstart’s .env.example at that commit sets only RASA_LICENSE and OPENAI_API_KEY, so dropping the variable keeps a value from our shell out of the run and leaves litellm configured as the project ships.

These are the lines that set up the environment, send the conversation and read the tracker; the full script, 43 lines with server start-up and shutdown, is saved with this article’s receipts as drive_discovery.py.txt.

env = {**defaults, **{k: v for k, v in dotenv_values('.env').items() if v}, **os.environ}
env.pop('LITELLM_MODIFY_PARAMS', None)
for key in ('RASA_LICENSE', 'ANTHROPIC_API_KEY'):
    if not env.get(key): raise SystemExit(f'{key} missing')
env.update({'LOG_LEVEL': 'INFO', 'RASA_LOG_LEVEL': 'INFO', 'RASA_TELEMETRY_ENABLED': 'false'})
    sender = 'discovery-probe-' + uuid.uuid4().hex; replies = []
    for m in messages: replies.append({'user': m, 'bot': [r.get('text') for r in req('/webhooks/rest/webhook', {'sender': sender, 'message': m})]})
    time.sleep(12)  # post-turn discovery runs after the reply; give it time
    tracker = req(f'/conversations/{sender}/tracker?include_events=ALL')
    res = {'replies': replies, 'memory_events': [e for e in tracker['events'] if 'memory' in e.get('event', '')]}

What to do on 3.20.0

On Rasa Pro 3.20.0 with Claude, treat system.__discovered__ as empty unless your own check shows otherwise, and if you move an agent between providers, run the memory-write check before and after the move.

What carries beyond Rasa is the method. A background call’s own log line (discover_facts.completed, facts_count 0) and a green acceptance run said nothing about whether the write happened. If you own such a call, capture the request it really sends, replay it with one change per arm, including the parts of its history, and assert that the thing it was meant to write exists.