Skip to content
RasaGet a free licence
Guides for AI teams

Guide · AI product manager

Is an MCP integration a drop-in replacement?

Is an MCP integration a drop-in replacement? A green proof shows the instructions held. Audit where each tool gets its customer id before you scope the swap.

by Rod Rivera

About 9 minutes

Key takeaways (5)
  • A byte-identical check proves the skill’s instruction text did not change. It says nothing about whether that text still describes what the tool underneath does, or about the arguments the model must now supply.
  • The companion proof checks text, and not all of it: moving a confirmation constraint off the write tool in a skill’s frontmatter still passes every check.
  • In the companion CRM agent, both tools that used to read the customer id from memory take it as an argument after the swap: one reads tickets, the other writes notes to a customer record.
  • Read in the Rasa Pro 3.20.0rc1 source, a mistyped remote tool name is not caught until the agent starts and connects to the MCP server. The companion proof tests its own loopback server, so it cannot stand in for that start.
  • Split the work by owner: the config swap is one ordinary ticket, and each tool whose customer id became an argument gets its own ticket with a named security reviewer.

An engineer on Meridian’s support team opens a pull request: move the CRM tools onto the vendor’s MCP server. The companion project makes that change with make mcp-swap (Makefile, lines 101–113). Each of the three skill files is replaced by an MCP twin whose only frontmatter change is the import: check_tickets and log_interaction gain an import_tools: key and one mcp/ entry, and identify_customer has its one import line rewritten to - mcp/hubspot_crm:find_contact_by_email. The proof below reports “2 changed line(s)” for each because it counts removed plus added lines. integrations.yml gains an mcp_servers: block, and three Python files are moved aside: tools/crm.py, skills/check_tickets/tools.py and skills/log_interaction/tools.py. As review evidence, the pull request carries the output of the companion’s own proof, make mcp-prove. This is that output, from an offline run at the pinned revision, with no model and no API key:

  Proving: the same three skills keep their instructions while
  tools/crm.py is replaced by import_tools: mcp/<server>:<tool>.
  No licence, no API key, no HubSpot account, loopback only.

  1. skill instructions are unchanged by the swap
  ✓ check_tickets: only import_tools differs     2 changed line(s)
  ✓ check_tickets: instruction body byte-identical 687 bytes
  ✓ identify_customer: only import_tools differs 2 changed line(s)
  ✓ identify_customer: instruction body byte-identical 804 bytes
  ✓ log_interaction: only import_tools differs   2 changed line(s)
  ✓ log_interaction: instruction body byte-identical 1226 bytes

  2. integrations.yml parses as MCP server configuration
  ✓ parse_mcp_servers accepts the file           servers: ['hubspot_crm']
  ✓ server url is loopback http                  http://127.0.0.1:8931/mcp

  3. each skill's import parses as mcp/<server>:<tool>
  ✓ check_tickets imports list_open_tickets      server=hubspot_crm llm sees 'list_open_tickets'
  ✓ identify_customer imports find_contact_by_email server=hubspot_crm llm sees 'find_contact_by_email'
  ✓ log_interaction imports add_timeline_note    server=hubspot_crm llm sees 'add_timeline_note'

  4. every imported server is configured
  ✓ no skill imports an unconfigured server      ['hubspot_crm'] ⊆ ['hubspot_crm']

  5. the MCP server exposes the imported tools
  ✓ server exposes list_open_tickets             imported by check_tickets
  ✓ server exposes find_contact_by_email         imported by identify_customer
  ✓ server exposes add_timeline_note             imported by log_interaction

  6. a tool call over MCP returns the CRM fact
  ✓ result is structured, not text blocks        outputSchema published
  ✓ find_contact_by_email over MCP               Dana Okafor at Okafor Logistics
  ✓ absent contact is not an error               error=contact_not_found

  The transport changed. The instructions did not.

The product manager approves it as a same-sprint config change. Eighteen green checks, and the last line says what the pull request claims.

Now read the 687 bytes that check 1 proved unchanged. The MCP variant of skills/check_tickets/skill.md still tells the model, at lines 12–13:

Call `list_open_tickets`. It reads the identified customer from project memory,
so it is the authority on whether the caller has been identified yet.

The memory read that sentence describes lived in skills/check_tickets/tools.py, one of the files the swap moves aside. The MCP tool that now answers to the name does not read memory at all (scripts/mcp_crm_server.py, lines 107–118):

@mcp.tool(description="List the support tickets on the identified customer's account.")
async def list_open_tickets(contact_id: str) -> CrmResult:
    """Read the caller's tickets from the CRM.

    Args:
        contact_id: HubSpot contact id of the identified customer.
    """
    # `contact_id` is a PARAMETER here, not a memory read. An MCP tool has no
    # ToolContext, so the value has to travel as an argument. See the README
    # section "What MCP does not do for you".
    if not contact_id:
        return CrmResult(ok=False, error="not_identified")

The proof checked that the sentence survived the swap. Nothing in it checked whether the sentence is still true.

Meridian, its pull request and the approval are an illustrative scenario. The code and output are real: tutorials/rasa-hubspot-crm-tutorial in RasaHQ/rasa-community-resources at revision 69e27b6, where the agent is called Ora (agent.yml, lines 5 and 13), run against Rasa Pro 3.20.0rc1. The MCP server in that project is a local mock, and nothing here describes any vendor’s real server. The engineer’s walkthrough is the CRM transport swap tutorial; this guide covers the approval, and it disagrees with that tutorial in places, listed chapter by chapter further down.

Is an MCP integration a drop-in replacement?

Check 1 is the evidence the approval leaned on, so read what it accepts. It diffs each REST skill file against its MCP twin and collects the changed lines (scripts/prove_mcp_swap.py, lines 89–118). A changed line passes if it is import_tools: or starts with - (lines 94–100):

        # Every changed line must be an import declaration: either the
        # `import_tools:` key itself or one of its list entries.
        offending = [
            line
            for line in changed
            if line != "import_tools:" and not line.startswith("- ")
        ]

It then compares the text below the frontmatter byte for byte (lines 109–114):

        # And the body below the frontmatter must be identical, byte for byte.
        rest_body = _read(f"skills/{skill}/skill.md").split("---\n", 2)[-1]
        mcp_body = _read(f"mcp_variant/skills/{skill}/skill.md").split("---\n", 2)[-1]
        check(
            f"{skill}: instruction body byte-identical",
            rest_body == mcp_body,

The body comparison is real: a reworded instruction would turn it red, as the script’s docstring says (lines 26–28). The frontmatter test is looser than its comment. Any list entry passes, whatever key it sits under, and tool_constraints: entries are list entries too. We changed one line in the MCP log_interaction skill, moving its confirmation constraint from the write tool to the ticket read:

@@ -6,7 +6,7 @@
 import_tools:
   - mcp/hubspot_crm:add_timeline_note
 tool_constraints:
-  - add_timeline_note:
+  - list_open_tickets:
       requires_confirmation:
         enabled: true
         utter_for_confirmation: utter_confirm_note

The proof still passed. Its line for that skill became ✓ log_interaction: only import_tools differs 4 changed line(s), every other check stayed green, the last line still read “The transport changed. The instructions did not.”, and make mcp-prove exited 0. The confirmation declaration had moved off the write tool, and the proof counted the two extra changed lines as import declarations.

So the diff and the green run answer a narrow question: did the instruction text move? The scoping decision turns on a different one: for each tool, did the value it relies on stay where the caller cannot influence it? A clean diff cannot stand in for that audit, and the audit, not the size of the diff, decides whether this is one ticket or several.

Which value moved, tool by tool

Chapter 3 of the tutorial was the first to name the move for the ticket read: list_open_tickets stopped reading the id from memory and started taking it as an argument. The same move happened to the write tool; none of the tutorial’s chapters says so.

In the REST version, the two tools that act on an identified customer read the customer id from project memory, where the lookup tool wrote it. The ticket read (skills/check_tickets/tools.py, lines 10–12):

async def list_open_tickets(context: ToolContext = None) -> ToolResult:
    """Read the caller's tickets from the CRM."""
    contact_id = context.memory.get("project.contact_id") if context else None

The note write (skills/log_interaction/tools.py, lines 10–16), whose only argument is the summary:

async def add_timeline_note(summary: str, context: ToolContext = None) -> ToolResult:
    """Write a note onto the customer's CRM record.

    Args:
        summary: One or two sentences describing what the caller wanted.
    """
    contact_id = context.memory.get("project.contact_id") if context else None

Over MCP there is no ToolContext, so the write tool takes the id as an argument too, and writes the note to whichever contact it is given (scripts/mcp_crm_server.py, lines 129–139):

async def add_timeline_note(contact_id: str, summary: str) -> CrmResult:
    """Write a note onto the customer's CRM record.

    Args:
        contact_id: HubSpot contact id of the identified customer.
        summary: One or two sentences describing what the caller wanted.
    """
    if not contact_id:
        return CrmResult(ok=False, error="not_identified")
    try:
        note_id = await log_note(str(contact_id), summary)

You do not need to read Python to check this for your own vendor. Ask the engineer to list each MCP tool’s required arguments. In a copy of the companion project, that is one command and three lines of output (offline, no model; the command discards a library warning on stderr):

$ uv run python -c 'import asyncio, scripts.mcp_crm_server as s
for t in asyncio.run(s.mcp.list_tools()): print(t.name, "requires", t.inputSchema["required"])' 2>/dev/null
find_contact_by_email requires ['email']
list_open_tickets requires ['contact_id']
add_timeline_note requires ['contact_id', 'summary']

A contact_id in the required list is a value the caller of the tool must supply. In both versions the caller chooses the email the lookup searches for. The REST lookup then wrote the id it found into memory (tools/crm.py, lines 42–45), and the REST tools act only on that id: one a lookup produced. The MCP lookup can only return the id in its result (scripts/mcp_crm_server.py, lines 99–104), and the MCP tools act on whatever id they are passed. Who passes it is stated by the companion README: the model “fills it from the conversation” (README.md, lines 264–266), and the caller influences the conversation.

Caller gives an email find_contact_by_email(email) same input on both sides project.contact_id (memory, REST only) REST: writes id Model fills contact_id (MCP) MCP: returns id list_open_tickets reads tickets REST add_timeline_note writes a note REST MCP MCP
FigureWhere each CRM tool gets the customer id, before and after the swap

The lookup’s input did not move. Both tools that act on a customer did.

The per-tool checklist

Take this table into the scoping meeting, with one row per tool the vendor’s server will serve. Every row asks the question from the README’s own rule: does the tool’s correctness depend on a value “the user must not be able to influence” (README.md, lines 267–269)?

ToolActs onInput before (REST)Input after (MCP)If the input is wrongMoved in the swap?
find_contact_by_emailLooks a customer upCaller’s emailCaller’s emailThe same in both versionsNo
list_open_ticketsReads ticketsId the lookup wrote to memoryId passed as an argumentThe agent reads out another person’s ticketsYes: needs a decision
add_timeline_noteWrites to a CRM recordId the lookup wrote to memoryId passed as an argumentA note lands on another person’s recordYes: needs a decision, and it is a write

The last two columns are where the approver’s attention goes. A wrong read discloses; a wrong write changes someone else’s system of record. Both need a decision the diff does not contain.

The write step, as shipped, does not fit the new schema

There is a second problem on the write side. log_interaction fixes its write in an ordered block, and the MCP variant keeps that block unchanged. The write step passes one parameter (mcp_variant/skills/log_interaction/skill.md, lines 33–36):

  - id: write_note
    execute_tool: add_timeline_note
    parameters:
      summary: session.log_interaction.note_summary

The MCP tool requires contact_id as well. In the Rasa Pro source, an execute step’s arguments are validated against the schema the tool publishes (rasa/mantle/orchestration/tool_execution/invoker.py, lines 63–81). On a failure the step does not call the tool; it hands the turn back to the model (rasa/mantle/orchestration/tool_execution/constraints.py, lines 695–706):

        validation_error = self._tool_invoker.validate_builder_tool_arguments(
            active_skill_id,
            step.execute_tool,
            tool_func,
            resolved_parameters,
        )
        if validation_error is not None:
            structlogger.warning(
                "mantle.skill_executor.advance_steps.tool_arguments_invalid",
                tool=step.execute_tool,
            )
            return ExecuteStepControl.YIELD_TO_LLM

We ran Rasa Pro’s own validator offline against the schema the MCP server publishes, converted the way MCPRuntime.prepare converts it, with the arguments the step passes and then with contact_id added:

write_note step as shipped: {'summary': 'Dana asked about invoice 4471.'}
  -> {"error": "Invalid tool arguments: Missing required argument(s): 'contact_id'.", "retryable": true, "code": "invalid_arguments"}
with contact_id added: {'contact_id': '101', 'summary': 'Dana asked about invoice 4471.'}
  -> None

The shipped step fails the check. That check sits before the confirmation logic, which starts at line 708, so on this path the caller is never asked. That ordering is a reading of the source; we did not run the agent against a model for this guide, so what a live run does next is untested. The model may stall on the step, or it may call add_timeline_note itself with a contact_id it chose. Either outcome needs a ticket, and neither is in the pull request.

Why a green build is not the release gate

The second thing the pull request cannot show is whether the vendor’s server offers the tools the skills import. What follows is a reading of the Rasa Pro 3.20.0rc1 source, not a run. When a model is loaded, parse_mcp_imports is called (rasa/mantle/model_archive/bootstrap.py, line 208). It parses each mcp/<server>:<tool> import with try_parse_mcp_tool_import (rasa/mantle/tools/mcp_import_spec.py, lines 34–55, called at line 86) but does not check whether the tool exists (lines 74–76):

    Remote existence is not checked here: that requires ``list_tools`` at
    connection time. Collisions with local/shared callables and reserved
    framework names are static and fail at model load.

The existence check runs when the agent starts. Loading an agent calls await agent.prepare_runtime_integrations() (rasa/core/agent.py, line 421). That method (lines 664–672) calls _prepare_runtime_integrations (lines 141–148), which calls the processor’s own prepare_runtime_integrations (rasa/mantle/processor.py, lines 546–558), which runs await mcp_runtime.prepare(...): MCPRuntime.prepare in rasa/mantle/tools/mcp_runtime.py, lines 87–131. That connects to each server, calls list_tools (line 124) and passes the result to _filter_imported_tools (line 131), which raises on a missing tool (lines 237–241):

            if tool_schema is None:
                raise ValueError(
                    f"MCP server {reference.server_id!r} does not expose imported "
                    f"tool {reference.tool_name!r} for skill {skill_id!r}."
                )

So a mistyped tool name is not caught until the agent starts. That is a loud failure rather than a quiet one, but it lands after the build and the proof are green.

The proof stays useful for what it covers: the skill text and the companion’s mock. The check that covers the vendor is the agent starting cleanly against the vendor’s endpoint, outside production, with the same server id and import lines. The proof needs uv (Makefile, lines 90–91) and no licence, key or account. These are the steps from a clone; we ran the proof once, on macOS, in a git archive export of the same revision rather than a fresh clone:

git clone https://github.com/RasaHQ/rasa-community-resources.git
cd rasa-community-resources
git checkout 69e27b6
cd tutorials/rasa-hubspot-crm-tutorial
make mcp-prove

We did not run it on Linux or Windows. The target’s command is uv run python scripts/prove_mcp_swap.py.

Split the ticket by owner

The config swap itself is ordinary work. The boundary decisions are not, and they need a reviewer who is not the engineer who wrote the diff. If you scoped the workflow with a written boundary, as in choosing an agent workflow worth building, this is where that boundary gets checked against the tools. A split for the companion project:

TicketCoversOwnerReviewerDone when
Config swapmcp_servers: block, import_tools linesEngineerAny engineermake mcp-prove passes on the loopback configuration, and the agent starts against the vendor endpoint outside production without an “imported tool” error
Ticket read boundarylist_open_tickets takes its contact_id as an argumentEngineerNamed security reviewerOne closure from the table below, signed off
Note write boundaryadd_timeline_note takes its contact_id as an argument, and the write step fails its argument checkEngineerNamed security reviewerOne closure from the table below, signed off
Instruction accuracycheck_tickets/skill.md lines 12–13 describe a memory readEngineer and product managerConversation designer or product managerThe sentence matches the tool that ships

Each boundary ticket closes one of two ways, and the reviewer picks one per tool:

ClosureWhat the reviewer acceptsDone whenKnock-on effects
Keep the tool localNothing new: the tool still reads the id from memoryThe tool and its REST skill frontmatter stay; the check_tickets sentence stays trueThe lookup must stay local too, because it is the only code that writes the id to memory. Keep both boundary tools local and none of the three CRM tools moves without new code. Record it as a scope reduction
Accept the argumentIn writing, what a wrong id costs for this tool: another person’s tickets read out, or a note on another person’s recordConversation tests where the caller names another customer; for the write, the write_note step fixed to satisfy the schema and tested in a live runThe check_tickets sentence is rewritten, so check 1 of the proof goes red and the proof is updated with it

The first closure is the README’s own rule. It is also, in this project, a decision about the whole swap rather than one tool.

Where this guide and the tutorial disagree

The tutorial’s chapters are the engineer’s reference for this swap. On these points the evidence above says something different:

ChapterThe tutorial saysWhat the evidence shows
3, What MCP takes awaySame name, keys and error string, “which is why the instructions still work”The instruction text is unchanged, but check_tickets/skill.md lines 12–13 now describe a memory read the tool does not do
3, What MCP takes awayThe tutorial “can afford the looser arrangement” for the ticket read, because it reads the caller’s own record and the data is fixture dataWhose record it reads is exactly what the argument no longer guarantees, and fixture data is a property of the tutorial, not of the swap. A wrong id reads out another person’s tickets, which is a decision for the reviewer
3, What MCP takes awayA missing remote tool is “the specific gap make mcp-prove closes”: its check 5 calls list_tools “before you spend a training run”Check 5 lists the tools of the proof’s own loopback mock (prove_mcp_swap.py, lines 147–154 and 193–207), not the server the agent will use. It cannot show that a vendor’s server offers the imported tools
4, What survives anywayThe write_note step “is otherwise unchanged”, and “the caller still sees this before anything is written”The step passes only summary, Rasa Pro’s validator rejects that for the MCP tool (“Missing required argument(s): ‘contact_id’”), and the check (constraints.py, lines 695–706) runs before the confirmation logic (line 708)
4, What survives anywayThe swap “preserves that interface exactly”Two tools’ argument lists changed: see the listing above
5, Shaping for the channelInstruction prose and tool constraints are “portable across the swap”; the one thing not portable is “the tool” reading project.contact_idTwo tools read it, one of them a write. And the proof would pass a constraint moved to the wrong tool

Chapter 3 is still the right place to learn why an MCP tool cannot read memory, and its rule for a tool that fails the check is the one this guide uses. Its claim that the proof closes the start-time gap is the one this guide does not accept.

Limits of this evidence

  • It shows where the customer id comes from in one companion project. It says nothing about how any vendor’s MCP server behaves.
  • The boundary question is one question. It is not a security review of the vendor, the transport or the data.
  • Neither version checks that the caller owns the email they give. That is identity verification, a separate decision this guide does not cover.
  • The Rasa Pro behaviour described here comes from reading the source. No live Mantle run was made, so whether a model passes the right id is untested.
  • It gives no estimate of hours or points. The split tells you who reviews what, not how long it takes.

Three questions an approver will hear

The engineer says the model always passes the id the lookup returned. Is that enough?

It is a claim about model behaviour, and the companion project has no test files to check it against. The REST version did not depend on it: the tools read the id the lookup had written to memory, not one taken from the conversation. If the argument is that the model will behave, ask for the conversation tests that show it, including a caller who names another customer, and decide whether a sampled result is enough for a write.

Can we move only the lookup to MCP and keep the other two tools local?

Not as the companion is written. The only code that writes the customer id to memory is the REST lookup (tools/crm.py, line 43). The project’s memory.yml declares contact_id without llm_settable (lines 8–10), and Rasa Pro 3.20.0rc1 defaults that setting to false (rasa/mantle/memory/field.py, line 52), so the model cannot set the id through set_fields either. A search of the whole project at 69e27b6 finds no other writer. Move the lookup to MCP and nothing writes the id: the local ticket and note tools find none and return not_identified (skills/check_tickets/tools.py, lines 12–14, and skills/log_interaction/tools.py, lines 16–18). Plan a partial swap as its own piece of work.

Does any of this change with a real vendor server?

The question stays the same, and the answers come from the vendor’s schema rather than the mock’s. Ask for the listing of required arguments for every tool the vendor publishes, and fill in the checklist from that. Any customer or account identifier in a required list is a value your agent will have to pass.