Guide · AI product manager
Is an MCP integration a drop-in replacement?
Is an MCP integration a drop-in replacement? A green proof shows the instructions held. Audit where each tool gets its customer id before you scope the swap.
Key takeaways (5)
- A byte-identical check proves the skill’s instruction text did not change. It says nothing about whether that text still describes what the tool underneath does, or about the arguments the model must now supply.
- The companion proof checks text, and not all of it: moving a confirmation constraint off the write tool in a skill’s frontmatter still passes every check.
- In the companion CRM agent, both tools that used to read the customer id from memory take it as an argument after the swap: one reads tickets, the other writes notes to a customer record.
- Read in the Rasa Pro 3.20.0rc1 source, a mistyped remote tool name is not caught until the agent starts and connects to the MCP server. The companion proof tests its own loopback server, so it cannot stand in for that start.
- Split the work by owner: the config swap is one ordinary ticket, and each tool whose customer id became an argument gets its own ticket with a named security reviewer.
An engineer on Meridian’s support team opens a pull request: move the CRM tools
onto the vendor’s MCP server. The companion project makes that change with
make mcp-swap (Makefile, lines 101–113). Each of the three skill files is
replaced by an MCP twin whose only frontmatter change is the import:
check_tickets and log_interaction gain an import_tools: key and one
mcp/ entry, and identify_customer has its one import line rewritten to
- mcp/hubspot_crm:find_contact_by_email. The proof below reports “2 changed
line(s)” for each because it counts removed plus added lines.
integrations.yml gains an mcp_servers: block, and three Python files are
moved aside:
tools/crm.py, skills/check_tickets/tools.py and
skills/log_interaction/tools.py. As review evidence, the pull request carries
the output of the companion’s own proof, make mcp-prove. This is that output,
from an offline run at the pinned revision, with no model
and no API key:
Proving: the same three skills keep their instructions while
tools/crm.py is replaced by import_tools: mcp/<server>:<tool>.
No licence, no API key, no HubSpot account, loopback only.
1. skill instructions are unchanged by the swap
✓ check_tickets: only import_tools differs 2 changed line(s)
✓ check_tickets: instruction body byte-identical 687 bytes
✓ identify_customer: only import_tools differs 2 changed line(s)
✓ identify_customer: instruction body byte-identical 804 bytes
✓ log_interaction: only import_tools differs 2 changed line(s)
✓ log_interaction: instruction body byte-identical 1226 bytes
2. integrations.yml parses as MCP server configuration
✓ parse_mcp_servers accepts the file servers: ['hubspot_crm']
✓ server url is loopback http http://127.0.0.1:8931/mcp
3. each skill's import parses as mcp/<server>:<tool>
✓ check_tickets imports list_open_tickets server=hubspot_crm llm sees 'list_open_tickets'
✓ identify_customer imports find_contact_by_email server=hubspot_crm llm sees 'find_contact_by_email'
✓ log_interaction imports add_timeline_note server=hubspot_crm llm sees 'add_timeline_note'
4. every imported server is configured
✓ no skill imports an unconfigured server ['hubspot_crm'] ⊆ ['hubspot_crm']
5. the MCP server exposes the imported tools
✓ server exposes list_open_tickets imported by check_tickets
✓ server exposes find_contact_by_email imported by identify_customer
✓ server exposes add_timeline_note imported by log_interaction
6. a tool call over MCP returns the CRM fact
✓ result is structured, not text blocks outputSchema published
✓ find_contact_by_email over MCP Dana Okafor at Okafor Logistics
✓ absent contact is not an error error=contact_not_found
The transport changed. The instructions did not.
The product manager approves it as a same-sprint config change. Eighteen green checks, and the last line says what the pull request claims.
Now read the 687 bytes that check 1 proved unchanged. The MCP variant of
skills/check_tickets/skill.md still tells the model, at lines 12–13:
Call `list_open_tickets`. It reads the identified customer from project memory,
so it is the authority on whether the caller has been identified yet.
The memory read that sentence describes lived in skills/check_tickets/tools.py,
one of the files the swap moves aside. The MCP tool that now answers to the
name does not read memory at all (scripts/mcp_crm_server.py, lines 107–118):
@mcp.tool(description="List the support tickets on the identified customer's account.")
async def list_open_tickets(contact_id: str) -> CrmResult:
"""Read the caller's tickets from the CRM.
Args:
contact_id: HubSpot contact id of the identified customer.
"""
# `contact_id` is a PARAMETER here, not a memory read. An MCP tool has no
# ToolContext, so the value has to travel as an argument. See the README
# section "What MCP does not do for you".
if not contact_id:
return CrmResult(ok=False, error="not_identified")
The proof checked that the sentence survived the swap. Nothing in it checked whether the sentence is still true.
Meridian, its pull request and the approval are an illustrative scenario. The
code and output are real: tutorials/rasa-hubspot-crm-tutorial in
RasaHQ/rasa-community-resources at revision 69e27b6, where the agent is
called Ora (agent.yml, lines 5 and 13), run against Rasa Pro 3.20.0rc1. The
MCP server in that project is a local mock, and nothing here describes any
vendor’s real server. The engineer’s walkthrough is the
CRM transport swap tutorial; this
guide covers the approval, and it disagrees with that tutorial in places,
listed chapter by chapter further down.
Is an MCP integration a drop-in replacement?
Check 1 is the evidence the approval leaned on, so read what it accepts. It
diffs each REST skill file against its MCP twin and collects the changed lines
(scripts/prove_mcp_swap.py, lines 89–118). A changed line passes if it is
import_tools: or starts with - (lines 94–100):
# Every changed line must be an import declaration: either the
# `import_tools:` key itself or one of its list entries.
offending = [
line
for line in changed
if line != "import_tools:" and not line.startswith("- ")
]
It then compares the text below the frontmatter byte for byte (lines 109–114):
# And the body below the frontmatter must be identical, byte for byte.
rest_body = _read(f"skills/{skill}/skill.md").split("---\n", 2)[-1]
mcp_body = _read(f"mcp_variant/skills/{skill}/skill.md").split("---\n", 2)[-1]
check(
f"{skill}: instruction body byte-identical",
rest_body == mcp_body,
The body comparison is real: a reworded instruction would turn it red, as the
script’s docstring says (lines 26–28). The
frontmatter test is looser than its comment. Any list entry passes, whatever
key it sits under, and tool_constraints: entries are list entries too. We
changed one line in the MCP log_interaction skill, moving its confirmation
constraint from the write tool to the ticket read:
@@ -6,7 +6,7 @@
import_tools:
- mcp/hubspot_crm:add_timeline_note
tool_constraints:
- - add_timeline_note:
+ - list_open_tickets:
requires_confirmation:
enabled: true
utter_for_confirmation: utter_confirm_note
The proof still passed. Its line for that skill became
✓ log_interaction: only import_tools differs 4 changed line(s), every other
check stayed green, the last line still read “The transport changed. The
instructions did not.”, and make mcp-prove exited 0. The confirmation
declaration had moved off the write tool, and the proof counted the two extra
changed lines as import declarations.
So the diff and the green run answer a narrow question: did the instruction text move? The scoping decision turns on a different one: for each tool, did the value it relies on stay where the caller cannot influence it? A clean diff cannot stand in for that audit, and the audit, not the size of the diff, decides whether this is one ticket or several.
Which value moved, tool by tool
Chapter 3 of the tutorial was the first to name the move for the ticket read:
list_open_tickets stopped reading the id from memory and started taking it as
an argument. The same move happened to the write tool; none of the tutorial’s
chapters says so.
In the REST version, the two tools that act on an identified customer read the
customer id from project memory, where the lookup tool wrote it. The ticket
read (skills/check_tickets/tools.py, lines 10–12):
async def list_open_tickets(context: ToolContext = None) -> ToolResult:
"""Read the caller's tickets from the CRM."""
contact_id = context.memory.get("project.contact_id") if context else None
The note write (skills/log_interaction/tools.py, lines 10–16), whose only
argument is the summary:
async def add_timeline_note(summary: str, context: ToolContext = None) -> ToolResult:
"""Write a note onto the customer's CRM record.
Args:
summary: One or two sentences describing what the caller wanted.
"""
contact_id = context.memory.get("project.contact_id") if context else None
Over MCP there is no ToolContext, so the write tool takes the id as an
argument too, and writes the note to whichever contact it is given
(scripts/mcp_crm_server.py, lines 129–139):
async def add_timeline_note(contact_id: str, summary: str) -> CrmResult:
"""Write a note onto the customer's CRM record.
Args:
contact_id: HubSpot contact id of the identified customer.
summary: One or two sentences describing what the caller wanted.
"""
if not contact_id:
return CrmResult(ok=False, error="not_identified")
try:
note_id = await log_note(str(contact_id), summary)
You do not need to read Python to check this for your own vendor. Ask the engineer to list each MCP tool’s required arguments. In a copy of the companion project, that is one command and three lines of output (offline, no model; the command discards a library warning on stderr):
$ uv run python -c 'import asyncio, scripts.mcp_crm_server as s
for t in asyncio.run(s.mcp.list_tools()): print(t.name, "requires", t.inputSchema["required"])' 2>/dev/null
find_contact_by_email requires ['email']
list_open_tickets requires ['contact_id']
add_timeline_note requires ['contact_id', 'summary']
A contact_id in the required list is a value the caller of the tool must
supply. In both versions the caller chooses the email the lookup searches for.
The REST lookup then wrote the id it found into memory (tools/crm.py, lines
42–45), and the REST tools act only on that id: one a lookup produced. The MCP
lookup can only return the id in its result (scripts/mcp_crm_server.py,
lines 99–104), and the MCP tools act on whatever id they are passed. Who
passes it is stated by the companion README: the model “fills it from the
conversation” (README.md, lines 264–266), and the caller influences the
conversation.
The lookup’s input did not move. Both tools that act on a customer did.
The per-tool checklist
Take this table into the scoping meeting, with one row per tool the vendor’s
server will serve. Every row asks the question from the README’s own rule: does
the tool’s correctness depend on a value “the user must not be able to
influence” (README.md, lines 267–269)?
| Tool | Acts on | Input before (REST) | Input after (MCP) | If the input is wrong | Moved in the swap? |
|---|---|---|---|---|---|
find_contact_by_email | Looks a customer up | Caller’s email | Caller’s email | The same in both versions | No |
list_open_tickets | Reads tickets | Id the lookup wrote to memory | Id passed as an argument | The agent reads out another person’s tickets | Yes: needs a decision |
add_timeline_note | Writes to a CRM record | Id the lookup wrote to memory | Id passed as an argument | A note lands on another person’s record | Yes: needs a decision, and it is a write |
The last two columns are where the approver’s attention goes. A wrong read discloses; a wrong write changes someone else’s system of record. Both need a decision the diff does not contain.
The write step, as shipped, does not fit the new schema
There is a second problem on the write side. log_interaction fixes its write
in an ordered block, and the MCP variant keeps that block unchanged. The write
step passes one parameter (mcp_variant/skills/log_interaction/skill.md,
lines 33–36):
- id: write_note
execute_tool: add_timeline_note
parameters:
summary: session.log_interaction.note_summary
The MCP tool requires contact_id as well. In the Rasa Pro source, an execute
step’s arguments are validated against the schema the tool publishes
(rasa/mantle/orchestration/tool_execution/invoker.py, lines 63–81). On a
failure the step does not call the tool; it hands the turn back to the model
(rasa/mantle/orchestration/tool_execution/constraints.py, lines 695–706):
validation_error = self._tool_invoker.validate_builder_tool_arguments(
active_skill_id,
step.execute_tool,
tool_func,
resolved_parameters,
)
if validation_error is not None:
structlogger.warning(
"mantle.skill_executor.advance_steps.tool_arguments_invalid",
tool=step.execute_tool,
)
return ExecuteStepControl.YIELD_TO_LLM
We ran Rasa Pro’s own validator offline against the schema the MCP server
publishes, converted the way MCPRuntime.prepare converts it, with the
arguments the step passes and then with contact_id added:
write_note step as shipped: {'summary': 'Dana asked about invoice 4471.'}
-> {"error": "Invalid tool arguments: Missing required argument(s): 'contact_id'.", "retryable": true, "code": "invalid_arguments"}
with contact_id added: {'contact_id': '101', 'summary': 'Dana asked about invoice 4471.'}
-> None
The shipped step fails the check. That check sits before the confirmation
logic, which starts at line 708, so on this path the caller is never asked.
That ordering is a reading of the source; we did not run the agent against a
model for this guide, so what a live run does next is untested. The model may stall on the step, or it may call
add_timeline_note itself with a contact_id it chose. Either outcome needs a
ticket, and neither is in the pull request.
Why a green build is not the release gate
The second thing the pull request cannot show is whether the vendor’s server
offers the tools the skills import. What follows is a reading of the Rasa Pro
3.20.0rc1 source, not a run. When a model is loaded, parse_mcp_imports is
called (rasa/mantle/model_archive/bootstrap.py, line 208). It parses each
mcp/<server>:<tool> import with try_parse_mcp_tool_import
(rasa/mantle/tools/mcp_import_spec.py, lines 34–55, called at line 86) but
does not check whether the tool exists (lines 74–76):
Remote existence is not checked here: that requires ``list_tools`` at
connection time. Collisions with local/shared callables and reserved
framework names are static and fail at model load.
The existence check runs when the agent starts. Loading an agent calls
await agent.prepare_runtime_integrations() (rasa/core/agent.py, line 421).
That method (lines 664–672) calls _prepare_runtime_integrations (lines
141–148), which calls the processor’s own prepare_runtime_integrations
(rasa/mantle/processor.py, lines 546–558), which runs
await mcp_runtime.prepare(...): MCPRuntime.prepare in
rasa/mantle/tools/mcp_runtime.py, lines 87–131. That connects to each server, calls list_tools (line 124) and passes
the result to _filter_imported_tools (line 131), which raises on a missing
tool (lines 237–241):
if tool_schema is None:
raise ValueError(
f"MCP server {reference.server_id!r} does not expose imported "
f"tool {reference.tool_name!r} for skill {skill_id!r}."
)
So a mistyped tool name is not caught until the agent starts. That is a loud failure rather than a quiet one, but it lands after the build and the proof are green.
The proof stays useful for what it covers: the skill text and the companion’s
mock. The check that covers the vendor is the agent starting cleanly against
the vendor’s endpoint, outside production, with the same server id and import
lines. The proof needs uv (Makefile, lines 90–91) and no licence, key or
account. These are the steps from a clone; we ran the proof once, on macOS, in
a git archive export of the same revision rather than a fresh clone:
git clone https://github.com/RasaHQ/rasa-community-resources.git
cd rasa-community-resources
git checkout 69e27b6
cd tutorials/rasa-hubspot-crm-tutorial
make mcp-prove
We did not run it on Linux or Windows. The target’s command is
uv run python scripts/prove_mcp_swap.py.
Split the ticket by owner
The config swap itself is ordinary work. The boundary decisions are not, and they need a reviewer who is not the engineer who wrote the diff. If you scoped the workflow with a written boundary, as in choosing an agent workflow worth building, this is where that boundary gets checked against the tools. A split for the companion project:
| Ticket | Covers | Owner | Reviewer | Done when |
|---|---|---|---|---|
| Config swap | mcp_servers: block, import_tools lines | Engineer | Any engineer | make mcp-prove passes on the loopback configuration, and the agent starts against the vendor endpoint outside production without an “imported tool” error |
| Ticket read boundary | list_open_tickets takes its contact_id as an argument | Engineer | Named security reviewer | One closure from the table below, signed off |
| Note write boundary | add_timeline_note takes its contact_id as an argument, and the write step fails its argument check | Engineer | Named security reviewer | One closure from the table below, signed off |
| Instruction accuracy | check_tickets/skill.md lines 12–13 describe a memory read | Engineer and product manager | Conversation designer or product manager | The sentence matches the tool that ships |
Each boundary ticket closes one of two ways, and the reviewer picks one per tool:
| Closure | What the reviewer accepts | Done when | Knock-on effects |
|---|---|---|---|
| Keep the tool local | Nothing new: the tool still reads the id from memory | The tool and its REST skill frontmatter stay; the check_tickets sentence stays true | The lookup must stay local too, because it is the only code that writes the id to memory. Keep both boundary tools local and none of the three CRM tools moves without new code. Record it as a scope reduction |
| Accept the argument | In writing, what a wrong id costs for this tool: another person’s tickets read out, or a note on another person’s record | Conversation tests where the caller names another customer; for the write, the write_note step fixed to satisfy the schema and tested in a live run | The check_tickets sentence is rewritten, so check 1 of the proof goes red and the proof is updated with it |
The first closure is the README’s own rule. It is also, in this project, a decision about the whole swap rather than one tool.
Where this guide and the tutorial disagree
The tutorial’s chapters are the engineer’s reference for this swap. On these points the evidence above says something different:
| Chapter | The tutorial says | What the evidence shows |
|---|---|---|
| 3, What MCP takes away | Same name, keys and error string, “which is why the instructions still work” | The instruction text is unchanged, but check_tickets/skill.md lines 12–13 now describe a memory read the tool does not do |
| 3, What MCP takes away | The tutorial “can afford the looser arrangement” for the ticket read, because it reads the caller’s own record and the data is fixture data | Whose record it reads is exactly what the argument no longer guarantees, and fixture data is a property of the tutorial, not of the swap. A wrong id reads out another person’s tickets, which is a decision for the reviewer |
| 3, What MCP takes away | A missing remote tool is “the specific gap make mcp-prove closes”: its check 5 calls list_tools “before you spend a training run” | Check 5 lists the tools of the proof’s own loopback mock (prove_mcp_swap.py, lines 147–154 and 193–207), not the server the agent will use. It cannot show that a vendor’s server offers the imported tools |
| 4, What survives anyway | The write_note step “is otherwise unchanged”, and “the caller still sees this before anything is written” | The step passes only summary, Rasa Pro’s validator rejects that for the MCP tool (“Missing required argument(s): ‘contact_id’”), and the check (constraints.py, lines 695–706) runs before the confirmation logic (line 708) |
| 4, What survives anyway | The swap “preserves that interface exactly” | Two tools’ argument lists changed: see the listing above |
| 5, Shaping for the channel | Instruction prose and tool constraints are “portable across the swap”; the one thing not portable is “the tool” reading project.contact_id | Two tools read it, one of them a write. And the proof would pass a constraint moved to the wrong tool |
Chapter 3 is still the right place to learn why an MCP tool cannot read memory, and its rule for a tool that fails the check is the one this guide uses. Its claim that the proof closes the start-time gap is the one this guide does not accept.
Limits of this evidence
- It shows where the customer id comes from in one companion project. It says nothing about how any vendor’s MCP server behaves.
- The boundary question is one question. It is not a security review of the vendor, the transport or the data.
- Neither version checks that the caller owns the email they give. That is identity verification, a separate decision this guide does not cover.
- The Rasa Pro behaviour described here comes from reading the source. No live Mantle run was made, so whether a model passes the right id is untested.
- It gives no estimate of hours or points. The split tells you who reviews what, not how long it takes.
Three questions an approver will hear
The engineer says the model always passes the id the lookup returned. Is that enough?
It is a claim about model behaviour, and the companion project has no test files to check it against. The REST version did not depend on it: the tools read the id the lookup had written to memory, not one taken from the conversation. If the argument is that the model will behave, ask for the conversation tests that show it, including a caller who names another customer, and decide whether a sampled result is enough for a write.
Can we move only the lookup to MCP and keep the other two tools local?
Not as the companion is written. The only code that writes the customer id to
memory is the REST lookup (tools/crm.py, line 43). The project’s
memory.yml declares contact_id without llm_settable (lines 8–10), and
Rasa Pro 3.20.0rc1 defaults that setting to false
(rasa/mantle/memory/field.py, line 52), so the model cannot set the id
through set_fields either. A search of the whole project at 69e27b6 finds
no other writer. Move the lookup to MCP and nothing writes the id: the local
ticket and note tools find none and return not_identified
(skills/check_tickets/tools.py, lines 12–14, and
skills/log_interaction/tools.py, lines 16–18). Plan a partial swap as its
own piece of work.
Does any of this change with a real vendor server?
The question stays the same, and the answers come from the vendor’s schema rather than the mock’s. Ask for the listing of required arguments for every tool the vendor publishes, and fill in the checklist from that. Any customer or account identifier in a required list is a value your agent will have to pass.