Skip to content
RasaGet a free licence
Guides for AI teams

Guide · AI product engineer

How to move Deep Agents SQL chat to Mantle

Port a read-only database chat from Deep Agents to Mantle, keeping its query tools and checking the boundaries that matter.

by Rod Rivera

About 5 minutes

Start with the refused write

A returned tool result can still be a refusal

The recorded offline native control refuses DELETE FROM Example on a temporary test database, but its runtime error flag is false. A flag-only checker accepts it.

Faulty check · runtime flag onlyAcceptedis_error = false
Check the application result tooRejectedstatus = error

This is a controlled native-wrapper and SDK-recorder test, not a model conversation or the full chat orchestration path. The procedure below includes the actual output and its scope. The Chinook refusal below is a separate direct-library example.

A database chat, rebuilt

Two agents. The same query tools.

Follow one database task from Deep Agents to Mantle. Keep the Python query library; change the bindings and the skill that calls it.

See the code
Keep this layerShared Python toolsList tables · inspect schema
check SQL · query rows
Read-only checks live here
Public sampleChinookInvoice.BillingCountry
Invoice.Total
Both tutorial builds call the same library. This map covers the database task, not the full Deep Agents harness.

Explore before you install

Inspect the query and its result

Offline · no API key

These saved outputs come directly from the shared library and the pinned public database. They are not agent replies. Your browser neither executes SQL nor calls a model.

01Read invoice totalsRead allowed

Request

Which five billing countries generated the most revenue?

SQL passed to the shared tool

                  SELECT BillingCountry, ROUND(SUM(Total),2) AS revenue FROM Invoice GROUP BY BillingCountry ORDER BY revenue DESC, BillingCountry ASC LIMIT 5
                
Inspect the raw tool result
                    {
  "status": "ok",
  "columns": [
    "BillingCountry",
    "revenue"
  ],
  "rows": [
    [
      "USA",
      523.06
    ],
    [
      "Canada",
      303.96
    ],
    [
      "France",
      195.1
    ],
    [
      "Brazil",
      190.1
    ],
    [
      "Germany",
      156.48
    ]
  ]
}
                  

Reference rows · invoice totals

Top five billing countries by summed invoice totals in the pinned Chinook sample
CountryTotal
USA523.06
Canada303.96
France195.10
Brazil190.10
Germany156.48

A fluent answer still needs to match these rows. The oracle runs without a model.

02Try a refused writeWrite refused

Request

Delete every invoice.

SQL passed to the shared tool

                  DELETE FROM Invoice
                
Inspect the raw tool result
                    {
  "status": "error",
  "reason": "not authorized",
  "next_step": "Use the schema to correct the query. Writes, attachments and PRAGMA are refused."
}
                  

The database refuses the write

status: error

not authorized

No invoice was deleted. A returned result is not proof that the requested task succeeded.

Reproduce these outputs with the independent oracle and the shared query library. The refusal check uses the same pinned public database.

Key takeaways (3)
  • Keep database access in shared tools, so changing the agent runtime does not change what the agent can do.
  • Move the task instructions into a Mantle skill, then test the tools through the actual runtime.
  • A working database chat proves one task can move. It does not prove that every Deep Agents feature has moved.

A refused SQL write still has is_error=false in this example’s offline control. A checker looking only at that flag accepts the refusal:

refused write | application=error | is_error=False | flag-only=True | both=False
valid read | application=ok | is_error=False | flag-only=True | both=True

The actual native query wrapper and deterministic SDK event recorder produced those lines with a controlled tracker. No model conversation was run. The SQL library reports a refusal through status and reason; the pinned runtime looks for error or errors. Checking both contracts rejects the refused write and still accepts the valid read. This does not demonstrate a failure of the full chat orchestration path.

When moving an agent, carry its definition of a successful result along with its tools. A tool that returned data is not necessarily a tool that completed the task. Keep that distinction through the binding, packaging and answer checks below.

This cookbook ports the official Deep Agents text-to-SQL example to Rasa Mantle. Both implementations are built to answer questions about Chinook, a fictional music store. They call the same Python functions over the same public sample database. The tool tests include attempts to delete invoices and attach another database. Those attempts fail before either model can change a record.

Start with the runnable companion. It contains both chat implementations, separate dependency locks and the shared query library. You need Python 3.12, an OpenAI key and, for Mantle, a Rasa Pro licence. These are text agents. No microphone or speech vendor is involved. The pinned Mantle development build labels itself experimental in the recorded training output; use the public sample data for this exercise.

Choose the task you are moving

The upstream example uses a model, SQL tools, task instructions and skills. The model explores the schema, writes a query, runs it and explains the result. That is the task we move. We do not port the whole Deep Agents harness.

Here is the migration map. Each row says what you must build or keep.

PartDeep Agents buildMantle buildWhat carries over
Records and query executionShared read-only Python libraryThe same libraryDatabase checks and returned rows
Tool bindingsLangChain tool decoratorsAsync Mantle tool decoratorsNames, arguments and descriptions
Task instructionsSystem prompt and a skill fileAgent persona and a skill fileThe task rules, with framework-specific references
Model selectionChat model configurationModel group in integrationsThe model and reasoning setting
Tool loopDeep Agents graphMantle runtimeThe task outcome, not the loop implementation
Planning notesQuery plan in the skill textQuery plan in the skill textPlan the query in the skill instructions
FilesDedicated Deep Agents workspaceNo agent filesystem in this exampleNothing; the answer is returned in chat
DelegationExplicitly disabled for this taskNo delegated agent in this taskNo claim about delegated workloads

Context management and persistent backends are outside this port. Moving this chat does not establish equivalent behaviour for those requirements. If your application needs background jobs, shell execution or delegated work, qualify them separately before deciding to migrate.

One detail is easy to miss. In the inspected upstream revision, subagents=[] means no custom subagents. The harness can still add its default general-purpose agent. Our build disables that agent through a harness profile. It also excludes the shell tool. The profile is part of the chosen scope, not evidence that Deep Agents lacks those features.

Both configurations request the same pinned OpenAI model and low reasoning. Deep Agents selects the Responses API explicitly. Check that your provider account supports the pinned model before starting billed chat sessions.

Invoice records BillingCountry · Total Group by BillingCountry SUM(Total) revenue requested metric COUNT(*) invoice count different question Sort by revenue keep five countries
FigureThe revenue question needs invoice totals

A valid query can calculate the wrong thing. The oracle sums invoice totals; counting invoices answers a different question. The visual explorer uses the oracle’s rows, so you can inspect the reference before running either agent.

Keep the database boundary outside the prompt

The shared library opens SQLite in read-only mode. It then enables query-only access and installs an authorizer. The authorizer permits reads, functions and query planning, while refusing writes, attachments and configuration changes.

connection = sqlite3.connect(
    database_path().as_uri() + "?mode=ro", uri=True
)
connection.enable_load_extension(False)
connection.execute("PRAGMA query_only=ON")
connection.execute("PRAGMA trusted_schema=OFF")
connection.set_authorizer(
    lambda action, *args:
        sqlite3.SQLITE_OK if action in ALLOWED else sqlite3.SQLITE_DENY
)

Both locks install the shared library as a local package. That keeps it available when Mantle loads a model archive in a fresh process. The model archive does not bundle that external package. Keep the package and its locked source with the worker; sending only the model archive is not a complete installation. The recorded archive listing shows exactly what was packaged. The same code runs beneath both agents. Removing “never update” from the prompt would not make an update possible through these tools. That is the property you want to preserve during a migration.

The library also bounds query work and refuses results above its row limit. It returns an error rather than silently cutting a result short. The work check counts SQLite VM progress; it is not a memory or wall-clock budget for every SQL function. A valid SQL statement can still be the wrong query: syntax checking does not establish whether the agent chose the right metric or joined the right records.

Run the tool checks without installing either framework:

git clone https://github.com/RasaHQ/rasa-community-resources.git
cd rasa-community-resources
git fetch origin 797cd552a94401f9bfcf1d374273e1e7a53d5e98
git checkout --detach 797cd552a94401f9bfcf1d374273e1e7a53d5e98
cd tutorials/deepagents-to-mantle
make test
make data

The data command downloads a pinned Chinook file and checks its hash before writing it locally. An existing file with the wrong hash is left untouched. The tests use a separate temporary database, so they cannot change the downloaded sample. In the database test file, test_write_attachment_and_sensitive_operations_refused checks the denied operations. test_schema_does_not_accept_injected_identifier checks schema names, test_row_bound_refuses_instead_of_truncating checks oversized results, and test_work_bound_interrupts_recursive_query checks recursive work. The separate data-preservation test checks the existing file without network access.

These are tool tests. They do not measure model accuracy, tokens or response speed. A passing direct function test also does not prove the runtime can call that function. Check that boundary next.

Before moving on, remove the “never update” sentence in a test copy of the prompt. Could that copy update an invoice through these tools? No: SQLite refuses the write below the model. Could it disclose a customer email through a read? Read-only access alone would allow that. A real service needs an access policy in the tool, even if its prompt sounds restrictive.

Replace the bindings, keep the functions

The query function returns column names and rows, or a structured error. Compare the bindings side by side. Both call the same query function; the return type and runtime context change.

Change the binding, keep the query

Deep Agents

@tool
async def sql_db_query(query: str) -> dict:
    """Run one read-only SELECT and return at most 100 rows, or an explicit error."""
    return database.query(query)

The decorator exposes the function as a LangChain tool. Its return value is the shared dictionary.

Rasa Mantle

@tool(description="Run one read-only SELECT and return at most 100 rows, or an explicit error.")
async def sql_db_query(
    query: str, context: ToolContext = None
) -> ToolResult:
    return ToolResult(llm_response=database.query(query))

Keep the async context argument. Wrap the shared dictionary in the native tool result.

Keep the binding async, and accept the runtime context even when this particular tool does not need it. The context is supplied by Mantle’s invoker. Calling the function directly without that argument can conceal a broken binding.

Run the native checks with the locked Mantle interpreter:

cd mantle
uv run --locked python -B -m unittest discover -s ../tests -v

The recorded offline log ends with this unedited excerpt. Its qualification record binds the installed SDK version, source revision and log hashes:

Ran 15 tests in 8.356s

OK

The suite includes nine database checks, one data-preservation check and five native checks. These are the three boundaries it tests:

BoundaryFaulty conditionRestored or complete check
Application resultA flag-only checker accepts a refused SQL writeChecking application status too rejects the write and accepts a valid read
Native bindingShared query checks pass while a wrapper without context failsThe actual invoker with the accepted context returns [[1]]
Dependency loadingA cached import hides the unavailable dependencyA cold process refuses the fault; the installed package restores normal loading

Inspect the native tests. test_actual_invoker_preserves_write_refusal generates the opening comparison. It invokes the real tool, then calls the deterministic SDK event recorder with a controlled tracker. Its flag-only check accepts the refused write; requiring application status=ok rejects it. The valid read is the positive control. That acceptance check does not establish the correct business query; the row oracle still matters.

test_actual_invoker_runs_the_query_binding checks the returned rows. test_removing_context_breaks_the_actual_invocation checks a separate binding mutation. Remove context and the actual invoker reports this unedited error:

NativeBindings.test_removing_context_breaks_the_actual_invocation.<locals>.broken_query() got an unexpected keyword argument 'context'

test_archive_loads_without_a_warm_shared_import removes inherited keys and Python paths before starting its child process. The invoker receives a controlled tracker, not a simulated invoker. No model is asked to answer a question in these tests, so their pass is integration evidence, not measured chat accuracy.

The dependency counter-test does not uninstall a package. It uses a Python import finder to make new imports of shared fail. One child imports that package first; the other starts cold. The recorded result is:

warm: cached dependency hides the fault
cold: missing dependency refused

That is why a warm process is a poor packaging check. It can keep working with an import that a newly started worker cannot make. The ordinary cold-load test checks the restored condition, with the declared package available. This is controlled fault injection, not a recorded production incident.

The source references distinguish the controlled recorder from the full LLM path.

The pinned recorder is rasa/mantle/orchestration/tool_execution/constraints.py, lines 919–954. The LLM path has its own hooks and event emission in rasa/mantle/orchestration/orchestrator.py, lines 1912–1948. This control does not execute that full path.

From the tutorial root, resume the remaining steps below. If you have just run the native checks in mantle, return there with cd .. first.

The other three tools follow the same pattern: list tables, read schema and check query syntax. No business query is rewritten in the Mantle wrapper. A failed query returns a reason that the agent can use to correct its SQL. The tool never returns a successful receipt for a refused write.

There is one deliberate adaptation from upstream. The official example configures its SQL toolkit with the same model as the agent. In the inspected checker implementation, calling the checker invokes that model’s chain. Our shared checker compiles SQL locally. Both builds use this checker. This changes the application, so it cannot establish a saving caused by switching runtimes.

StageUpstream exampleBoth tutorial builds
Check proposed SQLModel-based toolkit checkerShared checker compiles SQL with SQLite
Execute checked SQLDatabase query toolShared read-only query tool

Local compilation can reject invalid SQL, but cannot establish that it answers the user’s question. The upstream checker path is a code-inspection finding; we have not measured its requests here. Count the complete native request journal when measuring your own workload.

Put the procedure in a Mantle skill

The Mantle project has three parts: the agent’s task rules, integrations for model and chat, and a skill that exposes the query tools.

mantle/
├── agent.yml
├── integrations.yml
├── pyproject.toml
├── uv.lock
└── skills/query_store/
    ├── skill.md
    └── tools.py

The skill names the task and describes when to use it. Its procedure references the tools directly:

List tables with @tool.sql_db_list_tables. Inspect relevant schemas with
@tool.sql_db_schema. Check the SQL with @tool.sql_db_query_checker, then run
it with @tool.sql_db_query. For complex joins, decide which tables and keys
you need before executing. State the metric and time filter in the answer.

This preserves the SQL workflow. Neither selected build exposes a plan-management tool. If a visible plan is part of your user experience, define its format and persistence as a separate application requirement.

Both projects use the same base task instructions. That does not establish identical model requests. Each runtime assembles prompts and tool schemas; inspect those requests before attributing token differences to the instructions. In the pinned wheels, Mantle’s assembly is in rasa/mantle/orchestration/orchestrator.py, lines 710–786. LangChain’s langchain/agents/factory.py, lines 1050–1056 and 1429–1488, builds the system message and binds tools. The lock pins LangChain 1.4.4. Sharing a text file is not enough to claim identical input token counts.

For example, “most revenue” means summing invoice totals here, rather than counting tracks. The shared task instructions state that metric. The returned query and rows must prove the agent used it. Do not accept a fluent answer as evidence of the right calculation.

Run both versions against the same question

Install only the environment you plan to use. Keep the lock with its project; do not copy an environment into another checkout.

For Deep Agents:

cd deepagents
uv sync --locked --python 3.12
uv run --locked python inspect_chat.py \
  "Which five billing countries generated the most revenue?" \
  --out .local/billing-revenue.json

The file retains the native messages. Match a sql_db_query tool message to its assistant call by tool_call_id. The call’s arguments give the proposed query; the tool’s content gives the returned result. Parse that content and require application status=ok before accepting its rows. Inspect the final assistant message for the answer. LangChain Core 1.6.9 defines these message fields in langchain_core/messages/tool.py, lines 26–87 and 206–246. Its langchain_core/tools/base.py, lines 1390–1427 and 1471–1484, serialises returned dictionaries into tool messages. The companion saves their native JSON representation, rather than a reconstructed chat transcript. The command refuses an existing output file before calling the model, so a previous conversation is not silently replaced.

Set the model key explicitly in that terminal before running the command. The example does not search neighbouring repositories for credentials. Each invocation is a fresh session. Live calls use OpenAI API billing; offline checks do not exercise that provider.

For Mantle, set the model key and your Rasa Pro licence in the terminal. The pinned licence lookup is in rasa/utils/licensing.py, lines 276–296; it accepts RASA_LICENSE and the legacy RASA_PRO_LICENSE. Then:

cd ../mantle
make install
make validate
make test
make train
make run

Send a question from a second terminal:

mkdir -p .local
curl --fail-with-body http://localhost:5005/webhooks/rest/webhook \
  -H 'Content-Type: application/json' \
  -d '{"sender":"chinook-demo-1","message":"Which five billing countries generated the most revenue?"}' \
  -o .local/billing-answer.json
curl --fail-with-body \
  'http://localhost:5005/conversations/chinook-demo-1/tracker?include_events=ALL&start_session=false' \
  -o .local/billing-tracker.json

The response file holds the final chat answer. In the tracker file, inspect tool_executed events: arguments contains the query and result holds the raw returned data; parse result as JSON when it is text. Check both the runtime is_error flag and the application status inside that result. This wrapper returns refused SQL as {"status":"error","reason":...}. The pinned runtime parser looks for error or errors fields, so this application refusal can have is_error=false. A false runtime flag alone does not establish a successful query. The tracker request reads the existing session without starting another one. This follows the pinned SDK tracker route in rasa/server.py, lines 997–1025, and the tool event fields in rasa/shared/core/events.py, lines 4033–4118. These references are for Rasa Pro 3.21.0.dev5; its payload parser is in rasa/mantle/orchestration/tool_execution/payload.py, lines 77–98. Inspect the installed wheel when changing versions. Keep these files local; they are generated conversation records, not source to commit.

Use a new sender identifier to avoid existing conversation state. For a follow-up test, send the next question with the same identifier and inspect the tracker. Do not assume history survives a server restart. The REST channel passes sender into the incoming message in rasa/core/channels/rest.py, lines 150–183. rasa/core/processor.py, lines 1413–1444, retrieves or creates the tracker by conversation identifier. Do not compare a fresh Deep Agents invocation with a Mantle session that already holds earlier answers. The conversation history changes the task and its cost.

If validation passes but a tool returns a generic error, inspect the native invocation. Check the async signature and context argument first. Keep the original error and query when debugging; do not turn a failed call into an empty successful result.

Check correctness before comparing cost

Agree what an answer must contain before running either agent. For this store, you can compute the expected result independently in SQLite.

For billing-country revenue:

SELECT BillingCountry, ROUND(SUM(Total), 2) AS revenue
FROM Invoice
GROUP BY BillingCountry
ORDER BY revenue DESC, BillingCountry ASC
LIMIT 5;

From the tutorial root, run the offline oracle against the pinned database:

python3 oracle.py

Its expected rows are:

Billing countryInvoice revenue
USA523.06
Canada303.96
France195.10
Brazil190.10
Germany156.48

This is a SQLite result, not a model transcript. Use the captured native tool messages or tracker events to find sql_db_query, then inspect its SQL and rows.

Compare the agent’s returned rows with these rows. Then check its final answer: it must name the same countries and totals. Correct rows with an invented total in the answer are still a failed answer. Use a tie-breaker so equivalent queries do not appear different because SQLite returned tied rows in another order.

Add cases that change the decision. Ask for employee revenue through a join. Ask an ambiguous question and require clarification. Ask for an update and verify that the database bytes stay unchanged. Test a follow-up separately; it exercises retained conversation state rather than a fresh query.

Use the following measurements when deciding whether to move your own task. They need evidence from both native runtimes, not estimates from the visible application prompt.

MeasurementWhat to retainWhat it does not prove
Correct answersExpected rows, actual rows and final answer for each casePerformance on other datasets or tasks
Model requestsEvery provider dispatch, including helpers and retriesThe number of user turns
TokensRaw input, cached-input and output usage from each responseTokens inferred from a local text counter
Model costUsage priced against a dated tariff; missing usage remains unknownA reconciled invoice or total operating cost
Response timeStart and finish for each session, under the same measurement pathA general speed ranking from one run
ImplementationCounted application code and configuration, with shared code listed separatelyEngineering hours or the size of the SDK

Run repeated, interleaved sessions before claiming one version is faster or more reliable. Record the model, framework versions, data hash and date. A small tutorial run can show a migration works. It cannot establish a winner for every database chat.

Decide what to move next

Mantle is a useful fit for this task when you want a chat runtime with skills, model configuration and tool bindings around an existing business library. The database boundary remains yours to implement and test.

Keep the current harness if your workload depends on capabilities this example does not port. Files, long-running jobs and delegated context are concrete requirements, not details to drop because the first query worked. An ordinary Python function can perform several steps inside a tool, but that does not make it an isolated subprocess or a native delegated agent.

For your first migration, choose one read-only workflow, preserve its tool contract and run the same acceptance cases. Keep the existing endpoint available until the replacement passes those cases. Route a separate test client to the new endpoint before changing real traffic.

The next step is the Mantle quickstart for a new agent, or the existing migration guide for a tool-based LangGraph application. For the full Deep Agents feature set, consult its official documentation.

Start with the quickstart if you need the basic project setup before trying the Mantle port.