Guide · Platform / operations engineer
endpoints.yml said Ollama. Mantle still picked OpenAI
On Rasa Pro 3.20.0, Mantle takes the references embedder from agent.yml references.embeddings. Left empty, rasa train uses OpenAI and logs no provider.
3 chunks
of reference text sent to text-embedding-3-large after endpoints.yml was set to Ollama
0
lines in the train output naming the embedding provider or model
Key takeaways (5)
- On Rasa Pro 3.20.0 we pointed the companion voice agent's endpoints.yml embeddings group at ollama/all-minilm. rasa train still sent all three reference chunks to text-embedding-3-large, exited 0 and printed no line naming an embedding provider.
- That group ships in the committed project and no setting refers to it. It looked in charge because it lists text-embedding-3-large, the same model as the OpenAI default that an empty references.embeddings falls back to.
- build_references_embedder takes a model group id from agent.yml references.embeddings and looks it up only in integrations.yml. An empty id hands embedder_factory None and the OpenAI default.
- Every check on the key runs only when it is set: project validation, the raise in build_references_embedder and the model group resolver. An unknown group stops training; a misspelt embedding: key is dropped and the default is used without a word.
- Setting embeddings: local-embeddings, with that group declared in integrations.yml, trained with no request to the OpenAI stand-in. In any stack, change the setting to a value you can tell apart from the default and record what the provider receives.
We pointed the embeddings model group in the companion voice agent’s
endpoints.yml at a local model:
- id: openai-embeddings
models:
- - provider: openai
- model: text-embedding-3-large
- api_key_env: OPENAI_API_KEY
+ - provider: ollama
+ model: all-minilm
+ api_base: http://localhost:11434
Then we ran rasa train with OpenAI’s API replaced by a local stand-in that
records each request. This is everything the stand-in received:
exit code: 0
requests the OpenAI stand-in received: 1
{"method": "POST", "path": "/v1/embeddings", "model": "text-embedding-3-large", "inputs": 3, "authorization_header": "present"}
input head: "# Fixture data \u2014 entirely fictional\n\nEvery person, company, account, and event i"
input head: "# Horizon Travel FAQ\n\n## Baggage\n- Economy includes one cabin bag up to 7 kg and"
input head: "## Contact\n- Horizon Travel phone support is available daily from 7am to 10pm UK"
Training passed. One request went to the OpenAI embeddings endpoint for
text-embedding-3-large, carrying three chunks: two from the project’s travel
FAQ and one from the README.md that sits beside it in the references folder.
Horizon Travel and its FAQ are the companion project’s fictional fixture data,
as that README says. Nothing in the train output named an embedding provider or
model. What decided the model was a key the project never sets, in a different
file.
That output is unedited, from a real rasa train rather than an offline test.
The project is examples/mantle-voice-agent in
RasaHQ/rasa-community-resources at commit 01d6a6e, on Rasa Pro 3.20.0. We
recorded the runs on 29 September 2026 against that release. OPENAI_API_KEY
held a placeholder, and OPENAI_API_BASE pointed at the stand-in, which
answers with fake vectors and forwards nothing. The request is the one Rasa’s
OpenAI client sends. Re-running it needs a Rasa Pro licence in RASA_LICENSE; the
companion’s .env.example asks for the free Developer Edition licence.
We did not add the group we edited. This is how the committed endpoints.yml
declares it (lines 21 to 36, with the openai-llm group cut):
# ---------------------------------------------------------------------------
# Model groups (LLMs / embeddings referenced by endpoints features)
# ---------------------------------------------------------------------------
# https://rasa.com/docs/reference/config/components/llm-configuration/
model_groups:
...
- id: openai-embeddings
models:
- provider: openai
model: text-embedding-3-large
api_key_env: OPENAI_API_KEY
No setting anywhere in the project refers to openai-embeddings; the only
other match for the name is a copy of this file under tutorial/snippets. The
group looked in charge because it lists text-embedding-3-large, which is
also the model Rasa falls back to when agent.yml names none. Train the
project as committed and the request is identical, so the group appears to
work. Only changing it shows that it does nothing.
The model that embeds a Mantle project’s reference files is chosen by one key,
references.embeddings in agent.yml, and resolved only against
integrations.yml. Left empty, it is OpenAI: if your reference files are
private, every chunk of them goes to OpenAI on every training run, whatever
endpoints.yml says. Our position: a Mantle
project that ships reference files should set that key, even when the answer
is an OpenAI group, and prove the choice with one training run against a
recorder. The cost is a model group in integrations.yml to maintain, a
retrain every time the embedder changes, and a stand-in wherever the proof
runs.
Five trainings of one project, and which edits moved the request
We trained five separate copies of the committed project, each with the same
stand-in and placeholder key. Run a changes nothing; runs b and e edit one
file; runs c and d edit two, agent.yml and integrations.yml.
| Run | Edit to the committed project | Requests to the stand-in | Model requested | rasa train |
|---|---|---|---|---|
| a | none | 1, with 3 inputs | text-embedding-3-large | exit 0 |
| b | endpoints.yml group openai-embeddings set to ollama / all-minilm | 1, with 3 inputs | text-embedding-3-large | exit 0 |
| c | embedding: local-embeddings (no s) under references:; group added to integrations.yml | 1, with 3 inputs | text-embedding-3-large | exit 0 |
| d | embeddings: local-embeddings under references:; the same group added to integrations.yml | 0 | none | exit 0 |
| e | embeddings: no-such-group under references:; integrations.yml unchanged | 0 | none | exit 1, at project validation |
Runs a, b and c sent the same request, down to the first 80 characters of each input. Run b is the one the title describes.
Run c is a typo that nothing reports. ReferencesConfig is declared with
extra="ignore" (rasa/mantle/config/agent_spec.py, line 153), so an unknown
key under references: is dropped and embeddings keeps its empty default.
The companion’s own agent.yml warns at lines 17 to 19 that AgentSpec
ignores unknown keys; the class one level down is declared the same way.
Run e is the only loud one. Training stopped at project validation with one
finding, mantle.validation.config.unknown_model_group, whose message reads:
“‘references.embeddings’ in ‘agent.yml’ references model group
‘no-such-group’, which is not declared in ‘integrations.yml’.”
Where build_references_embedder gets the embedding model
The index is built when rasa train packages the model archive.
build_references_index loads every *.md file under a references/ folder
(line 99 globs **/*.md, which is how the README was embedded), splits them
into chunks, and then chooses the embedder on lines 201 and 202:
model_groups = resolve_integration_model_groups(project_root)
embedder = build_references_embedder(agent, model_groups)
resolve_integration_model_groups, in rasa/mantle/model_archive/bootstrap.py,
returns the model_groups list from the project’s integrations.yml, or an
empty list. endpoints.yml is not an argument. This
is the function those groups feed, whole, from the 3.20.0 wheel (sha256
a68fa7efff47dd035e3ff2b7b1cfdc467c394ff75981426a4377d4008207e9b5):
def build_references_embedder(
agent: AgentSpec, model_groups: list[dict[str, Any]]
) -> "EmbeddingClient":
"""Resolve the references embeddings model group into an embedding client.
``agent.references.embeddings`` names a ``model_groups`` id; empty → the
OpenAI default (mirrors ``EnterpriseSearchPolicy``). The resolved config is
handed to ``embedder_factory``, whose client ``FAISS_Store`` consumes
directly.
"""
group_id = agent.references.embeddings
embeddings_config: Optional[dict[str, Any]] = None
if group_id:
# mantle resolves model groups only from the packaged integrations.yml.
# resolve_model_client_config treats an empty model_groups list as falsy
# and silently falls back to endpoints.yml, which would resolve against
# the wrong (or missing) config. Fail closed with a clear message instead.
if not model_groups:
raise InvalidConfigException(
f"references.embeddings names model group '{group_id}', but no "
f"model_groups are defined in {V2_INTEGRATIONS_FILE}. Define the "
f"model group there, or clear references.embeddings to use the "
f"default embedder."
)
embeddings_config = resolve_model_client_config(
{MODEL_GROUP_CONFIG_KEY: group_id},
"references",
model_groups=model_groups,
)
return embedder_factory(embeddings_config, DEFAULT_EMBEDDINGS_CONFIG)- Line 62: the only input that names a model is
references.embeddingsfromagent.yml. Its default is the empty string. - Line 64: an empty string is falsy, so lines 65 to 80 never run, and nothing is looked up, checked or logged.
- Lines 65 to 68: the authors’ comment says model groups come only from
integrations.yml, and that an empty list would let resolution drift toendpoints.yml. The next line exists to stop that drift. - Lines 69 and 70: the function’s own
raise, for a set id whenintegrations.ymldefines no groups at all. - Line 81: for an empty id this is
embedder_factory(None, DEFAULT_EMBEDDINGS_CONFIG).
The default is two keys. DEFAULT_EMBEDDINGS_CONFIG in
rasa/core/policies/enterprise_search_policy_config.py names the provider and
the model, and embedder_factory in rasa/shared/utils/llm.py passes a None
config straight through to the client factory with it:
DEFAULT_EMBEDDINGS_CONFIG = {
PROVIDER_CONFIG_KEY: OPENAI_PROVIDER,
MODEL_CONFIG_KEY: DEFAULT_OPENAI_EMBEDDING_MODEL_NAME,
}
llm.py sets DEFAULT_OPENAI_EMBEDDING_MODEL_NAME = "text-embedding-3-large" on line 135. That is the model in every request runs a to c
recorded, and the model the committed endpoints.yml group lists. The
coincidence is what hid the dead group.
The fallback is documented. The field’s docstring in agent_spec.py says
“Empty → the built-in default embeddings (OpenAI), like
EnterpriseSearchPolicy”, and the docstring at the top of
references_index.py (lines 7 to 9) says it again. What is missing is any
sign of it at run time. The build path’s only info-level log call,
mantle.packaging.references_index.wrote, records the index path, document and
chunk counts and chunking settings, and no provider. In the full, unfiltered
log of run a it is the only line about the index, and no line names an
embedding provider or model. Rasa’s training telemetry code has the same blind
spot: _embeddings_attributes in training_telemetry.py reports
embeddings_provider as None whenever the key is empty.
Every check on references.embeddings waits for the key to be set
A bad references.embeddings can stop training in three places, and all
three start from the same test. Project validation runs first:
validate_agent in rasa/mantle/config/validation.py stops checking the key
when it is empty (lines 93 and 94) and otherwise looks the id up in
integrations.yml alone. That lookup produced run e’s finding. The second is
the raise on line 70 of build_references_embedder, inside if group_id:.
The third is resolve_model_client_config in llm.py, called from the same
branch. It raises its own InvalidConfigException when the id is missing from
a non-empty group list, and a default rasa train only reaches it with
validation off, because rasa/cli/train.py passes validate=not args.skip_validation. Its message says the group was not found “in
endpoints.yml”, although Mantle handed it the groups from integrations.yml.
No check of any kind covers runs a, b and c.
- Validation checks the key only when it is set, and only against
integrations.yml. - The build step’s
raiseand the resolver’s sit behind the sameif group_id:; with validation on, the resolver’s never fires for a missing id. - An empty or misspelt key reaches
embedder_factorywith only the OpenAI default, and no log line names the provider. rasa traindoes readendpoints.yml: run b’s output includes a tracing notice naming that file. None of it reaches the references path.
The code fails closed on a group it cannot find and falls open, to a hosted
provider, on a key it cannot see. The comment on lines 65 to 68 refuses to fall
back silently to the wrong configuration. An empty key falls back, documented
but unannounced, to a provider the project may never have chosen. The
docstring’s comparison with EnterpriseSearchPolicy says this default is
deliberate, and a maintainer could reasonably keep it. Two small changes would
have made run b visible: one info line at line 81 naming the provider and model
it chose, or a validation finding when a project ships reference files with no
embeddings key. Both are our suggestions, not behaviour we observed.
Name the embedder in agent.yml and declare it in integrations.yml
Run d is the change, and it needs two files: line 62 takes the id from
agent.yml, and line 201 decides that the id is looked up in
integrations.yml. The diff against the committed project, agent.yml
first, then integrations.yml:
references:
+ embeddings: local-embeddings
instructions: |
Answer from retrieved FAQ snippets only. Keep answers short for voice.
temperature: 0.0
+ - id: local-embeddings
+ models:
+ - provider: ollama
+ model: all-minilm
+ api_base: http://localhost:11434
+
channels:
Training passed and the stand-in received no request, with endpoints.yml
untouched. If you want OpenAI, name an OpenAI group the same way: the key then
records the choice, and validation checks it.
To audit a project before training it, read the key and look for anything
else under references:. This is the command we ran in the committed project
and in the copies for runs c and d. It is wrapped here with backslash
continuations, which the shell removes, so it is the same one-line program:
uv run --frozen python -c "import yaml; \
r = yaml.safe_load(open('agent.yml')).get('references') or {}; \
print('references.embeddings:', r.get('embeddings') or '(empty: OpenAI default)'); \
print('other keys under references:', \
sorted(set(r) - {'instructions', 'embeddings', 'chunking'}) or 'none')"
Its output, with uv’s install notices (a hardlink warning and the
Installed/Built lines) removed and nothing else changed:
$ cd examples/mantle-voice-agent
references.embeddings: (empty: OpenAI default)
other keys under references: none
$ cd runs/c-misspelt-key
references.embeddings: (empty: OpenAI default)
other keys under references: ['embedding']
$ cd runs/d-named-local
references.embeddings: local-embeddings
other keys under references: none
The label “(empty: OpenAI default)” is ours, taken from the source above. The command reads YAML; only a training run against a recorder shows what is sent.
The same fallback in stacks that are not Rasa
None of this needs Rasa to happen again. Four things lined up:
- A factory takes an optional user config and a hosted default, and
Nonemeans the default:embedder_factory(None, DEFAULT_EMBEDDINGS_CONFIG). - Validation runs only on values that are present, so absence is never an error.
- The schema ignores unknown keys, so a typo becomes absence.
- A config block elsewhere lists the same model as the default, so the fallback and the intended setting send the same request, and a green build proves nothing.
In code you do not own, search for calls where the config argument can be
None next to a module-level constant naming a hosted model, and for
extra="ignore" or its equivalent on the schema that holds the setting. The
log line that should exist is one at build time naming the provider and model
actually chosen. When it is missing, run the test we ran: point the
provider’s base URL at a recorder, change the setting to a value you can tell
apart from the default, and train. If the recorder still hears the default
model, the setting is dead.
The recorder is a plain HTTP server. It logs the path, the model, the input
count, the first 80 characters of each input and whether an Authorization
header was present (never its value), and it answers embeddings requests with
fake vectors.
Show the stand-in we trained against (29 lines of Python)
Start it with python openai_standin.py <port> requests.jsonl, then train
with
OPENAI_API_KEY=sk-standin-placeholder OPENAI_API_BASE=http://127.0.0.1:<port>/v1 uv run --frozen rasa train.
"""Local stand-in for api.openai.com. Records each request (path, model, input count,
first 80 characters of each input, auth header present or not; never its value) and
returns deterministic fake embeddings. Nothing is forwarded anywhere."""
import http.server, json, sys, datetime, hashlib
PORT, LOG = int(sys.argv[1]), sys.argv[2]
class H(http.server.BaseHTTPRequestHandler):
def do_POST(self):
body = json.loads(self.rfile.read(int(self.headers.get('content-length', 0))) or b'{}')
inputs = body.get('input', [])
inputs = [inputs] if isinstance(inputs, str) else inputs
rec = {'at': datetime.datetime.now().isoformat(timespec='seconds'), 'method': 'POST', 'path': self.path,
'authorization_header': 'present' if self.headers.get('authorization') else 'absent',
'model': body.get('model'), 'inputs': len(inputs),
'input_heads': [str(i)[:80] for i in inputs]}
with open(LOG, 'a') as f: f.write(json.dumps(rec) + '\n')
if self.path.rstrip('/').endswith('/embeddings'):
def vec(t):
h = hashlib.sha256(str(t).encode()).digest()
return [((h[i % 32] / 255.0) - 0.5) for i in range(1536)]
out = {'object': 'list', 'model': body.get('model'), 'data': [{'object': 'embedding', 'index': i, 'embedding': vec(t)} for i, t in enumerate(inputs)],
'usage': {'prompt_tokens': 0, 'total_tokens': 0}}
data, status = json.dumps(out).encode(), 200
else:
data, status = json.dumps({'error': {'message': 'stand-in serves /embeddings only'}}).encode(), 404
self.send_response(status); self.send_header('content-type', 'application/json')
self.send_header('content-length', str(len(data))); self.end_headers(); self.wfile.write(data)
def do_GET(self): self.do_POST()
def log_message(self, *a): pass
http.server.ThreadingHTTPServer(('127.0.0.1', PORT), H).serve_forever()We chose a recorder over an invalid key for this test. A call that fails tells you something was attempted; a recorder lets training finish, as it would in production, and keeps the request itself: the model, the input count and the start of every chunk.