Skip to content
RasaGet a free licence
Guides for AI teams

Guide · Platform / operations engineer

endpoints.yml said Ollama. Mantle still picked OpenAI

On Rasa Pro 3.20.0, Mantle takes the references embedder from agent.yml references.embeddings. Left empty, rasa train uses OpenAI and logs no provider.

by Rod Rivera

About 4 minutes

  • 3 chunks

    of reference text sent to text-embedding-3-large after endpoints.yml was set to Ollama

  • 0

    lines in the train output naming the embedding provider or model

Source: rasa train on Rasa Pro 3.20.0, companion commit 01d6a6e, OpenAI traffic sent to a local stand-in
Key takeaways (5)
  • On Rasa Pro 3.20.0 we pointed the companion voice agent's endpoints.yml embeddings group at ollama/all-minilm. rasa train still sent all three reference chunks to text-embedding-3-large, exited 0 and printed no line naming an embedding provider.
  • That group ships in the committed project and no setting refers to it. It looked in charge because it lists text-embedding-3-large, the same model as the OpenAI default that an empty references.embeddings falls back to.
  • build_references_embedder takes a model group id from agent.yml references.embeddings and looks it up only in integrations.yml. An empty id hands embedder_factory None and the OpenAI default.
  • Every check on the key runs only when it is set: project validation, the raise in build_references_embedder and the model group resolver. An unknown group stops training; a misspelt embedding: key is dropped and the default is used without a word.
  • Setting embeddings: local-embeddings, with that group declared in integrations.yml, trained with no request to the OpenAI stand-in. In any stack, change the setting to a value you can tell apart from the default and record what the provider receives.

We pointed the embeddings model group in the companion voice agent’s endpoints.yml at a local model:

   - id: openai-embeddings
     models:
-      - provider: openai
-        model: text-embedding-3-large
-        api_key_env: OPENAI_API_KEY
+      - provider: ollama
+        model: all-minilm
+        api_base: http://localhost:11434

Then we ran rasa train with OpenAI’s API replaced by a local stand-in that records each request. This is everything the stand-in received:

exit code: 0
requests the OpenAI stand-in received: 1
{"method": "POST", "path": "/v1/embeddings", "model": "text-embedding-3-large", "inputs": 3, "authorization_header": "present"}
    input head: "# Fixture data \u2014 entirely fictional\n\nEvery person, company, account, and event i"
    input head: "# Horizon Travel FAQ\n\n## Baggage\n- Economy includes one cabin bag up to 7 kg and"
    input head: "## Contact\n- Horizon Travel phone support is available daily from 7am to 10pm UK"

Training passed. One request went to the OpenAI embeddings endpoint for text-embedding-3-large, carrying three chunks: two from the project’s travel FAQ and one from the README.md that sits beside it in the references folder. Horizon Travel and its FAQ are the companion project’s fictional fixture data, as that README says. Nothing in the train output named an embedding provider or model. What decided the model was a key the project never sets, in a different file.

That output is unedited, from a real rasa train rather than an offline test. The project is examples/mantle-voice-agent in RasaHQ/rasa-community-resources at commit 01d6a6e, on Rasa Pro 3.20.0. We recorded the runs on 29 September 2026 against that release. OPENAI_API_KEY held a placeholder, and OPENAI_API_BASE pointed at the stand-in, which answers with fake vectors and forwards nothing. The request is the one Rasa’s OpenAI client sends. Re-running it needs a Rasa Pro licence in RASA_LICENSE; the companion’s .env.example asks for the free Developer Edition licence.

We did not add the group we edited. This is how the committed endpoints.yml declares it (lines 21 to 36, with the openai-llm group cut):

# ---------------------------------------------------------------------------
# Model groups (LLMs / embeddings referenced by endpoints features)
# ---------------------------------------------------------------------------
# https://rasa.com/docs/reference/config/components/llm-configuration/
model_groups:
  ...
  - id: openai-embeddings
    models:
      - provider: openai
        model: text-embedding-3-large
        api_key_env: OPENAI_API_KEY

No setting anywhere in the project refers to openai-embeddings; the only other match for the name is a copy of this file under tutorial/snippets. The group looked in charge because it lists text-embedding-3-large, which is also the model Rasa falls back to when agent.yml names none. Train the project as committed and the request is identical, so the group appears to work. Only changing it shows that it does nothing.

The model that embeds a Mantle project’s reference files is chosen by one key, references.embeddings in agent.yml, and resolved only against integrations.yml. Left empty, it is OpenAI: if your reference files are private, every chunk of them goes to OpenAI on every training run, whatever endpoints.yml says. Our position: a Mantle project that ships reference files should set that key, even when the answer is an OpenAI group, and prove the choice with one training run against a recorder. The cost is a model group in integrations.yml to maintain, a retrain every time the embedder changes, and a stand-in wherever the proof runs.

Five trainings of one project, and which edits moved the request

We trained five separate copies of the committed project, each with the same stand-in and placeholder key. Run a changes nothing; runs b and e edit one file; runs c and d edit two, agent.yml and integrations.yml.

RunEdit to the committed projectRequests to the stand-inModel requestedrasa train
anone1, with 3 inputstext-embedding-3-largeexit 0
bendpoints.yml group openai-embeddings set to ollama / all-minilm1, with 3 inputstext-embedding-3-largeexit 0
cembedding: local-embeddings (no s) under references:; group added to integrations.yml1, with 3 inputstext-embedding-3-largeexit 0
dembeddings: local-embeddings under references:; the same group added to integrations.yml0noneexit 0
eembeddings: no-such-group under references:; integrations.yml unchanged0noneexit 1, at project validation

Runs a, b and c sent the same request, down to the first 80 characters of each input. Run b is the one the title describes.

Run c is a typo that nothing reports. ReferencesConfig is declared with extra="ignore" (rasa/mantle/config/agent_spec.py, line 153), so an unknown key under references: is dropped and embeddings keeps its empty default. The companion’s own agent.yml warns at lines 17 to 19 that AgentSpec ignores unknown keys; the class one level down is declared the same way.

Run e is the only loud one. Training stopped at project validation with one finding, mantle.validation.config.unknown_model_group, whose message reads: “‘references.embeddings’ in ‘agent.yml’ references model group ‘no-such-group’, which is not declared in ‘integrations.yml’.”

Where build_references_embedder gets the embedding model

The index is built when rasa train packages the model archive. build_references_index loads every *.md file under a references/ folder (line 99 globs **/*.md, which is how the README was embedded), splits them into chunks, and then chooses the embedder on lines 201 and 202:

    model_groups = resolve_integration_model_groups(project_root)
    embedder = build_references_embedder(agent, model_groups)

resolve_integration_model_groups, in rasa/mantle/model_archive/bootstrap.py, returns the model_groups list from the project’s integrations.yml, or an empty list. endpoints.yml is not an argument. This is the function those groups feed, whole, from the 3.20.0 wheel (sha256 a68fa7efff47dd035e3ff2b7b1cfdc467c394ff75981426a4377d4008207e9b5):

rasa/mantle/model_archive/references_index.py, lines 52-81 (Rasa Pro 3.20.0)
def build_references_embedder(
    agent: AgentSpec, model_groups: list[dict[str, Any]]
) -> "EmbeddingClient":
    """Resolve the references embeddings model group into an embedding client.

    ``agent.references.embeddings`` names a ``model_groups`` id; empty → the
    OpenAI default (mirrors ``EnterpriseSearchPolicy``). The resolved config is
    handed to ``embedder_factory``, whose client ``FAISS_Store`` consumes
    directly.
    """
    group_id = agent.references.embeddings
    embeddings_config: Optional[dict[str, Any]] = None
    if group_id:
        # mantle resolves model groups only from the packaged integrations.yml.
        # resolve_model_client_config treats an empty model_groups list as falsy
        # and silently falls back to endpoints.yml, which would resolve against
        # the wrong (or missing) config. Fail closed with a clear message instead.
        if not model_groups:
            raise InvalidConfigException(
                f"references.embeddings names model group '{group_id}', but no "
                f"model_groups are defined in {V2_INTEGRATIONS_FILE}. Define the "
                f"model group there, or clear references.embeddings to use the "
                f"default embedder."
            )
        embeddings_config = resolve_model_client_config(
            {MODEL_GROUP_CONFIG_KEY: group_id},
            "references",
            model_groups=model_groups,
        )
    return embedder_factory(embeddings_config, DEFAULT_EMBEDDINGS_CONFIG)
  1. Line 62: the only input that names a model is references.embeddings from agent.yml. Its default is the empty string.
  2. Line 64: an empty string is falsy, so lines 65 to 80 never run, and nothing is looked up, checked or logged.
  3. Lines 65 to 68: the authors’ comment says model groups come only from integrations.yml, and that an empty list would let resolution drift to endpoints.yml. The next line exists to stop that drift.
  4. Lines 69 and 70: the function’s own raise, for a set id when integrations.yml defines no groups at all.
  5. Line 81: for an empty id this is embedder_factory(None, DEFAULT_EMBEDDINGS_CONFIG).

The default is two keys. DEFAULT_EMBEDDINGS_CONFIG in rasa/core/policies/enterprise_search_policy_config.py names the provider and the model, and embedder_factory in rasa/shared/utils/llm.py passes a None config straight through to the client factory with it:

DEFAULT_EMBEDDINGS_CONFIG = {
    PROVIDER_CONFIG_KEY: OPENAI_PROVIDER,
    MODEL_CONFIG_KEY: DEFAULT_OPENAI_EMBEDDING_MODEL_NAME,
}

llm.py sets DEFAULT_OPENAI_EMBEDDING_MODEL_NAME = "text-embedding-3-large" on line 135. That is the model in every request runs a to c recorded, and the model the committed endpoints.yml group lists. The coincidence is what hid the dead group.

The fallback is documented. The field’s docstring in agent_spec.py says “Empty → the built-in default embeddings (OpenAI), like EnterpriseSearchPolicy”, and the docstring at the top of references_index.py (lines 7 to 9) says it again. What is missing is any sign of it at run time. The build path’s only info-level log call, mantle.packaging.references_index.wrote, records the index path, document and chunk counts and chunking settings, and no provider. In the full, unfiltered log of run a it is the only line about the index, and no line names an embedding provider or model. Rasa’s training telemetry code has the same blind spot: _embeddings_attributes in training_telemetry.py reports embeddings_provider as None whenever the key is empty.

Every check on references.embeddings waits for the key to be set

A bad references.embeddings can stop training in three places, and all three start from the same test. Project validation runs first: validate_agent in rasa/mantle/config/validation.py stops checking the key when it is empty (lines 93 and 94) and otherwise looks the id up in integrations.yml alone. That lookup produced run e’s finding. The second is the raise on line 70 of build_references_embedder, inside if group_id:. The third is resolve_model_client_config in llm.py, called from the same branch. It raises its own InvalidConfigException when the id is missing from a non-empty group list, and a default rasa train only reaches it with validation off, because rasa/cli/train.py passes validate=not args.skip_validation. Its message says the group was not found “in endpoints.yml”, although Mantle handed it the groups from integrations.yml. No check of any kind covers runs a, b and c.

rasa train endpoints.yml openai-embeddings group reads validate_agent validation.py 93-108 tracing settings unknown_model_group exit 1 (run e) set, not declared build_references_embedder references_index.py 62-81 empty, or declared integrations.yml model_groups only if the key is set line 201 OpenAI default text-embedding-3-large (runs a, b, c) empty InvalidConfigException (only with --skip-validation) set, not found named group local-embeddings (run d) set, found
  1. Validation checks the key only when it is set, and only against integrations.yml.
  2. The build step’s raise and the resolver’s sit behind the same if group_id:; with validation on, the resolver’s never fires for a missing id.
  3. An empty or misspelt key reaches embedder_factory with only the OpenAI default, and no log line names the provider.
  4. rasa train does read endpoints.yml: run b’s output includes a tracing notice naming that file. None of it reaches the references path.
FigureThree checks, one condition: how each run left rasa train on Rasa Pro 3.20.0

The code fails closed on a group it cannot find and falls open, to a hosted provider, on a key it cannot see. The comment on lines 65 to 68 refuses to fall back silently to the wrong configuration. An empty key falls back, documented but unannounced, to a provider the project may never have chosen. The docstring’s comparison with EnterpriseSearchPolicy says this default is deliberate, and a maintainer could reasonably keep it. Two small changes would have made run b visible: one info line at line 81 naming the provider and model it chose, or a validation finding when a project ships reference files with no embeddings key. Both are our suggestions, not behaviour we observed.

Name the embedder in agent.yml and declare it in integrations.yml

Run d is the change, and it needs two files: line 62 takes the id from agent.yml, and line 201 decides that the id is looked up in integrations.yml. The diff against the committed project, agent.yml first, then integrations.yml:

 references:
+  embeddings: local-embeddings
   instructions: |
     Answer from retrieved FAQ snippets only. Keep answers short for voice.
         temperature: 0.0
 
+  - id: local-embeddings
+    models:
+      - provider: ollama
+        model: all-minilm
+        api_base: http://localhost:11434
+
 channels:

Training passed and the stand-in received no request, with endpoints.yml untouched. If you want OpenAI, name an OpenAI group the same way: the key then records the choice, and validation checks it.

To audit a project before training it, read the key and look for anything else under references:. This is the command we ran in the committed project and in the copies for runs c and d. It is wrapped here with backslash continuations, which the shell removes, so it is the same one-line program:

uv run --frozen python -c "import yaml; \
r = yaml.safe_load(open('agent.yml')).get('references') or {}; \
print('references.embeddings:', r.get('embeddings') or '(empty: OpenAI default)'); \
print('other keys under references:', \
sorted(set(r) - {'instructions', 'embeddings', 'chunking'}) or 'none')"

Its output, with uv’s install notices (a hardlink warning and the Installed/Built lines) removed and nothing else changed:

$ cd examples/mantle-voice-agent
references.embeddings: (empty: OpenAI default)
other keys under references: none

$ cd runs/c-misspelt-key
references.embeddings: (empty: OpenAI default)
other keys under references: ['embedding']

$ cd runs/d-named-local
references.embeddings: local-embeddings
other keys under references: none

The label “(empty: OpenAI default)” is ours, taken from the source above. The command reads YAML; only a training run against a recorder shows what is sent.

The same fallback in stacks that are not Rasa

None of this needs Rasa to happen again. Four things lined up:

  1. A factory takes an optional user config and a hosted default, and None means the default: embedder_factory(None, DEFAULT_EMBEDDINGS_CONFIG).
  2. Validation runs only on values that are present, so absence is never an error.
  3. The schema ignores unknown keys, so a typo becomes absence.
  4. A config block elsewhere lists the same model as the default, so the fallback and the intended setting send the same request, and a green build proves nothing.

In code you do not own, search for calls where the config argument can be None next to a module-level constant naming a hosted model, and for extra="ignore" or its equivalent on the schema that holds the setting. The log line that should exist is one at build time naming the provider and model actually chosen. When it is missing, run the test we ran: point the provider’s base URL at a recorder, change the setting to a value you can tell apart from the default, and train. If the recorder still hears the default model, the setting is dead.

The recorder is a plain HTTP server. It logs the path, the model, the input count, the first 80 characters of each input and whether an Authorization header was present (never its value), and it answers embeddings requests with fake vectors.

Show the stand-in we trained against (29 lines of Python)

Start it with python openai_standin.py <port> requests.jsonl, then train with OPENAI_API_KEY=sk-standin-placeholder OPENAI_API_BASE=http://127.0.0.1:<port>/v1 uv run --frozen rasa train.

"""Local stand-in for api.openai.com. Records each request (path, model, input count,
first 80 characters of each input, auth header present or not; never its value) and
returns deterministic fake embeddings. Nothing is forwarded anywhere."""
import http.server, json, sys, datetime, hashlib
PORT, LOG = int(sys.argv[1]), sys.argv[2]
class H(http.server.BaseHTTPRequestHandler):
    def do_POST(self):
        body = json.loads(self.rfile.read(int(self.headers.get('content-length', 0))) or b'{}')
        inputs = body.get('input', [])
        inputs = [inputs] if isinstance(inputs, str) else inputs
        rec = {'at': datetime.datetime.now().isoformat(timespec='seconds'), 'method': 'POST', 'path': self.path,
               'authorization_header': 'present' if self.headers.get('authorization') else 'absent',
               'model': body.get('model'), 'inputs': len(inputs),
               'input_heads': [str(i)[:80] for i in inputs]}
        with open(LOG, 'a') as f: f.write(json.dumps(rec) + '\n')
        if self.path.rstrip('/').endswith('/embeddings'):
            def vec(t):
                h = hashlib.sha256(str(t).encode()).digest()
                return [((h[i % 32] / 255.0) - 0.5) for i in range(1536)]
            out = {'object': 'list', 'model': body.get('model'), 'data': [{'object': 'embedding', 'index': i, 'embedding': vec(t)} for i, t in enumerate(inputs)],
                   'usage': {'prompt_tokens': 0, 'total_tokens': 0}}
            data, status = json.dumps(out).encode(), 200
        else:
            data, status = json.dumps({'error': {'message': 'stand-in serves /embeddings only'}}).encode(), 404
        self.send_response(status); self.send_header('content-type', 'application/json')
        self.send_header('content-length', str(len(data))); self.end_headers(); self.wfile.write(data)
    def do_GET(self): self.do_POST()
    def log_message(self, *a): pass
http.server.ThreadingHTTPServer(('127.0.0.1', PORT), H).serve_forever()

We chose a recorder over an invalid key for this test. A call that fails tells you something was attempted; a recorder lets training finish, as it would in production, and keeps the request itself: the model, the input count and the start of every chunk.