Skip to content
RasaGet a free licence
Guides for AI teams

Guide · Platform / operations engineer

Rasa 3.20.0 sent endpoints.yml's api_key_env; OpenAI refused

On Rasa Pro 3.20.0 an endpoints.yml model group sends api_key_env to OpenAI as a request parameter. With the health check on, OpenAI refused it; Rasa exited.

by Rod Rivera

About 7 minutes

  • exit 1

    server exit code on real OpenAI with the startup health check on, project unedited

  • 29 lines

    active api_key_env lines in 15 endpoints.yml files in the companion at 01d6a6e

  • 1 of 15

    of those files run against a real provider

Source: Live runs of examples/mantle-voice-agent at companion commit 01d6a6e on Rasa Pro 3.20.0
Key takeaways (5)
  • On Rasa Pro 3.20.0 the OpenAI and default litellm clients built from an endpoints.yml model group keep every key they do not recognise and pass it to litellm, which sent api_key_env to the provider as a request parameter.
  • Rasa already checks endpoints.yml model-group keys by name when it trains and when it loads a model: api_key and the other keys on its SENSITIVE_DATA list must be written as ${NAME}. api_key_env, a key Rasa's own Mantle client defines, is on no list, so it passes.
  • With LLM_API_HEALTH_CHECK=true the unedited companion voice agent exited 1 at startup after OpenAI answered "Unknown parameter: 'api_key_env'." In a run with the rephraser group switched to Anthropic, Anthropic answered "api_key_env: Extra inputs are not permitted".
  • The startup check tested a client no reply used. We recorded four REST turns on this Mantle project; in the three that could tell the clients apart, the rephrase ran on the orchestrator's client from integrations.yml, and only the startup test call used the endpoints.yml group.
  • Write api_key: ${NAME} in both endpoints.yml and integrations.yml. On 3.20.0 integrations.yml also accepts api_key_env, but the placeholder works in both. Then start once with the health check on against the real provider before you deploy.

The companion voice agent, unedited, starting on Rasa Pro 3.20.0 against the real OpenAI API:

$ LLM_API_HEALTH_CHECK=true uv run --frozen rasa run --enable-api --port 55931
INFO   Sending a test LLM API request for the component -
       ContextualResponseRephraser.
       config: {"model": "gpt-5.2", […], "api_key_env": "OPENAI_API_KEY"}
ERROR  Test call to the LLM API failed for component -
       ContextualResponseRephraser.
       […] litellm.BadRequestError: OpenAIException -
       Unknown parameter: 'api_key_env'.
(harness) server process exit code: 1; /status: connection refused (http 000)

Excerpt, re-wrapped to fit. Cut: timestamps, logger names, the event and level fields, the config except model and api_key_env, and the ProviderClientAPIException wrapper; […] marks cuts inside a line. The last line is our harness’s summary, prefix ours. The unedited lines are in the collapsed block below. One line changed in the rephraser’s model group in endpoints.yml starts the server:

   - id: openai-llm
     models:
       - provider: openai
         model: gpt-5.2
-        api_key_env: OPENAI_API_KEY
+        api_key: ${OPENAI_API_KEY}
         temperature: 0.0
(harness) POST /webhooks/rest/webhook  "Hi there"
reply: Hi. I’m Atlas, Horizon Travel’s voice assistant. I can help with
       things like viewing your itinerary, checking flight status, updating
       a booking, or reporting lost baggage. What would you like to do?

First line: our summary of the request. The reply is the response’s text field, re-wrapped, with its JSON escapes decoded.

Atlas and Horizon Travel are the companion project’s fictional fixture. The failure needs LLM_API_HEALTH_CHECK=true: Rasa’s default is false, the companion’s .env.example does not set it, and with it off the same unedited files started and answered with no ERROR line.

Here is the mechanism. Rasa checks the keys of endpoints.yml model groups by name when it trains and when it loads a model, but only names on its lists: api_key must be written as ${NAME}, for example. api_key_env is a key that Rasa’s own Mantle client defines for integrations.yml, and it is on no list, so it passes. On the OpenAI and default litellm paths we read, the client built from the group then keeps every key it does not recognise and hands it to litellm, which sends it to the provider. The first objection comes from the provider, and only when something calls it. Here the caller was the startup health check, and it tested a client that no reply used: in the recorded turns that could tell, replies were rephrased by the orchestrator’s client, built from integrations.yml.

endpoints.yml openai-llm group ContextualResponseRephraser created by the Agent nlg.llm.model_group health check if LLM_API_HEALTH_CHECK=true startup test call body carries api_key_env integrations.yml orchestrator group LLMClient.from_bootstrap pops api_key_env orchestrator _rephrase_and_send orchestrator.py 268 turn-time rephrase no api_key_env in body line 936
  1. nlg.llm.model_group names this group, and nothing on its path reads api_key_env.
  2. Run when the server starts; in our training run with the check on it sent no test call.
  3. OpenAI and Anthropic each refused this request, and the server did not come up.
  4. The one reader of api_key_env, fed from integrations.yml.
  5. _rephrase_and_send calls the orchestrator’s client; in the recorded turns that could tell, the replies used it.
FigureTwo clients, two files: which call uses which model group (Rasa Pro 3.20.0, Mantle)

Our position: Rasa should refuse api_key_env in endpoints.yml at load time with a dedicated rule whose message names api_key: ${NAME}. The cheaper fix is one entry in SENSITIVE_DATA, the list its load-time check already keeps. By our reading of the source that entry would reject both string forms, but with the wrong advice: api_key_env: NAME would be told it “must be set as an environment variable”, and api_key_env: ${NAME} would then be told it is “not allowed”, with a list of the keys that may use ${...}. Neither message says to write api_key instead. A dedicated rule is one more check in a loop that already walks every model entry, and it can name the fix. Either way, a project whose endpoints.yml carries the line, as the companion’s examples do, would stop at rasa train instead of running quietly with the check off. We would take that break. On 3.20.0, write api_key: ${NAME} in both files and make the provider answer once before every deploy; integrations.yml also accepts api_key_env: NAME there, but the placeholder is the form that works in both.

Show the unedited log lines from the OpenAI and Anthropic runs

OpenAI, the run above:

2026-09-29 11:50:21 INFO     rasa.shared.utils.health_check.health_check  - {"event_info": "Sending a test LLM API request for the component - ContextualResponseRephraser.", "config": {"model": "gpt-5.2", "api_base": null, "api_version": null, "api_type": "openai", "provider": "openai", "reasoning_effort": "none", "temperature": 0.0, "max_completion_tokens": 256, "timeout": 5, "api_key_env": "OPENAI_API_KEY"}, "event": "contextual_response_rephraser.init.send_test_llm_api_request", "level": "info"}
2026-09-29 11:50:21 ERROR    rasa.core.run  - {"event_info": "Test call to the LLM API failed for component - ContextualResponseRephraser.", "config": {"model": "gpt-5.2", "api_base": null, "api_version": null, "api_type": "openai", "provider": "openai", "reasoning_effort": "none", "temperature": 0.0, "max_completion_tokens": 256, "timeout": 5, "api_key_env": "OPENAI_API_KEY"}, "error": "ProviderClientAPIException(\"\\nOriginal error: litellm.BadRequestError: OpenAIException - Unknown parameter: 'api_key_env'.)\")", "event": "contextual_response_rephraser.init.send_test_llm_api_request_failed", "level": "error"}

Anthropic, the run in the provider section below:

2026-09-29 11:33:04 ERROR    rasa.core.run  - {"event_info": "Test call to the LLM API failed for component - ContextualResponseRephraser.", "config": {"model": "claude-haiku-4-5", "provider": "anthropic", "api_key_env": "ANTHROPIC_API_KEY", "temperature": 0.0}, "error": "ProviderClientAPIException('\\nOriginal error: litellm.BadRequestError: AnthropicException - {\"type\":\"error\",\"error\":{\"type\":\"invalid_request_error\",\"message\":\"api_key_env: Extra inputs are not permitted\"},\"request_id\":\"req_011CfXaYEuMPSm6AqRe3Du4h\"})')", "event": "contextual_response_rephraser.init.send_test_llm_api_request_failed", "level": "error"}

Live runs of examples/mantle-voice-agent (RasaHQ/rasa-community-resources at 01d6a6e) on Rasa Pro 3.20.0 with litellm 1.100.1, as the project’s uv.lock pins them. Runs recorded 29 September 2026 against the 3.20.0 release. Real-provider runs used the key in the project’s .env; the rest used placeholder keys and a local stand-in. Receipts are in editorial/receipts/endpoints-api-key-env-ignored/ in the site repository.

The line is ours. Companion commit 432d416, titled “Use api_key_env: NAME; api_key: ${VAR} never expanded”, made this change to every endpoints.yml credential line:

$ git -C rasa-community-resources show 432d416 -- 'examples/*/endpoints.yml' 'tutorials/*/endpoints.yml' 'community/*/endpoints.yml' | grep '^[-+] .*api_key' | sort | uniq -c
   1 -        api_key: ${GEMINI_API_KEY}
  28 -        api_key: ${OPENAI_API_KEY}
   1 -  #       api_key: ${OPENAI_API_KEY}
   1 +        api_key_env: GEMINI_API_KEY
  28 +        api_key_env: OPENAI_API_KEY
   1 +  #       api_key_env: OPENAI_API_KEY

On 3.20.0 the commit’s premise did not hold for endpoints.yml: the client expands ${NAME} when it calls the provider, which is why the edit above started the server and why a stand-in received the expanded placeholder key. At 01d6a6e, 29 active api_key_env lines remain in 15 endpoints.yml files. We ran one of those files against a real provider.

The request carried api_key_env as a body field

OpenAI’s error names a parameter; a recorder shows the request that carried it. The stand-in is a local HTTP server that OPENAI_API_BASE points at. It writes one JSON line per request (path, bearer value, top-level keys of the body, model) and answers with a fake completion. For two runs we gave the rephraser’s group its own variable, set OPENAI_API_KEY and REPHRASER_OPENAI_KEY to two placeholders, trained, started the server and sent “Hi there”. One run kept the api_key_env spelling; the other used api_key: ${...}.

What the stand-in received, by rephraser credential line (selected fields)

Avoid: api_key_env: REPHRASER_OPENAI_KEY

/v1/embeddings        sk-standin-default
/v1/chat/completions  sk-standin-default
    body keys: api_key_env, max_completion_tokens,
    messages, model, reasoning_effort, temperature
    api_key_env = REPHRASER_OPENAI_KEY
/v1/chat/completions  sk-standin-default
    body keys: messages, model,
    reasoning_effort, temperature

The test call goes out on the default key and carries the variable’s name as a body field.

Prefer: api_key: ${REPHRASER_OPENAI_KEY}

/v1/embeddings        sk-standin-default
/v1/chat/completions  sk-standin-rephraser
    body keys: max_completion_tokens,
    messages, model, reasoning_effort, temperature
/v1/chat/completions  sk-standin-default
    body keys: messages, model,
    reasoning_effort, temperature

The test call carries the named key and no stray field. The last request is unchanged.

Each pane shows the path, bearer and body keys from the stand-in’s raw lines, in the order received; the raw lines are in the receipts. The first request is an embeddings call made during training by Mantle’s default references embedder, not by an endpoints.yml group. The second is the rephraser’s startup test call, the only request with max_completion_tokens, matching the 256 in the health check’s logged config. The third is the reply to “Hi there”.

The stand-in accepts anything, so in the left pane the field simply arrives. The left pane also shows the test call going out on sk-standin-default, the value of OPENAI_API_KEY, while the YAML named REPHRASER_OPENAI_KEY; the litellm source below explains why. The stand-in is 25 lines of Python, and these two records are its output.

Show the stand-in and the commands for these two runs

Start the recorder, then train and start Rasa against it with two different placeholder keys (set the same variables for rasa train). The ports are yours to choose.

python openai_standin_chat.py <port> requests.jsonl &
LLM_API_HEALTH_CHECK=true OPENAI_API_KEY=sk-standin-default \
  REPHRASER_OPENAI_KEY=sk-standin-rephraser \
  OPENAI_API_BASE=http://127.0.0.1:<port>/v1 \
  uv run --frozen rasa run --enable-api --port <rasa-port>

Give the rephraser’s group api_key_env: REPHRASER_OPENAI_KEY for the left pane or api_key: ${REPHRASER_OPENAI_KEY} for the right. The script logs a bearer value only when it starts with sk-standin, so a real key sent by mistake is never written down.

"""Local stand-in for api.openai.com used only with placeholder keys. Records, per request:
path, the bearer value (placeholders set by the test, never a real key), top-level body keys,
model. Returns a fake chat completion or fake embeddings. Nothing is forwarded."""
import http.server, json, sys, datetime, hashlib
PORT, LOG = int(sys.argv[1]), sys.argv[2]
class H(http.server.BaseHTTPRequestHandler):
    def do_POST(self):
        body = json.loads(self.rfile.read(int(self.headers.get('content-length', 0))) or b'{}')
        auth = self.headers.get('authorization', '')
        bearer = auth.split(' ', 1)[1] if auth.startswith('Bearer ') else ('(none)' if not auth else '(non-bearer)')
        if not bearer.startswith('sk-standin') and bearer not in ('(none)', '(non-bearer)'): bearer = '(unexpected value, not logged)'
        with open(LOG, 'a') as f:
            f.write(json.dumps({'at': datetime.datetime.now().isoformat(timespec='seconds'), 'path': self.path, 'bearer': bearer,
                                'body_keys': sorted(body), 'model': body.get('model'), 'api_key_env_in_body': body.get('api_key_env')}) + '\n')
        if self.path.rstrip('/').endswith('/embeddings'):
            inputs = body.get('input', []); inputs = [inputs] if isinstance(inputs, str) else inputs
            out = {'object': 'list', 'data': [{'object': 'embedding', 'index': i, 'embedding': [0.01] * 1536} for i, _ in enumerate(inputs)], 'model': body.get('model'), 'usage': {'prompt_tokens': 0, 'total_tokens': 0}}
        else:
            out = {'id': 'chatcmpl-standin', 'object': 'chat.completion', 'created': 0, 'model': body.get('model'),
                   'choices': [{'index': 0, 'message': {'role': 'assistant', 'content': 'stand-in reply'}, 'finish_reason': 'stop'}],
                   'usage': {'prompt_tokens': 1, 'completion_tokens': 1, 'total_tokens': 2}}
        data = json.dumps(out).encode()
        self.send_response(200); self.send_header('content-type', 'application/json'); self.send_header('content-length', str(len(data))); self.end_headers(); self.wfile.write(data)
    def log_message(self, *a): pass
http.server.ThreadingHTTPServer(('127.0.0.1', PORT), H).serve_forever()

Rasa’s load-time check has a credential list, and api_key_env is not on it

Rasa does check these groups before any request. validate_model_group_configuration_setup in rasa/engine/validation.py runs over every endpoints.yml model group when rasa train builds a model (rasa/model_training.py, line 347) and when Rasa loads one (rasa/engine/loader.py, line 46). Two of its checks are about key names. The first allows a ${...} value only on these keys (lines 1351 to 1360):

    allowed_env_vars = {
        DEPLOYMENT_CONFIG_KEY,
        API_BASE_CONFIG_KEY,
        API_KEY,
        API_VERSION_CONFIG_KEY,
        AWS_REGION_NAME_CONFIG_KEY,
        AWS_ACCESS_KEY_ID_CONFIG_KEY,
        AWS_SECRET_ACCESS_KEY_CONFIG_KEY,
        AWS_SESSION_TOKEN_CONFIG_KEY,
    }

The second requires every key on this list, from rasa/shared/constants.py (lines 361 to 372), to be written as ${...}, and raises a ValidationError otherwise:

SENSITIVE_DATA = [
    API_KEY,
    AWS_ACCESS_KEY_ID_CONFIG_KEY,
    AWS_SECRET_ACCESS_KEY_CONFIG_KEY,
    AWS_SESSION_TOKEN_CONFIG_KEY,
    LANGFUSE_CONFIG_PUBLIC_KEY,
    LANGFUSE_CONFIG_PRIVATE_KEY,
    CLIENT_ID_CONFIG_KEY,
    CLIENT_SECRET_CONFIG_KEY,
    TOKEN_URL_CONFIG_KEY,
    A2A_JWT_SECRET_CONFIG_KEY,
]

api_key_env: OPENAI_API_KEY is neither a ${...} value nor a listed key, so it passes both, and rasa train exited 0 in each of the six runs where we recorded its exit code. Neither file mentions api_key_env. The list is where Rasa has already decided that key names matter for credentials; api_key_env is a credential name Rasa itself defines, in its Mantle client, and the list does not know it. In effect the two checks are a name-based schema for secrets, with one name missing.

Past the check, the two clients we read keep unknown keys

examples/mantle-voice-agent/endpoints.yml, lines 16-31 (comment block cut)
nlg:
  type: rephrase
  llm:
    model_group: openai-llm

model_groups:
  - id: openai-llm
    models:
      - provider: openai
        model: gpt-5.2
        api_key_env: OPENAI_API_KEY
        temperature: 0.0
  1. The rephraser’s group. ContextualResponseRephraser resolves it when it is created and runs the health check on it.
  2. provider: openai selects OpenAILLMClient. Any provider Rasa does not map, anthropic included, gets DefaultLiteLLMClient (rasa/shared/providers/mappings.py).
  3. The load-time check passes it, and nothing after it reads it. It stays in the config, is printed in the log’s config, and is the field OpenAI named.

The client config takes out the keys it knows and keeps the rest. Its comment says “The rest of parameters (e.g. model parameters) are considered as extra parameters (this also includes timeout)”. This is OpenAIClientConfig, in rasa/shared/providers/_configs/openai_client_config.py (lines 141 to 154):

        this = OpenAIClientConfig(
            # Required parameters
            model=config.pop(MODEL_CONFIG_KEY),
            # Pop the 'provider' key. Currently, it's *optional* because of
            # backward compatibility with older versions.
            provider=config.pop(PROVIDER_CONFIG_KEY, OPENAI_PROVIDER),
            # Optional parameters
            api_base=config.pop(API_BASE_CONFIG_KEY, None),
            api_version=config.pop(API_VERSION_CONFIG_KEY, None),
            api_type=config.pop(API_TYPE_CONFIG_KEY, OPENAI_API_TYPE),
            # The rest of parameters (e.g. model parameters) are considered
            # as extra parameters (this also includes timeout).
            extra_parameters=config,
        )

api_key_env is in “the rest”. The default litellm client’s config does the same after taking out only model and provider. Both clients then spread the extras into the litellm call, in rasa/shared/providers/llm/_base_litellm_client.py (lines 84 to 96):

    @property
    def _completion_fn_args(self) -> dict:
        return {
            # Since all providers covered by LiteLLM use the OpenAI format, but
            # not all support every OpenAI parameter, raise an exception if
            # provider/model uses unsupported parameter
            "drop_params": False,
            # All other parameters set through config, can override drop_params
            **self._litellm_extra_parameters,
            # Model name is constructed in the LiteLLM format from the provided config
            # Non-overridable to ensure consistency
            "model": self._litellm_model_name,
        }

The **self._litellm_extra_parameters spread carries api_key_env forward, beside temperature and reasoning_effort. drop_params: False is a separate setting; its comment says the aim is to “raise an exception if provider/model uses unsupported parameter”. In our runs the exception came from the provider, after the request was sent. The same file passes the arguments through resolve_environment_variables, which calls os.path.expandvars (rasa/shared/utils/io.py). That is where api_key: ${REPHRASER_OPENAI_KEY} became sk-standin-rephraser.

litellm 1.100.1 does the last step (litellm/utils.py, lines 4719 to 4745). For openai, a keyword it does not recognise goes into extra_body; for other providers it goes into the request parameters. With no api_key, litellm’s OpenAI path falls back to the OPENAI_API_KEY environment variable (litellm/main.py), which is where the left pane’s sk-standin-default came from. Both litellm points are source reading; the wire behaviour is in the records. This covers the OpenAI and default litellm clients, the two paths we read.

Only integrations.yml turns api_key_env into a key

A grep of the 3.20.0 wheel finds one reader of the YAML key, in rasa/mantle/llm/client.py. The Azure clients’ hits are a private attribute whose YAML key is api_key, and a docstring in rasa/cli/project_env.py names integrations.yml:

$ grep -rn api_key_env rasa/ --include='*.py'
rasa/mantle/llm/client.py:43:_API_KEY_ENV_CONFIG_KEY = "api_key_env"
rasa/mantle/llm/client.py:136:def _resolve_api_key_env(config: Dict[str, Any]) -> Dict[str, Any]:
rasa/mantle/llm/client.py:150:                        "mantle.llm.api_key_env.not_set",
rasa/mantle/llm/client.py:257:        Resolves ``api_key_env`` entries — where the value is an environment
rasa/mantle/llm/client.py:260:        ``api_key_env`` key (an engine convention unknown to LiteLLM).
rasa/mantle/llm/client.py:262:        llm_config = _resolve_api_key_env(dict(bootstrap.llm))
rasa/shared/providers/llm/azure_openai_llm_client.py:137:        self._api_key_env_var = (
rasa/shared/providers/llm/azure_openai_llm_client.py:138:            self._resolve_api_key_env_var() if not self._oauth else None
rasa/shared/providers/llm/azure_openai_llm_client.py:197:    def _resolve_api_key_env_var(self) -> str:
rasa/shared/providers/llm/azure_openai_llm_client.py:344:        elif self._api_key_env_var:
rasa/shared/providers/llm/azure_openai_llm_client.py:345:            auth_parameter = {LITE_LLM_API_KEY_FIELD: self._api_key_env_var}
rasa/shared/providers/embedding/azure_openai_embedding_client.py:98:        self._api_key_env_var = (
rasa/shared/providers/embedding/azure_openai_embedding_client.py:99:            self._resolve_api_key_env_var() if not self._oauth else None
rasa/shared/providers/embedding/azure_openai_embedding_client.py:104:    def _resolve_api_key_env_var(self) -> str:
rasa/shared/providers/embedding/azure_openai_embedding_client.py:244:        elif self._api_key_env_var:
rasa/shared/providers/embedding/azure_openai_embedding_client.py:245:            auth_parameter = {LITE_LLM_API_KEY_FIELD: self._api_key_env_var}
rasa/cli/project_env.py:64:    ``api_key_env`` in ``integrations.yml`` available even when the process has

_resolve_api_key_env walks a model group, pops api_key_env and writes the named variable’s value into api_key. Its one caller explains why, in words that fit this whole page (client.py, lines 254 to 262):

    def from_bootstrap(cls, bootstrap: ModelBootstrap) -> "LLMClient":
        """Build client from the project's integrations.yml LLM config.

        Resolves ``api_key_env`` entries — where the value is an environment
        variable *name* — into the corresponding ``api_key`` value so the
        config that reaches ``llm_factory`` never contains the raw
        ``api_key_env`` key (an engine convention unknown to LiteLLM).
        """
        llm_config = _resolve_api_key_env(dict(bootstrap.llm))

bootstrap.llm is the orchestrator’s group, resolved from the project’s integrations.yml (rasa/mantle/model_archive/bootstrap.py). In the voice agent that file carries the same line under a different group (lines 6 to 15):

llm:
  model_group: orchestrator

model_groups:
  - id: orchestrator
    models:
      - provider: openai
        model: gpt-5.2
        api_key_env: OPENAI_API_KEY
        temperature: 0.0

So one project holds the same spelling twice: resolved in integrations.yml, forwarded from endpoints.yml. The integrations.yml half rests on the source above and on the shape of the reply request, which carried no api_key_env in either stand-in record. No run gave the orchestrator’s group a second variable name to tell the two apart.

The startup check tested a client the recorded replies never used

The forwarded key reaches a provider only when something sends the rephraser’s group. perform_llm_health_check, in rasa/shared/utils/health_check/health_check.py (lines 77 to 149), sends a test call when the variable reads true (compared case-insensitively). With it off, the same function can still probe Azure deployment-only groups at inference; otherwise it logs this warning, quoted from our real-OpenAI run with the check off:

The LLM_API_HEALTH_CHECK environment variable is set to false, which will disable LLM health check. It is recommended to set this variable to true in production environments.

A failed test call is re-raised as a HealthCheckError (lines 268 to 295); on the first screen rasa.core.run logged it and the process exited 1. For this project’s openai group we checked the switch at the stand-in. rasa run alone with the check on sent one request, the test call, api_key_env included; with it off the stand-in received nothing. With the check on, rasa train exited 0 and the stand-in received only the embeddings request. In this project the check runs when the server starts, not when it trains.

With the check off, does anything else use the group? For one Mantle component this is on record: the guide on Mantle’s references embedder showed the embedder is chosen from agent.yml and integrations.yml, and that repointing the endpoints.yml embeddings group changed no request. The rephrase needs its own evidence; the source and three recorded turns give it.

The orchestrator builds its client with LLMClient.from_bootstrap(self._static_data), and _rephrase_and_send passes that client to self._call_llm (rasa/mantle/orchestration/orchestrator.py, lines 268 and 936). The endpoints.yml rephraser is created by the Agent’s NaturalLanguageGenerator.create (rasa/core/agent.py), and a grep of rasa/mantle finds no reference to ContextualResponseRephraser.

We recorded four REST turns, one in each of rows 2, 3, 5 and 6 of the table below. Row 2 cannot tell the clients apart, because both would carry the same key. The other three can. In row 3 the check was off on real OpenAI, the group still carried api_key_env, and the greeting came back reworded, logged as mantle.turn.completed with "source": "rephrase_text". OpenAI refused that field in row 1, and a failed rephrase falls back to the verbatim text (rasa/mantle/orchestration/responses.py, lines 93 to 98). So by inference the rephrase request did not carry the field. The first sentence of each, from the receipt:

responses.yml utter_greet:  Hello. I'm Atlas, your Horizon Travel voice assistant.
reply, check off:           Hi. I’m Atlas, Horizon Travel’s voice assistant.

At the stand-in, row 5’s rephrase request carried no api_key_env field although the group did. In row 6, with the group spelled api_key: ${REPHRASER_OPENAI_KEY}, the test call carried sk-standin-rephraser and the rephrase request carried sk-standin-default, the orchestrator’s key. These are REST turns on one project; we did not run other channels or a project without Mantle.

The provider is the first thing that objects

OpenAI’s reply is on the first screen. For the Anthropic run we switched the rephraser’s group to Anthropic and left api_key_env in it:

   - id: openai-llm
     models:
-      - provider: openai
-        model: gpt-5.2
-        api_key_env: OPENAI_API_KEY
+      - provider: anthropic
+        model: claude-haiku-4-5
+        api_key_env: ANTHROPIC_API_KEY
         temperature: 0.0

The orchestrator and embeddings went to the stand-in; only the rephraser’s test call went to Anthropic, and /status never answered. The error, as an excerpt re-wrapped to fit:

ERROR  Test call to the LLM API failed for component -
       ContextualResponseRephraser.
       […] AnthropicException - […]"type":"invalid_request_error",
       "message":"api_key_env: Extra inputs are not permitted"[…],
       "request_id":"req_011CfXaYEuMPSm6AqRe3Du4h"

Cut: the timestamp, the logger name, the config object, the event and level fields, the ProviderClientAPIException and litellm.BadRequestError wrapper, and the outer JSON of Anthropic’s reply; […] marks the cuts inside a line, and the backslash escapes on the quotes are decoded. The full line is in the collapsed block near the top.

Two providers, two client classes in Rasa, two litellm branches (extra_body and request parameters), and the same refusal. Neither reply names the credential problem: both describe an unexpected parameter, which is accurate and points away from the key.

Strictness is what made the key visible. The stand-in is the lenient case: it accepted the field and the run carried on. A provider that ignored unknown fields would do the same. We ran no such provider; the stand-in shows the shape of that case, not any provider’s behaviour.

Seven runs of this project in six rows; row 1 holds two. They are every run on this page except the two switch runs, which started the server and sent no message, the training-only run, and the quickstart run, which is a different project. Only the rephraser’s group and the environment change between rows. In row 3 the check was off, so the group was sent nowhere.

#Rephraser group credential lineProviderLLM_API_HEALTH_CHECKWhat happenedReceipt
1api_key_env: OPENAI_API_KEY (as committed)real OpenAItrueTest call rejected: Unknown parameter: 'api_key_env'. Exit 1, /status never answeredr320-openai-rephraser-committed.txt, and run 2 of r320-openai-committed-health-check-off.txt
2api_key: ${OPENAI_API_KEY}real OpenAItrueTest call sent, no ERROR line. /status answered; “Hi there” got the Atlas greetingr320-openai-rephraser-placeholder.txt
3api_key_env: OPENAI_API_KEY (as committed)real OpenAIfalseStarted, 0 ERROR lines. The greeting came back rewordedr320-openai-committed-health-check-off.txt
4provider: anthropic, claude-haiku-4-5, api_key_env: ANTHROPIC_API_KEYreal Anthropic, for the test call onlytrueTest call rejected: api_key_env: Extra inputs are not permitted. /status never answeredr320-rephraser-anthropic.txt
5api_key_env: REPHRASER_OPENAI_KEYstand-intrueTest call: bearer sk-standin-default, body field api_key_env. Reply request: bearer sk-standin-defaultr320-rephraser-own-var.txt and its raw .jsonl.txt
6api_key: ${REPHRASER_OPENAI_KEY}stand-intrueTest call: bearer sk-standin-rephraser, no api_key_env. Reply request: bearer sk-standin-defaultr320-rephraser-placeholder.txt and its raw .jsonl.txt

Rows 1 and 4 are one key refused by two providers in their own words. Rows 1 and 3 differ only in the switch. Rows 5 and 6 are the counter-test for the spelling. In rows 2, 4, 5 and 6, and in the first of row 1’s two runs, the test call was logged, so by the source the variable was true; where it came from was not recorded for those runs. The second run of row 1 set it on the command line, as the first screen shows.

Make the provider answer before you deploy

Find the lines. In the companion at 01d6a6e this matched 30 lines in 15 files, one of them inside a comment:

git grep -n api_key_env -- '*endpoints*.yml'

Rewrite each active match as api_key: ${NAME} with the same variable name; in endpoints.yml we ran that form on OpenAI only. Use the same form in integrations.yml. There 3.20.0 reads api_key_env too, and the placeholder is expanded at call time by the same base client. A live 3.20.0 run of this site’s quickstart project with api_key: ${ANTHROPIC_API_KEY} in integrations.yml trained and answered four conversations. One spelling in both files means one rule to check. Then start once with the check on and the real key, and require /status to answer:

LLM_API_HEALTH_CHECK=true uv run --frozen rasa run --enable-api

With the committed line it exits 1; with the edit it starts. In this project that is one real request per start. A strict provider tells you the request was malformed; it does not tell you which key it used. For that, run the stand-in above.

Added 29 September 2026: the prerelease refuses the key

Rasa Pro 3.21.0.dev3, a prerelease uploaded to PyPI on 28 September at 15:57 UTC, refuses api_key_env in an endpoints.yml model group at load time. validate_model_group_configuration_setup now calls validate_model_group_credentials, which raises api_key_env_not_supported with a message that ends “which is no longer supported. Replace it with ‘api_key: ${ENV_VAR_NAME}’.” That is a dedicated rule with the message argued for above. We read it in the prerelease source and did not run it; everything above this note is about 3.20.0. If you are on 3.20.0, write api_key: ${NAME} now: it works there in both files, and the prerelease refuses the other spelling in endpoints.yml.