Guide · Platform / operations engineer
Rasa 3.20.0 sent endpoints.yml's api_key_env; OpenAI refused
On Rasa Pro 3.20.0 an endpoints.yml model group sends api_key_env to OpenAI as a request parameter. With the health check on, OpenAI refused it; Rasa exited.
exit 1
server exit code on real OpenAI with the startup health check on, project unedited
29 lines
active api_key_env lines in 15 endpoints.yml files in the companion at 01d6a6e
1 of 15
of those files run against a real provider
Key takeaways (5)
- On Rasa Pro 3.20.0 the OpenAI and default litellm clients built from an endpoints.yml model group keep every key they do not recognise and pass it to litellm, which sent api_key_env to the provider as a request parameter.
- Rasa already checks endpoints.yml model-group keys by name when it trains and when it loads a model: api_key and the other keys on its SENSITIVE_DATA list must be written as ${NAME}. api_key_env, a key Rasa's own Mantle client defines, is on no list, so it passes.
- With LLM_API_HEALTH_CHECK=true the unedited companion voice agent exited 1 at startup after OpenAI answered "Unknown parameter: 'api_key_env'." In a run with the rephraser group switched to Anthropic, Anthropic answered "api_key_env: Extra inputs are not permitted".
- The startup check tested a client no reply used. We recorded four REST turns on this Mantle project; in the three that could tell the clients apart, the rephrase ran on the orchestrator's client from integrations.yml, and only the startup test call used the endpoints.yml group.
- Write api_key: ${NAME} in both endpoints.yml and integrations.yml. On 3.20.0 integrations.yml also accepts api_key_env, but the placeholder works in both. Then start once with the health check on against the real provider before you deploy.
The companion voice agent, unedited, starting on Rasa Pro 3.20.0 against the real OpenAI API:
$ LLM_API_HEALTH_CHECK=true uv run --frozen rasa run --enable-api --port 55931
INFO Sending a test LLM API request for the component -
ContextualResponseRephraser.
config: {"model": "gpt-5.2", […], "api_key_env": "OPENAI_API_KEY"}
ERROR Test call to the LLM API failed for component -
ContextualResponseRephraser.
[…] litellm.BadRequestError: OpenAIException -
Unknown parameter: 'api_key_env'.
(harness) server process exit code: 1; /status: connection refused (http 000)
Excerpt, re-wrapped to fit. Cut: timestamps, logger names, the event and
level fields, the config except model and api_key_env, and the
ProviderClientAPIException wrapper; […] marks cuts inside a line. The last
line is our harness’s summary, prefix ours. The unedited lines are in the
collapsed block below. One line changed in the rephraser’s model group in
endpoints.yml starts the server:
- id: openai-llm
models:
- provider: openai
model: gpt-5.2
- api_key_env: OPENAI_API_KEY
+ api_key: ${OPENAI_API_KEY}
temperature: 0.0
(harness) POST /webhooks/rest/webhook "Hi there"
reply: Hi. I’m Atlas, Horizon Travel’s voice assistant. I can help with
things like viewing your itinerary, checking flight status, updating
a booking, or reporting lost baggage. What would you like to do?
First line: our summary of the request. The reply is the response’s text
field, re-wrapped, with its JSON escapes decoded.
Atlas and Horizon Travel are the companion project’s fictional fixture. The
failure needs LLM_API_HEALTH_CHECK=true: Rasa’s default is false, the
companion’s .env.example does not set it, and with it off the same unedited
files started and answered with no ERROR line.
Here is the mechanism. Rasa checks the keys of endpoints.yml model groups by
name when it trains and when it loads a model, but only names on its lists:
api_key must be written as ${NAME}, for example. api_key_env is a key
that Rasa’s own Mantle client defines for integrations.yml, and it is on no
list, so it passes. On the OpenAI and default litellm paths we read, the
client built from the group then keeps every key it does not recognise and
hands it to litellm, which sends it to the provider. The first objection
comes from the provider, and only when something calls it. Here the caller
was the startup health check, and it tested a client that no reply used: in
the recorded turns that could tell, replies were rephrased by the
orchestrator’s client, built from integrations.yml.
nlg.llm.model_groupnames this group, and nothing on its path readsapi_key_env.- Run when the server starts; in our training run with the check on it sent no test call.
- OpenAI and Anthropic each refused this request, and the server did not come up.
- The one reader of
api_key_env, fed fromintegrations.yml. _rephrase_and_sendcalls the orchestrator’s client; in the recorded turns that could tell, the replies used it.
Our position: Rasa should refuse api_key_env in endpoints.yml at load
time with a dedicated rule whose message names api_key: ${NAME}. The cheaper
fix is one entry in SENSITIVE_DATA, the list its load-time check already
keeps. By our reading of the source that entry would reject both string
forms, but with the wrong advice: api_key_env: NAME would be told it “must
be set as an environment variable”, and api_key_env: ${NAME} would then be
told it is “not allowed”, with a list of the keys that may use ${...}.
Neither message says to write api_key instead. A dedicated rule is one more
check in a loop that already walks every model entry, and it can name the fix.
Either way, a project whose endpoints.yml carries the line, as the
companion’s examples do, would stop at rasa train instead of running quietly
with the check off. We would take that break. On 3.20.0, write
api_key: ${NAME} in both files and make the provider answer once before
every deploy; integrations.yml also accepts api_key_env: NAME there, but
the placeholder is the form that works in both.
Show the unedited log lines from the OpenAI and Anthropic runs
OpenAI, the run above:
2026-09-29 11:50:21 INFO rasa.shared.utils.health_check.health_check - {"event_info": "Sending a test LLM API request for the component - ContextualResponseRephraser.", "config": {"model": "gpt-5.2", "api_base": null, "api_version": null, "api_type": "openai", "provider": "openai", "reasoning_effort": "none", "temperature": 0.0, "max_completion_tokens": 256, "timeout": 5, "api_key_env": "OPENAI_API_KEY"}, "event": "contextual_response_rephraser.init.send_test_llm_api_request", "level": "info"}
2026-09-29 11:50:21 ERROR rasa.core.run - {"event_info": "Test call to the LLM API failed for component - ContextualResponseRephraser.", "config": {"model": "gpt-5.2", "api_base": null, "api_version": null, "api_type": "openai", "provider": "openai", "reasoning_effort": "none", "temperature": 0.0, "max_completion_tokens": 256, "timeout": 5, "api_key_env": "OPENAI_API_KEY"}, "error": "ProviderClientAPIException(\"\\nOriginal error: litellm.BadRequestError: OpenAIException - Unknown parameter: 'api_key_env'.)\")", "event": "contextual_response_rephraser.init.send_test_llm_api_request_failed", "level": "error"}Anthropic, the run in the provider section below:
2026-09-29 11:33:04 ERROR rasa.core.run - {"event_info": "Test call to the LLM API failed for component - ContextualResponseRephraser.", "config": {"model": "claude-haiku-4-5", "provider": "anthropic", "api_key_env": "ANTHROPIC_API_KEY", "temperature": 0.0}, "error": "ProviderClientAPIException('\\nOriginal error: litellm.BadRequestError: AnthropicException - {\"type\":\"error\",\"error\":{\"type\":\"invalid_request_error\",\"message\":\"api_key_env: Extra inputs are not permitted\"},\"request_id\":\"req_011CfXaYEuMPSm6AqRe3Du4h\"})')", "event": "contextual_response_rephraser.init.send_test_llm_api_request_failed", "level": "error"}Live runs of examples/mantle-voice-agent (RasaHQ/rasa-community-resources
at 01d6a6e) on Rasa Pro 3.20.0 with litellm 1.100.1, as the project’s
uv.lock pins them. Runs recorded 29 September 2026 against the 3.20.0
release. Real-provider runs used the key in the project’s .env; the rest
used placeholder keys and a local stand-in. Receipts are in
editorial/receipts/endpoints-api-key-env-ignored/ in the site repository.
The line is ours. Companion commit 432d416, titled “Use api_key_env: NAME;
api_key: ${VAR} never expanded”, made this change to every endpoints.yml
credential line:
$ git -C rasa-community-resources show 432d416 -- 'examples/*/endpoints.yml' 'tutorials/*/endpoints.yml' 'community/*/endpoints.yml' | grep '^[-+] .*api_key' | sort | uniq -c
1 - api_key: ${GEMINI_API_KEY}
28 - api_key: ${OPENAI_API_KEY}
1 - # api_key: ${OPENAI_API_KEY}
1 + api_key_env: GEMINI_API_KEY
28 + api_key_env: OPENAI_API_KEY
1 + # api_key_env: OPENAI_API_KEY
On 3.20.0 the commit’s premise did not hold for endpoints.yml: the client
expands ${NAME} when it calls the provider, which is why the edit above
started the server and why a stand-in received the expanded placeholder key.
At 01d6a6e, 29 active api_key_env lines remain in 15 endpoints.yml
files. We ran one of those files against a real provider.
The request carried api_key_env as a body field
OpenAI’s error names a parameter; a recorder shows the request that carried
it. The stand-in is a local HTTP server that OPENAI_API_BASE points at. It
writes one JSON line per request (path, bearer value, top-level keys of the
body, model) and answers with a fake completion. For two runs we gave the
rephraser’s group its own variable, set OPENAI_API_KEY and
REPHRASER_OPENAI_KEY to two placeholders, trained, started the server and
sent “Hi there”. One run kept the api_key_env spelling; the other used
api_key: ${...}.
Avoid: api_key_env: REPHRASER_OPENAI_KEY
/v1/embeddings sk-standin-default
/v1/chat/completions sk-standin-default
body keys: api_key_env, max_completion_tokens,
messages, model, reasoning_effort, temperature
api_key_env = REPHRASER_OPENAI_KEY
/v1/chat/completions sk-standin-default
body keys: messages, model,
reasoning_effort, temperatureThe test call goes out on the default key and carries the variable’s name as a body field.
Prefer: api_key: ${REPHRASER_OPENAI_KEY}
/v1/embeddings sk-standin-default
/v1/chat/completions sk-standin-rephraser
body keys: max_completion_tokens,
messages, model, reasoning_effort, temperature
/v1/chat/completions sk-standin-default
body keys: messages, model,
reasoning_effort, temperatureThe test call carries the named key and no stray field. The last request is unchanged.
Each pane shows the path, bearer and body keys from the stand-in’s raw lines,
in the order received; the raw lines are in the receipts. The first request is
an embeddings call made during training by Mantle’s default references
embedder, not by an endpoints.yml group. The second is the rephraser’s
startup test call, the only request with max_completion_tokens, matching the
256 in the health check’s logged config. The third is the reply to “Hi there”.
The stand-in accepts anything, so in the left pane the field simply arrives.
The left pane also shows the test call going out on sk-standin-default, the
value of OPENAI_API_KEY, while the YAML named REPHRASER_OPENAI_KEY; the
litellm source below explains why. The stand-in is 25 lines of Python, and
these two records are its output.
Show the stand-in and the commands for these two runs
Start the recorder, then train and start Rasa against it with two different
placeholder keys (set the same variables for rasa train). The ports are
yours to choose.
python openai_standin_chat.py <port> requests.jsonl &
LLM_API_HEALTH_CHECK=true OPENAI_API_KEY=sk-standin-default \
REPHRASER_OPENAI_KEY=sk-standin-rephraser \
OPENAI_API_BASE=http://127.0.0.1:<port>/v1 \
uv run --frozen rasa run --enable-api --port <rasa-port>Give the rephraser’s group api_key_env: REPHRASER_OPENAI_KEY for the left
pane or api_key: ${REPHRASER_OPENAI_KEY} for the right. The script logs a bearer value
only when it starts with sk-standin, so a real key sent by mistake is never
written down.
"""Local stand-in for api.openai.com used only with placeholder keys. Records, per request:
path, the bearer value (placeholders set by the test, never a real key), top-level body keys,
model. Returns a fake chat completion or fake embeddings. Nothing is forwarded."""
import http.server, json, sys, datetime, hashlib
PORT, LOG = int(sys.argv[1]), sys.argv[2]
class H(http.server.BaseHTTPRequestHandler):
def do_POST(self):
body = json.loads(self.rfile.read(int(self.headers.get('content-length', 0))) or b'{}')
auth = self.headers.get('authorization', '')
bearer = auth.split(' ', 1)[1] if auth.startswith('Bearer ') else ('(none)' if not auth else '(non-bearer)')
if not bearer.startswith('sk-standin') and bearer not in ('(none)', '(non-bearer)'): bearer = '(unexpected value, not logged)'
with open(LOG, 'a') as f:
f.write(json.dumps({'at': datetime.datetime.now().isoformat(timespec='seconds'), 'path': self.path, 'bearer': bearer,
'body_keys': sorted(body), 'model': body.get('model'), 'api_key_env_in_body': body.get('api_key_env')}) + '\n')
if self.path.rstrip('/').endswith('/embeddings'):
inputs = body.get('input', []); inputs = [inputs] if isinstance(inputs, str) else inputs
out = {'object': 'list', 'data': [{'object': 'embedding', 'index': i, 'embedding': [0.01] * 1536} for i, _ in enumerate(inputs)], 'model': body.get('model'), 'usage': {'prompt_tokens': 0, 'total_tokens': 0}}
else:
out = {'id': 'chatcmpl-standin', 'object': 'chat.completion', 'created': 0, 'model': body.get('model'),
'choices': [{'index': 0, 'message': {'role': 'assistant', 'content': 'stand-in reply'}, 'finish_reason': 'stop'}],
'usage': {'prompt_tokens': 1, 'completion_tokens': 1, 'total_tokens': 2}}
data = json.dumps(out).encode()
self.send_response(200); self.send_header('content-type', 'application/json'); self.send_header('content-length', str(len(data))); self.end_headers(); self.wfile.write(data)
def log_message(self, *a): pass
http.server.ThreadingHTTPServer(('127.0.0.1', PORT), H).serve_forever()Rasa’s load-time check has a credential list, and api_key_env is not on it
Rasa does check these groups before any request. validate_model_group_configuration_setup
in rasa/engine/validation.py runs over every endpoints.yml model group
when rasa train builds a model (rasa/model_training.py, line 347) and
when Rasa loads one (rasa/engine/loader.py, line 46). Two of its checks
are about key names. The first allows a ${...} value only on these keys
(lines 1351 to 1360):
allowed_env_vars = {
DEPLOYMENT_CONFIG_KEY,
API_BASE_CONFIG_KEY,
API_KEY,
API_VERSION_CONFIG_KEY,
AWS_REGION_NAME_CONFIG_KEY,
AWS_ACCESS_KEY_ID_CONFIG_KEY,
AWS_SECRET_ACCESS_KEY_CONFIG_KEY,
AWS_SESSION_TOKEN_CONFIG_KEY,
}
The second requires every key on this list, from rasa/shared/constants.py
(lines 361 to 372), to be written as ${...}, and raises a ValidationError
otherwise:
SENSITIVE_DATA = [
API_KEY,
AWS_ACCESS_KEY_ID_CONFIG_KEY,
AWS_SECRET_ACCESS_KEY_CONFIG_KEY,
AWS_SESSION_TOKEN_CONFIG_KEY,
LANGFUSE_CONFIG_PUBLIC_KEY,
LANGFUSE_CONFIG_PRIVATE_KEY,
CLIENT_ID_CONFIG_KEY,
CLIENT_SECRET_CONFIG_KEY,
TOKEN_URL_CONFIG_KEY,
A2A_JWT_SECRET_CONFIG_KEY,
]
api_key_env: OPENAI_API_KEY is neither a ${...} value nor a listed key, so
it passes both, and rasa train exited 0 in each of the six runs where we
recorded its exit code. Neither file
mentions api_key_env. The list is where Rasa has already decided that key
names matter for credentials; api_key_env is a credential name Rasa itself
defines, in its Mantle client, and the list does not know it. In effect the
two checks are a name-based schema for secrets, with one name missing.
Past the check, the two clients we read keep unknown keys
nlg:
type: rephrase
llm:
model_group: openai-llm
model_groups:
- id: openai-llm
models:
- provider: openai
model: gpt-5.2
api_key_env: OPENAI_API_KEY
temperature: 0.0- The rephraser’s group.
ContextualResponseRephraserresolves it when it is created and runs the health check on it. provider: openaiselectsOpenAILLMClient. Any provider Rasa does not map,anthropicincluded, getsDefaultLiteLLMClient(rasa/shared/providers/mappings.py).- The load-time check passes it, and nothing after it reads it. It stays in
the config, is printed in the log’s
config, and is the field OpenAI named.
The client config takes out the keys it knows and keeps the rest. Its comment
says “The rest of parameters (e.g. model parameters) are considered as extra
parameters (this also includes timeout)”. This is OpenAIClientConfig, in
rasa/shared/providers/_configs/openai_client_config.py (lines 141 to 154):
this = OpenAIClientConfig(
# Required parameters
model=config.pop(MODEL_CONFIG_KEY),
# Pop the 'provider' key. Currently, it's *optional* because of
# backward compatibility with older versions.
provider=config.pop(PROVIDER_CONFIG_KEY, OPENAI_PROVIDER),
# Optional parameters
api_base=config.pop(API_BASE_CONFIG_KEY, None),
api_version=config.pop(API_VERSION_CONFIG_KEY, None),
api_type=config.pop(API_TYPE_CONFIG_KEY, OPENAI_API_TYPE),
# The rest of parameters (e.g. model parameters) are considered
# as extra parameters (this also includes timeout).
extra_parameters=config,
)
api_key_env is in “the rest”. The default litellm client’s config does the
same after taking out only model and provider. Both clients then spread
the extras into the litellm call, in
rasa/shared/providers/llm/_base_litellm_client.py (lines 84 to 96):
@property
def _completion_fn_args(self) -> dict:
return {
# Since all providers covered by LiteLLM use the OpenAI format, but
# not all support every OpenAI parameter, raise an exception if
# provider/model uses unsupported parameter
"drop_params": False,
# All other parameters set through config, can override drop_params
**self._litellm_extra_parameters,
# Model name is constructed in the LiteLLM format from the provided config
# Non-overridable to ensure consistency
"model": self._litellm_model_name,
}
The **self._litellm_extra_parameters spread carries api_key_env forward,
beside temperature and reasoning_effort. drop_params: False is a
separate setting; its comment says the aim is to “raise an exception if
provider/model uses unsupported parameter”. In our runs the exception came
from the provider, after the request was sent. The same file passes the
arguments through resolve_environment_variables, which calls
os.path.expandvars (rasa/shared/utils/io.py). That is where
api_key: ${REPHRASER_OPENAI_KEY} became sk-standin-rephraser.
litellm 1.100.1 does the last step (litellm/utils.py, lines 4719 to 4745).
For openai, a keyword it does not recognise goes into extra_body; for
other providers it goes into the request parameters. With no api_key,
litellm’s OpenAI path falls back to the OPENAI_API_KEY environment variable
(litellm/main.py), which is where the left pane’s sk-standin-default came
from. Both litellm points are source reading; the wire behaviour is in the
records. This covers the OpenAI and default litellm clients, the two paths we
read.
Only integrations.yml turns api_key_env into a key
A grep of the 3.20.0 wheel finds one reader of the YAML key, in
rasa/mantle/llm/client.py. The Azure clients’ hits are a private attribute
whose YAML key is api_key, and a docstring in rasa/cli/project_env.py
names integrations.yml:
$ grep -rn api_key_env rasa/ --include='*.py'
rasa/mantle/llm/client.py:43:_API_KEY_ENV_CONFIG_KEY = "api_key_env"
rasa/mantle/llm/client.py:136:def _resolve_api_key_env(config: Dict[str, Any]) -> Dict[str, Any]:
rasa/mantle/llm/client.py:150: "mantle.llm.api_key_env.not_set",
rasa/mantle/llm/client.py:257: Resolves ``api_key_env`` entries — where the value is an environment
rasa/mantle/llm/client.py:260: ``api_key_env`` key (an engine convention unknown to LiteLLM).
rasa/mantle/llm/client.py:262: llm_config = _resolve_api_key_env(dict(bootstrap.llm))
rasa/shared/providers/llm/azure_openai_llm_client.py:137: self._api_key_env_var = (
rasa/shared/providers/llm/azure_openai_llm_client.py:138: self._resolve_api_key_env_var() if not self._oauth else None
rasa/shared/providers/llm/azure_openai_llm_client.py:197: def _resolve_api_key_env_var(self) -> str:
rasa/shared/providers/llm/azure_openai_llm_client.py:344: elif self._api_key_env_var:
rasa/shared/providers/llm/azure_openai_llm_client.py:345: auth_parameter = {LITE_LLM_API_KEY_FIELD: self._api_key_env_var}
rasa/shared/providers/embedding/azure_openai_embedding_client.py:98: self._api_key_env_var = (
rasa/shared/providers/embedding/azure_openai_embedding_client.py:99: self._resolve_api_key_env_var() if not self._oauth else None
rasa/shared/providers/embedding/azure_openai_embedding_client.py:104: def _resolve_api_key_env_var(self) -> str:
rasa/shared/providers/embedding/azure_openai_embedding_client.py:244: elif self._api_key_env_var:
rasa/shared/providers/embedding/azure_openai_embedding_client.py:245: auth_parameter = {LITE_LLM_API_KEY_FIELD: self._api_key_env_var}
rasa/cli/project_env.py:64: ``api_key_env`` in ``integrations.yml`` available even when the process has
_resolve_api_key_env walks a model group, pops api_key_env and writes the
named variable’s value into api_key. Its one caller explains why, in words
that fit this whole page (client.py, lines 254 to 262):
def from_bootstrap(cls, bootstrap: ModelBootstrap) -> "LLMClient":
"""Build client from the project's integrations.yml LLM config.
Resolves ``api_key_env`` entries — where the value is an environment
variable *name* — into the corresponding ``api_key`` value so the
config that reaches ``llm_factory`` never contains the raw
``api_key_env`` key (an engine convention unknown to LiteLLM).
"""
llm_config = _resolve_api_key_env(dict(bootstrap.llm))
bootstrap.llm is the orchestrator’s group, resolved from the project’s
integrations.yml (rasa/mantle/model_archive/bootstrap.py). In the voice
agent that file carries the same line under a different group (lines 6 to
15):
llm:
model_group: orchestrator
model_groups:
- id: orchestrator
models:
- provider: openai
model: gpt-5.2
api_key_env: OPENAI_API_KEY
temperature: 0.0
So one project holds the same spelling twice: resolved in integrations.yml,
forwarded from endpoints.yml. The integrations.yml half rests on the
source above and on the shape of the reply request, which carried no
api_key_env in either stand-in record. No run gave the orchestrator’s group
a second variable name to tell the two apart.
The startup check tested a client the recorded replies never used
The forwarded key reaches a provider only when something sends the
rephraser’s group. perform_llm_health_check, in
rasa/shared/utils/health_check/health_check.py (lines 77 to 149), sends a
test call when the variable reads true (compared case-insensitively). With
it off, the same function can still probe Azure deployment-only groups at
inference; otherwise it logs this warning, quoted from our real-OpenAI run
with the check off:
The LLM_API_HEALTH_CHECK environment variable is set to false, which will disable LLM health check. It is recommended to set this variable to true in production environments.
A failed test call is re-raised as a HealthCheckError (lines 268 to 295);
on the first screen rasa.core.run logged it and the process exited 1. For this project’s openai group we
checked the switch at the stand-in. rasa run alone with the check on sent
one request, the test call, api_key_env included; with it off the stand-in
received nothing. With the check on, rasa train exited 0 and the stand-in
received only the embeddings request. In this project the check runs when the
server starts, not when it trains.
With the check off, does anything else use the group? For one Mantle
component this is on record: the
guide on Mantle’s references embedder
showed the embedder is chosen from agent.yml and integrations.yml, and
that repointing the endpoints.yml embeddings group changed no request. The
rephrase needs its own evidence; the source and three recorded turns give it.
The orchestrator builds its client with
LLMClient.from_bootstrap(self._static_data), and _rephrase_and_send
passes that client to self._call_llm
(rasa/mantle/orchestration/orchestrator.py, lines 268 and 936). The
endpoints.yml rephraser is created by the Agent’s
NaturalLanguageGenerator.create (rasa/core/agent.py), and a grep of
rasa/mantle finds no reference to ContextualResponseRephraser.
We recorded four REST turns, one in each of rows 2, 3, 5 and 6 of the table
below. Row 2 cannot tell the clients apart, because both would carry the same
key. The other three can. In row 3 the check was off on real OpenAI, the group
still carried api_key_env, and the greeting came back reworded, logged as
mantle.turn.completed with "source": "rephrase_text". OpenAI refused that
field in row 1, and a failed rephrase falls back to the verbatim text
(rasa/mantle/orchestration/responses.py, lines 93 to 98). So by inference
the rephrase request did not carry the field. The first sentence of each, from
the receipt:
responses.yml utter_greet: Hello. I'm Atlas, your Horizon Travel voice assistant.
reply, check off: Hi. I’m Atlas, Horizon Travel’s voice assistant.
At the stand-in, row 5’s rephrase request carried no api_key_env field
although the group did. In row 6, with the group spelled
api_key: ${REPHRASER_OPENAI_KEY}, the test call carried
sk-standin-rephraser and the rephrase request carried sk-standin-default,
the orchestrator’s key. These are REST turns on one project; we did not run
other channels or a project without Mantle.
The provider is the first thing that objects
OpenAI’s reply is on the first screen. For the Anthropic run we switched the
rephraser’s group to Anthropic and left api_key_env in it:
- id: openai-llm
models:
- - provider: openai
- model: gpt-5.2
- api_key_env: OPENAI_API_KEY
+ - provider: anthropic
+ model: claude-haiku-4-5
+ api_key_env: ANTHROPIC_API_KEY
temperature: 0.0
The orchestrator and embeddings went to the stand-in; only the rephraser’s
test call went to Anthropic, and /status never answered. The error, as an
excerpt re-wrapped to fit:
ERROR Test call to the LLM API failed for component -
ContextualResponseRephraser.
[…] AnthropicException - […]"type":"invalid_request_error",
"message":"api_key_env: Extra inputs are not permitted"[…],
"request_id":"req_011CfXaYEuMPSm6AqRe3Du4h"
Cut: the timestamp, the logger name, the config object, the event and
level fields, the ProviderClientAPIException and litellm.BadRequestError
wrapper, and the outer JSON of Anthropic’s reply; […] marks the cuts inside
a line, and the backslash escapes on the quotes are decoded. The full line is in the collapsed block near the top.
Two providers, two client classes in Rasa, two litellm branches (extra_body
and request parameters), and the same refusal. Neither reply names the
credential problem: both describe an unexpected parameter, which is accurate
and points away from the key.
Strictness is what made the key visible. The stand-in is the lenient case: it accepted the field and the run carried on. A provider that ignored unknown fields would do the same. We ran no such provider; the stand-in shows the shape of that case, not any provider’s behaviour.
Seven runs of this project in six rows; row 1 holds two. They are every run on this page except the two switch runs, which started the server and sent no message, the training-only run, and the quickstart run, which is a different project. Only the rephraser’s group and the environment change between rows. In row 3 the check was off, so the group was sent nowhere.
| # | Rephraser group credential line | Provider | LLM_API_HEALTH_CHECK | What happened | Receipt |
|---|---|---|---|---|---|
| 1 | api_key_env: OPENAI_API_KEY (as committed) | real OpenAI | true | Test call rejected: Unknown parameter: 'api_key_env'. Exit 1, /status never answered | r320-openai-rephraser-committed.txt, and run 2 of r320-openai-committed-health-check-off.txt |
| 2 | api_key: ${OPENAI_API_KEY} | real OpenAI | true | Test call sent, no ERROR line. /status answered; “Hi there” got the Atlas greeting | r320-openai-rephraser-placeholder.txt |
| 3 | api_key_env: OPENAI_API_KEY (as committed) | real OpenAI | false | Started, 0 ERROR lines. The greeting came back reworded | r320-openai-committed-health-check-off.txt |
| 4 | provider: anthropic, claude-haiku-4-5, api_key_env: ANTHROPIC_API_KEY | real Anthropic, for the test call only | true | Test call rejected: api_key_env: Extra inputs are not permitted. /status never answered | r320-rephraser-anthropic.txt |
| 5 | api_key_env: REPHRASER_OPENAI_KEY | stand-in | true | Test call: bearer sk-standin-default, body field api_key_env. Reply request: bearer sk-standin-default | r320-rephraser-own-var.txt and its raw .jsonl.txt |
| 6 | api_key: ${REPHRASER_OPENAI_KEY} | stand-in | true | Test call: bearer sk-standin-rephraser, no api_key_env. Reply request: bearer sk-standin-default | r320-rephraser-placeholder.txt and its raw .jsonl.txt |
Rows 1 and 4 are one key refused by two providers in their own words. Rows 1 and 3 differ only in the switch. Rows 5 and 6 are the counter-test for the spelling. In rows 2, 4, 5 and 6, and in the first of row 1’s two runs, the test call was logged, so by the source the variable was true; where it came from was not recorded for those runs. The second run of row 1 set it on the command line, as the first screen shows.
Make the provider answer before you deploy
Find the lines. In the companion at 01d6a6e this matched 30 lines in 15
files, one of them inside a comment:
git grep -n api_key_env -- '*endpoints*.yml'
Rewrite each active match as api_key: ${NAME} with the same variable name;
in endpoints.yml we ran that form on OpenAI only. Use the same form in
integrations.yml. There 3.20.0 reads api_key_env too, and the
placeholder is expanded at call time by the same base client. A live 3.20.0
run of this site’s quickstart project with api_key: ${ANTHROPIC_API_KEY} in
integrations.yml trained and answered four conversations. One spelling in
both files means one rule to check. Then start once with the check on and the
real key, and require /status to answer:
LLM_API_HEALTH_CHECK=true uv run --frozen rasa run --enable-api
With the committed line it exits 1; with the edit it starts. In this project that is one real request per start. A strict provider tells you the request was malformed; it does not tell you which key it used. For that, run the stand-in above.
Added 29 September 2026: the prerelease refuses the key
Rasa Pro 3.21.0.dev3, a prerelease uploaded to PyPI on 28 September at 15:57
UTC, refuses api_key_env in an endpoints.yml model group at load time.
validate_model_group_configuration_setup now calls
validate_model_group_credentials, which raises api_key_env_not_supported
with a message that ends “which is no longer supported. Replace it with
‘api_key: ${ENV_VAR_NAME}’.” That is a dedicated rule with the message argued
for above. We read it in the prerelease source and did not run it; everything
above this note is about 3.20.0. If you are on 3.20.0, write api_key: ${NAME}
now: it works there in both files, and the prerelease refuses the other
spelling in endpoints.yml.