# A correct device correction was rejected by our parser

Source: https://rasa.community/library/casebook/retail-guided-selling/
Author: Rod Rivera
Published: 2026-10-09T09:00:00.000Z

The shopper corrected “Lumen 7 Pro” to “the regular Lumen 7”. The model sent the right words to the tool. Our original parser rejected them, so the assistant asked for a device the shopper had already named.

This happened in a recorded Willow Shop tutorial run on 30 September 2026. The catalogue and shopper are fictional. The correction turn contained:

```text
Shopper: Wait, it's actually the regular Lumen 7, not the Pro.
record_requirements argument: {"device_model": "regular Lumen 7"}
Tool result: blocked / requirements_ambiguous
Assistant: Just to confirm, which exact model do you have?
```

The [recorded result](https://github.com/RasaHQ/rasa-community-resources/blob/4aa0c4419dc193fef7a969c12d59edcf720f2606/examples/mantle-text-retail-guided-selling-gpt/case-build/results/2026-09-30-gpt-5.5-reasoning-low/results.json) preserves that exchange in `correction-pro-to-standard-case`. No new model run was needed to investigate it.

## Find the failure between the argument and the result

The tool asked for the model exactly as the shopper said it. Its original resolver then required those words to equal a complete catalogue name or alias. “Regular” made the whole string unequal to “Lumen 7”. The model followed the tool contract, but the parser could not accept the resulting argument.

This was a correction failure, not an unsafe recommendation. Two assertions in the saved result make that distinction inspectable:

| Check after the correction            | Original run              | Resolver-fix rerun        |
| ------------------------------------- | ------------------------- | ------------------------- |
| Record the standard phone as `DEV-L7` | Failed: no matching call  | Passed: one matching call |
| Do not recommend the earlier Pro case | Passed: no forbidden call | Passed: no forbidden call |

The [rerun result](https://github.com/RasaHQ/rasa-community-resources/blob/4aa0c4419dc193fef7a969c12d59edcf720f2606/examples/mantle-text-retail-guided-selling-gpt/case-build/results/2026-09-30-resolver-fix-rerun/results.json) contains the same `device_model: "regular Lumen 7"` argument. It was recorded successfully. Comparing the tool arguments avoids mistaking a parser change for a better model answer. These are two saved tutorial traces, not an estimate of correction accuracy.

## Replay the parser change without paying for another conversation

We extracted the original resolver from Git and ran it alongside the fixed resolver against the same fixture. Only the original function is loaded from the old revision. The catalogue and helper functions come from the pinned fixed revision. This isolates that function's change; it does not replay the entire old application.

| Argument                  | Original resolver | Fixed resolver |
| ------------------------- | ----------------- | -------------- |
| `regular Lumen 7`         | Unresolved        | `DEV-L7`       |
| `Lumen 7 Pro`             | `DEV-L7P`         | `DEV-L7P`      |
| `Lumen 7 Lite`            | `DEV-L7L`         | `DEV-L7L`      |
| `Lumen`                   | Unresolved        | Unresolved     |
| `Lumen 7 and Lumen 7 Pro` | Unresolved        | Unresolved     |
| `Lumen 7 Charging Dock`   | Unresolved        | Unresolved     |

The last five rows are controls. Accepting an extra word must not collapse the Pro and Lite models into the standard phone, choose between two named devices, or turn a product name into the shopper's device. All six assertions passed in the offline replay on 7 October 2026.

[The original function](https://github.com/RasaHQ/rasa-community-resources/blob/040fce090b5be1143f01c25dd6f98df610b6719a/examples/mantle-text-retail-guided-selling-gpt/lib/willowshop.py#L154-L169) compares the entire normalised string. The [fixed function and helpers](https://github.com/RasaHQ/rasa-community-resources/blob/4aa0c4419dc193fef7a969c12d59edcf720f2606/examples/mantle-text-retail-guided-selling-gpt/lib/willowshop.py#L106-L187) reuse device recognition from shopper messages. They match longer names first, mask full catalogue product names, and accept an argument only when it names one distinct device. This is bounded catalogue matching, not fuzzy guessing.

::::solution{title="Replay the original and fixed resolvers"}

From the pinned companion tests directory shown below, save this as `resolver-replay.py`, then run `python3 resolver-replay.py`. Git reads a historical function. Nothing changes on disk and no model or provider is called.

```python
"""Replay the original resolver and its fix. Stdlib only, no model calls.
Run from examples/mantle-text-retail-guided-selling-gpt/tests at the pinned revision.
"""
import ast
import json
import subprocess
import test_guard as t

old_revision = "040fce090b5be1143f01c25dd6f98df610b6719a"
source = subprocess.check_output([
    "git", "show", old_revision + ":examples/mantle-text-retail-guided-selling-gpt/lib/willowshop.py"
], text=True)
node = next(n for n in ast.parse(source).body if isinstance(n, ast.FunctionDef) and n.name == "resolve_device")
namespace = dict(vars(t.ws))
exec(compile(ast.Module(body=[node], type_ignores=[]), "original-resolve-device.py", "exec"), namespace)
original = namespace["resolve_device"]
inputs = (
    ("regular Lumen 7", None, "DEV-L7"),
    ("Lumen 7 Pro", "DEV-L7P", "DEV-L7P"),
    ("Lumen 7 Lite", "DEV-L7L", "DEV-L7L"),
    ("Lumen", None, None),
    ("Lumen 7 and Lumen 7 Pro", None, None),
    ("Lumen 7 Charging Dock", None, None),
)
for argument, expected_old, expected_fixed in inputs:
    old_id, _ = original(argument)
    fixed_id, _ = t.ws.resolve_device(argument)
    assert (old_id, fixed_id) == (expected_old, expected_fixed)
    print(json.dumps({"argument": argument, "original": old_id, "fixed": fixed_id}, sort_keys=True))
```

The first output line reproduces the rejected argument:

```text
{"argument": "regular Lumen 7", "fixed": "DEV-L7", "original": null}
```

::::

::::checkpoint{id="retail-guided-selling-decision" question="The tool receives the correct device name but rejects it. What should you inspect first?" options="Replace the model|Compare the received argument with the resolver|Assume the shopper changed their mind" answer="1"}

Read the received argument and the resolver’s match rule. The saved argument names the requested device. Its rejection came from the exact-string rule.

::::

## Accept the correction without carrying over the old fit

A parser can recognise a device without proving that the shopper has it. `record_requirements` also checks the latest device named in the shopper's messages. The recommendation tool checks that recorded device against the latest message again. An earlier Pro requirement cannot justify a recommendation after the shopper names the standard phone.

The saved rerun exposes the next two boundaries. After recording `DEV-L7`, it attempts a standard-phone case with an old stock observation and a folio case without sourced fit evidence. The results are:

```text
record_requirements("regular Lumen 7"): recorded
recommend_product("WS-CS-L7"): blocked / availability_stale
recommend_product("WS-CS-L7-FOLIO"): blocked / compatibility_unverified
```

The fix did not manufacture a replacement recommendation. It let the conversation reach the stock and fit checks with the right device. A test that required any recommendation after a correction would be wrong for this fixture. Pair the positive assertion “record the corrected device” with the negative assertion “discard the earlier device's result”.

::::callout{type="info" title="Keep the parser and service assertions separate"}

The model argument can be correct while the tool refuses it. A correctly recorded correction can still leave every available product unsuitable or unverified. Inspect each boundary before changing a prompt or counting the conversation as a successful sale.

::::

## Label evidence before you score recommendations

For an evaluation set, “not recommended” is too broad. It hides the difference between a known mismatch and absent evidence. Use a label for each verdict, then attach the attribute that justifies it.

| Evaluation label | Evidence needed                                       | Customer-facing meaning                |
| ---------------- | ----------------------------------------------------- | -------------------------------------- |
| Compatible       | Required attributes match with sources                | This documented fit can be recommended |
| Not compatible   | At least one sourced attribute contradicts the device | Name the mismatch                      |
| Unverified       | A required attribute has no source                    | Ask the catalogue owner                |

A product name is not one of those sources. Search returns names, prices and identifiers without a fit verdict. That keeps a relevant search hit from silently becoming a compatibility claim.

::::callout{type="info" title="Do not score uncertainty as a mismatch"}

A missing measurement does not prove that a product will fail. Keep unverified separate in your labels and support queue. Keep the missing field visible so the catalogue owner knows what evidence is needed.

::::

## Inspect the branches behind the labels

The source returns a known mismatch before it evaluates missing evidence. That order makes a concrete contradiction visible without filling in unknown fields.

These excerpts come from the recommendation function. The sourced comparison produces matched, mismatched and unknown attributes; the two negative branches preserve different meanings.

[Source excerpt](https://github.com/RasaHQ/rasa-community-resources/blob/4aa0c4419dc193fef7a969c12d59edcf720f2606/examples/mantle-text-retail-guided-selling-gpt/lib/willowshop.py#L353-L384):

```python
base.update(device_id=device_id, device_name=device_name)
if result["mismatched"]:
    # A sourced attribute contradicts the device: a known mismatch.
    return dict(
        base,
        status="not_compatible",
        recommended=False,
        mismatched_attributes=result["mismatched"],
        note=NAME_IS_NOT_EVIDENCE,
        next_step=(
            "Tell the shopper it does not fit, naming the attribute and its "
            "source. Offer to look for a product that matches."
        ),
    )
facts["compatibility_source_present"] = not result["unknown"]
as_of = datetime.fromisoformat(data["as_of"])
observed = product["stock"].get("observed_at")
window = timedelta(hours=data["stock_freshness_hours"])
facts["availability_current"] = bool(observed) and as_of - datetime.fromisoformat(observed) <= window
reason = evaluate(facts)
if reason == "compatibility_unverified":
    return _blocked(
        reason,
        **base,
        facts=facts,
        matched_attributes=result["matched"],
        unknown_attributes=[
            {"attribute": r["attribute"], "required": r["required"], "status": "unknown"}
            for r in result["unknown"]
        ],
        note=NAME_IS_NOT_EVIDENCE,
    )
```

Only a sourced contradiction supports a mismatch. Missing evidence cannot establish either fit or failure.

## Check the latest device again at recommendation time

This branch rejects an earlier recorded device when the shopper has named another one. It is separate from the parser fix.

[Source excerpt](https://github.com/RasaHQ/rasa-community-resources/blob/4aa0c4419dc193fef7a969c12d59edcf720f2606/examples/mantle-text-retail-guided-selling-gpt/lib/willowshop.py#L333-L350):

```python
product = data["products"].get(number)
if product is None:
    return {"status": "unknown_product", "product_id": number, "next_step": "Use search_catalogue for valid ids."}
declared = latest_declared_devices(user_texts, data)
device_id = (requirements or {}).get("device_id")
requirements_ok = device_id in data["devices"] and device_id in declared
facts: dict[str, Any] = {"requirements_confirmed": requirements_ok}
base = {"product_id": number, "product_name": product["name"]}
if not requirements_ok:
    detail = dict(base, facts=facts)
    if device_id and declared and device_id not in declared:
        detail["note"] = (
            "The shopper has named a different device since requirements were "
            "recorded. The earlier result is discarded; record the new device."
        )
    return _blocked("requirements_ambiguous", **detail)

uses_case = bool(requirements.get("uses_case"))
```

::::diagram{title="A search hit needs evidence before it becomes a recommendation"}

```dot
search [label="Search: name, price and ID"]
req [label="Shopper names current device"]
attrs [label="Compare sourced attributes"]
yes [label="Evidence matches: recommend"]
no [label="Evidence contradicts: does not fit", class="blocked"]
unknown [label="Evidence absent: ask catalogue owner"]
search -> attrs
req -> attrs
attrs -> yes
attrs -> no
attrs -> unknown
```

::::

## Keep availability out of the fit label

A compatible item can be out of stock or have an old stock observation. The fixture checks freshness separately. A successful recommendation carries supporting attributes and sources, stock evidence and unresolved fit questions. For a dock, whether it charges with a case on can remain unanswered.

The specialist route preserves a question and returns no compatibility answer. It is not a simulated expert. The catalogue owner needs to supply the missing measurement or connect that route to a real specialist. This creates data and support work that a name-based guess avoids, but the resulting recommendation can be audited.

## Build a panel that includes mismatches, missing evidence and corrections

Run these checks from the pinned companion project. They use Python's standard library and make no model calls. The outputs below were recorded on 7 October 2026; elapsed times can differ.

```bash
git clone https://github.com/RasaHQ/rasa-community-resources.git
cd rasa-community-resources
git checkout 4aa0c4419dc193fef7a969c12d59edcf720f2606
cd examples/mantle-text-retail-guided-selling-gpt/tests
```

::::run{cmd="python3 -m unittest test_guard.RecommendationTests test_guard.RequirementTests -q"}
::::

```text
----------------------------------------------------------------------
Ran 19 tests in 0.004s

OK
```

The [companion tests](https://github.com/RasaHQ/rasa-community-resources/blob/4aa0c4419dc193fef7a969c12d59edcf720f2606/examples/mantle-text-retail-guided-selling-gpt/tests/test_guard.py) contain the assertions for these cases. To print the additional fixture states yourself, run this script from the same directory:

::::solution{title="Show the script that printed the fixture states"}

Paste it into a file such as `trace.py`, then run `python3 trace.py`. It prints selected fields from the actual tool results; it does not invent replies.

```python
"""Run from the pinned companion project tests directory. Stdlib only."""
import json
import test_guard as t

case = t.RecommendationTests()
for product, device in (("WS-DK-L7L", "Lumen 7"), ("WS-DK-L7", "Lumen 7 Pro"), ("WS-CM-GRIP", "Lumen 7 Pro")):
    result = case.rec(product, device)
    print(json.dumps({"product": product, "device": device, "status": result["status"], "reason": result.get("reason")}, sort_keys=True))
```

::::

These are fixture decisions, not customer return rates or an inventory reservation. Give the catalogue owner the three labelled pairs and ask for a source for every required attribute. For evaluation design, [build a set around observable outcomes](/library/guides/build-an-evaluation-set/). Include the exact rejected argument, the correction acceptance assertion, and the assertion that prevents an earlier fit from surviving. Keep missing evidence separate from a known mismatch.

::::cta{href="/library/guides/build-an-evaluation-set/" label="Build the recommendation evaluation set"}
::::

The [complete fixture implementation](https://github.com/RasaHQ/rasa-community-resources/blob/4aa0c4419dc193fef7a969c12d59edcf720f2606/examples/mantle-text-retail-guided-selling-gpt/lib/willowshop.py#L1-L439) defines the service state and remaining branches cited here.