Skip to content
RasaGet a free licence
AI team casebook

Evaluation / data specialist · Retail & e-commerce (external)

A correct device correction was rejected by our parser

Replay a recorded shopping correction through the original and fixed parser, then test that accepting the new device still discards the old recommendation.

by Rod Rivera

About 10 minutesCase study

Key takeaways (3)
  • Check the tool argument before blaming the model for a failed correction.
  • Accept one complete device name without confusing Pro, Lite and product names.
  • Test that the correction is recorded and the old recommendation is discarded.

The shopper corrected “Lumen 7 Pro” to “the regular Lumen 7”. The model sent the right words to the tool. Our original parser rejected them, so the assistant asked for a device the shopper had already named.

This happened in a recorded Willow Shop tutorial run on 30 September 2026. The catalogue and shopper are fictional. The correction turn contained:

Shopper: Wait, it's actually the regular Lumen 7, not the Pro.
record_requirements argument: {"device_model": "regular Lumen 7"}
Tool result: blocked / requirements_ambiguous
Assistant: Just to confirm, which exact model do you have?

The recorded result preserves that exchange in correction-pro-to-standard-case. No new model run was needed to investigate it.

Find the failure between the argument and the result

The tool asked for the model exactly as the shopper said it. Its original resolver then required those words to equal a complete catalogue name or alias. “Regular” made the whole string unequal to “Lumen 7”. The model followed the tool contract, but the parser could not accept the resulting argument.

This was a correction failure, not an unsafe recommendation. Two assertions in the saved result make that distinction inspectable:

Check after the correctionOriginal runResolver-fix rerun
Record the standard phone as DEV-L7Failed: no matching callPassed: one matching call
Do not recommend the earlier Pro casePassed: no forbidden callPassed: no forbidden call

The rerun result contains the same device_model: "regular Lumen 7" argument. It was recorded successfully. Comparing the tool arguments avoids mistaking a parser change for a better model answer. These are two saved tutorial traces, not an estimate of correction accuracy.

Replay the parser change without paying for another conversation

We extracted the original resolver from Git and ran it alongside the fixed resolver against the same fixture. Only the original function is loaded from the old revision. The catalogue and helper functions come from the pinned fixed revision. This isolates that function’s change; it does not replay the entire old application.

ArgumentOriginal resolverFixed resolver
regular Lumen 7UnresolvedDEV-L7
Lumen 7 ProDEV-L7PDEV-L7P
Lumen 7 LiteDEV-L7LDEV-L7L
LumenUnresolvedUnresolved
Lumen 7 and Lumen 7 ProUnresolvedUnresolved
Lumen 7 Charging DockUnresolvedUnresolved

The last five rows are controls. Accepting an extra word must not collapse the Pro and Lite models into the standard phone, choose between two named devices, or turn a product name into the shopper’s device. All six assertions passed in the offline replay on 7 October 2026.

The original function compares the entire normalised string. The fixed function and helpers reuse device recognition from shopper messages. They match longer names first, mask full catalogue product names, and accept an argument only when it names one distinct device. This is bounded catalogue matching, not fuzzy guessing.

Replay the original and fixed resolvers

From the pinned companion tests directory shown below, save this as resolver-replay.py, then run python3 resolver-replay.py. Git reads a historical function. Nothing changes on disk and no model or provider is called.

"""Replay the original resolver and its fix. Stdlib only, no model calls.
Run from examples/mantle-text-retail-guided-selling-gpt/tests at the pinned revision.
"""
import ast
import json
import subprocess
import test_guard as t

old_revision = "040fce090b5be1143f01c25dd6f98df610b6719a"
source = subprocess.check_output([
    "git", "show", old_revision + ":examples/mantle-text-retail-guided-selling-gpt/lib/willowshop.py"
], text=True)
node = next(n for n in ast.parse(source).body if isinstance(n, ast.FunctionDef) and n.name == "resolve_device")
namespace = dict(vars(t.ws))
exec(compile(ast.Module(body=[node], type_ignores=[]), "original-resolve-device.py", "exec"), namespace)
original = namespace["resolve_device"]
inputs = (
    ("regular Lumen 7", None, "DEV-L7"),
    ("Lumen 7 Pro", "DEV-L7P", "DEV-L7P"),
    ("Lumen 7 Lite", "DEV-L7L", "DEV-L7L"),
    ("Lumen", None, None),
    ("Lumen 7 and Lumen 7 Pro", None, None),
    ("Lumen 7 Charging Dock", None, None),
)
for argument, expected_old, expected_fixed in inputs:
    old_id, _ = original(argument)
    fixed_id, _ = t.ws.resolve_device(argument)
    assert (old_id, fixed_id) == (expected_old, expected_fixed)
    print(json.dumps({"argument": argument, "original": old_id, "fixed": fixed_id}, sort_keys=True))

The first output line reproduces the rejected argument:

{"argument": "regular Lumen 7", "fixed": "DEV-L7", "original": null}

Accept the correction without carrying over the old fit

A parser can recognise a device without proving that the shopper has it. record_requirements also checks the latest device named in the shopper’s messages. The recommendation tool checks that recorded device against the latest message again. An earlier Pro requirement cannot justify a recommendation after the shopper names the standard phone.

The saved rerun exposes the next two boundaries. After recording DEV-L7, it attempts a standard-phone case with an old stock observation and a folio case without sourced fit evidence. The results are:

record_requirements("regular Lumen 7"): recorded
recommend_product("WS-CS-L7"): blocked / availability_stale
recommend_product("WS-CS-L7-FOLIO"): blocked / compatibility_unverified

The fix did not manufacture a replacement recommendation. It let the conversation reach the stock and fit checks with the right device. A test that required any recommendation after a correction would be wrong for this fixture. Pair the positive assertion “record the corrected device” with the negative assertion “discard the earlier device’s result”.

Label evidence before you score recommendations

For an evaluation set, “not recommended” is too broad. It hides the difference between a known mismatch and absent evidence. Use a label for each verdict, then attach the attribute that justifies it.

Evaluation labelEvidence neededCustomer-facing meaning
CompatibleRequired attributes match with sourcesThis documented fit can be recommended
Not compatibleAt least one sourced attribute contradicts the deviceName the mismatch
UnverifiedA required attribute has no sourceAsk the catalogue owner

A product name is not one of those sources. Search returns names, prices and identifiers without a fit verdict. That keeps a relevant search hit from silently becoming a compatibility claim.

Inspect the branches behind the labels

The source returns a known mismatch before it evaluates missing evidence. That order makes a concrete contradiction visible without filling in unknown fields.

These excerpts come from the recommendation function. The sourced comparison produces matched, mismatched and unknown attributes; the two negative branches preserve different meanings.

Source excerpt:

base.update(device_id=device_id, device_name=device_name)
if result["mismatched"]:
    # A sourced attribute contradicts the device: a known mismatch.
    return dict(
        base,
        status="not_compatible",
        recommended=False,
        mismatched_attributes=result["mismatched"],
        note=NAME_IS_NOT_EVIDENCE,
        next_step=(
            "Tell the shopper it does not fit, naming the attribute and its "
            "source. Offer to look for a product that matches."
        ),
    )
facts["compatibility_source_present"] = not result["unknown"]
as_of = datetime.fromisoformat(data["as_of"])
observed = product["stock"].get("observed_at")
window = timedelta(hours=data["stock_freshness_hours"])
facts["availability_current"] = bool(observed) and as_of - datetime.fromisoformat(observed) <= window
reason = evaluate(facts)
if reason == "compatibility_unverified":
    return _blocked(
        reason,
        **base,
        facts=facts,
        matched_attributes=result["matched"],
        unknown_attributes=[
            {"attribute": r["attribute"], "required": r["required"], "status": "unknown"}
            for r in result["unknown"]
        ],
        note=NAME_IS_NOT_EVIDENCE,
    )

Only a sourced contradiction supports a mismatch. Missing evidence cannot establish either fit or failure.

Check the latest device again at recommendation time

This branch rejects an earlier recorded device when the shopper has named another one. It is separate from the parser fix.

Source excerpt:

product = data["products"].get(number)
if product is None:
    return {"status": "unknown_product", "product_id": number, "next_step": "Use search_catalogue for valid ids."}
declared = latest_declared_devices(user_texts, data)
device_id = (requirements or {}).get("device_id")
requirements_ok = device_id in data["devices"] and device_id in declared
facts: dict[str, Any] = {"requirements_confirmed": requirements_ok}
base = {"product_id": number, "product_name": product["name"]}
if not requirements_ok:
    detail = dict(base, facts=facts)
    if device_id and declared and device_id not in declared:
        detail["note"] = (
            "The shopper has named a different device since requirements were "
            "recorded. The earlier result is discarded; record the new device."
        )
    return _blocked("requirements_ambiguous", **detail)

uses_case = bool(requirements.get("uses_case"))
Search: name, price and ID Compare sourced attributes Shopper names current device Evidence matches: recommend Evidence contradicts: does not fit Evidence absent: ask catalogue owner
FigureA search hit needs evidence before it becomes a recommendation

Keep availability out of the fit label

A compatible item can be out of stock or have an old stock observation. The fixture checks freshness separately. A successful recommendation carries supporting attributes and sources, stock evidence and unresolved fit questions. For a dock, whether it charges with a case on can remain unanswered.

The specialist route preserves a question and returns no compatibility answer. It is not a simulated expert. The catalogue owner needs to supply the missing measurement or connect that route to a real specialist. This creates data and support work that a name-based guess avoids, but the resulting recommendation can be audited.

Build a panel that includes mismatches, missing evidence and corrections

Run these checks from the pinned companion project. They use Python’s standard library and make no model calls. The outputs below were recorded on 7 October 2026; elapsed times can differ.

git clone https://github.com/RasaHQ/rasa-community-resources.git
cd rasa-community-resources
git checkout 4aa0c4419dc193fef7a969c12d59edcf720f2606
cd examples/mantle-text-retail-guided-selling-gpt/tests
Any system
python3 -m unittest test_guard.RecommendationTests test_guard.RequirementTests -q
----------------------------------------------------------------------
Ran 19 tests in 0.004s

OK

The companion tests contain the assertions for these cases. To print the additional fixture states yourself, run this script from the same directory:

Show the script that printed the fixture states

Paste it into a file such as trace.py, then run python3 trace.py. It prints selected fields from the actual tool results; it does not invent replies.

"""Run from the pinned companion project tests directory. Stdlib only."""
import json
import test_guard as t

case = t.RecommendationTests()
for product, device in (("WS-DK-L7L", "Lumen 7"), ("WS-DK-L7", "Lumen 7 Pro"), ("WS-CM-GRIP", "Lumen 7 Pro")):
    result = case.rec(product, device)
    print(json.dumps({"product": product, "device": device, "status": result["status"], "reason": result.get("reason")}, sort_keys=True))

These are fixture decisions, not customer return rates or an inventory reservation. Give the catalogue owner the three labelled pairs and ask for a source for every required attribute. For evaluation design, build a set around observable outcomes. Include the exact rejected argument, the correction acceptance assertion, and the assertion that prevents an earlier fit from surviving. Keep missing evidence separate from a known mismatch.

The complete fixture implementation defines the service state and remaining branches cited here.