Skip to content
RasaGet a free licence
AI team casebook

Platform / operations engineer · Travel & hospitality (external)

A retry dropped an unresolved booking from our repair queue

A harmless replay removes open work when a repair queue treats attempt counters as booking state.

by Rod Rivera

About 4 minutesCase study

Key takeaways (3)
  • Zero new errors does not mean earlier work was repaired.
  • Key open work by stable operation references, not attempts.
  • Reconcile the queue with authoritative booking and reversal state.

The retry debits no extra points. Our test dashboard still removes a booking that has no ticket and no completed reversal:

{"attempt": "same_hold_retry", "authoritative_unresolved": 2, "naive_queue_size": 1, "new_mismatch_counter": 0, "queue_covers_open_work": "FAIL"}

We deliberately wrote a defective queue reducer around the Horizon Rewards fixture: add work when an attempt reports a mismatch; remove it when the next attempt reports zero. The first attempt and retry both behave as the fixture specifies. The dashboard makes the mistake.

This is an offline counterexample using invented members and an authored partner failure. It measures neither real airline failures nor model reliability.

Run the reducer that loses the repair

Use the pinned rewards example. Its records already contain one unresolved redemption. The new hold adds another when points are debited but ticketing fails. We seed the queue from the existing records, then apply the defective reducer:

Save this as /tmp/travel-redemption-check.py. From the example’s tests directory, run PYTHONPATH=. python3 -B /tmp/travel-redemption-check.py:

import json
import test_guard as t
svc = t.fresh()
hold = t.held(svc, "RW-OPO-3302")
def open_records():
    return {ref for ref, record in svc.redemptions.items()
            if record["points"]["state"] == "debited"
            and record["booking"]["state"] == "not_ticketed"
            and record["reversal"]["state"] in ("unresolved", "requested")}
queue = open_records()
for attempt in ("first_commit", "same_hold_retry"):
    result = t.redeem(svc, hold["hold_id"])
    reference = svc.holds[hold["hold_id"]]["redemption_reference"]
    # Deliberately defective dashboard reducer: replace state with each attempt's event.
    if result["unmatched_points_booking"]:
        queue.add(reference)
    else:
        queue.discard(reference)
    unresolved = open_records()
    print(json.dumps({"attempt": attempt, "naive_queue_size": len(queue), "authoritative_unresolved": len(unresolved), "queue_covers_open_work": "PASS" if queue == unresolved else "FAIL", "new_mismatch_counter": result["unmatched_points_booking"]}, sort_keys=True))
review = t.hz.rewards_desk_review(svc, t.ME, reference)
print(json.dumps({"attempt": "desk_review", "reversal_state": review["reversal"]["state"], "authoritative_unresolved": len(open_records()), "requested_repair_stays_open": reference in open_records()}, sort_keys=True))

Actual offline output:

{"attempt": "first_commit", "authoritative_unresolved": 2, "naive_queue_size": 2, "new_mismatch_counter": 1, "queue_covers_open_work": "PASS"}
{"attempt": "same_hold_retry", "authoritative_unresolved": 2, "naive_queue_size": 1, "new_mismatch_counter": 0, "queue_covers_open_work": "FAIL"}
{"attempt": "desk_review", "authoritative_unresolved": 2, "requested_repair_stays_open": true, "reversal_state": "requested"}

The first observation covers both outstanding records. The retry removes the new redemption, leaving one queue entry while two authoritative records remain unresolved. The set comparison fails. Stable references let us compare the actual work items, rather than relying on equal counts.

The final call requests desk review and verifies that the same reference remains open. The oracle includes both unresolved and requested reversals; routing does not silently clear the repair.

This reducer is new demonstration code. It is not a monitor shipped by Rasa or the companion project.

Read what the replay actually promises

The pinned replay branch returns the saved redemption and explicitly resets attempt fields:

record = service.redemptions[reference]
return {
    **service.public_redemption(reference, record),
    "effects": 0,
    "replay": True,
    "unmatched_points_booking": 0,
    # The remaining next_step text asks callers to report the stored state.
}

The selected HoldTests and RedeemTests receipt contains 14 passing checks without skips. The original replay experiment keeps the balance at 76,900 with one debit ledger entry. Its pending booking and unresolved reversal remain unchanged:

QuestionEvidenceAnswer on replay
Did this attempt create a mismatch?Attempt counterNo
Did this attempt debit again?Effects and ledgerNo
Is the member still owed repair?Saved points, booking and reversal stateYes

Maintain a set of outstanding operations

A production repair queue should use a stable redemption reference. Repeated attempts update evidence for that reference instead of creating or removing work based on an event counter. Periodically reconcile queue membership against authoritative records. This is a proposed operating contract; the sample does not implement a production queue.

Plan a way to query authoritative records and reconcile the queue. Keep the event counter for newly created failures; it cannot replace the state check. These are recommended operating responsibilities, not measured production costs. After a restart, reconstruct outstanding work from records rather than assuming a quiet counter means an empty queue.

Clear a reference only on evidence of completed ticketing or a reconciled reversal under the owner’s contract. The fixture’s rewards-desk route marks the reversal requested. Requested is still open work; it does not return the points. This fixture has no completed reversal implementation.

:::

These checks do not prove durable storage, partner integration, or successful repair. Test those separately before replacing the fixture.