Trust is earned, not given

A different perspective

2026-09-29 · Projects

AI-Enabled RMA, part 6: The wizard end to end — state, paid repairs, and the tests

Part 6from the AI-Enabled RMA series · 6 parts in all

The last piece of AIEnabledRMA is the part a customer actually sees: the return wizard, and behind it the workflow that moves a request from draft to awaiting shipment. This article covers the six steps, the two outcomes that took the most thought — a paid repair and a partially created return — and the test suite that proves the steps are wired together rather than merely present.

Six steps, addressed by an id

RmaWizardController walks through Start (customer plus serials, one per line), Triage (describe the fault), Problems (per-line detail for the bench), Shipping, Payment and Confirm. The payment step only exists when money is due. Route-id steps from triage onward are addressed by the request GUID, not by a session.

That is a deliberate, documented trade-off. The GUID is the capability: a 122-bit unguessable token plays the role a session would play, so a bookmarked step works across browsers and a shareable link lets a second device finish a return. The accepted cost is that anyone holding a request URL can drive that return until it is confirmed, and the README says so explicitly, along with the trigger to revisit: before the wizard gains authenticated pages, an operator portal, or multiple users per return. Every wizard POST is protected by antiforgery, because each one changes the return's state.

Outcome is a first-class return value

The workflow returns a StepResult carrying a StepOutcome — and one member of that enum earns its own paragraph. PartiallyCreated means a return was created, but not for every item supplied. It is deliberately distinct from Continue, so a caller that ignores the detail list still cannot mistake a partial return for a complete one.

Mixed batches are the normal case, not the edge case: customers return a working accessory alongside a broken unit. A single ineligible item does not sink the batch. Every device is evaluated, everything that clears policy becomes a line, out-of-warranty units become paid lines rather than refusals, and only the items that cannot proceed at all are excluded. The whole batch is refused only when nothing in it can proceed.

Exclusions are persisted, in an rma_exclusions table with one row per dropped item and the reason stored as an enum name. The triage page rebuilds its warning from the database, so a refresh or a new browser session cannot silently make the notice that part of the submission was refused disappear. That is the kind of decision that looks like extra work until the first support call about a missing item.

The governing reason for the request follows a rule worth stating: a paid line governs a paid return. If anything in the batch is not covered, the request is filed as PaidRepair under the out-of-warranty reason rather than under a covered line's reason — otherwise the audit trail would say a return was covered when the customer is about to be charged.

The paid-repair path, and a bug about tiers

An out-of-warranty unit is neither refused nor free. When the return is a PaidRepair, the shipping step computes the fee as the sum of its lines' product prices once per request, stores it as the deposit, and the payment step authorises it with manual-capture semantics: authorised now, captured once the item is received and inspected. Shipping is zero — the customer is paying for the repair, not for freight, and the payment page's copy stays true because of it. In-warranty returns skip the step entirely.

The bug is instructive. WarrantyTier was originally left null on the request, and because the shipping step resolves the tier by name, every return was quoted at zero regardless of policy. The fix sets the governing device's tier at creation, and the reason it is worth writing down is that a null field and a zero price are indistinguishable at the point of failure. A multi-item return spanning two tiers is genuinely ambiguous — the schema charges once per request, not per line — so the governing device sets the tier and the trace records which one.

One transaction per step

Repositories track changes but do not commit them, so a step that touches several repositories is a single transaction. Every mutating step ends by calling the unit of work, and the helper logs a warning when a commit wrote zero rows — usually meaning the step was a no-op or the request it changed was not tracked. That warning has caught more bugs than any assertion in the suite, because "the write silently did nothing" is the failure mode that looks like success.

References look like RMA-2026-00042 and come from a single global PostgreSQL sequence with CACHE 1; the year in the string is display-only. The obvious MAX(rma_number) + 1 is wrong under concurrency — two requests read the same maximum and one collides — so the sequence allocates, the migration path can adopt an existing database, and the seeder realigns the sequence afterwards so seeded demo rows and the sequence can never disagree.

The tests that make it credible

The suite is roughly 190 tests across ten files, and the interesting ones are the ones that refuse to fake their dependencies. RmaNumberSequenceTests and WebWizardTests need a real PostgreSQL server: concurrency semantics for sequence allocation and end-to-end persistence for the wizard cannot be faked usefully. They skip rather than fail when RMA_TEST_CONNECTION is unset or the server is unreachable, so dotnet test with no environment still goes green — which is what makes it plausible to run in CI at all.

WebWizardTests boots the real web host with WebApplicationFactory against a throwaway database and walks the wizard through the paid-repair path, landing on the payment step and billing the per-product price. Booting the real host is the only layer that proves the steps, the workflow and persistence are wired together; a controller unit test with mocked services would pass while the application failed at step four.

One small detail from that file is worth stealing: the project declares a WebAppAnchor class, because three projects in the solution have a top-level Program and the test project references all of them, making the generated entry point name ambiguous. A distinct anchor type keeps WebApplicationFactory<WebAppAnchor> unambiguous.

Health, and setting it up

GET /health reports the two things that fail independently and both look like "the assistant is broken" from the customer's side: the loaded knowledge base with its per-file errors, and the configured AI provider. It deliberately does not check the database, and the code says not to wire it into a liveness probe and read green as "the app can serve returns" — the kind of note that saves someone an outage.

docker compose up -d
dotnet run --project tools/DbAdmin -- reset --force
dotnet run --project src/AIEnabledRma.Web

From there, the shipped demo catalog gives you a device for every outcome, including one that walks the paid-repair path and one that is genuinely unknown.

What I would take from this

The pattern generalises past returns. Any customer-facing flow with a model in it benefits from the same three moves: make the consequential decision a deterministic function you can unit-test in isolation, make the model's contribution advisory and attribute every claim it makes to something you retrieved, and make the failure mode a human rather than an answer. The rest — the persistence of exclusions, the partial-created outcome, the sequence instead of a maximum, the tests that boot the real host — is ordinary software engineering doing the unglamorous work that makes the first three believable.

Repository: github.com/bobhuang1/AIEnabledRMA