AI-Enabled RMA, part 1: The AI advises, the policy decides
Part 1from the AI-Enabled RMA series · 6 parts in all
AIEnabledRMA is an end-to-end returns (RMA) reference implementation. A customer types the fault, a language model helps troubleshoot it, and a deterministic, auditable policy decides whether the return is covered, what it costs, and what happens next. The model can suggest a fix. It can quote nothing and approve nothing. This first article covers why the split is drawn there, and what falls out of it for the rest of the system.
The problem with an AI that approves returns
A return touches three things that are not negotiable: a customer's money, a company's stock, and a warranty obligation. Give a model the ability to approve one and you have given it the ability to be talked into approving one — by a persuasive customer, by text embedded in a support ticket, or by an instruction sitting in a knowledge-base article. The failure is not that the model is usually wrong. It is that when it is wrong, there is no rule to point at afterwards, and no way to answer "why was this approved?" other than pasting a chat log.
The README states the position in one line: the AI can suggest a fix — it can quote nothing and approve nothing. Everything in the architecture is a consequence of taking that sentence literally rather than aspirationally.
One evaluator, one vocabulary of outcomes
Eligibility lives in a single class,
RmaPolicyEvaluator,
which is ordinary testable C# with no model anywhere near the call path. It answers with one of
four outcomes, and the vocabulary is deliberately small:
- Eligible — covered by an active warranty; a free repair or replacement is owed.
- EligibleForPaidRepair — not covered, but the device is real and repairable, so the customer may still be quoted a fee.
- RequiresHumanReview — the system must not decide; a specialist does.
- NotEligible — there is no repair path: excluded cause, outside the return window, unknown device.
Every decision also carries a RmaDecisionReason, a customer-safe
Explanation, and — the part that makes the system auditable — a
MatchedRuleIds list naming the rules that actually fired. When someone asks in
March why a specific unit was refused in January, the answer is a stored list of rule
identifiers plus the same code path re-run against the same inputs, not an interrogation of a
model's mood.
Two abstractions sit over that one implementation: IRmaPolicy, which belongs to
the domain, and IEligibilityService, which is the seam the web and MCP layers
resolve when they need to ask "may this proceed?" without loading a whole RMA aggregate. They
are satisfied by the same class, so the wizard and an agent's tool call cannot disagree about
an outcome. That is a recurring theme: the same registration is reused by all three hosts so
that a threshold, a price or a rule can never drift between the customer's browser and an
agent's answer.
Fail closed, and only narrow
Three invariants do most of the work.
The AI may only narrow. The configuration flag
RmaPolicyOptions.AiIsAdvisoryOnly defaults to true and stays true in
production. When it is true, the model's confidence cannot make an ineligible request eligible.
It is configurable only so the deterministic path can be regression-tested in isolation.
Every failure resolves to a human. A model timeout, a parse failure, an empty
retrieval, or an out-of-range confidence all land on
FallbackEligibility, which is RequiresHumanReview. There is no code
path where a model failure becomes an approval. The fallback also never claims resolution and
never claims high confidence.
Troubleshooting success cancels, it does not approve. When the model resolves
the fault, the RMA is marked Cancelled with the reason "Resolved during automated
troubleshooting" — not approved. A return the customer no longer needs is not the same object as
a return that was granted, and conflating them would put a fake approval in the audit trail.
What the split buys you
A testable core. The policy evaluator has 23 test cases and needs no model, no network and no database. The AI path has its own tests with a fake chat model. Neither suite has to simulate the other.
Determinism as a feature. The same serial number, the same date and the same
region produce the same answer whether the caller is the wizard, the MCP server or a unit test.
Dates are supplied explicitly (the MCP tools make today a required parameter), so
an answer is reproducible rather than dependent on when it was asked.
A small, boring blast radius. Because the model is confined to a troubleshooting verdict stored as advisory text on the line, the worst outcome of a model misbehaving is a bad suggestion — a support cost, not a financial or compliance incident. That is what makes it safe to put a language model in a customer-facing flow at all.
The rest of this series is the implementation: the project layout and its dependency rule, the nine ordered policy rules and how money is priced, the retrieval and lockdown machinery around the model, the MCP server that exposes the same policy to an agent, and the wizard with its state machine and test suite.
Repository: github.com/bobhuang1/AIEnabledRMA