AI-Enabled RMA, part 3: The policy evaluator — nine rules, a rule trace, and the money
Part 3from the AI-Enabled RMA series · 6 parts in all
The whole eligibility story of AIEnabledRMA
lives in one class,
RmaPolicyEvaluator.
This article walks its nine rules in order, because the order is the design, and then
covers the part everyone gets wrong first: where the money comes from.
A closed input record
The evaluator takes a single record, RmaPolicyRequest, and it is deliberately
closed: adding a field forces every rule's behaviour to be explicit at the call site. It carries
the device, today's date, an optional region code, optional problem text, whether the device
already has an open request, a nullable proof-of-purchase flag (nullable so "not required" is an
explicit state rather than a missing one), and an optional AI confidence that is ignored unless
the deployment has relaxed the lockdown.
Nothing else is reachable. There is no repository, no clock, no HTTP client. The date comes in
as a parameter, which is why the same question asked twice about the same request produces the
same answer — and why the MCP tools make today a required argument.
The nine rules, in order
- Unknown device. If the identifier did not resolve, there is nothing to
decide against:
DeviceNotFound, not eligible. The alternative — treating an unknown serial as "probably covered" — is how a returns system ends up shipping replacements for units it has never sold. - Suspected serial tampering. A configured prefix match escalates to
RequiresHumanReviewrather than denying. A false positive must not silently discard a legitimate customer's return; the fraud team looks at it. - One open request per device. Escalate so the customer is routed to the existing return instead of creating a second one.
- Excluded causes. Matched against the problem text, and checked before warranty, because liquid ingress is out of scope even on a unit bought last week. Each rule declares a match mode: any term, all terms, or substring, with whole-word token matching as the default so "screen" does not fire on "screened".
- Return window. Measured from the ship date when it is known. This is an absolute window, independent of warranty — a two-year-old unit can be well inside its warranty and still outside the window the business accepts returns for.
- Proof of purchase. Only when the caller says it is missing, and then it escalates rather than denies.
- Warranty, including the grace period. If no warranty governs, the outcome is
EligibleForPaidRepairrather than a refusal — an expired unit is still repairable. The same path handles a tier that is configured as uncovered, which is how the sample'sno-servicetier becomes the demonstration of a paid repair. - Region restriction. Configured regions always go to a human before approval.
- AI confidence — last, and only if the lockdown is off. Kept as the final rule
so it can never bypass a denial above it. When the flag is off and the confidence is below the
threshold, the outcome is the configured fallback, which defaults to
RequiresHumanReview.
Two details that took a rewrite to get right
The grace period is read from the same options object as everything else. The
warranty selector used to ignore a configurable grace period, so a deployment that set one still
had expired-but-recent units denied outright. It now filters on
EndDate.AddDays(GracePeriodDays) >= today, and when several records overlap the
one that runs latest wins — so a superseded record on the same unit cannot shadow the live one.
Out of warranty is not a dead end. The original code denied an expired unit
outright, which turned a perfectly repairable device into a stop. The enum documentation had
always said a paid repair was available. Routing it to
EligibleForPaidRepair instead is a one-line behaviour change with a large effect on
the shape of the product: the wizard grows a payment step, and the money becomes explicit.
Every branch appends to a fired list, and the returned decision carries that list
as MatchedRuleIds — identifiers like warranty.expired,
warranty.tier:standard, region:EU, fraud.prefix:XX. That
is the audit trail, and it is also what the workflow uses to pick a governing reason for a
multi-item return.
Where the money lives
Pricing is the part most implementations get subtly wrong, so the rule is stated in the code
comments and enforced by the types: a tier describes the plan, a price describes the
product. Turnaround time, replacement tier and coverage come from
WarrantyTierPolicy; the repair fee comes from a
RepairPrices table matched by SKU first, then product name, then the configured
DefaultRepairFee so an unlisted product still has a quotable price.
Two consequences fall out of that. The wizard's bill and the MCP get_fee_quote
tool read the same table, so an agent's quote and the actual charge cannot disagree. And a
request with two out-of-warranty units charges the sum of their own product prices once
per request, rather than a tier-wide rate that happens to be wrong for one of them.
There is a matching subtlety in the tier lookup. GetTier is lenient on purpose: a
legacy or mistyped tier on a device row must not make every return unevaluable, so it falls back
to the default tier. TryGetTier is exact and returns null. The evaluator uses the
lenient one because its job is to decide. Anything answering a question about money uses
the exact one, because silently answering a "premium" question with standard-tier numbers is a
false answer about money rather than a rounding error.
Testing a decision engine
Twenty-three cases cover the evaluator with no model, no network and no database: each rule in isolation, the boundary conditions around the return window and the grace period, the ordering guarantees (an in-warranty liquid-damage report is still excluded), and the advisory-only behaviour in both positions of the flag. Because the input is a closed record and the output carries the fired rules, a failing test tells you which rule is wrong in one line of the assertion.
Next: the machinery around the model — an in-process lexical index, a scope gate that runs before any model call, and what happens to untrusted output.
Repository: github.com/bobhuang1/AIEnabledRMA