Trust is earned, not given

A different perspective

2026-08-06 · Projects

AI-Enabled RMA, part 3: The policy evaluator — nine rules, a rule trace, and the money

Part 3from the AI-Enabled RMA series · 6 parts in all

The whole eligibility story of AIEnabledRMA lives in one class, RmaPolicyEvaluator. This article walks its nine rules in order, because the order is the design, and then covers the part everyone gets wrong first: where the money comes from.

A closed input record

The evaluator takes a single record, RmaPolicyRequest, and it is deliberately closed: adding a field forces every rule's behaviour to be explicit at the call site. It carries the device, today's date, an optional region code, optional problem text, whether the device already has an open request, a nullable proof-of-purchase flag (nullable so "not required" is an explicit state rather than a missing one), and an optional AI confidence that is ignored unless the deployment has relaxed the lockdown.

Nothing else is reachable. There is no repository, no clock, no HTTP client. The date comes in as a parameter, which is why the same question asked twice about the same request produces the same answer — and why the MCP tools make today a required argument.

The nine rules, in order

  1. Unknown device. If the identifier did not resolve, there is nothing to decide against: DeviceNotFound, not eligible. The alternative — treating an unknown serial as "probably covered" — is how a returns system ends up shipping replacements for units it has never sold.
  2. Suspected serial tampering. A configured prefix match escalates to RequiresHumanReview rather than denying. A false positive must not silently discard a legitimate customer's return; the fraud team looks at it.
  3. One open request per device. Escalate so the customer is routed to the existing return instead of creating a second one.
  4. Excluded causes. Matched against the problem text, and checked before warranty, because liquid ingress is out of scope even on a unit bought last week. Each rule declares a match mode: any term, all terms, or substring, with whole-word token matching as the default so "screen" does not fire on "screened".
  5. Return window. Measured from the ship date when it is known. This is an absolute window, independent of warranty — a two-year-old unit can be well inside its warranty and still outside the window the business accepts returns for.
  6. Proof of purchase. Only when the caller says it is missing, and then it escalates rather than denies.
  7. Warranty, including the grace period. If no warranty governs, the outcome is EligibleForPaidRepair rather than a refusal — an expired unit is still repairable. The same path handles a tier that is configured as uncovered, which is how the sample's no-service tier becomes the demonstration of a paid repair.
  8. Region restriction. Configured regions always go to a human before approval.
  9. AI confidence — last, and only if the lockdown is off. Kept as the final rule so it can never bypass a denial above it. When the flag is off and the confidence is below the threshold, the outcome is the configured fallback, which defaults to RequiresHumanReview.

Two details that took a rewrite to get right

The grace period is read from the same options object as everything else. The warranty selector used to ignore a configurable grace period, so a deployment that set one still had expired-but-recent units denied outright. It now filters on EndDate.AddDays(GracePeriodDays) >= today, and when several records overlap the one that runs latest wins — so a superseded record on the same unit cannot shadow the live one.

Out of warranty is not a dead end. The original code denied an expired unit outright, which turned a perfectly repairable device into a stop. The enum documentation had always said a paid repair was available. Routing it to EligibleForPaidRepair instead is a one-line behaviour change with a large effect on the shape of the product: the wizard grows a payment step, and the money becomes explicit.

Every branch appends to a fired list, and the returned decision carries that list as MatchedRuleIds — identifiers like warranty.expired, warranty.tier:standard, region:EU, fraud.prefix:XX. That is the audit trail, and it is also what the workflow uses to pick a governing reason for a multi-item return.

Where the money lives

Pricing is the part most implementations get subtly wrong, so the rule is stated in the code comments and enforced by the types: a tier describes the plan, a price describes the product. Turnaround time, replacement tier and coverage come from WarrantyTierPolicy; the repair fee comes from a RepairPrices table matched by SKU first, then product name, then the configured DefaultRepairFee so an unlisted product still has a quotable price.

Two consequences fall out of that. The wizard's bill and the MCP get_fee_quote tool read the same table, so an agent's quote and the actual charge cannot disagree. And a request with two out-of-warranty units charges the sum of their own product prices once per request, rather than a tier-wide rate that happens to be wrong for one of them.

There is a matching subtlety in the tier lookup. GetTier is lenient on purpose: a legacy or mistyped tier on a device row must not make every return unevaluable, so it falls back to the default tier. TryGetTier is exact and returns null. The evaluator uses the lenient one because its job is to decide. Anything answering a question about money uses the exact one, because silently answering a "premium" question with standard-tier numbers is a false answer about money rather than a rounding error.

Testing a decision engine

Twenty-three cases cover the evaluator with no model, no network and no database: each rule in isolation, the boundary conditions around the return window and the grace period, the ordering guarantees (an in-warranty liquid-damage report is still excluded), and the advisory-only behaviour in both positions of the flag. Because the input is a closed record and the output carries the fired rules, a failing test tells you which rule is wrong in one line of the assertion.

Next: the machinery around the model — an in-process lexical index, a scope gate that runs before any model call, and what happens to untrusted output.

Repository: github.com/bobhuang1/AIEnabledRMA