Trust is earned, not given

A different perspective

2026-09-10 · Projects

AI-Enabled RMA, part 5: One policy, three hosts — the MCP tool server

Part 5from the AI-Enabled RMA series · 6 parts in all

An agent asked "AX2-0001-0005 is dead on arrival and out of warranty — what happens if the customer returns it, and how much would it cost?" should not answer from its own opinion. In AIEnabledRMA it calls two tools and gets the same answer the wizard would compute: a paid repair, at $69.00 for the AX-200-HS. This article is the MCP server that makes that true.

A stdio server, and the logging trap

src/AIEnabledRma.Mcp is a stdio MCP server built on ModelContextProtocol, with tools discovered from the assembly by attribute. Two lines of setup matter more than they look:

builder.Logging.AddConsole(o => o.LogToStandardErrorThreshold = LogLevel.Trace);
builder.Services.AddMcpServer().WithStdioServerTransport().WithToolsFromAssembly();

stdout is the JSON-RPC channel. The default console logger writes to stdout, which corrupts the stream and makes the host fail to parse the first response frame. Redirecting every log level to stderr is the fix, and it is the kind of bug that presents as "the MCP server does not work" with no error message anywhere useful.

The server then registers the same data layer the web host uses — AddRmaDbContext and AddRmaData — which is what guarantees that the fuzzy-match thresholds a customer-facing wizard applies and the ones an agent's lookup tool applies cannot drift apart.

The tool catalog

check_warranty_coverage — the headline tool. It resolves an identifier, then evaluates the same RmaPolicyRequest the workflow builds, returning the eligibility outcome, the reason, the customer-safe explanation and the fired rule ids. Its description in the source tells the agent plainly: use this to answer a coverage question; do not answer coverage questions from the model.

get_fee_quote — what a repair would cost. The price comes from the per-product repair price table, which is the same table the wizard charges against, so a quote and a charge can never disagree. It also reports coverage, because a covered unit owes nothing and the quoted figure is what the repair would cost without it.

find_device — lookup by serial, MAC, IMEI or asset tag, with fuzzy matching so a mistyped or partially transcribed serial still resolves. It returns the unit together with the warranty that governs it today.

search_devices — for when the customer has no identifier to hand. Best-first, capped at 25 results, and the description explicitly prefers find_device when any identifier exists, because an identifier is exact and this is not.

get_open_returns_for_device — so an agent does not open a second return for a unit that already has one.

Two more tools round out the surface: customer lookup and resolve_shipping_address, which matches free text to a customer's stored addresses and refuses to pretend a weak match is confident.

Read-only by construction

Every tool is a pure function of stored data and policy configuration. There is no write path, no state, and no way to create, approve or modify a return. The category is enforced by what the tools are allowed to depend on: PolicyTools takes repositories, the eligibility service and the options object — not the workflow. There is no code path from an agent's tool call to IRmaWorkflow.

The response shape reinforces it. CoverageResponse carries an AllowsRmaCreation field that is always false, documented as such, so an agent reading it can never mistake an evaluation for an approval. Every response also carries a note saying the same thing in words: this is a policy evaluation, not an approval.

The sharpest design decision is a parameter that does not exist. The policy has an AiConfidence field, and the MCP tool never passes it. The comment says why: the policy ignores AI confidence while the advisory-only lockdown is on, the tool has no way to obtain a trustworthy value, so it does not pretend to — a caller must not be able to set that field through a tool.

Two bugs worth repeating

A tool that states a false fact is worse than one that asks. An earlier version of check_warranty_coverage defaulted hasOpenReturn to true, which made the tool tell customers "there is already an open return for this item" for units that had none. It now reads the database, which is authoritative, while still letting a caller that holds fresher information — an in-flight session, say — supply its own value.

An unknown serial is not "probably covered". When an identifier resolves to nothing, coverage cannot be confirmed and the answer is a refusal with an explicit Identified: false flag. The alternative is a system that ships replacements for units it has never sold.

Reproducibility, PII, and narrow projections

today is a required string parameter on the coverage and quote tools, parsed as ISO yyyy-MM-dd. That is deliberate: warranty and return windows are date-sensitive, so an answer that depends on the wall clock is not reproducible, and a support conversation three weeks later would get a different answer to the same question.

What leaves the process is deliberately narrow. Tools return projection records — DeviceSummary, ProductSummary, WarrantySummary — rather than domain entities, on the theory that serial numbers, MAC addresses and IMEIs are customer-identifying and an agent needs no more than enough to confirm it has the right unit. find_device goes further: a unit that already has a return in progress is suppressed, with Suppressed reported distinctly from Found so a caller can tell "no such device" apart from "found, but not shown to you," and re-run with an explicit opt-in if it has a reason to.

Matching itself is a deterministic scorer in the domain — token overlap weighted 60%, a Levenshtein similarity weighted 40%, with exact, prefix and whole-token cases short-circuiting at high scores. The tokenizer squashes everything that is not a letter or digit, which is what makes "O'Brien", "o brien" and "OBRIEN" compare equal. PostgreSQL trigram ranking sits on top of it in the data layer for large tables, but the in-memory scorer is what runs in tests and in the fallback path, so behaviour is never only true against a database.

Eighteen tests cover the tools: unknown identifiers, the suppression path, the proof-of-purchase override, date parsing, and the invariant that AllowsRmaCreation is false on every response.

Next: the wizard — the state machine, the paid-repair path, and the test suite that holds the whole thing together.

Repository: github.com/bobhuang1/AIEnabledRMA