AI-Enabled RMA, part 5: One policy, three hosts — the MCP tool server
Part 5from the AI-Enabled RMA series · 6 parts in all
An agent asked "AX2-0001-0005 is dead on arrival and out of warranty — what happens if the
customer returns it, and how much would it cost?" should not answer from its own opinion. In
AIEnabledRMA it calls two tools and gets
the same answer the wizard would compute: a paid repair, at $69.00 for the AX-200-HS.
This article is the MCP server that makes that true.
A stdio server, and the logging trap
src/AIEnabledRma.Mcp
is a stdio MCP server built on ModelContextProtocol, with tools discovered from the
assembly by attribute. Two lines of setup matter more than they look:
builder.Logging.AddConsole(o => o.LogToStandardErrorThreshold = LogLevel.Trace);
builder.Services.AddMcpServer().WithStdioServerTransport().WithToolsFromAssembly();
stdout is the JSON-RPC channel. The default console logger writes to stdout, which corrupts the stream and makes the host fail to parse the first response frame. Redirecting every log level to stderr is the fix, and it is the kind of bug that presents as "the MCP server does not work" with no error message anywhere useful.
The server then registers the same data layer the web host uses — AddRmaDbContext
and AddRmaData — which is what guarantees that the fuzzy-match thresholds a
customer-facing wizard applies and the ones an agent's lookup tool applies cannot drift apart.
The tool catalog
check_warranty_coverage — the headline tool. It resolves an
identifier, then evaluates the same RmaPolicyRequest the workflow builds, returning
the eligibility outcome, the reason, the customer-safe explanation and the fired rule ids.
Its description in the source tells the agent plainly: use this to answer a coverage
question; do not answer coverage questions from the model.
get_fee_quote — what a repair would cost. The price comes from
the per-product repair price table, which is the same table the wizard charges against, so a
quote and a charge can never disagree. It also reports coverage, because a covered unit owes
nothing and the quoted figure is what the repair would cost without it.
find_device — lookup by serial, MAC, IMEI or asset tag, with
fuzzy matching so a mistyped or partially transcribed serial still resolves. It returns the unit
together with the warranty that governs it today.
search_devices — for when the customer has no identifier to
hand. Best-first, capped at 25 results, and the description explicitly prefers
find_device when any identifier exists, because an identifier is exact and this is
not.
get_open_returns_for_device — so an agent does not open a second
return for a unit that already has one.
Two more tools round out the surface: customer lookup and
resolve_shipping_address, which matches free text to a customer's stored addresses
and refuses to pretend a weak match is confident.
Read-only by construction
Every tool is a pure function of stored data and policy configuration. There is no write path,
no state, and no way to create, approve or modify a return. The category is enforced by what the
tools are allowed to depend on: PolicyTools takes repositories, the eligibility
service and the options object — not the workflow. There is no code path from an agent's tool
call to IRmaWorkflow.
The response shape reinforces it. CoverageResponse carries an
AllowsRmaCreation field that is always false, documented as such,
so an agent reading it can never mistake an evaluation for an approval. Every response also
carries a note saying the same thing in words: this is a policy evaluation, not an approval.
The sharpest design decision is a parameter that does not exist. The policy has an
AiConfidence field, and the MCP tool never passes it. The comment says why: the
policy ignores AI confidence while the advisory-only lockdown is on, the tool has no way to
obtain a trustworthy value, so it does not pretend to — a caller must not be able to set that
field through a tool.
Two bugs worth repeating
A tool that states a false fact is worse than one that asks. An earlier
version of check_warranty_coverage defaulted hasOpenReturn to
true, which made the tool tell customers "there is already an open return for this
item" for units that had none. It now reads the database, which is authoritative, while still
letting a caller that holds fresher information — an in-flight session, say — supply its own
value.
An unknown serial is not "probably covered". When an identifier resolves to
nothing, coverage cannot be confirmed and the answer is a refusal with an explicit
Identified: false flag. The alternative is a system that ships replacements for
units it has never sold.
Reproducibility, PII, and narrow projections
today is a required string parameter on the coverage and quote tools,
parsed as ISO yyyy-MM-dd. That is deliberate: warranty and return windows are
date-sensitive, so an answer that depends on the wall clock is not reproducible, and a support
conversation three weeks later would get a different answer to the same question.
What leaves the process is deliberately narrow. Tools return projection records —
DeviceSummary, ProductSummary, WarrantySummary — rather
than domain entities, on the theory that serial numbers, MAC addresses and IMEIs are
customer-identifying and an agent needs no more than enough to confirm it has the right unit.
find_device goes further: a unit that already has a return in progress is
suppressed, with Suppressed reported distinctly from Found so
a caller can tell "no such device" apart from "found, but not shown to you," and re-run with an
explicit opt-in if it has a reason to.
Matching itself is a deterministic scorer in the domain — token overlap weighted 60%, a Levenshtein similarity weighted 40%, with exact, prefix and whole-token cases short-circuiting at high scores. The tokenizer squashes everything that is not a letter or digit, which is what makes "O'Brien", "o brien" and "OBRIEN" compare equal. PostgreSQL trigram ranking sits on top of it in the data layer for large tables, but the in-memory scorer is what runs in tests and in the fallback path, so behaviour is never only true against a database.
Eighteen tests
cover the tools: unknown identifiers, the suppression path, the proof-of-purchase override, date
parsing, and the invariant that AllowsRmaCreation is false on every response.
Next: the wizard — the state machine, the paid-repair path, and the test suite that holds the whole thing together.
Repository: github.com/bobhuang1/AIEnabledRMA