Trust is earned, not given

A different perspective

2026-05-07 · Projects

AI Frontiers, part 63: Interoperability — standards, protocols, and portable context

Part 63from the AI Frontiers series · 65 parts in all

In November 2024, a tool-and-context protocol appeared with an open specification and a set of provider implementations, and within a year it was the assumed interface for connecting models to external systems (part 12 covered that arrival). In April 2025, a second protocol addressed the other direction — agent-to-agent communication between separately built systems — and by 2026 there were protocols and draft specifications for tool description, telemetry, identity, and dataset metadata. This is what a market looks like when the integration cost finally exceeds the differentiation value of being proprietary, and it happens on a recognizable schedule.

This entry is about that layer: which protocols have earned adoption, what "portable context" can and cannot mean, how to build for interoperability without hollowing out your product, and why the most portable asset in your stack is the one most teams treat as an internal artifact.

Why protocols appear on schedule

The pattern is old enough to be predictive. A layer gets standardized when three conditions hold: everyone implements it, nobody can differentiate on it, and the cost of pairwise integration has become intolerable. HTTP and HTML standardized the web's transport and document model while leaving content proprietary. SQL and its driver interfaces standardized database access while leaving query engines fiercely competitive. OpenAPI standardized service descriptions — and the specification work that produced it, and its predecessor, was largely a response to the combinatorial explosion of clients and servers needing to talk to each other. The Open Container Initiative standardized image formats precisely because every orchestrator reimplementing them was waste, and the container runtime market only got more competitive afterward (Open Container Initiative). The web's own architecture was itself a standardization argument about which constraints produce interoperability and which produce fragility (Fielding).

Each time the same shape appears: the standardized layer is the one where the buyer's switching cost is the vendor's leverage, and vendors accept it once the market is large enough that interoperability expands the pie more than it erodes the moat. In AI systems, that layer is now the interface between models and everything else — tools, data, other agents, and the telemetry that describes what happened.

The landscape, with honest assessments

Tool and context protocols. A specification for exposing tools, resources, and prompts to a model in a uniform way, with a transport and a capability handshake. Its value is directly proportional to how many independent implementations exist, and it has the property that made earlier protocols succeed: a server written once works with any compliant client. The reason it caught on so fast is that it solved a problem every team had, in the direction where integration cost was highest.

Agent-to-agent protocols. The complementary direction: describing an agent's capabilities so another agent can discover and invoke it, with task lifecycle and message formats. This is genuinely harder than tool invocation because the semantics of a task are richer than the semantics of a function call, and the early specifications are correspondingly broader. My expectation is that the durable parts will be the discovery and task-state elements, while the parts that try to standardize reasoning or negotiation will be ignored — the same fate that met earlier attempts to standardize business process.

Tool description in the existing idiom. The most effective agent tool interface in practice is still an OpenAPI document, because it is already the world's most widely deployed machine-readable description of an HTTP interface, and models read it well. Protocols that require rewriting existing services will lose to the approach of projecting existing interfaces into a form a model can use.

Telemetry. The semantic conventions for generative AI work in the OpenTelemetry project are unglamorous and among the most valuable efforts in the stack, because they make traces comparable across vendors — the artifact part 49 argued is the most valuable thing an AI system produces is only shareable if its fields are standard.

Identity and authorization. OAuth remains the substrate (Hardt), and the extending profile work for delegation — letting an agent act on behalf of a user with a scoped, revocable, auditable credential — is the piece most required for enterprise adoption and least settled. Verifiable-credential formats provide the verifiable-claims half of the problem and are mature enough to use (W3C). Read part 57 alongside any protocol decision, because interoperability with an untrusted party is a security problem first.

Metadata for models and data. Model cards established the practice of describing a model's intended use, evaluation, and limitations in a structured document (Mitchell et al.), and dataset metadata formats such as Croissant extended the idea to training data (Akhtar et al.). This is the least glamorous category and the one most likely to matter to regulators, because it is the machine-readable evidence that a system was documented before it was deployed.

What "portable context" can actually mean

The phrase suggests that everything a system knows about a user or a task can be lifted and carried to a different model or vendor. In practice, portability decomposes into five components with very different transfer properties.

Fully portable: the raw interaction history, the documents, the tool definitions, the trace of what happened. These are your data in an open format, and if you store them that way, no vendor owns them. This is also why a trace schema is a strategic asset: it is the container for everything else.

Mostly portable: the evaluation set and its judgments. A test case with a rubric and a known-correct answer transfers to any model; only the scores change. This is the reason I keep arguing that evaluation is the highest-leverage artifact in an applied system — it is the only asset that is both vendor-independent and directly controlling of quality.

Partially portable: prompts. A prompt encodes assumptions about one model's failure modes and idiosyncrasies. The structure and the content travel; the performance does not. Any migration is a re-tuning project, and teams that budget for it as a copy-paste operation are consistently surprised.

Weakly portable: memory and accumulated state. Summaries generated by one model's worldview, embeddings from one encoder's space, learned preferences from one model's interactions. The data is yours; its meaning is model-specific. Part 43 made the point about embeddings and drift, and the migration consequence is the same.

Not portable: fine-tuned weights, adapters, prompt-cache state, and every optimization that depends on a specific runtime. That is fine — it is the part where specialization creates advantage — provided the reason it is not portable is a deliberate choice rather than an accident.

There is a real tension here worth stating plainly. Portability and efficiency pull in opposite directions: prompt caching rewards byte-identical prefixes, locality-aware routing rewards sticking with one provider, and provider-specific features such as structured outputs or context caching deliver measurable cost and quality benefits that vanish the moment you abstract over them. The resolution is not to pick a side but to layer: keep the portable artifacts authoritative, keep the optimizations in an adapter that can be swapped, and know at each layer which one you are paying for. A useful heuristic is that anything you would want to hand to a regulator, a customer's auditor, or your successor should be portable; anything that exists to make a request cheaper should not be.

Building for interop without hollowing out the product

Interoperability has costs, and the two classic ones are worth naming. The lowest-common-denominator API problem: designing to the intersection of what every implementation supports means never using the features that differentiate. And Postel's law's dark side: being liberal in what you accept produces a de facto dialect that breaks clients the day they meet a stricter implementation (Postel). The web's own history is a long argument about this.

Four practices avoid both. Standardize the boundary, specialize the core. The protocol defines how a tool is described and invoked; the implementation behind it can be as idiosyncratic as it likes. Version explicitly and refuse ambiguously. Validate strictly, reject malformed input, and document the profile you implement rather than claiming full conformance. Write contract tests against a second implementation. The only real evidence of interoperability is another team's client working against your server without a phone call, and the only way to keep that true is to test it on every change. Keep the adapter thin and the semantics in your own model. A translation layer that maps a standard concept onto your domain model is maintainable; one that leaks the standard's vocabulary into your core is a rewrite waiting to happen.

The procurement conversation

Portability shows up in every vendor negotiation, usually framed as lock-in, and it is worth being precise about where the lock-in actually lives. It is not in the API, which is cheap to replace. It is in the accumulated, model-specific artifacts: the tuned prompts, the fine-tunes, the memory state, the cached prefixes, and above all the evaluation history that tells you whether a migration improved anything. Those are the things that make leaving expensive, and every one of them is under your control if you decided so early.

So the questions worth asking a vendor, and worth answering about your own platform, are narrow: can I export my data and interactions in a documented format; is there an OpenAPI-compatible description of every capability; do you emit OpenTelemetry-compatible traces with the generative-AI conventions; can I bring my own identity provider and scoped credentials; and can I run your interface against a different backend for testing. A vendor that answers all five reasonably is not necessarily the best vendor, but it is the one whose advantage has to be earned every quarter rather than collected as rent — and in a field moving this fast, that is the property that matters most.

Migration: the project nobody scopes

Interoperability is tested when you actually move something, and the work is predictable even though teams keep being surprised by it. A migration between providers or between a provider and self-hosted infrastructure has five phases, and the cost is dominated by the ones that do not involve the model.

1. Freeze the portable artifacts. Export interactions, documents, tool definitions, traces, and the evaluation set in documented formats, and verify the export by loading it into a second system. Export paths that have never been exercised are export paths that do not work. This is a one-time cost of a few days and it should be paid before a migration is needed, not during one.

2. Re-baseline the evaluation. Run the suite against the new backend and compare per-task rather than in aggregate, because the differences are the project plan. Anything the new configuration does worse becomes a work item; anything it does better becomes an opportunity. This step is cheap exactly in proportion to how good your evaluation was, which is the recurring theme of this series.

3. Re-tune the interface layer. Prompts, output schemas, tool descriptions, and retry policies. Assume all of it needs revisiting and budget accordingly; the parts that turn out to work unchanged are a pleasant surprise rather than a baseline.

4. Rebuild the model-specific state. Embeddings under a new encoder, memory summaries under a new model, fine-tunes against a new base. This is the phase with the longest tail, and the reason part 43's advice about versioning the embedding model is really advice about migration cost.

5. Run both, then switch. A shadow period where the new path processes recorded traffic, followed by a canary on live traffic, followed by a cutover with the old path kept warm for a defined window. The old path being warm is what turns a bad discovery in week two into an afternoon rather than a quarter.

Two closing observations about the whole subject. First, the organizations best positioned to exploit a better or cheaper model are the ones whose systems were built to be portable, and that position is accumulated through many small decisions rather than purchased at migration time. Second, the standardization wave is not a threat to differentiation — it is a reallocation of where differentiation lives. When everyone can connect tools and agents the same way, the remaining competitions are over the quality of the tools, the data behind them, and the evaluation discipline that tells you which combination actually works. All three of those are yours.

Works Cited

Akhtar, Mubashara, et al. "Croissant: A Metadata Format for ML-Ready Datasets." arXiv, 2024, arxiv.org/abs/2403.19546. Accessed 7 May 2026.

Anthropic. "Introducing the Model Context Protocol." Anthropic, 25 Nov. 2024, anthropic.com/news/model-context-protocol. Accessed 7 May 2026.

Fielding, Roy T. "Architectural Styles and the Design of Network-Based Software Architectures." Dissertation, University of California, Irvine, 2000. Accessed 7 May 2026.

Google. "Announcing the Agent2Agent Protocol (A2A)." Google for Developers Blog, 9 Apr. 2025, developers.googleblog.com/en/a2a-a-new-era-of-agent-interoperability/. Accessed 7 May 2026.

Hardt, Dick, editor. "The OAuth 2.0 Authorization Framework." RFC 6749, Internet Engineering Task Force, 2012, datatracker.ietf.org/doc/html/rfc6749. Accessed 7 May 2026.

Mitchell, Margaret, et al. "Model Cards for Model Reporting." Proceedings of the Conference on Fairness, Accountability, and Transparency, 2019, pp. 220–229. Accessed 7 May 2026.

Open Container Initiative. "OCI Image Format Specification." Open Container Initiative, 2017, github.com/opencontainers/image-spec. Accessed 7 May 2026.

OpenAPI Initiative. "OpenAPI Specification Version 3.1.0." OpenAPI Initiative, 2021, spec.openapis.org/oas/v3.1.0. Accessed 7 May 2026.

OpenTelemetry. "Semantic Conventions for Generative AI Systems." OpenTelemetry Documentation, 2024, opentelemetry.io/docs/specs/semconv/gen-ai/. Accessed 7 May 2026.

Postel, Jonathan B. "DoD Standard Transmission Control Protocol." RFC 761, Internet Engineering Task Force, 1980, datatracker.ietf.org/doc/html/rfc761. Accessed 7 May 2026.

World Wide Web Consortium. "Verifiable Credentials Data Model 1.1." W3C Recommendation, 2022, w3.org/TR/vc-data-model/. Accessed 7 May 2026.