Trust is earned, not given

A different perspective

2025-12-04 · Projects

AI Frontiers, part 58: Data residency and sovereignty — where inference is allowed to run

Part 58from the AI Frontiers series · 65 parts in all

The question used to be simple: choose a cloud region, tick a box, sign a data processing agreement. By 2025 it had become one of the first questions in an enterprise AI deal and one of the hardest to answer, because inference does not behave like storage. It reads your data, produces derived data, may log both, may be served by a provider that is itself a sub-processor, and may route the request across continents for capacity. Every one of those steps is a data flow with a legal address, and the dropdown that says "Region: Frankfurt" does not describe any of them.

This entry maps the problem: the legal stack that governs where computation may happen, the places prompts actually travel that teams do not expect, the deployment options that resolve the issue honestly versus the ones that merely relabel it, and what to build if you are going to serve customers in regulated jurisdictions. It is the compliance-shaped companion to the on-device argument in part 46 and the licensing material in part 37.

The legal stack, in one place

Four layers govern the answer, and they interact.

Data protection law. For European personal data, the General Data Protection Regulation (European Parliament and Council, "Regulation (EU) 2016/679") sets the baseline: lawful basis, purpose limitation, minimization, processor contracts, and — the operative restriction for cloud AI — limits on transfers of personal data to third countries. The Court of Justice's Schrems II decision invalidated the previous adequacy arrangement for United States transfers and required an assessment of whether the destination jurisdiction's law undermines the protections, including the possibility of government access (Court of Justice of the European Union). The European Data Protection Board's supplementary-measures guidance followed, effectively requiring organizations to document a transfer impact assessment and, in some cases, to apply technical measures such as encryption with keys held outside the receiving jurisdiction (European Data Protection Board). The newer standardization of contractual clauses gave teams a template but not a free pass: the assessment still has to be done.

AI-specific regulation. The EU AI Act adds obligations layered on top of data protection, including requirements that apply to providers and deployers of higher-risk systems and expectations about record-keeping, transparency, and human oversight (European Parliament and Council, "Regulation (EU) 2024/1689"). Its residency implications are indirect but real: obligations attach to roles, and a provider that cannot demonstrate where processing happened struggles to demonstrate compliance.

National localization and access laws. Some jurisdictions require data to remain within borders or impose conditions on export; others grant government access to data held by entities they regulate, which is precisely the conflict Schrems II identified. The general academic literature on data localization describes a spectrum from permissive to strictly restrictive regimes and is the fastest way to build a mental map of the variation (Chander and Lê).

Sector and contractual rules. Health, financial services, and public-sector procurement each add their own regimes, and enterprise contracts add more: a customer may require that no personnel outside a specified set access their data at any point. These contractual terms are frequently stricter than the law and are what actually bind implementation.

Where prompts actually go

The gap between a data-flow diagram and reality is where most residency incidents originate. Seven flows to enumerate explicitly, in rough order of how often they are missed.

Retention and abuse monitoring. Human review of flagged interactions is standard practice at API providers, which means the data can reach people you did not think of as processing it, in jurisdictions you may not have considered. Ask specifically: is there an option for zero retention, and what does it exclude?

Observability and tracing. The trace store from part 49 is, by design, the richest repository of user data in the system. If it ships to a third-party platform in another region, the residency property of your model deployment is irrelevant. This is the most common self-inflicted violation I have seen.

Embeddings and retrieved content. A vector index built over customer documents is derived personal data sitting in a database that someone chose for latency reasons. Vector stores are frequently overlooked in the residency review because they are infrastructure, not "data."

Evaluation sets and trace-derived training data. The loop in part 49 that turns production traces into test cases and fine-tuning datasets is excellent engineering and a residency hazard, because it copies production data into places with different controls.

Support access. Break-glass engineering access does not respect regions unless the platform enforces it. Ask who can read the logs and from where.

Backups and disaster recovery. A cross-region replica for availability is a cross-border transfer, whether or not it is ever restored.

Sub-processor chains. The model provider may use an inference host, a GPU cloud, a logging service, and an annotation vendor. Each is a downstream processor with its own location, and the provider's list of them changes without asking you.

The options, ranked by honesty

On-device or customer-premises inference. The only architecture that answers the question completely, because the data never leaves the customer's control. Covered in part 46, including its capability ceiling and operational cost. For the highest-sensitivity workloads this is increasingly the standard answer, and the small-model improvements of the last two years are what made it viable rather than token.

Self-hosted open weights in a region you control. You choose the hardware location, the retention policy, the encryption keys, and the access controls, and you can produce evidence for each. The cost is real: a platform team, capacity planning, and a capability lag behind the frontier that you must decide is acceptable. This is the configuration most regulated deployments end up at, and the licensing term that makes it legal is the subject of part 37.

Sovereign cloud offerings. Providers built regional offerings with in-region operations, in-region support staff, and often third-party attestations — Germany's C5 attestation and France's SecNumCloud qualification are the two most commonly cited — to make the jurisdictional claim verifiable rather than aspirational. These are genuine products with genuine premium pricing, and their principal appeal is that a customer's legal team has something to point at. The engineering caveat is that the service catalog in a sovereign region lags the global one, sometimes by a year, and the feature you need may not be there.

Regional endpoints with contractual guarantees. A model provider offering a regional endpoint with commitments on processing location, zero retention, and published sub-processors. This is the pragmatic mainstream answer, and its quality depends entirely on how specific the commitments are. "Data is processed in the EU" is weak if support staff and the observability pipeline are elsewhere. "Data is stored and processed in the EU; no human review; sub-processor list published; zero retention available" is a claim you can assess.

Region-agnostic with consent. Sometimes the right answer is to tell the user where their data goes and get explicit consent. That works for consumer products and fails for procurement-driven enterprise deals, which is worth knowing before designing around it.

What to build

Four artifacts turn this from a recurring argument into a routine.

A data-flow map with locations. Every store, every third-party call, every log sink, with a jurisdiction and a retention period. It should be generated from infrastructure configuration rather than written by hand, because hand-written maps are wrong within a quarter.

Residency as a configuration property. Region selection, endpoint choice, logging destination, key location, and encryption boundary modelled as configuration with a single tenant-level setting. Retrofitting this after launch is the expensive version of the work, and I have watched three teams do it the expensive way.

Keys held where the data must stay. Customer-managed keys, held in a key management service under the customer's control or in the required jurisdiction, are the strongest technical measure available for the transfer problem, because the ciphertext is useless to whoever holds it. This is also the specific measure European regulators have pointed to as a supplement to contractual protections.

Evidence, produced continuously. Logs of where each request was processed, attestation records, sub-processor change notifications, and access reviews. Regulated customers do not accept architecture diagrams; they accept evidence, and the teams that win these deals are the ones that can produce it in an afternoon.

The honest trade-off

Residency costs capability, latency, and money, and pretending otherwise wastes everyone's time. The frontier model you want may not be available in the region you need; the sovereign region may lack the vector store you standardized on; the latency budget may not survive a regional round trip; and the operational burden of self-hosting falls on one or two people who become a single point of failure.

What makes the trade manageable is deciding it per workload rather than per company. A support assistant summarizing public documentation has different constraints from a clinical transcription pipeline, and a single platform-wide answer is usually either over-restrictive for the first or non-compliant for the second. Classify the data, classify the workload, and route accordingly — which is, incidentally, the same routing discipline that pays for itself in cost and quality, as the next entry discusses. The places where computation happens are increasingly part of the product, and the teams that treat location as a first-class design variable rather than an infrastructure afterthought are the ones that can sell into every market rather than one.

A worked example: clinical documentation in the EU

Concrete cases make the abstract stack legible, so here is one I have watched designed. The product drafts clinical notes from a recorded consultation for a European hospital group. The data is special-category personal data under the regulation, the customer is a public hospital with a procurement process, and the failure mode that loses the deal is a data flow the customer's legal team did not approve.

The first decision was architectural rather than contractual: speech recognition and note drafting run on hardware inside the hospital network, using open-weight models sized to the available accelerators. That choice, which the on-device and self-hosting material in parts 46 and 37 made feasible, removes the transfer question for the content itself and converted a two-year legal negotiation into a deployment project. What remained was everything else, and everything else is where the work was.

The telemetry had to be rebuilt. The tracing platform that had been defaulting to a hosted collector was replaced with a self-managed one inside the same boundary, and the sampled traces stored at full fidelity were reduced to a redacted subset with a defined retention window. The evaluation pipeline had to stop shipping raw clinical text to a hosted judge model; the replacement used a locally hosted model for routine grading with human review on a sample, which improved both the compliance posture and, unexpectedly, the relevance of the rubric. Embeddings for the retrieval index over clinical guidelines moved into the same boundary, along with the vector store's backups. Audit logs stayed in-region and became an artifact the customer could inspect, which turned out to be the single most persuasive item in the procurement pack.

The capability cost was real and worth naming honestly. The local models were roughly a generation behind the frontier on the hardest summarization cases, the customer accepted that explicitly in exchange for the residency property, and a routing layer sent anything that fell outside the local models' competence to a human rather than to a remote model. The engineering team spent more effort on the redaction, retention, and evidence machinery than on the models, which is the pattern in every regulated deployment I have seen.

Three transfers from that project. Residency is cheapest to satisfy when it is an architectural property rather than a contractual one. Your observability stack is a data flow and must be in scope from day one. And the deliverable that closes deals is not an architecture diagram but a set of evidence a customer's counsel can audit — which, conveniently, is the same set of evidence your own incident response will want.

Works Cited

Chander, Anupam, and Uyên P. Lê. "Data Nationalism." Emory Law Journal, vol. 64, no. 3, 2015, pp. 677–739. Accessed 4 Dec. 2025.

Court of Justice of the European Union. "Data Protection Commissioner v. Facebook Ireland Ltd and Maximillian Schrems, Case C-311/18." Court of Justice of the European Union, 16 July 2020. Accessed 4 Dec. 2025.

European Data Protection Board. "Recommendations 01/2020 on Measures That Supplement Transfer Tools to Ensure Compliance with the EU Level of Protection of Personal Data." European Data Protection Board, 2021, edpb.europa.eu. Accessed 4 Dec. 2025.

European Parliament and Council. "Regulation (EU) 2016/679 on the Protection of Natural Persons with Regard to the Processing of Personal Data and on the Free Movement of Such Data (General Data Protection Regulation)." Official Journal of the European Union, 2016, eur-lex.europa.eu/eli/reg/2016/679/oj. Accessed 4 Dec. 2025.

European Parliament and Council. "Regulation (EU) 2024/1689 Laying Down Harmonised Rules on Artificial Intelligence (Artificial Intelligence Act)." Official Journal of the European Union, 2024, eur-lex.europa.eu/eli/reg/2024/1689/oj. Accessed 4 Dec. 2025.

Microsoft. "EU Data Boundary." Microsoft Learn, 2025, learn.microsoft.com/en-us/privacy/eudb/eu-data-boundary-learn. Accessed 4 Dec. 2025.