Trust is earned, not given

A different perspective

2025-11-06 · Projects

AI Frontiers, part 57: Security of agentic systems — permissions, sandboxes, and blast radius

Part 57from the AI Frontiers series · 65 parts in all

Part 34 argued that prompt injection cannot be patched, for a structural reason: instructions and data arrive on the same channel, so a system cannot reliably distinguish "content to reason about" from "orders to obey." That entry was written about a chatbot with a plugin. Two years later the same vulnerability is sitting inside systems that hold production credentials, write to databases, send email, and call APIs on behalf of real users, and its severity has changed by an order of magnitude even though the mechanism has not.

What changed is capability. A chatbot that can be tricked into saying something embarrassing is a PR problem. An agent that can be tricked into exfiltrating a customer list, wiring money, or deleting a production table is an incident. Which means agentic security is not an extension of the alignment conversation or a matter for prompt engineering; it is ordinary systems security with an unusually persuasive and unusually literal-minded component in the middle of it, and the usable defenses are the boring ones.

Why the model's judgment cannot be a security boundary

Every secure system rests on a boundary that cannot be talked across. A database enforces access control regardless of what the query intends. A shell separates code from data. A browser's same-origin policy holds whether or not a page is malicious. Language models dissolve those boundaries because they are trained to be maximally helpful and maximally responsive to whatever text is in front of them, and because the channel that carries their instructions is the same channel that carries their input.

The empirical record makes the point unambiguously. System prompts can be extracted and overridden. Instructions embedded in retrieved documents, web pages, calendar invites, source code comments, and image alt text are followed at meaningful rates (Greshake et al.). Defenses that looked promising individually — paraphrasing, spotlighting, delimiters, structured queries, instruction hierarchy — have each reduced attack success under some evaluations while remaining breakable by adaptive adversaries (Hines et al.; Chen et al., "StruQ"; Wallace et al.). The general taxonomy of adversarial risks for machine learning systems now has a formal treatment that reaches much the same conclusion (Vassilev et al., "Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations").

The correct inference is not that prompt-injection defenses are worthless. It is that they belong in the same category as input validation: useful, worth doing, and never the thing you rely on to prevent a breach. If a security review rests on the model recognizing a malicious instruction, the review has failed before it starts.

The failure taxonomy worth designing against

Indirect injection. The attacker does not talk to your agent; they plant text your agent will read. A support ticket, a shared document, a webpage the agent fetches, a dependabot diff, a poisoning of the memory store. This is the dominant real-world vector because it requires no access to the system, only to any content the system reads.

The confused deputy. The agent holds privileges the requester does not, and acts on their behalf using its own authority. Triggering an exfiltration through the agent's API credentials is confused-deputy abuse, and it is the reason per-user authorization must live in the tool layer, not in the prompt.

Data exfiltration over channels you forgot you had. An agent with a fetch tool can leak data by requesting a URL that encodes it in the path. An agent that renders images or markdown can leak it through a URL in the output. An agent that can call an arbitrary webhook can leak it anywhere. The exfiltration surface is everything the agent can reach that leaves the trust boundary, and the list is longer than the obvious tools.

Memory and state poisoning. Anything the agent persists — conversation summaries, learned preferences, a vector store of past interactions — becomes an injection carrier that survives the session and can steer future behavior. Persistent memory turns a one-shot injection into a durable implant.

Supply chain. Tool servers, plugin packages, and third-party integrations execute code and receive credentials. Treating an MCP server or an agent framework plugin as trusted because it is in the ecosystem is the same mistake that led to the npm and PyPI incidents, with a worse blast radius because the component holds your auth tokens.

Blast radius compounding. Individually benign capabilities compose into dangerous ones. Read email plus send email is phishing from the user's own account. Read files plus HTTP access is exfiltration. Read database plus write database is data corruption at machine speed. The security question is not whether each tool is safe but what the set of tools can collectively do — which is exactly the point part 47 made from the design side.

The defenses that work

None of these are novel, which is the reassuring part. They are the principles in Saltzer and Schroeder's 1975 paper — least privilege, fail-safe defaults, complete mediation, separation of privilege — restated for a component that cannot be trusted to follow instructions.

Least privilege, enforced outside the model. Each tool gets the narrowest scope that does the job: a read-only token, a single table, a specific directory, a rate limit, an amount cap. The model never sees a credential it does not need, because a credential the model can see is a credential the model can be persuaded to use.

Separate read from write, and write from irreversible. Reading is how injection enters; writing is how it becomes damage. Consequential or irreversible actions get an explicit confirmation path — a human approval, a two-step commit, a dry-run diff. The industry phrase for this is human-in-the-loop, and the design discipline behind it is the subject of a later entry.

Egress control. The most underrated control in agentic systems. An agent that can only reach an allowlisted set of hosts cannot exfiltrate to an attacker-controlled endpoint, regardless of what it decides. This is also the single control most often missing from prototypes, because development wants an open fetch tool.

Provenance labeling and channel separation. Mark retrieved content as untrusted data in a structured channel; never concatenate it into the instruction region. Spotlighting and delimiting do not make injection impossible, but they raise the cost, and they make the architecture legible to a reviewer who needs to see where instructions can come from.

Sandboxing with real boundaries. Execute generated code in a container with no network, no credentials, and a filesystem snapshot. Run the browsing agent in a session with no access to the authenticated context. This is where the LM-emulated sandbox line of work is useful both as a research environment and as a design pattern: if you cannot articulate exactly what your agent can reach, you have not designed its sandbox (Ruan et al.).

Instrumentation and detection. Log every tool call with arguments and authorization context — the trace discipline of part 49 — and alert on anomalies: a retrieval-heavy session that suddenly starts making network calls, an agent reading the same document three times, a burst of failed authorization attempts. You will not prevent every injection; you can notice the exploitation.

Testing it like an adversary

The evaluation tooling caught up in 2024 and 2025, and it is usable. AgentDojo provides a dynamic environment with realistic tools, a corpus of injected content, and a suite of attack and defense combinations measured together, which is the property that matters: a defense that reduces attack success by breaking utility is not a defense (Debenedetti et al.). InjecAgent catalogs injection cases across tool-integrated agents (Zhan et al.). The design-pattern work from mid-2025 goes further and proposes architectural templates — action-selector, plan-then- execute, dual-model separation — that constrain what an injected instruction can even reach (Beurer-Kellner et al.).

What a team should actually run: a red-team suite of at least a hundred injection cases covering each tool, each data channel, and each persistence mechanism; an adaptive attacker that knows your defenses, because the static cases will stop working once you add delimiters; and a utility measurement alongside every attack-success number. Then a quarterly rerun, because the attack surface grows every time someone adds a tool.

What to promise a customer

The governance question is harder than the technical one, because customers and regulators want a guarantee. The honest position is that injection is an accepted risk that is mitigated by containment rather than eliminated by detection, and that the mitigation is strongest where permissions are narrow, data flows are auditable, and irreversible actions have human confirmation. Promising more than that — "our system cannot be prompt-injected" — is a claim no architecture currently supports, and making it is worse than making no claim at all, because it transfers risk to the customer without their knowledge.

Two further governance habits are worth building early, because they are cheap now and expensive later. Keep a written register of every capability the system has, who approved it, and what it can reach — the register is what makes a quarterly review possible and what answers a customer's security questionnaire without a scramble. And decide in advance what the response to a suspected injection incident is: which logs are pulled, who is notified, what is disabled first, and how a customer or regulator is told. Incident response plans written in the middle of an incident are written badly. The reassuring half of the story: the standard toolchain already covers most of what is needed. Identity, scoped tokens, network policy, sandboxes, audit logs, approval workflows, and an allowlist are all mature and all familiar to any competent infrastructure team. The work of agentic security is mostly resisting the temptation to let a very capable text generator hold the keys, and then applying twenty-year-old practice to a new component.

Threat modeling an agent in one afternoon

Security work tends to arrive as a document nobody reads. What actually changes a design is a short, specific exercise, and for an agent it fits in a two-hour session with the engineers who built it. Six questions, answered in writing, with the answers attached to the pull request that adds a tool.

1. What untrusted text can this system read? Enumerate every source: user input, retrieved documents, web pages, email bodies, attachments, ticket comments, code comments, filenames, image metadata, and anything a tool returns. Every one of those is an injection channel, and the list is almost always longer than the team expected. Mark which ones cross a trust boundary — content authored by someone outside the organization.

2. What credentials does it hold, and what can each one reach? Not "what tools does it have" but the underlying authority: whose token, what scope, what resource limits, what irreversible operations are reachable through that credential. The check is uncomfortable and worth doing: for each credential, ask what the worst possible action is, and who could trigger it.

3. Where can data leave? Every outbound channel: HTTP requests, webhooks, email, rendered images and links, error messages, logs shipped to a third party. The allowlist should cover all of them, and the question to answer is whether a crafted instruction could encode sensitive data into a destination the attacker controls.

4. Which actions are irreversible? Deletions, transfers, external communications, writes to a system of record, anything with a physical consequence. Each one should either require explicit human confirmation or be guarded by a mechanism outside the model, such as a staging table and a separate commit step.

5. What does it remember, and where does that memory live? Persistent state is a durable injection carrier. Ask whether a poisoned memory can influence a future session, who can write to the store, and how a bad entry is removed.

6. If it is compromised right now, what is the blast radius, and how would you know? This is the question that produces the useful design changes. If the honest answer to the second half is "we would find out from the customer," that is the finding: add the detection before adding the capability.

Attach the six answers to the tool-provisioning review, and re-run the exercise whenever a tool is added or a permission is widened. The exercise is cheap, it is specific enough to produce changes, and it surfaces the composed-capability problems that no individual tool review will catch.

Works Cited

Beurer-Kellner, Luca, et al. "Design Patterns for Securing LLM Agents against Prompt Injections." arXiv, 2025, arxiv.org/abs/2506.08837. Accessed 6 Nov. 2025.

Chen, Sizhe, Julien Piet, Chawin Sitawarin, and David Wagner. "StruQ: Defending against Prompt Injection with Structured Queries." arXiv, 2024, arxiv.org/abs/2402.06363. Accessed 6 Nov. 2025.

Debenedetti, Edoardo, et al. "AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents." arXiv, 2024, arxiv.org/abs/2406.13352. Accessed 6 Nov. 2025.

Greshake, Kai, et al. "Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection." arXiv, 2023, arxiv.org/abs/2302.12173. Accessed 6 Nov. 2025.

Hines, Keegan, et al. "Defending against Indirect Prompt Injection Attacks with Spotlighting." arXiv, 2024, arxiv.org/abs/2403.14720. Accessed 6 Nov. 2025.

OWASP Foundation. "OWASP Top 10 for Large Language Model Applications." OWASP, 2025, owasp.org/www-project-top-10-for-large-language-model-applications/. Accessed 6 Nov. 2025.

Ruan, Yangjun, et al. "Identifying the Risks of LM Agents with an LM-Emulated Sandbox." arXiv, 2023, arxiv.org/abs/2309.15817. Accessed 6 Nov. 2025.

Saltzer, Jerome H., and Michael D. Schroeder. "The Protection of Information in Computer Systems." Proceedings of the IEEE, vol. 63, no. 9, 1975, pp. 1278–1308. Accessed 6 Nov. 2025.

Vassilev, Apostol, et al. "Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations." NIST AI 100-2, 2025, nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-2e2025.pdf. Accessed 6 Nov. 2025.

Wallace, Eric, et al. "The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions." arXiv, 2024, arxiv.org/abs/2404.13208. Accessed 6 Nov. 2025.

Zhan, Qiusi, et al. "InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents." arXiv, 2024, arxiv.org/abs/2403.02691. Accessed 6 Nov. 2025.