Prompt Injection Is a Supply Chain Problem
Treating prompt injection as an AI quirk leads to prompt patches. Treating it as untrusted input reaching a privileged interpreter leads to the architecture you already know how to build.
I spent an afternoon trying to write a system prompt that would make an agent refuse instructions found inside a retrieved document. I produced about four hundred words of careful prose. Then I read a page it had fetched and found the phrase "ignore previous instructions" sitting in a comment near the bottom, and the four hundred words did nothing at all.
That is the moment the framing changes. Prompt injection is not a model deficiency you can talk your way past. It is untrusted data reaching an interpreter with privileges, and we have a mature discipline for exactly that.
Why the prompt patch cannot work
An LLM has no principled boundary between "instructions I was given at setup" and "text I just read." Both arrive as tokens in one sequence. There is no privilege level attached to the system prompt that the model can enforce, because the model is not the one enforcing anything — your application is.
trusted: system prompt written by your team
untrusted: document contents, web pages, issue text, email bodies, tool output
the model sees: one continuous string
your app must enforce: the separation it cannot
Every "ignore instructions embedded in documents" rule is a request to the model to distinguish those two by appearance. An adversary only has to make theirs look like yours.
Reframe it as data reaching an interpreter
Compare the threat to a web application that concatenates a request parameter into a SQL statement. The fix was never "write a smarter query." The fix was a boundary: parameterised queries, least-privilege database accounts, network rules that stop the database from reaching the internet.
The LLM architecture has the same four controls available.
Parameterisation means the document is passed as data in a clearly delimited slot, never spliced into the instruction channel. The agent's own instructions live in code and configuration the operator controls.
Least privilege means the agent's tools do what the task needs and nothing more. An agent that reads tickets should not also have a tool that sends email or executes shell commands. If retrieval-only agents cannot exfiltrate, most injection stories end at the read step.
Egress control is the one teams skip. If the agent's network access is restricted to an allowlist of your own endpoints, a successful injection has nowhere to send what it stole. This control is unaffected by how persuasive the injected text was.
Provenance means marking which parts of the context are untrusted and logging them, so that when an agent does something odd you can see which retrieved chunk was in the window at the time.
None of these require the model to understand trust.
The supply chain analogy is not decorative
Retrieved documents, MCP tool responses, issue comments, and vendor API payloads are all dependencies: they arrive from outside your build, are updated without your review, and execute with your privileges. So apply the same practices — pin what you consume, prefer an allowlist of sources over an open crawl, scan before ingest, and treat a change in a tool's output format as a dependency bump rather than a runtime surprise.
An agent that can fetch arbitrary URLs is an application with an unsandboxed package manager and outbound network access. We would not ship that.
Detecting is weaker than containing
Classifiers for injected text help at the margin, but they are a tripwire, not a boundary — an attacker optimises directly against them, and a false negative lands in the privileged channel with no second check behind it. Use them to alert, not as the control. The reliable version is architectural: if the worst successful injection can only produce a malformed ticket comment, the system is safe regardless of detection quality.
What to build first
Inventory which text reaches your agent and which tools it can call. The answer to "could a malicious document do damage here?" is decided entirely by that second list. Then cut tools until a retrieved document cannot, on its own, cause an irreversible action.
That is the same instinct behind supply-chain scanning of container images: you do not make the base image trustworthy, you make an untrustworthy base image unable to become a running system without passing gates.
Summary
Prompt injection cannot be patched at the prompt, because the model cannot enforce a boundary between setup instructions and retrieved text. Enforce it in your application: pass documents as data, restrict tools to the task, control egress so exfiltration has nowhere to go, and log provenance. Treat every external text as an unpinned dependency, and design so that a successful injection produces a logged failure rather than a privileged action.
SDP Clouds Team
DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations — every article is based on real incidents and real pipelines, not docs-page rewrites.
More about us →