How AI governance works
What is prompt injection, and how do you defend against it?
Short answerPrompt injection is an attack where text an AI system reads, typed by a user or hidden in an email, web page, or document, changes what the system does. OWASP ranks it first among risks to LLM applications (LLM01:2025). Language models read instructions and data as one stream of text, so no filter stops every attempt. The defense is to limit what a fooled model can do: least-privilege tools, human approval for consequential actions, separated untrusted content, checks on output, and logs.
Prompt injection is entry LLM01 in the OWASP Top 10 for LLM Applications 2025, the first item on the list. OWASP describes it as prompts that alter a model's behavior or output in unintended ways, and notes the text does not need to be visible to a person, only readable by the model. MITRE ATLAS, the public knowledge base of attacks on AI systems, catalogs it as technique AML.T0051, LLM Prompt Injection: inputs that cause the model to ignore parts of its original instructions and follow the attacker's instead.
What is the difference between direct and indirect prompt injection?
The difference is who writes the instruction and how it reaches the model. In a direct injection the person typing is the attacker. In an indirect injection the attacker plants instructions in content the assistant reads later on someone else's behalf: an email it summarizes, a web page it browses, a shared document, or a file pulled in by a search. NIST observes that indirect attacks are mounted by a third party, and that the assistant's own user is often the one harmed.
| Type | Who writes the instruction | How it reaches the model | Typical goal |
|---|---|---|---|
| Direct (ATLAS AML.T0051.000) | The user of the assistant | Typed into the chat or API request | Get output the system is set up to refuse, or reveal its configuration |
| Indirect (ATLAS AML.T0051.001) | A third party | Hidden in an email, web page, document, or database record the assistant retrieves | Steer the assistant against its own user: leak data, plant false answers, trigger actions |
| Triggered (ATLAS AML.T0051.002) | A third party, planted in advance | Activated by a user action or event inside the victim's environment, often aimed at agents | Wait for the right moment to act with the agent's permissions |
Is a jailbreak the same as prompt injection?
A jailbreak is one kind of prompt injection. OWASP defines jailbreaking as prompt injection that makes the model disregard its safety protocols, and NIST classes a jailbreak as a specific kind of direct prompting attack. In practice the target differs. A jailbreak goes after the model's safety training, to get content it would normally decline to produce. Prompt injection in general goes after the application around the model: its data, its tools, and the actions it is allowed to take. A jailbroken chatbot says something harmful; an injected agent with email access does something harmful. This page describes the categories and leaves out working attack text.
Why can't filtering stop prompt injection?
SQL injection was brought under control with parameterized queries, which keep data and commands in separate channels so the database never runs user input as a command. The UK National Cyber Security Centre (NCSC) explained in December 2025 why that fix has no equivalent here: a language model draws no boundary between instructions and data inside a prompt. Everything is text, and the model predicts the next token from all of it. NCSC calls the model an inherently confusable deputy and advises reducing the likelihood and impact of injection, since the risk itself does not go away.
The other primary sources agree. OWASP writes that it is unclear whether fool-proof prevention exists. NIST AI 100-2 E2025 states that current mitigations do not offer full protection and recommends designing systems on the assumption that injection succeeds whenever a model reads untrusted input. Microsoft's Security Response Center describes indirect injection as an inherent risk of probabilistic models and defends against it in layers. A filter is a classifier, and an attacker can rephrase until the wording falls outside what it learned. Filters belong in the stack as one layer among several.
What are real examples of prompt injection attacks?
- EchoLeak, Microsoft 365 Copilot (June 2025). Researchers at Aim Security showed that a crafted email sitting in an Outlook inbox could carry hidden instructions. When the user later asked Copilot an ordinary question, Copilot retrieved the email alongside internal data and could send that data out through a link. Microsoft rated it CVE-2025-32711, critical at CVSS 9.3, fixed it on the service side, and reported no evidence of exploitation.
- Slack AI (August 2024). A researcher showed that a message posted in a public channel could instruct Slack AI, when it answered another user's question, to surface data from that user's private channels behind a phishing link. Slack deployed a patch on August 20, 2024 and reported no evidence of unauthorized access to customer data.
Both follow the same pattern: the assistant read an attacker's text in the same context as private data, and it had a route to send something outward. Removing either condition breaks the attack, which is why the defenses below concentrate on access and outbound routes.
How do you defend against prompt injection?
Every framework cited here reaches the same design rule: assume some injections succeed, and limit what a successful one can do.
| Defense | What it means in practice | Where it comes from |
|---|---|---|
| Least privilege for tools | Give the model only the tools and data the task needs. Keep credentials in application code, out of the model's reach. When the model reads outside content, drop its permissions to match the author of that content. | OWASP LLM01, NCSC |
| Human approval for actions | A person confirms anything consequential: sending email, moving money, deleting records, changing access. | OWASP LLM01, Microsoft MSRC |
| Separate untrusted content | Mark external content as data and tell the model to ignore instructions inside it. Route untrusted sources through a model with fewer permissions or a narrow interface. | Microsoft MSRC (spotlighting), NIST AI 100-2 |
| Output checks | Define the expected output format and validate it with ordinary code. Block known exfiltration routes such as auto-loading images and untrusted links. | OWASP LLM01, Microsoft MSRC |
| Logging and monitoring | Record prompts, responses, tool calls, and API calls. A run of failed tool calls can be the first sign of an attack. | NCSC |
| Adversarial testing | Test the system regularly as an attacker would, treating the model as an untrusted user. | OWASP LLM01 |
The first two rows carry the most weight. NCSC recommends deterministic safeguards, meaning ordinary code and permissions that do the same thing every time, to constrain what the system can do. A model that can read a stranger's email and also send email on your behalf needs either fewer permissions or a person in between. More on where those controls sit is in what AI enforcement at runtime looks like.
What is a system prompt, and does it protect against injection?
A system prompt is the set of instructions a developer gives the model before any user message: its role, its rules, and what to refuse. OWASP lists a well-written system prompt as a first mitigation. It sits in the same stream of text as the attack, so it shapes behavior and enforces nothing. Treat the system prompt as guidance and put enforcement in code outside the model: permissions, approvals, and output checks.
What is an LLM firewall?
LLM firewall is a market term for a service that inspects prompts and responses and blocks ones that look malicious. Microsoft's Prompt Shields is one documented example of the classifier approach. It is a useful detection layer with the same limit as any filter: a new phrasing gets through. Pair it with the permission and approval controls above. Definitions of the surrounding terms are in the AI glossary.
Frequently asked questions
- What is prompt injection in simple terms?
- Prompt injection is when text an AI reads makes it do something its owner did not intend. The text can come from the person typing or be hidden in an email, web page, or file the AI reads for them.
- What is the OWASP definition of prompt injection?
- OWASP lists prompt injection as LLM01:2025, the first risk in its Top 10 for LLM Applications. It defines the vulnerability as prompts that alter a model's behavior or output in unintended ways, and splits it into direct and indirect injection.
- What is MITRE ATLAS AML.T0051?
- AML.T0051 is the MITRE ATLAS technique ID for LLM Prompt Injection. Its sub-techniques are Direct (AML.T0051.000), Indirect (AML.T0051.001), and Triggered (AML.T0051.002), as listed on October 8, 2026.
- How does prompt injection through an untrusted email work?
- An attacker sends an email containing instructions aimed at the AI assistant, often formatted so a person skims past them. When the assistant later reads the inbox to answer a question, it treats those instructions as part of its task. The EchoLeak flaw in Microsoft 365 Copilot (CVE-2025-32711) worked this way.
- Why does "ignore previous instructions" work on AI models?
- A language model reads its developer's instructions and the incoming text as one stream, with no hard boundary between them. Blocking that one phrase does little, because the same request can be reworded endlessly.
- What does jailbroken AI mean?
- A jailbroken AI is a model that has been talked out of its safety rules and produces content it is trained to decline. OWASP treats jailbreaking as a form of prompt injection, and NIST classes it as a direct prompting attack.
- Can prompt injection be fully prevented?
- No. OWASP, NIST, the UK NCSC, and Microsoft each say current defenses do not stop every attempt. The working approach is to limit what a fooled model can reach and to require a person's approval for consequential actions.
Sources
- OWASP GenAI Security Project, LLM01:2025 Prompt Injection. Read October 8, 2026.
- MITRE ATLAS, LLM Prompt Injection (AML.T0051) and sub-techniques AML.T0051.000, .001, and .002. Read October 8, 2026.
- NIST, AI 100-2 E2025, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (March 2025), sections 3.3 and 3.4. Read October 8, 2026.
- UK National Cyber Security Centre, Prompt injection is not SQL injection (it may be worse), December 8, 2025. Read October 8, 2026.
- Microsoft Security Response Center, How Microsoft defends against indirect prompt injection attacks, July 29, 2025. Read October 8, 2026.
- CVE Program, CVE-2025-32711, M365 Copilot Information Disclosure Vulnerability. Read October 8, 2026.
- The Hacker News, Zero-Click AI Vulnerability Exposes Microsoft 365 Copilot Data Without User Interaction, June 12, 2025. Read October 8, 2026.
- Slack, Slack security update, August 21, 2024. Read October 8, 2026.
Reviewed