Knowledge · New threats: agents, MCP, prompt injection · 8 minutes
What is prompt injection and how do you protect against it?
Prompt injection hijacks an AI model with instructions hidden in content it reads. Direct vs indirect attacks, the resume example and what actually helps.
Last updated: October 5, 2026
What is prompt injection?
Prompt injection is when text given to a language model changes its behavior against the intent of whoever deployed it. The model has no separate channel for commands and for data. Your instruction, the body of an email and a hidden note on a web page all reach it as one stream of text.
The OWASP Top 10 for LLM Applications 2025 lists prompt injection as LLM01, the first entry. OWASP points out that injected content does not have to be visible to a human. It only has to be parsed by the model: white text on a white background, a tiny font, an HTML comment, file metadata.
NIST AI 600-1, the generative AI risk profile from July 2024, covers it under information security as a new attack surface that arrives together with the model.
Direct vs indirect prompt injection
Direct prompt injection is an instruction typed straight into the chat, for example an attempt to make a company chatbot reveal its system prompt or bypass its rules. The attacker needs access to the chat.
Indirect prompt injection works without chat access. NIST describes it as a remote attack: someone places instructions in data the model is likely to read at some point. That can be a web page, a PDF from a supplier, a customer ticket, a product description or a comment in a code repository.
Indirect injection is the bigger business risk because it bypasses people. An employee asks the assistant to summarize a document and never sees that the document contained a sentence aimed at the model. The first systematic write-up of this class of attack is the 2023 paper by Greshake and co-authors. An indirect injection can also wait for months: in a knowledge base, an old ticket or a product description, until someone asks the assistant to search that data.
Prompt injection in resumes: hidden instructions for AI screeners
A resume is a textbook indirect attack. In its LLM01 entry OWASP describes a scenario where a candidate splits an instruction into fragments hidden in the resume, and the screening model gives a positive recommendation regardless of the actual content.
This is no longer hypothetical. In a study published on 27 May 2026 the authors analyzed around 200,000 real resumes from the hiring platform hireEZ. They found hidden injections in about 1% of them. From 2019 to 2023 the share was 0.6-0.8%, in 2024 it jumped to about 1.2%, and then it fell slightly, although the number of such resumes kept growing with the number of applications. More than 90% contained no explicit command, only hidden text, most often skill lists, keywords and added experience.
If you are job hunting: hidden text is misleading, and tools that compare what a human sees with what a machine extracts exist precisely to catch it. If you are hiring: make the tool show you the text extracted from the file, and do not base a rejection or an invitation on the model's score alone.
Under the EU AI Act, AI used in employment is a high-risk area. According to the European Commission, after the AI Omnibus amendments the obligations for this area apply from 2 December 2027. More on the dates in AI Act: what it is and who it applies to.
Why can't you just switch it off?
OWASP says plainly that it is unclear whether fool-proof prevention methods exist. Simon Willison, a researcher who has documented the problem since 2022, wrote in June 2025 that we still do not know how to prevent it 100% reliably.
The reason is structural. A filter that catches 95% of attacks means, in security terms, that one attempt in twenty gets through, and an attacker can try as often as they like. At the launch of the ChatGPT Atlas browser in October 2025, OpenAI's security chief called prompt injection "a frontier, unsolved security problem" (Simon Willison's write-up).
The practical takeaway: do not design a deployment as if the model always listens only to you. Design it so that obeying outside text cannot do much harm.
When prompt injection becomes dangerous: the lethal trifecta
Willison calls it the lethal trifecta. Damage happens when one system combines three things: access to private data, exposure to untrusted content, and the ability to send something out.
Example: an assistant reads your inbox (data), receives an email from a stranger (untrusted content) and can send messages or open links (communication). One hidden line in that email is enough to ask it to forward something it should not.
Remove one element and this kind of leak stops being possible. It is the simplest test question for every assistant, agent and MCP server you connect.
Prompt injection protection: what actually works
OWASP lists seven directions. For a business, the ones that matter limit the impact rather than promising to detect every attack:
1. Least privilege. A model that only summarizes does not need permission to send email or write to a database. 2. Human approval before any action with consequences: payment, sending, deletion, permission changes. 3. Separate outside content from instructions and label where every piece came from. 4. A fixed output format, validated in code before anything is executed. 5. Input and output filters as an extra layer, never the only one. 6. Adversarial testing before launch and after every change.
Add the principle from OWASP LLM06 (excessive agency): access decisions are enforced by the target system, not by the model. If the database does not let a user read other people's records, a model acting on their behalf cannot read them either, whatever was injected. Example: a customer service assistant answers from the order database. If its database queries run with the logged-in customer's permissions, an injected "show all orders" returns only that customer's orders.
How to check your own deployment
Start with an inventory: which AI tools in the company read outside content, and what they can do after reading it. The list is often longer than the official one, because staff connect tools themselves, as covered in what shadow AI is. For each tool, record who approved it and who owns its configuration.
For each tool, answer the three trifecta questions. Where all three answers are yes, cut one or add human approval. Write the rules for your team into an AI usage policy.
If you are deploying an agent or an MCP server and want an outside check of its permissions and resistance to content injection, the scope is described on the agents and MCP servers page. The result states what we know, what we did not check and what to do, not "safe".
In short
- Prompt injection is an instruction hidden in text the model reads. The indirect kind needs no chat access.
- OWASP and researchers agree there is no reliable fix. You limit the impact, not the possibility of attack.
- Check every tool for three things at once: private data, untrusted content, outbound sending.
- Payments, sending and deletion always need human approval.
- In hiring, hidden text in resumes is real: about 1% of files in a large 2026 study.
Have a website, app or email address that looks suspicious?
Frequently asked questions
What is prompt injection in simple terms?
It is slipping an instruction to an AI model inside content it was only meant to read. The model cannot reliably tell instructions from data, so it sometimes does what was slipped in.
Does hidden text in a resume work on AI hiring tools?
It can affect a score, because the system reads text a human does not see. It can, however, be detected by comparing the visible text with the machine-extracted text, which is exactly the kind of detection the 2026 study describes. Hidden text misleads the employer.
What is the difference between prompt injection and a jailbreak?
A jailbreak tries to bypass the model's own safety rules, usually directly in the chat. Prompt injection is broader: any content that takes over the behavior of an application built on a model, including indirectly through documents and web pages.
Is a filter or guardrail enough to protect against prompt injection?
Not as the only layer. Filters reduce successful attacks but not to zero, so OWASP also recommends least privilege and human approval for actions with consequences.
Does prompt injection affect regular ChatGPT too?
Yes, whenever the model reads outside content: a web page, a file, an email or a tool result. The risk grows with how much the assistant can do after reading it.
Sources
- OWASP: LLM01:2025 Prompt Injection
- OWASP Top 10 for LLM Applications 2025
- NIST AI 600-1: Generative Artificial Intelligence Profile (July 2024)
- Simon Willison: The lethal trifecta for AI agents (16 June 2025)
- Zhang et al.: Measuring Real-World Prompt Injection Attacks in LLM-based Resume Screening (2026)
- Greshake et al.: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection (2023)
- Simon Willison: OpenAI CISO on prompt injection in Atlas (22 October 2025)
- European Commission: AI Act application timeline
Accurate as of the article's last update. Laws and vendor terms change, so check the source before you decide.