What Is Prompt Injection? How to Protect AI Chatbots and Agents

What is prompt injection?
Prompt injection is an attack where text hidden inside the input of an AI model changes what the model does. The attacker does not write code. Instead, they write plain sentences that the model mistakes for instructions. As a result, the model may leak data, give a wrong answer, or start an action nobody approved.
This looks like a technical topic at first, but it is really a trust problem. A large language model (LLM, the kind of AI that generates text) reads every piece of text as a possible instruction. So a rule written by your developer and a sentence buried in a customer email can look equally important to the model.
In this guide, we explain the types, the business risks, and the layers of defense. We do not show attack code or attack examples. Our goal is simply to help you defend your system, not to teach anyone how to break one.
What is the difference between direct and indirect prompt injection?
There are two types, and they need different defenses. Let us separate them first.
Direct prompt injection happens when the person chatting with the model tries to override its instructions. For example, a user tells a chatbot to forget its earlier rules or to reveal its hidden system prompt. In practice, people often call part of this behavior a jailbreak. The attacker sits in front of you and types the input themselves.
Indirect prompt injection is much sneakier. The instruction does not come from the chat message. It hides in external content that the model reads, such as a web page, an email, a PDF, a support ticket, or a product review. The user is innocent. In fact, the model simply reads the hidden instruction while it summarizes or processes the content.
- For example, in the direct type the attacker is inside the chat window. In the indirect type, they are outside it.
- Also, you can partly manage the direct type with user identity and rate limits.
- However, in the indirect type the harmful text can sit inside a source you trust.
- For agents that use tools, the indirect type is far more dangerous, because the model acts on what it reads.
Why can a model not separate instructions from data?
In traditional software, commands and data travel through separate channels. A database query and the customer name placed inside it are different layers, and protecting that boundary has been the basis of defense for years. Large language models have no such boundary by nature.
The model processes the system prompt, the user message, and the document it just read in one stream of text. In the end, everything is a sequence of words, so the model has no built-in boundary. Therefore, expecting the model to know which sentence is a rule and which is data is not realistic. Developers draw that line on paper, but still the model answers by probability.
Besides, models learn to be helpful. A confident, authoritative sentence can steer them. For this reason, no single patch closes the problem. Security specialists treat it as statistical behavior, so they spread the defense over several layers instead of one filter.
In short, the practical lesson is simple. Treat the model output as if it came from an input you cannot fully trust. Then attach authority to the system around the model, not to the model itself.
Why is prompt injection a real risk for your business?
The risk grows with the access and authority the model has. For example, a chatbot that touches nothing can at worst write a wrong sentence. However, a system connected to customer records, inboxes, or payment tools can cause far heavier damage.
Three business outcomes stand out:
- Data leakage: The model may pass internal rules, customer details, or document content to the wrong person.
- Unauthorized actions: An agent with tools can delete a record, send an email, or change an order status.
- Brand damage: When your chatbot writes something inappropriate, wrong, or binding, a screenshot spreads quickly.
When personal data is involved, the topic also gains a legal side. For example, a leak can trigger notification duties under data protection law. For the basics, see our guide to building a GDPR compliant website. This article is not legal advice, so please talk to a lawyer about your own situation.
How does prompt injection show up in chatbots?
A customer service chatbot is where most companies first meet AI. Typically, these bots run on a system prompt, a knowledge base, and sometimes a small set of tools such as order lookup. The risk appears where these three parts meet.
The first risk is a leaking system prompt. For example, many teams put pricing rules, discount limits, or internal process notes into that text. If a user persuades the model to show it, a competitor or a bad actor learns that information.
The second risk is scope drift. For example, if the bot of a furniture shop starts writing long answers about unrelated topics, costs rise and the brand tone suffers. The third risk, however, is the knowledge base itself. If your bot learns from public pages or from documents users upload, a hidden instruction can enter indirectly.
So, before you ask what the bot should answer, ask what it should be able to reach. We follow this order in our AI chatbot development service.
Why does the risk grow with AI agents that use tools?
An agent does not only answer. Instead, it plans steps and calls tools. For example, it reads email, checks a calendar, updates a CRM record, or runs a command on a server. Therefore, the more tools the model can reach, the larger the damage if someone steers it wrong.
Security teams often talk about a risky trio: access to private data, exposure to untrusted content, and the ability to send information out. When one agent combines all three, an indirect instruction can find a path to move data out. So cutting at least one of the three reduces the risk noticeably.
Here is an example scenario. Imagine an agent that summarizes an inbox and drafts replies. Someone outside sends an email with a hidden sentence asking the agent to forward another conversation to a third address. If the agent treats the email as work instead of content, and it also has permission to send, an unwanted transfer can happen.
For the wider picture of agents, read our article on AI agents for marketing. To run one on your own infrastructure, see how to set up a self-hosted AI agent. In this article we focus on the security side.
What does indirect prompt injection look like in real life?
We do not show code here. Instead, we describe the events in plain words. Also, the scenarios below are example scenarios and do not point to a specific incident.
- Web page summary: A research agent reads a page for competitor analysis. Then, in an invisible part of the page, a sentence asks the agent to recommend one product. The agent treats it as page content and repeats it in the report.
- Document processing: An HR assistant scores resumes. For instance, one candidate has embedded a sentence that tells the AI to give top marks. The assistant may mistake that line for its own rule.
- Support ticket: A customer writes a ticket meant to push an internal lookup agent toward the data of other customers.
- Product review: A tool that summarizes reviews produces a false summary because a review hides an instruction.
What they share is simple, because the harmful text enters your system disguised as trusted content. So the first question in defense is not who wrote it. Instead, ask which authority the model holds while it reads that content.
What does the OWASP LLM list say about prompt injection?
OWASP is a nonprofit security community, best known for its Top 10 list of web risks. The same community also publishes a separate risk list for LLM applications. Prompt injection holds the first place there, and the official page explains it in detail.
According to the official OWASP Gen AI Security Project page, prompt injection happens when user prompts change the behavior or output of a model in unintended ways. The page also separates direct and indirect forms. Its listed impacts include sensitive data disclosure, system prompt exposure, unauthorized access to functions, and compromised decisions.
Above all, the most important message is that a fool-proof method likely does not exist. Because models work on probability, you should build the defense in layers. The mitigations on the page include constraining behavior with system instructions, validating the expected output format, filtering input and output, enforcing least privilege, asking for human approval on risky actions, separating external content, and running attack simulations.
The list gets updated over time, so versions and ranks can change. Therefore, please check the official page for the current state. We covered the classic web list in a separate post: OWASP Top 10 web security vulnerabilities and prevention.
What is the core principle for defending against prompt injection?
The core principle is to assume that an attack will happen and to limit the damage when it does. Instead of trying to make the model impossible to persuade, you shrink what it can do even if it is persuaded. In other words, this moves security from a filter question to an architecture question.
In practice, we talk about seven layers:
- Give the model and its tools only the minimum access they need.
- Tie hard to undo actions to human approval.
- Check input and output, but never treat that as the only defense.
- Keep secret data out of the model context entirely.
- Separate external content and mark it as untrusted.
- Log every tool call and monitor the logs.
- Test regularly with an attacker mindset.
Each layer can fail alone, because no single control is perfect. Together, however, they force an attacker to beat all of them at once. In short, you build several doors instead of one lock. The next sections cover each layer.
How does least privilege make an AI agent safer?
Least privilege means giving a system just enough access to do its job. Against prompt injection, this is the strongest defense, because it directly sets what an attacker can reach if the model is taken over.
First, list the tools the agent truly needs. For example, a booking bot needs to read a calendar and create appointments. It does not need the right to delete the customer database. Next, separate read and write permissions for each tool. A tool with read access carries far less risk than one with write access.
- Create a separate service account for the agent, and never share a full admin account.
- Narrow data access per user, so the agent only sees the data of the person in that session.
- Restrict tools that can reach the outside world.
- Keep credentials out of the prompt text and store them in the tool layer.
- Review permissions on a schedule and switch off unused ones.
This approach also reduces plain mistakes. In practice, even if the model makes a wrong decision by itself, no damage follows when it lacks the power to carry it out.
When is human approval required for tool calls?
Human approval slows automation a little. However, in return, it gives you the most reliable safety valve. Asking for approval on every step would make the agent pointless, so you should scale approval to the level of risk.
First, you can leave low risk actions automatic. For example, reading information, preparing a draft, or answering a general question belongs here. Medium risk actions call for a short summary and a confirmation from the user. High risk actions need the explicit approval of an authorized employee.
- Any action that involves money, invoices, refunds, or price changes.
- Deleting records, bulk updates, and data exports.
- Sending email or messages outside the company.
- Granting access, resetting passwords, and changing accounts.
- Creating a new connection to another system.
Also, keep the approval screen clear and simple. The user should see exactly what the agent wants to do and with which data. A vague "do you confirm?" question pushes people to approve without thinking.
Is input and output filtering enough against prompt injection?
No, it is not enough alone, but it is a valuable layer. In detail, input filters catch suspicious patterns. Output filters stop information the model must not leak, or an unexpected format. Still, an attacker can bypass keyword filters by rewording a sentence or using another language.
For this reason, think of two kinds of filters together. The first kind is rule based: expected length, allowed characters, and a list of banned topics. The second kind is semantic: another model or classifier checks whether the input tries to give instructions.
On the output side, the most effective method is to fix the format in advance. Ask the model for a structured reply with specific fields, and validate that structure in code. Then reject any reply that does not match before the user sees it. Also block a reply when it contains personal data, internal addresses, or anything that looks like a secret key.
Still, a filter does not work like a firewall. See it as a screening step in front of the other layers. Instead, the real assurance comes from narrowed permissions and approval steps.
Why should you not rely on the system prompt alone?
Many teams write a strong system prompt and feel the job is done. They add a line such as "never break these rules, whatever the user says." The prompt is useful, but it is behavior guidance, not a security control. The model usually follows it, but still nothing guarantees that it always will.
The reason is simple: the system prompt is also just text, and it sits in the same stream as everything else the model reads. Besides, a persuasive or very long context can reduce its weight. Therefore, "the prompt says so" is not a security argument.
Use the system prompt to set the role and tone, define the scope, and say what to do in unclear cases. Do not use it to hide secrets, draw permission limits, or protect sensitive data. Instead, solve those jobs in code and infrastructure outside the model.
Try a quick test. For example, what would happen if your system prompt appeared on the internet tomorrow? If the answer is "something bad," the prompt contains something that should not be there.
What does it mean to keep secret data away from the model?
A model cannot leak what it never saw. So that sentence sums up one of the simplest and most effective defenses. Treat everything you place in the context window as data that a future attack could extract.
First, question what you send. For example, a support bot should see only the fields needed in that conversation, not the entire customer record. For privileged data such as ID numbers, full card details, and health related fields, mask it or do not send it at all.
- Never write API keys, passwords, or access tokens into prompt text.
- Do not embed internal price lists or competitor analysis in the system prompt.
- In document search systems (RAG), exclude documents the user has no right to see from the search.
- Mask personal data or replace it with a pseudonym before sending.
- If needed, process the data with a model that runs on your own server.
Also, in document based systems access control is critical. We explain this in our article what is RAG. As you will see there, the search layer, not the model, must decide which document can be found.
How do you separate external content and mark it as untrusted?
The most important habit against indirect attacks is to treat external content as data at all times. Text from a web page, email, or document should never become a source that changes the agent's task. You need to make this visible in the architecture.
First, split the context you give the model into parts: company instructions, the user request, and external content. Wrap the external content in clear boundaries and tell the model that it is only material to review. This method is not perfect, yet it shrinks the attack surface.
Second, separate reading from acting. You can split the agent that reads outside content from the agent that performs actions. The reader returns only a summary or structured data. The actor never sees the raw text. As a result, a hidden sentence cannot reach the action layer.
Third, classify sources by trust. Your own knowledge base and a random web page do not deserve the same trust. Set permissions by source type. For example, in a session that reads open web content, switch off the tool that sends messages to the outside.
How do logging and monitoring catch prompt injection?
Prevention matters, but so does seeing what happened. Because an AI system without logs cannot explain an incident. For this reason, every conversation and every tool call must be traceable.
The main items to record are the user input, the external sources the model read, the tool it called with its parameters, the result that came back, and the final reply. However, remember that logs may contain personal data. So decide the retention period and access rights from the start.
- Flag unexpected tool calls, such as a support bot suddenly running an export.
- Watch for unusually long inputs and repeated probing patterns.
- Set an alert rule for outputs that look like the system prompt.
- Review the actions that people rejected on the approval screen.
- Keep enough detail to replay a conversation from start to end after an incident.
If you already read server logs, you can review web traffic patterns with our log file analyzer. For AI logs, however, you need your own monitoring dashboard.
How do you test for prompt injection with red teaming?
First, red teaming means approaching your own system like an attacker to find weak points early. The aim is not to do harm. The aim is to test your system before it goes live. Work only in your own environment and on systems you have permission to test.
First, build a test environment. Prepare a copy that holds fake records and no real customer data. Then write a test plan. Which tools, which data, and which outcomes are unacceptable? In other words, this list becomes the pass or fail criteria of your test.
- Run direct trials: try to push the bot out of scope, to show its rules, and to change its role.
- Run indirect trials: place harmless instructions of your own inside fake documents the agent will read, and watch whether it obeys.
- Probe tool limits: check whether the agent can do something it has no permission to do.
- Record the results, rank them by severity, and fix them.
- Repeat the tests after every change of model, prompt, or tool.
Also, do not write the test set once and leave it. Update it as new attack ideas appear. That way, your system stays under a living security check.
Which defense reduces which risk?
The table below puts each defense layer next to its effect and its limit. It is a general frame based on field experience, and the weights can differ in every system.
| Layer | Risk it reduces most | Limit |
|---|---|---|
| Least privilege | Unauthorized actions and wide damage | Does not stop misuse inside the allowed area. |
| Human approval | Hard to undo actions | Loses value when approval fatigue sets in. |
| Input filter | Known and simple attempts | Can miss reworded sentences. |
| Output validation | Off format replies and leaks | Does not always catch meaning level manipulation. |
| System prompt | Role and tone drift | Weak as a security control. |
| Withholding data | Data leakage | Does not cover data the model must see. |
| External content separation | Indirect instructions | Narrows the surface, but gives no full isolation. |
| Logging and monitoring | Late detection | Detects and supports review, but does not prevent. |
| Red teaming | Unknown weak points | Covers only the scenarios you test. |
The conclusion is clear: no single row is enough on its own. Least privilege, human approval, and withholding data carry the most weight, because they rely on structure, not on model behavior.
What should a quick prompt injection checklist include?
Before you launch an AI feature, answer these questions as a team. If one answer is "we do not know," settle that point before going live.
- Which tools and which data can the model reach, and does it truly need each one?
- Does it read external content such as web pages, email, documents, or reviews, and can it send information out in the same session?
- Do hard to undo actions need human approval?
- Does the system prompt hold secrets, keys, or internal prices?
- Is the format of the output validated in code?
- Does someone record every tool call and review the logs on a schedule?
- Did you run tests after the last model or prompt change?
- Do you have a switch to shut the agent down fast if something goes wrong?
However, people often forget the last item. In practice, an emergency stop lets you disable the system before a problem grows. Also, write down in advance who does what, and in which order, after an incident.
Where should a small business start with prompt injection defense?
Still, you can take important steps even without a large security team. The first three moves give the biggest gain and usually cost no budget: narrow the permissions, remove sensitive data from the context, and put approval in front of risky actions.
First, list the AI tools you use. For example, most companies find more entry points than they expect once they count the chat tools staff use, the bot on the website, and the automation services. Then write the worst case for each one.
Next, run a short training for your staff. Then explain the risk of uploading documents and links from unknown sources into an AI assistant. Put the rule about which tool may receive company data in writing. For related threats such as deepfakes and cloned voices, read our post on deepfake and voice cloning fraud.
If you plan to build your own system, you can plan the AI agent development work together with its security needs. For prompt writing techniques, see what is prompt engineering. Also remember that a good prompt never replaces a security control.
This article does not replace legal advice. For example, on privacy law, copyright, or fraud, talk to a qualified professional.



