Knowledge · New threats: agents, MCP, prompt injection · 8 minutes
What is an AI agent and how do you use one safely at work?
An AI agent takes actions, not just answers. How it differs from an assistant, OWASP risks, agentic browsers, permissions, human approval and logs.
Last updated: October 5, 2026
What is an AI agent?
An AI agent is a program in which a language model gets a goal, tools and permissions, and then decides on the next steps by itself to reach that goal. It can search an inbox, open a web page, fill in a form, save a file or call a company system without you approving every move.
The OWASP Top 10 for Agentic Applications, published on 9 December 2025, describes agents as systems that plan, hold memory, call tools and act with delegated authority. Agents often get their tools through MCP servers.
AI agent vs AI assistant: what is the difference?
An AI assistant answers and a human acts. You ask for a draft email, you get text, you send it yourself. An assistant's mistake ends as bad text you can fix.
An AI agent acts by itself. You ask it to reply to overdue emails and the agent reads, writes and sends them. An agent's mistake is a sent email, an executed payment or a deleted file.
The line is blurry, because the same product can work in both modes. For security one question matters: can the system do something in the world without your click? If yes, treat it as an agent. Workplace examples: a bot that answers customers and issues refunds on its own, a code editor assistant that runs terminal commands, an automation that reads invoices and prepares transfers for approval.
AI agent security: what can go wrong
The first three entries of the OWASP agentic list are a good risk map. ASI01, agent goal hijack: text the agent reads changes what it is trying to achieve. This is prompt injection with consequences. ASI02, tool misuse: the agent uses tools it is allowed to use in harmful ways, in loops or at scale. ASI03, identity and privilege abuse: the agent acts with broader rights than anyone meant to give it.
The earlier OWASP LLM list (LLM06, excessive agency) names three sources of harm: too much functionality, too many permissions and too much autonomy. Each one can be reduced through configuration before anything goes wrong.
The key point: an agent does not need to be "hacked". It only has to read a malicious document, page or email and treat a sentence in it as an instruction. So the question at deployment is not "is the agent safe" but "what happens if it does the worst thing its permissions allow".
Agentic browsers: Comet, ChatGPT Atlas and web content
An agentic browser is one in which the model can click, scroll and fill in forms by itself on sites where you are logged in. It is the most direct contact between an agent and untrusted content, because every page it visits is input for the model.
On 20 August 2025 Brave described a vulnerability in Perplexity's Comet browser: instructions hidden in a Reddit comment, read when the user asked for a page summary, made the agent fetch the user's email address, trigger a one-time code, read it from the logged-in Gmail inbox and post both in a reply. Brave recommended separating user instructions from page content, human confirmation for sensitive actions, and isolating agent mode from regular browsing.
For ChatGPT Atlas, OpenAI's help center describes agent mode safeguards: no running code or downloading files, a logged-out mode that does not use your accounts, and a mode where on sensitive sites such as banks you must keep the tab in front and watch. In December 2025 the company wrote that prompt injection, like scams on the web, is unlikely ever to be fully solved (OpenAI).
In practice: use logged-out mode for tasks on unfamiliar sites. Do not leave the agent alone on your bank, email or admin panels.
Least privilege in practice
Least privilege means the agent gets exactly what a specific task requires and nothing extra. Review those permissions regularly, because agents usually gain new tools faster than anyone removes old ones. In practice:
1. A separate account or token for the agent, not an admin account and not your full account. 2. Read instead of write wherever the task allows. 3. An allowlist of addresses and recipients instead of "the whole internet". 4. Acting in the context of the user who requested the task, so the agent sees no more than that person. OWASP LLM06 also recommends that access is decided by the target system, not by the model.
Check against Simon Willison's lethal trifecta: private data, untrusted content and the ability to send things out in one agent is a recipe for a leak. Cut at least one.
Human in the loop: payments, sending and deletion
Human in the loop means the agent only prepares certain actions and carries them out after a human approves them. OWASP (LLM01 and LLM06) and Brave both recommend it.
A minimum list of actions to approve: every payment and change to payment details, anything sent outside (email, message, post, file share), data deletion, changes to permissions and security settings, purchases and accepting terms.
Approval only makes sense when the human sees exactly what they approve: the full recipient, the amount, the content. A dialog saying "the agent wants to take an action, OK?" with no details teaches people to click without reading.
Logs and a kill switch
Agent logs should answer three questions: what it did, when, and on whose instruction, together with the content that led it there. Without that you cannot investigate an incident or tell a model error from an attack. Keep logs out of the agent's own reach so it cannot change or delete them.
A kill switch is the ability to stop the agent immediately and revoke its tokens without hunting for settings. Before launch, check who can do it and how long it takes. Also plan what happens to tasks in progress after a stop, so the agent does not leave an operation half done.
AI agents at work: where to start
1. List the agents, assistants with an action mode and agentic browsers already in use, including those staff connected on their own (shadow AI). 2. For each, record the permissions requested and needed. 3. Agree on the list of actions that need approval. 4. Write the rules into an AI usage policy. 5. Test the agent on a document with a hidden instruction before it gets real data.
If you want an outside check of permissions, approvals, action logs and resistance to content injection, the scope is described on the agents and MCP servers page. The result states what we checked, what we did not and what to fix.
In short
- An AI agent acts, an assistant only answers. If a system can do something without your click, treat it as an agent.
- An agent's safety is set by its permissions, because content it reads can take it over.
- Payments, outbound sending and deletion always need approval from a human who sees the details.
- Use agentic browsers on unfamiliar sites without being logged in to your accounts.
- Log every agent action and have a way to stop it immediately.
Have a website, app or email address that looks suspicious?
Frequently asked questions
What is an AI agent in simple terms?
It is an AI assistant that does not just suggest but carries out tasks itself: opening pages, writing and sending messages, saving files, using company systems.
What is the difference between an AI agent and an AI assistant?
An assistant gives an answer and a human takes the action. An agent takes the action itself, so its mistakes have immediate real-world effects.
Are AI agent browsers safe?
It depends on the settings and where you use them. Researchers have shown that page content can take over the agent, and OpenAI says prompt injection is unlikely ever to be fully solved. Use logged-out mode and do not leave the agent alone on banking or email sites.
What permissions should a company give an AI agent?
Only those a specific task requires, on a separate account or token. Read instead of write, an allowlist of recipients, and human approval for payments, sending and deletion.
What does human in the loop mean?
It means the agent prepares selected actions and carries them out only after approval by a human who sees the full details.
Sources
- OWASP Top 10 for Agentic Applications for 2026 (9 December 2025)
- OWASP: LLM06:2025 Excessive Agency
- Brave: Agentic Browser Security, Indirect Prompt Injection in Perplexity Comet (20 August 2025)
- OpenAI Help Center: ChatGPT Atlas, agent mode
- OpenAI: Hardening ChatGPT Atlas against prompt injection (December 2025)
- Simon Willison: The lethal trifecta for AI agents (16 June 2025)
- OWASP: LLM01:2025 Prompt Injection
Accurate as of the article's last update. Laws and vendor terms change, so check the source before you decide.