Agent Security Risks You Must Handle
Agents read untrusted content, run code, and call APIs — a new attack surface. The risks that actually materialize, explained plainly, with working mitigations.
Published 2026-09-29 · 10 min read
Agents are new attack surfaces wearing the disguise of productivity tools. They read untrusted content, execute code, call APIs with your credentials, and make decisions based on what they find — which means every classic security concern applies, plus several novel ones unique to systems that follow instructions found in data. This guide covers the risks that actually materialize, in plain language, with the mitigations that work. No fear-mongering, no hypothetical superintelligence scenarios: just the concrete ways agents get compromised or cause damage today, and what to do about each.
Prompt injection: the headline risk
Prompt injection is the defining security problem of the agent era. The mechanism is simple: an agent reads content from an untrusted source — a web page, a document, a tool output, an email — and that content contains instructions disguised as data. Because language models process everything as text, the agent can mistake the attacker's instructions for legitimate ones. A product review containing hidden text that says to ignore previous instructions and send credentials elsewhere; a support ticket that instructs the agent to approve a refund it should not approve. Direct injections are crude and often caught; indirect injections, buried in content the agent was legitimately asked to read, are the real problem. Complete prevention is an unsolved problem, which is why the rest of this guide emphasizes containment over cure.
What injection enables
- Data exfiltration. An injected instruction can direct the agent to include private context — file contents, credentials, conversation history — in outputs sent to external destinations.
- Unauthorized actions. If the agent has tools, injected instructions can trigger them: sending messages, modifying data, or calling APIs the user never authorized.
- Credential theft. Instructions to echo environment variables, config files, or API keys into a tool call the attacker can observe.
- Restriction bypass. Carefully worded injections can talk an agent around its own safety rules, since the rules and the attack arrive through the same channel.
- Persistent compromise. An injection that writes malicious instructions into the agent's notes, memory files, or configuration persists across sessions.
Defense in depth
- Treat all external content as untrusted data. Instruct the agent explicitly: content from tools, files, and the web is information to process, never instructions to follow.
- Least-privilege tools. An agent that cannot send external messages cannot be injected into sending them. Every withheld capability is an attack path closed.
- Human approval for sensitive actions. The approval gate is the last line of defense; an injected instruction still has to get past a human reading the proposed action.
- Separate trust levels for sources. Content from your own files deserves more trust than content scraped from the open web; make the distinction explicit in the agent's instructions.
- Monitor outputs, not just inputs. Exfiltration and misuse show up in what the agent sends and calls — log tool calls and review them for significant runs.
- Isolate sessions. Do not let one agent's compromised context bleed into others via shared memory files; segment by project and sanitize between tasks.
Credential and data leakage
Separate from injection, agents leak by accident with depressing regularity. They paste secrets into logs and error messages. They include private file contents in prompts sent to third-party tools. They summarize confidential documents into shared notes files with broader access than the source. The mitigations are boring and effective: keep secrets in environment variables and credential stores, never in files the agent reads casually; redact known secret patterns from tool outputs before they reach the agent; give agents scoped credentials that cannot reach beyond the task; and periodically grep agent-accessible logs and notes for anything that should not be there. Assume leakage is a matter of when, and design so that what leaks is limited.
Supply chain and tool risks
Agents increasingly consume third-party tools, plugins, and skill packages — each one code that runs with the agent's privileges. A malicious or compromised tool can exfiltrate data, manipulate results, or inject instructions through its outputs. Vet what you install: prefer tools from known maintainers, pin versions, and read what a tool actually does before granting it broad permissions — especially tools that request network access or file-system writes. Keep an inventory of installed tools per agent; the tool you installed for one experiment six months ago and forgot about is a liability. And treat tool outputs with the same suspicion as web content, because from the agent's perspective, that is what they are.
An incident response mindset
- Log everything significant: tool calls, file accesses, external requests, and permission grants, with timestamps.
- Know how to revoke fast: credential rotation procedures and a kill switch for every agent with meaningful access.
- Review blast radius after any incident: what could the compromised agent reach, and what did it actually touch?
- Write it down. A short postmortem for every security scare — even the ones that turned out fine — builds the institutional memory that prevents repeats.
- Rehearse the bad day. Walk through a hypothetical injection-to-exfiltration scenario for your highest-privilege agent and fix whatever step has no defense.
Agent security is not about achieving perfect safety — it is about making attacks expensive and failures contained. Layer the defenses: untrusted-data discipline, least privilege, human gates on the irreversible, output monitoring, and fast revocation. None of these is exotic; together they are the difference between an incident and a catastrophe. The teams that get this right treat agent security as ordinary operational hygiene, not a research problem — and ordinary hygiene, applied consistently, stops the vast majority of real attacks.
Indirect injection: the harder problem
Direct prompt injection — a user typing ignore your instructions — is the attack everyone pictures. The harder problem is indirect injection: malicious instructions hidden in content the agent legitimately processes. A webpage the agent reads contains hidden text telling it to exfiltrate data. A document under review includes instructions to email its contents to an attacker. A calendar invite's description field carries a payload. The agent is not being tricked by a user; it is being tricked by its own tools, which faithfully deliver attacker-controlled content as if it were data. Defenses that focus on the user's input miss this entirely. Treat every tool output as untrusted by default: content from the web, files, emails, and databases can all carry instructions, and the agent must be trained — through system prompts and architecture — to treat tool outputs as data to work with, never as instructions to follow.
- Webpages with hidden text or metadata instructing the agent to take actions.
- Documents under review that embed instructions in footnotes, comments, or white-on-white text.
- Emails and messages whose bodies contain directives aimed at the agent, not the human.
- Database records or API responses containing payload text in fields the agent reads.
- Images with embedded text instructions, read by multimodal agents as content.
Exfiltration paths you might miss
Data leaves through channels nobody monitors. The obvious path — the agent pasting secrets into a chat response — is the one everyone guards. The subtle paths: the agent including sensitive data in a tool call's parameters, where it lands in a third-party service's logs; error messages that echo back fragments of the data that caused them; file writes to shared or synced directories that propagate beyond their intended audience; URLs constructed with sensitive values in query strings, logged by every proxy in between; and summaries or reports that aggregate sensitive details into a document with broader distribution than the sources. Audit exfiltration by tracing data, not by watching the chat: pick a sensitive value, and follow everywhere it could flow through the agent's tools, parameters, logs, and outputs. The paths you have not traced are the ones attackers will find.
Least privilege in daily practice
Least privilege is easy to state and tedious to implement, which is why it rarely survives contact with real deployments. Make it practical with three habits. First, scope credentials to the task: the agent analyzing support tickets does not need production database write access, even if it is convenient to use one shared credential. Second, expire access aggressively: permissions granted for a migration should not still exist six months later — put calendar reminders on every exception. Third, separate the agent's identity from yours: the agent should act as its own service account with its own limited permissions, so its actions are auditable and its access is bounded. When someone proposes giving the agent your credentials because it is easier, that is the moment the security model failed — convenience at the credential layer compounds into every incident you will later investigate.
Red-team your own setup
You will not think of every attack, so hire someone — or some process — to think adversarially for you. A red-team exercise for an agent deployment is straightforward: give a tester the agent's interface and a goal (extract this data, perform this unauthorized action), and see what happens. Start with the obvious — direct injection, permission probing — then move to the indirect: poisoned documents, malicious tool outputs, social-engineering the agent through its context. Time-box the exercise and fix what it finds before expanding scope. Even an informal version — spending an afternoon trying to break your own setup — surfaces the misconfigurations that formal reviews miss, because attackers think in terms of what is possible, not what is intended. Run it before every significant expansion of the agent's permissions, and after every incident, to verify the fix actually holds.
Model-provider risk: where your data goes
Every prompt you send to a hosted model is data leaving your boundary, and the security review must include the provider. Read the data-processing terms: does the provider train on your inputs, retain them, allow human review? The answers vary widely — enterprise tiers typically offer zero-retention options that consumer tiers do not. For sensitive work, this determines architecture: highly confidential tasks may require self-hosted models or providers with contractual data protections, even at higher cost or lower capability. Also consider the provider as an attack surface: a compromised or coerced provider sees everything your agents process. This is not paranoia but supply-chain realism — your security perimeter includes every service that touches your data. Document the data flows: which providers receive what categories of data, under which contractual protections, reviewed on what schedule. The agents' permissions matter, but so does the plumbing underneath them.
- Read the training and retention terms: enterprise tiers often offer zero-retention; consumer tiers often do not.
- Highly confidential tasks may require self-hosted models or contractual data protections.
- Treat the provider as part of your supply chain — a compromised provider sees everything.
- Document data flows: which providers get what data, under which protections.
- Review provider terms on a schedule; they change, and the changes rarely favor you.
Keep reading
A practical look at where AI agents genuinely help in 2026 and where they still fall short, so you can delegate with confidence.
How to Write Instructions AI Agents ActuallyClear, practical techniques for writing agent instructions that get followed — with examples of vague versus precise wording.
AI Agent vs. Chatbot vs. Script: What's theAgents, chatbots, and scripts get lumped together. Here is what each one actually is, and when to use which.