AI-agent threat-model workflow
- Threat model an AI agent by first defining its goals, scope, assets, permissions, and architecture; then map the agentic loop and trust boundaries; identify AI- and agent-specific threats across prompts, model, memory, tools, orchestration, and connected systems; prioritize risks; and implement layered preventive, detective, and reactive controls. Conventional data-flow threat modeling remains useful, but it must be extended to cover autonomy, dynamic reasoning, tool chains, persistent memory, and inter-agent interactions.
- Create a system and data-flow diagram, using a standard threat-modeling tool where helpful.
- Combine conventional techniques such as STRIDE with agent-specific attack trees, abuse cases, or a framework such as MAESTRO or OWASP guidance.
- Constrain autonomy and tool permissions; use separate non-human identities and narrowly scoped credentials.
- Require human confirmation for irreversible or high-impact actions and implement circuit breakers or quarantine paths.
What an AI-agent threat model must cover
Threat model an AI agent by defining its goals, scope, assets, permissions, and architecture; mapping its agentic loop and trust boundaries; identifying threats across prompts, model, memory, tools, orchestration, and connected systems; prioritizing risks; and applying layered preventive, detective, and reactive controls. For a broader foundation, see our AI threat modeling guide.
Conventional data-flow threat modeling remains useful, but an AI-agent threat model must also address autonomy, dynamic reasoning, tool chains, persistent memory, chained actions, dependencies, human approvals, and inter-agent interactions. Model the complete agentic system—not only the language model or API.
The workflow is: define the system, map its architecture and data flows, identify threats and trust-boundary failures, prioritize risks, control high-impact paths, then validate, monitor, and reassess the model.
Step 1: Define the agent’s purpose, scope, and operating context
Start with what the AI agent is intended and authorized to do. Record its purpose, users, operating context, autonomy, capabilities, decision points, permitted actions, and excluded components.
- List the data the agent can read, create, retain, retrieve, or modify.
- Inventory its models, prompts, tools, APIs, external services, frameworks, and dependencies.
- Document human approval points and actions requiring explicit authorization.
- Identify affected assets, including data, business processes, accounts, and external services.
- Record the identities, authorization limits, and consequences associated with each action.
Make controls proportional to actual permissions and business impact. An agent that produces drafts has a different risk profile from one that can change records or invoke consequential external actions. For the broader risk-assessment process, see our AI agent risk assessment framework.
Step 2: Map the architecture, agent loop, and data flows
Create a system and data-flow diagram before enumerating threats. Show how information enters the system, influences decisions, reaches tools, and returns as an output or action. For tools that can support this security work, see our guide to AI agent security tools.
Include these components
- User inputs, external content, system instructions, prompts, retrieved context, and outputs.
- The foundation model or other reasoning and decision-making components.
- Orchestration, planning, workflow logic, and action execution.
- Short-term or persistent memory, vector stores, and other retained data.
- Tools, APIs, databases, external services, and dependencies.
- Human review, approval points, downstream systems, and multi-agent communication.
Trace the agentic loop
Follow the path from user input and prompt construction through reasoning, planning, tool selection, execution, returned results, memory updates, and final output. Show whether retrieved or retained information can influence a later decision or action.
Mark where data is transformed, stored, validated, trusted, or passed to another component. A standard diagramming or threat-modeling tool can help document the system, but the diagram must reflect the agent’s actual workflow.
Step 3: Mark trust boundaries, identities, permissions, and control points
Mark every location where data, authority, or responsibility changes directly on the architecture and data-flow diagram. For a control model built around verifying each boundary and request, see Zero Trust for AI agents.
Track untrusted content and retained data
Identify where untrusted content can influence:
- Prompts or instructions.
- Retrieved context and memory.
- Agent decisions, plans, or reasoning paths.
- Tool parameters and external actions.
- Outputs passed to users, systems, or other agents.
Document who or what can write to memory, retrieve it, validate it, correct it, or remove it. In multi-agent systems, treat agent-to-agent communication as another trust boundary: record which agents can send instructions, delegate work, access data, or trigger actions.
Also mark tool and API boundaries, user and service identities, privilege changes, approval gates, validation points, and logging points. Use narrowly scoped credentials and separate non-human identities where appropriate. For a deeper treatment of agent identities and permissions, see our guide to identity and privilege abuse.
Step 4: Identify threats across the agentic system
Assess conventional software threats alongside agent-specific threats. For each threat, trace the affected component, data flow, decision point, possible action, and consequence. Keep threats separate from the controls that address them.
Prompts, data, and memory
- Direct prompt injection: supplied instructions attempt to alter the agent’s intended behavior.
- Indirect prompt injection: external or retrieved content influences instructions or decisions.
- Data poisoning: manipulated or unreliable data affects reasoning or outputs.
- Memory poisoning: malicious or incorrect information enters retained memory and influences later decisions.
- Model weaknesses: the reasoning component produces an unsafe, incorrect, or unauthorized result.
Tools, permissions, dependencies, and agents
- Unsafe tool use: the agent invokes a tool or service incorrectly or without sufficient authorization.
- Authorization failures and excessive permissions: the agent accesses or changes more than its intended scope.
- Identity spoofing or impersonation: an incorrect identity is trusted during an agent or service interaction.
- Data leakage: sensitive information reaches an unauthorized user, tool, service, or output.
- Dependency and supply-chain risks: models, tools, libraries, frameworks, or connected protocols introduce unwanted exposure or behavior.
- Multi-agent manipulation: a rogue or compromised agent influences another agent or delegates an unintended task.
Use abuse cases or attack paths to follow threats through chained actions. For example, assess whether poisoned retained information could influence a later plan or whether an incorrect tool decision could produce an unintended external action. Keep testing within authorized, controlled environments.
Step 5: Use suitable threat-modeling methods
Use each method as a starting point, then complete the AI-agent threat model with system-specific analysis.
- Data-flow diagrams expose components, information movement, trust boundaries, and control locations.
- STRIDE can organize conventional software and service threats.
- MITRE ATT&CK can organize relevant adversary behaviors where applicable.
- The Microsoft Threat Modeling Tool can help document a system-specific model.
- OWASP resources and the CSA MAESTRO framework can focus analysis on autonomy, memory, tools, and interactions.
Use these resources alongside architecture-specific analysis, abuse cases, and attack paths. Extend conventional modeling to cover dynamic reasoning, changing workflows, persistent memory, tool chains, autonomy, and inter-agent communication.
Step 6: Prioritize and document risks
Create a threat register that connects every identified threat to a component, consequence, control, owner, and validation record.
Threat-register fields
- Affected component, data flow, trust boundary, or decision point.
- Threat description, possible attack path, consequence, and affected asset.
- Existing safeguards, remaining exposure, and proposed preventive, detective, or reactive controls.
- Responsible owner, status, validation evidence, and reassessment date.
Prioritize by likelihood, impact, exploitability, autonomy, and permission level. Give early attention to sensitive assets, irreversible actions, broad permissions, weak approval boundaries, and threats that cross multiple chained systems.
Use a consistent organizational method for ranking risks, but do not treat a numerical score or threshold as universally sufficient. Explain each priority and revisit it when the agent’s capabilities or context changes.
Step 7: Design layered mitigations
Connect each threat to controls that match the agent’s architecture, permissions, data, and business impact. Use defense in depth rather than relying on a single prompt rule or model safeguard.
Preventive and constraining controls
- Apply least privilege, narrowly scoped credentials, and separate identities for distinct responsibilities where appropriate.
- Restrict tools, APIs, data access, and delegation to what the task requires.
- Validate and filter inputs, retrieved content, tool parameters, and outputs.
- Use guardrails and runtime authorization checks at consequential control points.
- Use isolation or sandboxing where appropriate to the agent’s risk and execution environment.
Detection, approval, and response controls
- Require human confirmation for irreversible or high-impact actions.
- Log prompts, relevant decisions or traces, memory access, tool calls, privilege changes, approvals, and outputs as appropriate.
- Monitor policy violations, anomalous actions, unexpected tool activity, and unusual agent interactions.
- Provide blocking, quarantine, rollback, suspension, or shutdown paths appropriate to the system’s risk.
| Threat category | Relevant controls | Validation or monitoring signal |
|---|---|---|
| Prompt injection or poisoned content | Input validation, filtering, guardrails, constrained instructions, and human review for high-impact actions | Unexpected instruction changes, policy violations, or unusual outputs |
| Memory poisoning | Controlled memory writes, validation, scoped access, and correction paths | Unexpected memory updates or decisions based on questionable retained data |
| Unsafe tool use or excessive permissions | Least privilege, scoped credentials, runtime authorization, and approval gates | Unexpected tool calls, denied actions, privilege changes, or anomalous activity |
| Identity, dependency, or multi-agent risks | Identity controls, dependency review, communication boundaries, logging, and monitoring | Unexpected agent communication, delegation, or service interaction |
Step 8: Validate with authorized adversarial scenarios
Validate whether the documented controls work in controlled, authorized environments. Test scenarios involving:
- Direct and indirect prompt attacks.
- Data or memory poisoning.
- Incorrect or unauthorized tool use.
- Authorization failures and excessive permissions.
- Agent-to-agent manipulation or unintended delegation.
- Unexpected behavior during chained or dynamic workflows.
Record the scenario, affected flow, observed behavior, control response, residual risk, and corrective action. Tailor scope and success criteria to the agent’s actual permissions, workflow, and impact level.
Step 9: Monitor, govern, and reassess continuously
Threat modeling is a lifecycle activity, not a one-time document. Monitor prompts, relevant decisions or traces, outputs, policy violations, tool calls, memory activity where appropriate, privilege changes, and anomalous behavior.
Revisit the model when the agent’s tools, prompts, models, memory, permissions, dependencies, or workflows change. Also reassess when it gains autonomy, connects to another system, adds an identity, or enters a multi-agent workflow.
Use monitoring findings to update the architecture diagram, threat register, control mapping, validation record, and approval boundaries. Controls should evolve with actual behavior rather than remain fixed because the original design was approved.
Practical AI-agent threat-model deliverables
A complete AI-agent threat-modeling review should produce:
| Artifact | Purpose |
|---|---|
| System inventory and scope statement | Defines purpose, users, autonomy, assets, permissions, dependencies, and boundaries. |
| Architecture, agent-loop, and data-flow diagram | Shows inputs, prompts, reasoning, orchestration, memory, tools, outputs, approvals, services, and agent communication. |
| Trust-boundary, identity, permission, and control-point map | Shows where validation, authorization, approval, and monitoring are required. |
| Threat register and control mapping | Connects threats and consequences to safeguards, owners, status, and residual risk. |
| Validation record | Documents adversarial scenarios, observed results, control responses, and corrective actions. |
| Monitoring and reassessment plan | Defines relevant signals, review responsibilities, change triggers, and reassessment records. |
A framework, template, or tool can organize these artifacts, but it does not replace system-specific analysis. Update the model whenever autonomy, tools, memory, permissions, dependencies, or workflows change, and keep controls proportional to the actions the agent can actually take.



