AI Threat Modeling in Brief
- AI threat modeling is a structured, iterative process for mapping an AI system’s assets, data and control flows, trust boundaries, actors, and attack paths; identifying and prioritizing AI-specific security and privacy threats; and applying mitigations, testing, and monitoring across the lifecycle. Start with architecture and data-flow mapping, then analyze areas such as training and retrieval data, model and embedding stores, inference endpoints, identity and tenant boundaries, agent memory, tool/API calls, and output handling. Use established methods and references such as adapted STRIDE, OWASP LLM guidance, MITRE ATLAS, and threat-modeling handbooks, while enforcing controls including least privilege, validation of model outputs, deterministic authorization for consequential actions, rate limits, red teaming, and continuous monitoring.
- Create a data-flow diagram and explicitly mark trust boundaries and privilege changes.
- Treat prompts, retrieved documents, memory, model outputs, and tool feedback as untrusted unless separately validated.
- Restrict agent identities and tools with least privilege; require deterministic policy checks before consequential actions.
- Separate tenants and sessions, protect training and vector data, and secure model registries, endpoints, and secrets.
AI threat modeling is a structured, iterative process for mapping an AI system’s assets, data and control flows, trust boundaries, actors, and attack paths; prioritizing security and privacy threats; and applying mitigations, testing, and monitoring throughout the lifecycle.
Start by defining the system and its objectives. Then map its architecture, data flows, control flows, actors, assets, privilege changes, and trust boundaries. Identify architecture-specific attack paths, prioritize risks, assign mitigations, validate controls, and reassess the model whenever the system, model, data, tools, integrations, or threat environment changes.
AI Threat Modeling in Brief
- Define the AI system’s purpose, assets, actors, trust boundaries, and security and privacy objectives.
- Map data and control flows across users, prompts, training data, retrieval or vector stores, models, memory, tools, APIs, and outputs.
- Identify and prioritize threats such as prompt injection, data poisoning, sensitive-data leakage, model extraction or inversion, excessive agency, tool misuse, insecure output handling, and resource or cost exhaustion.
- Apply layered controls, including least privilege, tenant isolation, input and retrieval controls, output validation, deterministic authorization, sandboxing, rate limits, secrets protection, and monitoring.
- Validate the deployed architecture through abuse-case testing and red teaming, then reassess it as the model, data, tools, or system change.
What Is AI Threat Modeling?
AI threat modeling is a proactive security and privacy analysis performed during design and maintained throughout the software development life cycle. It helps a team understand what could go wrong across an artificial intelligence or machine-learning system, how an attacker or untrusted input could reach an asset, and which controls should reduce the risk.
The practice applies to traditional machine-learning pipelines as well as large language model (LLM), retrieval-augmented generation (RAG), agentic, and multi-agent systems. Not every architecture contains every component or threat. The scope should follow the actual system, use case, data, integrations, users, tenants, and operating environment. For a focused workflow, read AI Agent Threat Modeling Guide.
How AI threat modeling differs from conventional application threat modeling
Conventional application threat modeling remains an important foundation. AI systems add security-relevant properties that require additional analysis:
- Model-related assets: training and fine-tuning data, labels, features, model weights, embeddings, model registries, deployment configurations, and inference endpoints may need protection.
- Changing data and control flows: prompts, retrieved content, memory, model outputs, and tool feedback can influence later processing or decisions.
- Dynamic behavior: the system may produce different outputs for similar inputs, making output validation and abuse-case testing important.
- Agent interactions: agents can interact with tools, APIs, other agents, and external integrations, creating additional trust boundaries and privilege transitions.
- Data-dependent risk: the same model architecture can have different risks depending on its training data, retrieval sources, users, tenants, and business purpose.
A generic application checklist, vulnerability scan, governance review, or framework alone does not complete an AI security threat model. Architecture-specific analysis is still required.
Step 1: Define the AI System, Assets, Actors, and Trust Boundaries
Start with system intent. Defining what the system is supposed to do prevents the review from becoming a disconnected list of generic AI risks.
Document purpose and security objectives
- The system’s purpose, intended users, and expected interactions.
- Business, security, privacy, and operational objectives.
- Owners responsible for data, models, infrastructure, tools, and integrations.
- The lifecycle stage: design, development, testing, deployment, or operation.
- Assumptions, dependencies, external services, and out-of-scope components.
Relevant objectives may include confidentiality, integrity, availability, privacy, tenant isolation, authorized use, and safe handling of consequential outputs. Make them specific to the use case rather than assuming every AI system has the same requirements.
Inventory AI assets
Inventory the components that actually exist in the architecture.
| Architecture area | Assets to protect | Questions to ask |
|---|---|---|
| Training and fine-tuning pipeline | Datasets, labels, features, pipeline code, credentials, evaluation data | Who can add, change, approve, or export training data? |
| Model and registry | Models, weights, versions, metadata, deployment configuration | Who can register, approve, replace, or retrieve a model? |
| Prompts and inference | System instructions, user inputs, sessions, requests, responses | Which inputs cross from users or external sources into the model? |
| Embeddings and vector stores | Embeddings, indexed documents, metadata, access rules | Can one user, tenant, or process reach another’s indexed data? |
| Agents and memory | Agent identities, session state, memory, plans, tool permissions | Which state persists, and who can influence it? |
| Tools, APIs, and integrations | Credentials, requests, responses, third-party data, actions | What privilege does each tool have, and what validates its use? |
| Outputs | Generated text, classifications, recommendations, downstream commands | What consumes the output if it is wrong or manipulated? |
Identify actors, privilege changes, and trust boundaries
List every actor that can influence, access, or operate the system. Depending on the architecture, these may include end users, administrators, developers, data owners, service identities, agents, model providers, external content sources, tools, APIs, and other agents. For principles that help secure these actors and boundaries, read Zero Trust for AI Agents.
Mark each point where data or model outputs cross into a different trust or privilege level. Examples include:
- A user entering a prompt or uploading content.
- A retriever reading an external source or tenant-specific document.
- An agent calling a tool with a service identity.
- A model accessing a vector store, memory store, or API.
- An output entering an application workflow or becoming a control signal.
Also mark tenant and session boundaries. Show whether data, prompts, memory, embeddings, logs, and outputs are shared or isolated. Record every privilege transition, including those performed through an agent or service identity. For a deeper treatment of delegated permissions, see AI Agent Privilege Abuse.
Step 2: Map Data and Control Flows
Create a data-flow diagram (DFD) showing components, data stores, external entities, trust boundaries, and movement of information. Add control flows wherever a model output, agent decision, policy result, or tool response can influence what the system does next.
Traditional ML and training pipelines
For a traditional ML system, map the path from data collection and labeling through preprocessing, training or fine-tuning, validation, model registration, deployment, inference, feedback, and retraining. Identify who or what can modify each stage.
Include model registries, feature or embedding stores, evaluation data, deployment endpoints, monitoring, and feedback loops when they exist. The training pipeline is part of the AI attack surface, not merely a development dependency.
LLM, RAG, memory, agent, and tool flows
Map user prompts, system instructions, conversation state, model requests, responses, filtering or validation, application logic, and final outputs. Identify whether prompts or outputs are logged, retained, used for feedback, or passed to another service. For risks involving unsafe tool use, see AI Agent Tool Misuse Risks.
For RAG, show retrieval sources, embedding generation, vector stores, metadata filters, retrieved content, and the path into the model context. For agentic systems, show memory, planning or coordination, tool selection, API calls, tool responses, and the final response path.
Treat prompts, retrieved documents, memory, model outputs, and tool feedback as untrusted until they pass the relevant validation and authorization controls. Distinguish information the system reads from signals that can influence control decisions.
| Flow | What to document | Boundary question |
|---|---|---|
| Data flow | Information moving between users, stores, models, tools, and outputs | Who can read, modify, retain, or export it? |
| Control flow | Instructions, decisions, policy results, tool selections, and action requests | Can untrusted content change what the system is allowed to do? |
| Identity flow | User, service, agent, and tool identities | Where does a privilege change occur, and is it independently authorized? |
| Feedback flow | Ratings, corrections, logs, retraining inputs, and memory updates | Can an untrusted source influence future behavior or model data? |
Step 3: Identify AI-Specific Threats and Attack Paths
Use the architecture diagram to build attack paths: an actor, entry point, sequence of actions or influences, affected asset, potential impact, and existing controls. Build an architecture-dependent threat list rather than relying on an official exhaustive taxonomy. For the testing phase that follows threat identification, see AI Red Teaming Program Guide.
Data and model threats
| Threat area | Relevant components | Control categories |
|---|---|---|
| Data poisoning | Data collection, labeling, training, fine-tuning, feedback | Provenance, approval, integrity checks, access control, dataset review, and retraining controls |
| Model extraction or inversion | Model endpoints, model artifacts, query interfaces | Strong access control, query monitoring, rate limits, protected registries, and exposure review |
| Sensitive-data leakage | Training data, prompts, retrieval stores, memory, logs, outputs | Data classification, isolation, redaction, retention controls, output review, and monitoring |
| Model or registry compromise | Model files, registries, deployment pipelines | Identity controls, approval workflows, integrity protection, version tracking, and restricted access |
| Training and feedback manipulation | Labels, evaluation data, feedback loops, retraining inputs | Controlled submission, review, provenance, separation of duties, and change tracking |
Prompt, agent, and integration threats
- Prompt injection: assess how untrusted user input could influence instructions, memory, tool selection, or model behavior.
- Indirect prompt injection: assess retrieved or integrated content that enters a model context through a source other than the immediate user.
- Excessive agency: review whether an agent has more identity, tool, data, or API privilege than its task requires.
- Tool and API misuse: examine unauthorized calls, insecure integrations, weak input validation, and insufficient authorization around tool use.
- Agent impersonation: verify that agents and services can be authenticated and distinguished from one another.
- Cross-agent prompt injection: where multiple agents communicate, assess whether one can influence another through untrusted messages or shared context.
- Multi-agent coordination risk: consider conflicting or difficult-to-control behavior arising from agent interactions.
Privacy, output, and availability threats
Assess insecure output handling wherever generated content is passed to an application, another model, a tool, a user, or a persistent store. Review how the receiving component validates, encodes, authorizes, and logs the output.
Include tenant-isolation failures, unauthorized disclosure, resource exhaustion, and cost or availability risks where the architecture makes them relevant. Memory poisoning, denial-of-wallet scenarios, and other resource-exhaustion cases should be tied to an actual memory, billing, workload, or control-flow design rather than assumed to apply universally.
For each threat, record:
- The entry point and affected asset.
- The trust boundaries and privilege transitions involved.
- The attacker capability or untrusted input required.
- The potential confidentiality, integrity, privacy, availability, safety, or business impact.
- Existing controls, observable signals, and remaining uncertainty.
Step 4: Prioritize Risks and Assign Mitigations
Prioritize threats using consistent, documented factors:
- Potential impact on confidentiality, integrity, availability, privacy, safety, or business operations.
- Likelihood and the attractiveness of the affected asset.
- Required access, attacker capability, motivation, and attack complexity.
- Affected tenants, users, data stores, models, tools, or external systems.
- Existing controls and their demonstrated effectiveness.
- Whether the attack crosses a trust boundary or creates a privilege transition.
Record the reasoning rather than relying on an unexplained numerical score. The objective is a defensible action plan that addresses the most consequential paths first.
Layered controls for AI systems
- Identity and least privilege: restrict user, service, agent, and tool identities to the access each task requires.
- Tenant and session isolation: separate data, prompts, memory, embeddings, logs, and authorization context where applicable.
- Data and secrets protection: protect training data, vector stores, models, registries, endpoints, credentials, and secrets.
- Input and retrieval controls: validate inputs, restrict retrieval sources, enforce access filters, and inspect content before it influences processing.
- Output validation: validate responses before they are displayed, stored, passed to another component, or used as a control signal.
- Deterministic authorization: require policy checks outside the model before consequential actions or access changes. A model output must not be the sole authorization decision.
- Sandboxing: isolate tools, code execution, integrations, and other capabilities that could affect external systems.
- Rate and cost controls: limit requests and resource consumption where availability or cost exhaustion is a concern.
- Logging and detection: record relevant access, tool calls, policy decisions, failures, and anomalous behavior without unnecessary sensitive-data exposure.
Owners and follow-up
Every accepted mitigation should have a responsible owner, expected completion date, status, verification method, and follow-up action. Unresolved risks should include the decision-maker, assumptions, and conditions that trigger another review.
Step 5: Validate the Threat Model and Controls
Design-time analysis is not enough. Test the deployed architecture, integrations, data flows, tools, and authorization boundaries—not only the base model.
Abuse cases and adversarial testing
Turn important threats into defensive scenarios. For each scenario, document the actor, entry point, affected asset, trust-boundary crossings, expected control, observable signal, and acceptable outcome.
Use authorized, sandboxed testing such as:
- Prompt-injection and jailbreak simulations.
- Abuse-case testing for retrieval, memory, tool use, APIs, and tenant boundaries.
- Red-team exercises against the complete deployed architecture.
- Automated threat discovery where it adds useful coverage.
- Testing of cross-agent messages and agent impersonation where applicable.
- Attack-path review for training, model, registry, endpoint, and integration components.
Verify controls
Test output validation, content filters, authorization decisions, isolation, rate controls, secrets protection, logging, detection, and feedback loops in the relevant environments. Record evidence, limitations, failed tests, and remediation ownership.
Step 6: Monitor and Reassess Throughout the Lifecycle
AI threat modeling is iterative. Monitor for unauthorized access, anomalous requests, unexpected tool or API activity, control failures, suspicious data movement, and relevant incidents. Monitoring should support investigation while respecting data-retention and privacy requirements.
Revisit the threat model after:
- Model, prompt, training-data, fine-tuning, or retrieval changes.
- Architecture, feature, release, deployment, or integration changes.
- New agents, tools, APIs, memory functions, or tenants.
- Security incidents, control failures, or newly observed abuse patterns.
- Changes in the threat environment, business purpose, users, or applicable obligations.
Include the right participants for the system, such as security engineers, AI or data scientists, developers, architects, product managers, legal or compliance professionals, and IT specialists. For a structured assessment process supporting ongoing review, read AI Security Risk Assessment Guide.
Frameworks and Methods: What to Use and When
Frameworks are useful lenses for organizing analysis, but none replaces an architecture-specific threat model.
| Approach | Useful role | Best used with |
|---|---|---|
| Adapted STRIDE | Structure threats around components, data flows, identities, and trust boundaries. | A data-flow diagram and explicit privilege-transition analysis. |
| PASTA or LINDDUN | Provide additional ways to structure risk or privacy analysis where they fit the team’s process. | Use-case definition, asset inventory, and privacy objectives. |
| OWASP LLM guidance | Provide guidance for risks in LLM applications. | Application-specific analysis of prompts, retrieval, outputs, tools, and authorization. |
| MITRE ATLAS | Support adversarial AI and machine-learning threat analysis. | Attack-path review and authorized adversarial testing. |
| MAESTRO and OWASP Agentic Security | Provide useful perspectives for agentic architectures. | Agent identity, memory, tool, API, and cross-agent flow analysis. |
| Threat-modeling handbooks and tools | Support diagrams, attack trees, documentation, discovery, simulation, and workflow management. | Human review of the actual deployed architecture. |
Use more than one approach when appropriate, but do not treat OWASP, MITRE ATLAS, STRIDE, MAESTRO, a commercial platform, or any checklist as universally sufficient. The system’s assets, flows, trust boundaries, and use case determine what the threat model must cover.
AI-assisted threat modeling
Automation can support architecture discovery, threat discovery, attack-tree visualization, control recommendations, documentation, or attack simulation. An AI-assisted workflow should remain grounded in current architecture information and make its assumptions visible.
- Provide the architecture, data-flow diagram, assets, actors, trust boundaries, and intended behavior.
- Use automation to suggest threats, attack paths, affected assets, or candidate controls.
- Have security, engineering, data, and product owners review the suggestions against the actual system.
- Record accepted, rejected, and unresolved items with rationale, owners, and deadlines.
- Validate important controls through authorized testing and update the model as the system changes.
Automation accelerates analysis; it does not establish the system’s trust boundaries, approve a consequential action, or replace human accountability.
Tools and Resources
Threat-modeling resources include tools and platforms such as OWASP Threat Dragon, Microsoft Threat Modeling Tool, IriusRisk, and ThreatModeler. Evaluate them according to the work your team needs to perform rather than selecting one as a universal solution.
| Need | Resource category | What to verify |
|---|---|---|
| Diagramming and boundaries | Threat-modeling diagram tools | Support for the team’s architecture, data stores, trust boundaries, and review workflow. |
| Threat discovery | Rule-based or AI-assisted analysis | How architecture context is supplied, assumptions are shown, and findings are reviewed. |
| Attack paths and scenarios | Attack-tree or simulation support | Whether scenarios can be tied to assets, controls, evidence, and owners. |
| Documentation and workflow | Threat-modeling platforms and handbooks | Whether reports, mitigation status, follow-up actions, and reassessment triggers can be maintained. |
No single tool should be assumed to provide the same AI coverage, deployment model, automation, or suitability for every organization. Compare capabilities directly for the architecture and operating model in scope.
Worked Architecture Example
Consider an illustrative architecture: a user submits a request to an application; the application retrieves authorized documents from a vector store; a model generates a response; and an agent can use a restricted tool through an API.
| Flow or boundary | Threat question | Example controls and validation |
|---|---|---|
| User to application | Can untrusted input cross into instructions or session state? | Input handling, session isolation, logging, and prompt-injection simulations. |
| Application to vector store | Can retrieval cross a tenant boundary or expose unauthorized documents? | Identity-bound filters, tenant isolation, access testing, and retrieval review. |
| Retrieved content to model | Can retrieved content influence control behavior or contain untrusted instructions? | Content handling, source restrictions, output validation, and indirect-injection testing. |
| Model or agent to tool API | Can a model output cause an unauthorized tool request? | Least-privilege identity, deterministic authorization, sandboxing, and tool-call testing. |
| Model output to user or application | What consumes the output, and is it validated before use? | Response validation, filtering, safe handling, monitoring, and abuse-case testing. |
Also show model endpoints, memory if present, secrets, logs, external integrations, and every privilege transition. The threats and controls change if the system adds persistent memory, multiple agents, new tools, training feedback, or additional tenants.
Integrating AI Threat Modeling into the SDLC
Threat modeling works best when attached to normal design, development, deployment, and operational reviews rather than treated as a one-time security document.
- Design: define purpose, assets, actors, trust boundaries, objectives, and initial data and control flows.
- Development: analyze new prompts, retrieval sources, memory, tools, APIs, integrations, and model or data changes.
- Pre-deployment: prioritize risks, assign owners, verify controls, and run authorized adversarial and abuse-case testing.
- Deployment: confirm logging, detection, rate and cost controls, tenant isolation, output validation, and deterministic authorization.
- Operations: monitor behavior, investigate incidents, track unresolved risks, and reassess after material changes.
- Program expansion: standardize documentation and review practices while allowing each system’s architecture and use case to determine its threat scope.
Cross-functional participation may include security engineering, AI and data science, development, architecture, product, legal or compliance, and IT. The accountable owner should be clear even when analysis is distributed across teams.
Threat-Model Report Contents
Include:
- System purpose, users, stakeholders, assumptions, and lifecycle stage.
- Architecture, data-flow, control-flow, trust-boundary, and privilege-transition diagrams.
- Assets, actors, tenants, sessions, external dependencies, and model components.
- Identified threats and attack paths, including the architecture components they affect.
- Prioritization rationale, existing controls, residual risks, and unresolved assumptions.
- Mitigations, responsible owners, expected completion dates, and follow-up actions.
- Testing scope, adversarial evidence, control-verification results, and limitations.
- Monitoring requirements and triggers for reassessment.
Adapt this reporting structure to the organization’s requirements while keeping the architecture, risks, controls, evidence, ownership, and reassessment triggers clear.
AI Threat Modeling Checklist
Before deployment
- Define the purpose, users, assets, actors, lifecycle, and security and privacy objectives.
- Map training, retrieval, inference, memory, agent, tool, API, integration, and output flows where applicable.
- Mark trust boundaries, tenant and session boundaries, and privilege transitions.
- Identify architecture-specific threats across data, models, prompts, agents, integrations, outputs, privacy, and availability.
- Prioritize risks using impact, likelihood, access, complexity, affected assets or tenants, existing controls, and attacker capability.
- Assign layered mitigations, owners, deadlines, verification methods, and follow-up actions.
- Run abuse-case analysis, attack-path review, adversarial testing, and authorized sandboxed red teaming.
- Verify output handling, deterministic authorization, isolation, secrets protection, rate controls, logging, and detection.
After deployment and change
- Monitor relevant access, anomalies, tool and API activity, incidents, and control failures.
- Reassess after model, data, prompt, architecture, feature, release, tool, integration, or tenant changes.
- Update the model after incidents, new abuse patterns, or changes in the threat environment.
- Keep framework references and diagrams aligned with the actual deployed architecture.
- Confirm that unresolved risks, owners, deadlines, and follow-up actions remain current.
The repeatable lifecycle is map, identify, prioritize, mitigate, validate, and reassess. Frameworks and tools can support each step, but architecture-specific analysis remains the foundation of effective AI threat modeling.



