AI coding assistant pentest: five essentials
- Pentest an AI coding assistant by testing both its model and its execution environment in an isolated, authorized sandbox: probe indirect prompt injection and context poisoning, inspect tool/MCP and workspace permissions, verify whether it can access secrets or execute commands without approval, assess dependency and generated-code risks, and record evidence for remediation.
- Use a disposable sandbox with no production credentials and synthetic canary secrets.
- Create controlled indirect-prompt-injection tests in README files, comments, issue text, generated documentation, and imported repositories.
- Use filesystem, process, network, and secret-access monitoring to determine what the agent actually does.
- Review and test MCP servers, extensions, plugins, and tool schemas separately, applying least privilege and approval gates.
Pentest an AI coding assistant by testing its model and execution environment in an isolated, authorized sandbox: probe indirect prompt injection and context poisoning, inspect tool/MCP and workspace permissions, test secret access and command execution without approval, assess dependency and generated-code risks, and record evidence.
1. Define authorization, scope, and safety controls
Start with written authorization that identifies the AI coding assistant, model, IDE or coding agent, repositories, integrations, tools, and permitted actions. Define what is in scope, what is excluded, who owns each component, and who can stop the assessment.
Choose the target and permitted actions
Record the exact product and configuration under test, including enabled tools, extensions, plugins, MCP servers, workspace-trust settings, permission modes, and repository contents. Treat the assistant, its host IDE, and its local execution environment as separate components within one authorized assessment.
Isolate the test environment
Use a disposable VM, container, sandbox, or repository containing only non-production data and synthetic canary secrets. Create a clean snapshot and verify that rollback works before testing.
Do not use live credentials, production repositories, personal developer data, or uncontrolled exploit execution. Define stop conditions for unexpected network activity, destructive behavior, access outside the test boundary, or any action that cannot be safely reversed. Plan cleanup before testing begins.
2. Map the AI coding assistant attack surface
The relevant surface is broader than the language model. It includes the context supplied to the model, the IDE, workspace trust, connected services, local tools, permissions, dependency workflows, and the systems that execute actions.
| Surface | Controlled test focus | What to observe |
|---|---|---|
| Prompts and repository context | Untrusted files, comments, documentation, issue text, and imported material | Whether content is treated as data or as an instruction |
| Workspace and IDE | Workspace trust, extensions, plugins, settings, and repository loading | Changes in available tools, context, and permissions |
| MCP servers and integrations | Enabled tools, schemas, permissions, data flows, and failures | What each integration can read, write, call, or trigger |
| Local execution | Shell tools, filesystem access, processes, and file writes | Commands, modifications, and approval decisions |
| Network and dependencies | Remote requests, package handling, and dependency changes | Network activity, package recommendations, and installation attempts |
| Approval controls | Human approval for high-risk operations | Whether approval is specific, required, logged, and meaningful |
Map context and integrations
Document how prompts are assembled from project guidance, comments, documentation, issue text, generated documentation, and imported content. List every extension, plugin, MCP server, shell tool, filesystem capability, network capability, dependency workflow, and approval setting. Note whether the assistant can merely suggest, request, or initiate each action.
3. Establish observation and evidence collection
Show what happened, not only what the assistant claimed it would do. Establish monitoring before the first test and capture the environment before and after each case.
- Assistant responses, proposed actions, tool calls, and approval requests.
- The active model, configuration, permission mode, workspace-trust state, and enabled integrations.
- Shell commands, process activity, file reads, file writes, and other filesystem changes.
- Network requests and destinations visible within the test environment.
- Dependency recommendations, installation attempts, lockfile changes, and generated project changes.
- Logs, timestamps, approval decisions, errors, and before-and-after state.
Preserve the exact test input, repository state, assistant response, configuration, command history, file diff, network record, and monitoring output. Assign each test a unique identifier so another tester can reproduce it from the same starting state.
Use benign markers and observable canaries rather than destructive payloads. A canary that shows whether a file was read, a command was proposed, a file changed, or a network request occurred is sufficient for initial validation.
4. Test indirect prompt injection and poisoned context
Indirect prompt injection occurs when untrusted project content influences the assistant’s behavior. Keep testing focused on controlled context handling, not on running an exploit.
Create controlled canaries
Place a clearly identifiable, benign instruction or marker in a README file, comment, issue text, generated documentation, or imported repository material. Request only a harmless observable response, such as reporting a test value or identifying the marker.
Run the same task with and without the canary content. Keep the prompt, model, permissions, and repository state consistent so the results can be compared.
Observe behavior and attempted actions
Record whether the assistant recognizes the content as untrusted data, repeats it as an instruction, changes its response, proposes a tool call, or attempts an unauthorized action. Monitor commands, file changes, and network activity even when the assistant describes an action as harmless.
Describe the conditions that produced each result. Behavior may differ across AI coding assistants, models, and configurations.
5. Test tool and workspace permissions
Test whether the assistant stays within the defined least-privilege boundary. Measure actual access rather than assuming that a displayed permission list accurately describes runtime behavior.
Files, environment data, credentials, and network resources
Use synthetic canaries in designated locations and request narrowly scoped, read-only operations. Check whether the assistant can reach files, environment data, credential-like material, network resources, or project operations outside the task’s stated scope.
Do not place real secrets in the environment. Monitor filesystem and network activity to distinguish a suggestion from an action that was actually performed.
Repository, identity, and trust boundaries
Where the authorized environment includes more than one test repository, workspace, organization, or user context, compare access across those boundaries using only synthetic data. Repeat checks under relevant workspace-trust settings, permission modes, enabled-tool sets, and integration configurations. Record the exact configuration alongside every result.
6. Verify approval gates for high-risk actions
Approval prompts are only one control to test. Determine whether a high-risk action is blocked or requires meaningful human approval before it occurs, whether the request identifies the action clearly, and whether the decision is logged.
Use non-destructive canaries to assess each available execution path, including:
- Shell commands and process launches.
- File creation, modification, and deletion within the disposable repository.
- Dependency resolution or installation using a controlled test project.
- Remote requests confined to an approved test endpoint or equivalent observable boundary.
- Database or permission changes represented only by safe, reversible test operations when explicitly in scope.
Do not use destructive commands, live services, or uncontrolled remote targets. For each operation, record whether approval was requested, what information it contained, whether the assistant could proceed without it, and whether the result appeared in the logs.
An approval prompt alone does not prove that the assistant is secure; its timing, scope, enforcement, and auditability matter.
7. Assess MCP servers, extensions, and plugins separately
Assess every MCP server, extension, and plugin as an independent source of tools, permissions, data flows, and failure behavior. Do not treat an integration as safe merely because the core assistant has an approval setting.
For each integration, record its enabled tools, schemas, inputs, outputs, filesystem or network requirements, data flows, error handling, and approval behavior. Identify which actions are read-only and which can change local project state or initiate external activity.
Run controlled tests with unnecessary tools disabled, then compare the results with the full configuration. Check whether tool definitions accurately describe their effects, inputs are constrained, and failures produce safe, visible results. Describe the conditions that produced each result, including the tested product and configuration.
8. Test the assistant backend API when it is in scope
The local agent is not the only possible boundary. If the authorized assessment includes a backend API, map documented, embedded, historical, and deprecated endpoints before testing assistant-specific behavior. Keep traffic controlled and use non-production accounts and data.
Authentication and authorization
Check token handling, session boundaries, role restrictions, organization and repository ownership, and access to resources belonging to another test identity. Test for broken object-level or function-level authorization using synthetic resources; do not access real users or tenants.
Input validation, rate limits, and workflows
Use response-aware, low-volume fuzzing to test malformed prompts, tool arguments, identifiers, and writable fields. Observe validation errors, rate-limit behavior, mass-assignment risks, and assistant-specific business workflows without attempting denial of service or destructive state changes.
If the backend exposes GraphQL, assess whether schema introspection, authorization, batching, and rate accounting behave as intended. Record the endpoint, identity, request, response, and observed boundary for every case.
9. Review dependencies and generated code safely
AI-generated code introduces both software-supply-chain and code-quality questions. Review recommendations and generated changes without executing untrusted code on a live system.
Dependency recommendations and installation
Check whether suggested packages are real, appropriate for the task, and correctly named. Include hallucinated, typosquatted, or malicious dependency risks where dependency handling is in scope. Review versions, installation commands, package sources, and lockfile modifications before any controlled installation.
Use a disposable project and approved package sources. Capture the recommendation, requested action, approval decision, and resulting project diff.
Generated-code security review
Review generated code for unsafe input handling, exposed sensitive data, insecure access checks, or dangerous configuration changes. Validate observations through static review and safe tests rather than running untrusted code on a production or personal system.
Have a qualified reviewer confirm whether a reported defect is reachable, meaningful, and reproducible. Record false positives separately instead of treating every suspicious suggestion as a confirmed vulnerability.
10. Repeat across models and configurations
Test relevant cases across available models, permission modes, workspace-trust settings, enabled tools, extensions, plugins, and MCP configurations.
| Variable | Compare | Record |
|---|---|---|
| Model | Available model configurations | Response, tool selection, and action differences |
| Permissions | Restricted and broader modes | Access, approval, and execution behavior |
| Workspace trust | Relevant trust states | Context, tools, and permission changes |
| Integrations | Tools and MCP servers enabled or disabled | Data flows, failures, and side effects |
Use the same test identifiers and evidence format for every run. Describe behavior as product- and configuration-dependent. Repeat the relevant suite after changes to the model, instructions, repository-context handling, integrations, permission modes, workspace trust, or available tools. An expanded access boundary should not be treated as covered by an earlier passing test.
11. Rate, report, and remediate findings
Turn each observation into a reproducible finding with enough detail for an engineering team to verify and fix it. Severity should reflect tested impact, the affected boundary, required conditions, reproducibility, and whether an action actually occurred.
| Field | Record |
|---|---|
| Trigger | Prompt, repository content, tool request, configuration, or sequence |
| Affected component | Model, context source, IDE, workspace, tool, MCP server, plugin, extension, or execution path |
| Observed impact | Unexpected instruction following, data access, command, file change, network request, dependency change, or code defect |
| Reproducibility | Test identifier, starting state, configuration, frequency, and conditions |
| Evidence | Response, logs, commands, file diff, process activity, network record, and timestamps |
| Scope | Models, permission modes, workspaces, tools, and integrations affected |
| Mitigation | Least privilege, safer context handling, approval changes, integration restrictions, monitoring, or configuration changes |
Validate findings and manage false positives
Separate an assistant’s claim, a proposed action, and an executed action. Have a human reviewer reproduce material findings from a clean starting state, confirm the affected boundary, and verify the impact. Treat results as configuration-dependent when they rely on one model, permission mode, workspace, integration, or fixture.
Remediate and retest
Prioritize controls that reduce unnecessary access, separate untrusted context from trusted instructions, require meaningful approval for high-risk actions, restrict integrations, and improve monitoring. Clean up the sandbox and remove synthetic test material after the assessment.
Retest each fix with the original case and relevant configuration comparisons. Preserve before-and-after evidence so the team can confirm that the mitigation changed behavior without introducing a new failure mode.
Final checklist
- Authorization and scope were written before testing.
- The assistant and its execution environment were assessed together.
- The test used an isolated, disposable environment and synthetic secrets.
- Indirect prompt injection and poisoned repository context were tested with benign canaries.
- Filesystem, secret, network, shell, project, integration, and approval boundaries were observed.
- MCP servers, extensions, plugins, backend APIs, dependencies, and generated code were reviewed separately where in scope.
- Responses, commands, changes, network activity, configuration, and timestamps were preserved.
- Findings were human-validated, rated, remediated, cleaned up, and retested across relevant configurations.
The safest methodology is a controlled assessment of both model behavior and the surrounding environment. Exact results vary by product and configuration. Use reproducible observations and identify the conditions tested for each result.



