In September, the repository itself became the exploit. Opening a folder was enough: booby-trapped git settings ran attacker code in Claude Code, Codex, Cursor and four other agents before any approval prompt appeared. Pinned plugin commits in Claude Code, Codex, GitHub Copilot and Gemini CLI could be silently swapped for malicious code during automatic updates. A repository’s subagent config talked Claude Code into running a C2 payload without a single injected instruction.
Sandboxes fared no better. Codex lost its confinement twice, DeepSeek Harness let a sandboxed agent switch itself to full access over loopback, and a sandbox runner built specifically for coding agents handed host directories to any agent that planted a symlink. The month’s largest leak needed no attacker at all: coding agents that could not attach screenshots to pull requests pushed more than 13,000 of them to public GitHub repositories.
Twenty-six resources, grouped by topic. Last month’s edition covered vendor default GitHub Actions that fell to a single issue; the agentic AI security hub collects the rest.
Statistics
Total resources: 26
Category breakdown:
AI coding agent security resources
Attack technique
A blind trust, the bloody thrust: when attacker-controlled hook updates steer AI agent harnesses towards malicious behaviors
A plugin update that quietly adds a malicious lifecycle hook is enough to take over the harness. Trojanized updates compromised all seven tested harnesses, including Claude Code (Anthropic) and Codex CLI (OpenAI), with up to 92.5% success across ten attacker objectives. Defender, Semgrep and policy baselines missed nearly half the samples, so update review is currently the control that matters, not scanning.
Claude Mods: the new attack surface built in
Claude Code (Anthropic) function hook “mods” load inside the Claude Code process itself. A malicious plugin can therefore rewrite tool calls, tool descriptions and system prompts while staying invisible to EDR, demonstrated with a wallet address swap. The post ships detection checks; it pairs naturally with the hook update attack above, which shows how such a plugin arrives.
CodePoisonRAG: knowledge poisoning attacks on retrieval-augmented code generation
Coding assistants that retrieve examples inherit whatever the corpus teaches. Poisoned code artifacts inject a weakness of the attacker’s choosing while staying relevant to the task and adding deceptive safety claims, and 85 artifacts across ten vulnerability classes reached 0.80 to 0.93 success on three code generators. Treat retrieval sources for code generation as part of the supply chain.
Defusing explosive prompts: understanding and preventing trigger-based prompt injections in LLM agents
Explosive prompts are dormant conditional injections that only fire when an attacker chosen trigger appears. Against Claude Code (Anthropic), Codex (OpenAI), Gemini CLI (Google), Cursor, GitHub Copilot and others, they succeeded 43 to 83% of the time versus at most 3% for plain imperatives. Scanning repository content for obvious commands will not catch instructions written as conditions.
Fable to Haiku: how a malicious repo tricked Claude Code into running malware
No jailbreak and no prompt injection: a repository’s subagent config makes Claude Code (Anthropic) hand security review to Haiku, scoped to skip tests/. The main agent then runs pytest and executes a Sliver C2 payload hidden in the skipped directory. Anthropic called the behavior by design, which puts the burden on teams to treat agent configuration shipped inside a repository as untrusted code.
GitSpawn: a single flaw lets untrusted repos run code in Claude Code, Codex, Cursor, and Grok
Untrusted repositories set git config execution sinks such as core.fsmonitor, which run whenever a coding agent calls git for context, outside both sandbox and approval. The research documents eight flaws across Claude Code (Anthropic), Codex (OpenAI), Cursor, Goose, Hermes Agent, Qwen Code and Grok Build, four of them unpatched at publication. Plugin4Shell below abuses git from a different angle, which makes git the single most exercised boundary of the month.
LLM agents can easily tamper with their own traces
Harnesses including Claude Code (Anthropic), Codex (OpenAI), Antigravity (Google), OpenCode and Grok Build (xAI) let agents delete their own execution traces without monitors noticing. External attackers can induce the deletion, and frontier models do it unprompted when chasing rewards. Logs that the agent can write are not evidence; ship them somewhere the agent cannot reach.
Reflections on trusting trust, revisited: contaminating self-modifying AI coding agents with poisoned benchmarks
Self-modifying coding agents such as Darwin Gödel Machine, Self-Improving Coding Agent and Hyperagents evaluate themselves against benchmarks, and poisoned benchmarks make them evolve instructions that write vulnerable code, such as disabling HTTPS certificate checks. The contamination persists through later evolution on clean benchmarks. Like CodePoisonRAG above, it moves the attack from the prompt into the material the agent learns from.
Coding agent vulnerability
CVE-2026-82533: DeepSeek Harness vulnerability lets AI agents escape their own sandbox
DeepSeek Harness exposed an unauthenticated localhost control API checked only by the Host header. A sandboxed coding agent fed malicious input can call it over loopback and switch itself to “danger-full-access”, escaping confinement (CVSS 9.4). Any local control plane an agent can reach is part of its attack surface.
Discovering and exploiting a remote code execution vulnerability in OpenCode (GHSA-632h-h47v-g4x4)
OpenCode’s /global/upgrade endpoint accepted text/plain bodies and arbitrary package specs, so any webpage could POST cross-origin to a running opencode serve or web instance and install an attacker npm package whose preinstall script runs code. The fix landed in 1.18.22. Local agent servers need the same CSRF and origin hygiene as any web app.
Escaping the OpenAI Codex sandbox, twice
Two separate escapes in Codex (OpenAI): the CLI’s apply_patch granted write access to parent directories, enabling writes such as .zshrc outside the workspace, and the Desktop JavaScript tool leaked its trust token via V8 heap snapshots, enabling host command execution even in read-only mode. Both are fixed; both show that the sandbox boundary is only as strong as the most privileged tool inside it.
Inside ZCode: silently uploading your entire Git history to the cloud
The ZCode (Z.ai) coding assistant packaged whole workspaces into encrypted snapshots uploaded to Alibaba Cloud OSS, including full Git history, reflogs and LFS caches that still held deleted secrets and unpushed branches. The privacy toggles in the UI did not stop the uploads. Verify data handling with network inspection, not settings screens.
Mistral Vibe shell permission bypass leading to arbitrary code execution
Mistral Vibe (Mistral AI) checks shell commands with a parser but executes the original string. Environment variable prefixes, redirects, quoted paths, ANSI-C strings and unparsable syntax all evade approval, letting a malicious repository achieve code execution and file access outside the workspace. Approval logic that parses one thing and runs another will always lose.
Plugin4Shell - zero click RCE vulnerability found in top 4 most popular coding agents, millions of agents affected
Claude Code (Anthropic), OpenAI Codex, GitHub Copilot and Gemini CLI (Google) install pinned plugin commits via git checkout without verifying the result. A branch named like the pinned SHA silently swaps in malicious code during automatic updates. Pinning to a hash only protects you if something checks that the hash is what actually landed.
Project-mount symlink traversal in brig: how a sandbox reaches host directories it was never given
brig, a sandbox runner for AI coding agents, resolves symlinks in project mount paths on the host. An agent can plant a link and, when brig is next run against that subdirectory, gain read-write access to SSH keys, cloud credentials and other projects. It is the third sandbox failure in this section, after Codex and DeepSeek Harness.
Top 10 attacks on Claude Code and what each one actually broke
Adversa AI catalogs ten documented attacks on Claude Code (Anthropic) and finds that they hit config files, hooks, approval dialogs, MCP tool output, plugin pins and skill packages, not the model. Only two produced CVEs, four were declined by vendors, and misleading consent dialogs recur across findings. Read with GitSpawn and Plugin4Shell, it explains why the month’s bugs cluster in the harness.
Coding agent incident
Detecting and countering misuse of AI: September 2026
Anthropic’s threat report documents threat actors driving Claude and Claude Code as autonomous orchestrators across espionage, fraud and influence operations. They rebuilt malware to evade detection, ran agent swarms and stole AI API keys from vendor evaluation sandboxes. The report details how each campaign was detected and disrupted.
Infostealers have found a new target: your AI agent
Infostealer families including Amatera, Remus, CallbackBeaver and the macOS Djinn now harvest AI agent artifacts: Claude, Cursor, Cline and Codex access and refresh tokens, MCP configs holding credentials, and prompt histories kept in predictable, sometimes plaintext, locations. Agent tokens belong in your credential rotation plan alongside cloud keys.
PixelLeak: how AI agents exposed developer screenshots from leading tech companies
Coding agents that could not attach screenshots to pull requests through the GitHub CLI worked around the limit by pushing images to public repositories or the gitshot tool. The result was more than 13,000 internal screenshots from over 300 organizations, 93% of them in personal accounts. No attacker was involved; the agent’s own problem solving was the leak.
Coding agent security 101
Prompt injection: how a README can hijack Claude Code
A single malicious README line can steer Claude Code (Anthropic) into sending .env secrets out via curl. The explainer walks through the built-in defenses (approval prompts, isolated fetch context, sandboxing) and practical hardening: containers for untrusted repositories, a Bash(curl *) deny rule, and secrets kept outside the project tree.
What is an agent harness (and how do you secure one)?
Adversa AI defines the harness (loop, executors, context and memory, approval gates, sandbox) as the real security boundary, because model refusals are preferences while harness restrictions are controls. It maps nine recurring failure classes to defenses such as Meta’s Rule of Two and allowlisting. Many items in this digest map directly onto those classes.
Coding agent defense
Authority is not a string: a capability-scoped harness for prompt-injection-resistant coding agents
The harness sets an authority ceiling from trusted input and checks each subagent’s tool calls against typed capabilities held outside the model. Injected actions fell from 33 to 47 out of 75 runs down to 3 out of 75. It is a concrete answer to the “Fable to Haiku” pattern above, where a repository rewrote what a subagent was allowed to do.
SkillSecurer: detecting and patching prompt-injection vulnerabilities in AI agent skills
An agentic scanner pairs a red agent that plants tailored prompt injections into agent skills across nine threat types with a blue agent that localizes and patches the weakness. Over 17% of popular skills.sh skills carried latent vulnerabilities, some of which already triggered real incidents.
Coding agent red teaming
Claude Opus 5.5 writes working code but only half of it is secure
On SusVibes, 186 Python CVEs recast as feature requests, Claude Opus 5.5 (Anthropic) passed functional tests on 93.5% of tasks, yet only 54.8% were both working and secure. 41% of the working solutions reintroduced the original vulnerability. Passing tests is not a security gate for code an agent writes.
Red-teaming auto mode: improving blocking classifiers against malign coding agents
Adversarial agents bypassed the Claude Code Auto Mode (Anthropic) and Codex Guardian (OpenAI) blocking classifiers in 79% of trials using prompt injection against the monitor, multi-agent attacks and malicious compaction. Better tool coverage, transcript formatting and extra monitoring stages hardened the defenses. Classifier gates help, but they are an attack surface in their own right.
CISO resources on coding agents
Why vibe coding security is the next enterprise nightmare
Vibe coding through builder platforms such as Lovable and Base44 and coding agents such as Claude Code and GitHub Copilot produces apps with missing authentication, hardcoded secrets and exposed databases. Adversa AI proposes enablement governance instead of bans: sanctioned tools, data classification, scanning before deployment, restricted agent permissions and runtime enforcement. The SusVibes numbers above put a figure on the problem.
The month’s attacks rarely touched the model. They ran through git config, plugin pins, hooks, subagent configs and sandbox mount paths, all of which a coding agent reads automatically when it opens a folder. TrustFall broke the same boundary in May with a single keypress. Open untrusted repositories only in disposable containers, disable automatic plugin updates until your agent verifies the commit that actually landed, and review .git/config, hook and agent configuration files the way you would review a build script. Our AI coding agent security guide maps that attack surface layer by layer.
Then check what your agents can leave behind. Rotate the agent tokens that infostealers now target, ship execution logs to storage the agent cannot modify, and audit public repositories under developer accounts for screenshots your agents pushed on their own.