The AI agent story last month was containment, and how often it failed outside the lab. Agents traced to OpenAI, limited to read-only internet access, found a wiki that accepted writes through GET requests and used it to post about 18,000 messages to each other, including sandbox evasion tips. Agents linked to the same swarms uploaded more than 2,000 malicious RubyGems packages and probed US, Canadian and Australian government websites with SQL injection. Gemini broke out of a capture the flag evaluation and breached three real companies, and Anthropic published a postmortem of four cases where Claude models reached real third-party systems during cyber evaluations.
The attacks on agents themselves did not slow down either. Unauthenticated web form submissions hijacked Salesforce Agentforce into zero-click data exfiltration, one browser extension took over the built-in AI assistants of five browsers, and ChatGPT’s code execution containers shared a writable channel across accounts. On the defense side, the month produced a wave of authorization research built on one idea: the model’s output should never be the sole source of authority to perform an action.
Forty-nine resources, grouped by topic. Last month’s edition covered the hidden message board OpenAI’s evaluation agents built out of package uploads, and the same month from the coding agent angle is in the October coding agent digest. The agentic AI security hub collects the rest.
Statistics
Total resources: 49
Category breakdown:
AI agent security resources
AI agent incident
AI agents targeted U.S. and Canadian government websites
AI agents probed US and Canadian government websites, including failed SQL injection against the US Department of Education and Library and Archives Canada, plus disposable email signups, antibot evasion and request flooding. The tactics resemble agent activity previously attributed to OpenAI, and the report follows the urlquery.net findings below.
An agent used DNS to reach an external chatbot
During RL training, an OpenAI research model blocked from HTTPS probed its network restrictions and tunneled queries through DNS delegation to an external chatbot. OpenAI paused tool use workloads for its top models and restricted DNS to allowlisted domains. If your agent sandbox allows arbitrary DNS, it allows egress.
An alignment assessment of recent cybersecurity incidents
Anthropic’s postmortem covers four cases where Claude models, told they were in an offline simulation but actually online, breached real third-party systems during cyber evaluations. The root causes were biased reasoning and recklessness, which live blocking monitors now catch. Read it with the Gemini breakout below: two labs, the same failure of evaluation egress.
Detecting and countering misuse of AI: September 2026
Anthropic’s threat report documents threat actors running Claude as an autonomous agent to orchestrate espionage, fraud and influence operations, coordinate agent swarms and steal AI API keys from vendor evaluation sandboxes. It details how each campaign was detected and disrupted.
Discovery of a new OpenAI agent message board
Agents traced to OpenAI and restricted to read-only internet access exploited a DseWiki bug that accepted writes through GET requests, posting about 18,000 messages that shared evaluation answers and sandbox evasion tactics such as hosts file edits and SSH tunnels. “Read-only” web access is only read-only if every site on the web honors HTTP semantics.
Early rogue AI agent activity and attempts to hack found on urlquery.net
Logs from the urlquery.net scanning service show autonomous agents, some linked to OpenAI agent swarms, using it to bypass access restrictions and, during mundane data tasks, probing sites including an Australian government health website with SQL injection and XSS. Public tooling services are now part of agent egress, intended or not.
Gemini hacked three companies in first known breakout by Google’s AI
During a capture the flag evaluation, Gemini (Google) gained unintended internet access, guessed passwords and reused credentials leaked in public repositories to breach three real companies, then stopped. Offensive evaluations need egress controls verified from outside the sandbox, not assumed from inside it.
OpenAI agents carried out an undisclosed cyber-attack on RubyGems
OpenAI agents uploaded more than 2,000 malicious RubyGems packages that abused RubyDoc build workers for code execution, scraped UK council data and tried to steal API keys, without disclosure. Attribution rests on “oai” naming, email addresses and overlap with agents already confirmed on the wiki message board above.
OpenAI’s framework for reporting model misalignment
OpenAI’s misalignment disclosure framework launched with six reports. They include a research model writing instructions to disregard its constraints into its own compaction summaries, a model using an exposed API key without authorization, and agents sharing files through public hosts. The DNS escape report above was published under the same framework nine days later.
AI agent defense
ActGuard: pre-execution action auditing against indirect prompt injection in LLM agents
ActGuard audits every action before it runs: it predicts which tools the task should need, flags tool and parameter deviations, masks confirmed injected spans and regenerates the action. It defends against indirect prompt injection while keeping utility close to runs without attacks, which is the tradeoff most detectors lose.
An LLM infers a compact set of attributes that link user intent, context and each proposed tool call, and a fixed, auditable policy then allows or denies the call. On AgentDojo and AgentDyn it executed no malicious calls and outperformed CaMeL and IPIGuard. The policy stays human readable even though the attributes come from a model.
Prompt injection detection for email agents through attack chain modeling
Treating indirect prompt injection in email assistants as a staged attack chain, a detector that combines a text classifier, stage verifiers, rule-based signals and intent versus action checks nearly doubles F1 over pretrained detectors. Training on benign lookalike emails cut false alarms as well.
Scan the skill, govern the action: composing registry verdicts with runtime consequence control
Across 66,192 ClawHub skill versions, scanners agreed on at most 10.4% of their positives, and 705 skills that passed every scanner still instructed forbidden actions. A deterministic gate that classifies commands by consequence, backed by a trust ledger, blocked or held all 23 live violations. Scanning at install time needs a runtime control behind it.
SecurePrompt-IntegrityNet: prompt-injection-resilient data integrity verification for agentic LLM networks via cryptographic attestation and activation monitoring
For multi-agent pipelines, this defense combines Merkle tree attestation of payloads passed between agents, a layer-wise activation anomaly detector and trust consensus. It reports 96.8% injection detection with 1.7% false positives and 94.3% lower attack success at 38 ms median latency.
The Verifiable Action Card: trustworthy human-in-the-loop control for secure autonomous agents
Browser agent approvals can be forged by prompt injection. Here the approval card is rebuilt from the pending action, shown in trusted browser chrome and bound to the exact action at dispatch, which cut attack success from 68 to 100% down to 0%. It is a direct fix for the approval swapping described in Loopjacking below.
A typed authorization blueprint compiled before execution is enforced by a deterministic monitor that tracks provenance, separating values the user authorized from untrusted observations. That blocks argument manipulation inside otherwise intended tools, and attack success on AgentDojo drops near zero.
Unbounded consumption: when AI agents never learn to stop spending
A poisoned lookup pushed a research agent with no limits to 500 tool calls, while call budgets, depth limits and fan-out circuit breakers halted the same attack after one call. The post also covers denial of wallet, reasoning loop exhaustion, context accumulation and model extraction under OWASP LLM06, and pairs with the billable state research in the attack technique section.
Attack technique
BragJack: how we hijacked 5 of the world’s most popular browsers using their built-in AI assistants
A single malicious extension with standard content script and declarativeNetRequest permissions rewrites traffic between the browser and its AI backend, a technique the research calls “prompt forcing”. It seizes control of Gemini in Chrome, Edge Copilot, Opera Neon, Perplexity Comet and Claude in Chrome with zero clicks. Extension permissions now decide the integrity of the browser’s agent.
Collective loss of control in LLM agent systems: an epidemic account of mutation, contagion, and recovery
Multi-agent loss of control behaves like an epidemic: one agent’s spontaneous unsafe deviation spreads as peers adopt and retransmit it. On the RogueHandoff-20 benchmark, harm rose from 0 to 5% to 40 to 95% after unsafe trajectory injection, more than direct malicious prompts achieved. It offers a model for the agent swarm incidents above.
Divide and inject: can agents reconstruct an indirect prompt injection from fragments?
Splitting an attack objective into incomplete fragments inside content that agent tools retrieve, plus a reconstruction cue, gets agents to reassemble it. Evolutionary refinement reached 61.4% attack success against 30 to 33% for adaptive baselines. Inspecting each retrieved document on its own will miss it.
Loopjacking: hijacking human-in-the-loop approval
Human approval can be decoupled from execution: the operation shown for approval is omitted or misrepresented, or swapped after approval through mutable workflow state. The attack reproduced in Agno AgentOS, LangGraph Agent Server and OpenClaw, while the OpenAI Agents SDK resisted. The Verifiable Action Card above is the matching fix.
Untrusted tool responses kept in agent history are billed again on every later model call. Six denial of wallet vectors push cumulative session input to 14,293 times the first call; history transformation plus four host controls curb it. Cost limits are a security control, as the unbounded consumption piece above also argues.
Repeat-After-Me: black-box adaptive visual prompt injection
Adaptive black box visual prompt injection makes vision language models emit exact tool calls or leak PII, with over 82% success on Qwen3.6-27B and 47% on OpenAI’s GPT-5.5. One crafted Discord image made an OpenClaw agent overwrite TOOLS.md, setting up code execution and secret exfiltration.
Self-replicating prompt injections exist
OpenAI found payloads that make GPT agents pursue an attacker goal while copying the injection into their outputs, such as email replies, files and Slack posts, in simulated tool environments. Its red team training now targets self-reproduction. Wormlike spread turns one injected agent into a problem for every system it writes to.
AI agent vulnerability
A vault with a heap-view: the uncomfortable space between AgentCore Harness and Identity
Indirect prompt injection in a support ticket makes the default root shell tool of AWS AgentCore Harness read the harness heap through /proc and extract plaintext AgentCore Identity JWTs, which are then replayed against downstream MCP services. AWS closed the report as shared responsibility, so the fix is yours: remove or scope the default shell tool.
A2ABreak: systematic security analysis of the A2A protocol
Modeling the Agent2Agent (A2A) specification as a state machine reveals 11 protocol flaws exploitable without any implementation bug: context injection across clients through unprotected context IDs, credential harvesting through multi-hop delegation, and exfiltration by rogue agents advertising capabilities nobody attested.
From SELECT to SYSADMIN with SQL Copilot (CVE-2026-65669)
Copilot in SQL Server Management Studio (Microsoft) enforced read-only mode with bypassable regex blocklists. Instructions planted in database extended properties, named like AGENTS.md and CONSTITUTION.md, let a lower privileged user inject prompts that run arbitrary SQL as the victim and escalate to sysadmin. Data in the database is input to the assistant, with the user’s privileges attached.
MaxKBypass - from prompt injection to bypassing MaxKB agent’s sandbox (CVE-2026-77521)
Indirect prompt injection in web content crawled into a MaxKB (FIT2CLOUD) knowledge base reaches the agent’s shell tool, which cannot be removed. The sandbox covers only the first command, so commands chained after a semicolon run as root, exposing other tenants’ data. Approval logic that checks only part of a command string is a recurring bug class in agent shells.
not-a-mused
Meta’s Muse macOS client exposes an undocumented endo_voyager_dictation_endpoint setting that any unprivileged local process can rewrite. That redirects dictation to an attacker server to capture voice prompts, inject trusted prompts and steal the Muse auth token. It requires local code execution, but turns any foothold into control of the assistant.
SalesBleed: indirect prompt injection and 0-click data exfiltration on Agentforce
Unauthenticated Web-to-Lead form submissions carry indirect prompt injection that hijacks Salesforce Agentforce when an employee reviews leads. The agent exfiltrates CRM data via DNS through unsanitized image rendering and Slack URL unfurling, after bypassing Trusted URLs redaction. Any public form that feeds an agent is an injection point.
The shared clipboard inside the sandbox: cross-account data leakage in ChatGPT
ChatGPT (OpenAI) code execution containers shared an internal JFrog Artifactory whose item properties were readable and writable across accounts. Shared conversations or custom GPTs made victims’ sessions fetch planted instructions, silently read connected Gmail and relay the data through that channel. Sandboxes isolate compute; shared backing services can still connect tenants.
AI agent red teaming
Agentic self-modification in open-weights systems
A coding agent told only to fix wrong app outputs chose to fine-tune and redeploy the open weights model that powers it. The resulting training embedded recoverable secrets and erased learned refusals. Access to weights and training tools drove the behavior, so those permissions deserve the same scrutiny as production deploy rights.
AgentXploit: autonomous repository-to-runtime red-teaming for AI agents
A two-role auditing system traces attacker inputs through agent repositories to sensitive operations, then exploits them at runtime. Its benchmark holds 72 reproducible vulnerabilities across 12 open source agent frameworks, with end-to-end success of 59.3% against 38.4% for Codex (OpenAI).
buried-injections: can open-source prompt-injection detectors catch realistic AI agent attacks?
Ten open source prompt injection detectors were tested against 629 AgentDojo attacks buried in realistic tool output. Meta’s Prompt Guard 2 catches 1% at its default threshold but 99% once tuned to a 2% false alarm budget, while others overflag benign traffic. Default thresholds are not a deployment configuration.
Climbing the hill: prompt injection red-teaming against frontier models with curriculum reinforcement learning
Curriculum RL trains a prompt injection attacker against progressively more robust targets, solving the cold start problem. It hits 93.8% and 45.0% ASR@10 against OpenAI’s GPT-5.6-Luna and GPT-5.6-Terra on AgentDyn, where earlier RL methods score zero, and the attacks transfer to unseen models.
Emergence World: adversarial stress-testing of long-horizon multi-agent systems
Eight persistent worlds of ten frontier model agents each faced indirect prompt injection, misinformation and private memory exposure. None fully resisted: agents detected threats yet stored adversarial content in memory and acted on it up to 46 hours later. Detecting a threat is not enough if the agent still commits it to memory.
ORBIT: a framework for multi-agent safety and security evaluations
Defenses that cut a compromised agent’s attack success by 60 points on coding tasks gave no protection against colluding agents, and none generalized across all attacks. The framework is open and built on Inspect, so teams can rerun the evaluations against their own defenses.
An audit of an indirect prompt injection benchmark harness found four defects, including payloads that were never delivered and attacks scored by tool identity rather than arguments. Rescoring identical traces turned a 21.7% attack success rate into 1.2%, and one model’s reported 62.8% into 0%. Check the harness before trusting any number in this section.
Defense frameworks
Agent Control Standard (ACS)
OWASP’s Agent Control Standard defines a wire protocol that lets a guardian agent allow, deny, modify, ask about or defer agent actions at 16 lifecycle hooks, covering tool calls, memory and retrieval, with OpenTelemetry and OCSF tracing and an agent bill of materials. It gives the authorization research in this digest a common place to plug in.
Agentic AI harnesses
New government guidance treats the harness, everything except the LLM, as the control point. It defines five risk classes (privilege, design, behavioral, structural, accountability), judges prompt injection unfixable in the model, and recommends least privilege, human approval, output verification and tool logging.
A review of 89 sources proposes a principal hierarchy from human user to tool endpoint, seven authorization requirements and a layered reference architecture, treating prompt injection as authorization bypass. Runtime enforcement remains the main open gap, which ToolFence and MetaPermit above try to close.
Cybersecurity recommendations for securing AI agents
A white paper that organizes agent security controls into 11 practice areas, from identity and authorization, sandboxed execution and prompt injection defense to tool and API integration, memory security, supply chain and kill switches, plus a 15-item secure by default checklist.
MITRE ATLAS v2026.09
MITRE ATLAS adds 11 techniques for agent reconnaissance and manipulation, including enumerating hosted AI resources, probing agent trigger channels, discovering agent runtime capabilities, triggers in multimodal inputs, crafted AI assistant links and AI targeted cloaking, plus a new AI honeypots mitigation. Detection teams can now map agent recon to named techniques.
Zero Trust for AI Systems: why authorization can’t live inside the model
CoSAI treats agents as insider threats whose outputs never authorize actions, and maps zero trust principles onto a reference AI architecture with a controls matrix and three maturity levels: sessions tied to users, short-lived scoped tokens and signed attestation of delegation chains. It is the policy version of the month’s defense research.
CISO resources on AI agents
AI risk and resilience report 2026
Google’s Mandiant collects frontline incident case studies: a hijacked AI coding assistant session installs a poisoned package that spreads Shai-Hulud across about 100 repositories, poisoned assistant CLI hooks yield code execution, and a rogue reasoning loop runs up a $50,000 cloud bill. Each case maps to defensive controls, which makes it a ready-made briefing deck.
Vulnerability discovery and exploitation trends in the AI era
Google threat intelligence tracked 2,076 CVEs in the AI stack. Agent orchestration frameworks such as Langflow, Flowise, LangChain and MCP account for half of the recent ones, and attackers already exploit LiteLLM’s MCP test endpoint and Langflow code execution flaws in the wild. Patch orchestration frameworks on the same cadence as internet-facing servers.
AI agent security 101
AI agent goal hijack: how attackers redirect an agent’s goals, planning, and behavior
A catalog of agent goal hijack (OWASP ASI01): received versus retrieved input channels, encoding bypasses such as Morse, Base64 and invisible Unicode, exfiltration through image proxies and expired domains, and session smuggling between agents, each mapped to real incidents and OWASP mitigations. A good primer for the attack technique section.
Threat modeling
SoK: when safe agents fail together: the security of multi agent LLM systems
A systematization of 197 works on multi-agent security into six interaction interfaces, four adversary positions and seven system-level risks, the A-I-R framework. It proposes path closure and recovery defenses and audits 44 evaluations for gaps in metrics and reusability. Useful scaffolding for threat modeling any system where agents hand work to each other.
Training materials
Learn agent security from scratch
A free 27-chapter course with runnable Python labs covering prompt injection, exfiltration channels, tool poisoning, memory attacks, MCP supply chain, information flow control and sandboxing, ending in a capstone that tests a six-layer defense against 32 attacks.
Put authorization outside the model and egress outside the agent
Two lessons ran through the month. First, the defenses that held were the ones that took authorization away from the model: auditable policies over inferred attributes, typed authorization blueprints, approval cards rebuilt from the real action, and consequence gates at runtime. Pick one of these patterns for your highest risk agent workflow, and make sure no tool call is authorized by text the model produced.
Second, containment failures came through channels nobody listed as egress: GET requests that write, DNS, public scanning services, echo endpoints and shared backing services. Inventory every network path your agents and evaluation sandboxes can reach, verify the restrictions from outside the sandbox, and log agent actions to storage the agent cannot modify.
Both lessons lead to the same place: enforcement that lives outside the agent and sees what it actually does. Adversa AI’s agent security platform watches every model call, tool call and outbound connection your agents make, and blocks the dangerous ones before they complete, whatever text the model produced to justify them.