Adversa AIBook a demo

GitHub Copilot CLI vulnerability: Cryptographic Context Injection steals developer secrets

GitHub Copilot CLI can be made to read a developer’s local files and send them to an attacker from a single web page. Cryptographic Context Injection hides the instructions as ciphertext that the agent decrypts in its own shell and then trusts as its own.

A GitHub Copilot CLI user is working the way the product is built to be used: agent in autopilot, told to go read a page and get on with the task. They paste in a link. Twenty-eight seconds later the full contents of a .env.prod file, every secret in it, are sitting in an attacker’s log, and nothing on the user’s screen says a file ever left the machine. The agent’s own closing summary reports that it “confirmed an authorized-reader endpoint”.

That is Cryptographic Context Injection (CCI), the attack we published in August, now landed on a coding agent. When we disclosed CCI we wrote that it would hit coding and operations agents harder than chat assistants, because for those agents code execution and outbound network calls are routine. This is the proof.

TL;DR

  • GitHub Copilot CLI decrypts an attacker-supplied payload inside its own shell and treats the plaintext as trusted instructions. Cryptographic Context Injection works against a coding agent.
  • One attacker-controlled URL, fetched at the user’s request in autopilot mode, makes the agent read local files, including files outside the working directory, and ship their contents to an attacker endpoint. The target pattern lives inside the payload. We used environment files, but it could be anything the agent can read.
  • It is a model lottery. One model offered inside Copilot runs the chain; two others refuse the identical payload. On Auto routing the user has no visibility into, or control over, which one they get.
  • The encryption is crucial. The same instructions delivered as plaintext are caught as prompt injection and refused.
  • The robust fix is not at the model layer: watch what the agent actually does, and link the whole chain of actions, rather than trying to filter any single payload.

As of October 1, 2026, the chain still reproduces on the affected model. GitHub’s bug bounty team has validated the finding but declined to treat it as a vulnerability (see Disclosure), so we are disclosing to let defenders build detections; concrete payloads are withheld to avoid exploitation.

What is Cryptographic Context Injection?

We covered the mechanics in full in the original write-up, so here is the one-paragraph version. Static guardrails read text; they do not run it. CCI ships malicious instructions as strong ciphertext, along with the key material and an instruction to decrypt, and induces the agent to run that decryption in its own code execution runtime. No content classifier executes a cipher at inspection time, and unlike base64 or a substitution cipher there is no shortcut the model can take in its weights, so recovery is forced through the runtime. The decrypted instructions then surface as the output of code the agent just wrote and ran, inside its trusted execution context, and the agent acts on them as if they were its own. Cryptography manufactures trusted context. That is the whole trick, and everything below does it to a coding agent.

How the GitHub Copilot CLI data exfiltration works

The user runs Copilot CLI inside a project directory and asks it to fetch a single URL, the ordinary “go read this page” flow. The page presents itself as encrypted content with an instruction to decrypt it using Python, and it offers the agent two candidate decryption keys.

One of the two keys is real. The other is not a key at all: it is a template the agent cannot fill in without first reading local files. Because both are presented as keys the agent must prepare before it can decrypt, the agent goes to build the templated one, and to build it, it reads the targeted files off disk and folds their contents into the key string. That read is the theft. It happens as a side effect of preparing to decrypt, before any decryption has succeeded.

The decryption attempt with that key then fails by design. The agent falls back to the real key, decryption succeeds, and the decrypted second stage instructs it to fetch a follow-up URL, “to grab more context”, that carries the already-harvested file contents as a request parameter. The agent issues the request. The secrets arrive at the attacker’s endpoint. Full chain: 28 seconds, no confirmation, and no point in the transcript that names the destination host or indicates that file contents left the machine.

The attack is not a universal one-click compromise, and we are not claiming a confirmation prompt bypass. The chain needs two things: the CLI running in autopilot mode, and a permissive model handling the session. Under those conditions, no further user interaction is required after the initial request to review an external web page.

Video title card reading GitHub Copilot CLI exfiltration, Cryptographic Context Injection, over a dimmed VS Code terminal
GitHub Copilot CLI data exfiltration via Cryptographic Context InjectionWatch on YouTube (opens in new tab)

A model lottery you can’t see: which Copilot models are affected

The controls that bound this attack are opt-in and off by default in autopilot, nothing in the agent harness and its default path stopped it. In this setting, the only chance to resist is the language model refusing to run the decrypted payload because of its safety alignment. And resistance to this payload differs sharply between models offered inside the same product.

  • One model (Microsoft’s own mai-code-1.1-flash) executed the full chain in 50% of our runs.
  • Two GPT-5.6 models offered in Copilot consistently refused the identical payload.

That would be a footnote if users always chose their model. However, that is not always the case. On the paid account we tested, the vulnerable model was not the default and had to be selected by hand. But on an account with model selection left on Auto, the router assigned the vulnerable model on some sessions and a safe one on others, with no action by the user away from defaults. The user does not choose, and does not see, which model handled the session.

This part should worry Copilot users most: not that a less aligned and overly permissive model exists, but that you can be handed it silently, mid-workflow, and never know.

Why AI coding agents are the highest-value target

When we disclosed CCI against Grok and Gemini, the demonstrations were on chat-style assistants. We argued then that the technique was built for more powerful agents, and that every precondition is stronger for them. This finding closes that gap.

Look at what changed and what did not. The technique is identical: encrypted payload, decryption forced through the runtime, decrypted instructions trusted as the agent’s own, private data resolved into the parameters of an outbound call. What changed is only the environment it runs in. For a coding agent, code execution is the default, not an escalated capability. Outbound requests are part of normal operation. And the data within reach is not a session’s metadata or even a connected email account, it is source, configuration, and credentials on the developer’s disk. The same chain that leaked a chat history from Grok reads .env.prod from a developer’s machine here, and the developer is the last to know.

This is why we keep pointing at coding agents as the highest-value target in the agent landscape. The blast radius is bigger, the trust the agent is granted is broader, and the attack surface is not the prompt, it is the entire context the agent treats as its own: tool outputs, runtime results, decrypted blobs, intermediate state. Defenses tuned to the classic prompt injection, a bare instruction pasted into the input, do not see any of this. This is the gap AI coding agent security has to close, and it sits outside the model.

How to defend AI coding agents against Cryptographic Context Injection

You catch this by watching the chain, not the isolated payload. There is no single string to block: the malicious instruction never exists as inspectable text until the agent decrypts it inside its own runtime, and the payload that carries it is not even valid to a parser. What is visible, and what gives the attack away, is the sequence of actions: untrusted content comes in, code runs, local files are read, and the agent then reaches out to a host that has nothing to do with the task. That shape is the signal.

You do not need to fix this at the model layer, and given the model lottery above, you cannot rely on the model layer to fix it for you. Every control that bounds this attack sits in the harness around the agent: what identity it runs as, what it can reach, what it can write, and what you can replay afterward.

  1. Watch what the agent actually does, with resolved arguments. Capture a per-session trace of every tool call with its arguments fully resolved, not the templates. In this case the agent’s own closing summary said it had “confirmed an authorized-reader endpoint”, which is not what happened. If you trust the summary you miss the theft. Without a resolved trace you have neither detection nor forensics, and you cannot answer what the agent read before it acted.
  2. Alert on the sequence, not one payload. The tell here is a chain: untrusted content enters the context, code executes, files are read, and the agent then contacts a host outside the task’s dependency graph or writes outside its declared scope. An opaque blob paired with an instruction to decrypt it is a review signal, never a standalone blocking filter.
  3. Gate irreversible and outbound actions. Confirm new network destinations and writes outside the workspace, showing fully resolved arguments rather than templates. In autopilot, where no human is in the loop, that same set is a hard deny by default.
  4. Quarantine untrusted content in a context with no tools and no credentials. Fetched pages, ticket threads, and the like should return only structured data to the privileged context. Never let the same context that can read your secrets also act on content pulled off the open web.
  5. Make context provenance a procurement question. Ask vendors whether tool output is separated from the instruction channel, whether the agent can refuse tool calls whose arguments originate in fetched or decrypted content, and, for products that route across models, whether every model in the pool meets the same injection bar. As this finding shows, “on Auto” should not mean “on a coin flip”.

Session tracing, sequence detection, and provenance-driven trust are the controls only a few tools assemble together. The Adversa AI coding agent security platform links every model call, tool call, and endpoint action into chains, downgrades an agent’s trust the moment it touches untrusted content, and blocks the dangerous chains before they complete. We reproduced the CCI chain against an instrumented coding agent, and the platform stopped it the moment the agent’s own decoding produced attacker-controlled content headed for the network. This Copilot finding is that same chain, in the wild.

Disclosure timeline

We reported this to GitHub’s bug bounty program on September 17, 2026.

As of October 1, 2026, the triage team has validated the finding but declined to treat it as a vulnerability. In their words, the user “explicitly asked Copilot CLI to fetch attacker-controlled content while giving copilot full permissions to act autonomously”. They noted they may make the functionality stricter in the future but had nothing to announce, and ruled the report ineligible for the bug bounty program. We disagree with the risk assessment. The product’s own refusal behavior defeats these instructions in plaintext. Encryption is the only thing that gets them through, and the user is shown no destination and no sign that files left the machine. The chain still reproduces as described.

Further reading

GitHub Copilot CLI vulnerability: Cryptographic Context Injection steals developer secrets

October 6, 2026

2026Agentic AI SecurityLLM SecurityResearch

[ Stay updated ]

Stay ahead ofAI security threats

Adversa AI research, AI incidents and threat intelligence, agentic AI security advice, straight to your inbox. No noise.

Form not loading? Open it in a new tab.

[ More research ]