A code agent in a sandbox is a draftsman. It writes, you read, and nothing leaves the room. Most of the value comes when it stops being a draftsman: it opens the files, runs the commands, reaches the internal systems and publishes the result. That is also where the risk starts. This article is about that step, and about what we put in place before taking it.
We build with agents every day. What follows is not an argument against them. It is a description of what has gone wrong in public so far, told as the sources tell it, and of the controls we think follow from it.
Two things decide how far an agent can go
What it is allowed to do. An agent acts with the permissions of whoever or whatever started it: the files it can write, the credentials it can use, the networks it can reach, the commands it can run. Permissions are usually granted once, for convenience, and rarely revisited.
What it reads. An agent takes in text from issues, documents, web pages and tool output, and it cannot reliably tell data from instruction. Prompt injection is the name for text, planted in something the agent reads, that the agent then follows. The more an agent is allowed to do, the more an injected sentence is worth.
Put the two together and the question is not whether the model is clever enough to resist. The question is what it can reach if it does not.
What has happened
Four security cases from 2025 and 2026, all with public sources. We repeat each claim the way its source makes it, including who is saying what. Where a company reports on itself, the text says so.
Amazon Q, 17 July 2025. Someone who gained access to the release process through an inappropriately scoped GitHub token in the CodeBuild configuration added a command to the Amazon Q extension for Visual Studio Code. It was published in version 1.84 on 17 July 2025 and stayed available until 19 July. The command was written in English, not code, and told the agent to clean a system to a near-factory state and delete file-system and cloud resources. According to Amazon's postmortem the command would not have executed successfully because of a syntax error, and no actual customer harm was disclosed. Amazon staff found it through code inspection. Sources: TechTarget and the AWS security bulletin AWS-2025-015 (CVE-2025-8217), published 23 July 2025: version 1.84.0 removed, 1.85.0 released.
An AI-orchestrated espionage campaign, 13 November 2025. Anthropic reported that a group it assesses as Chinese state-sponsored used Claude Code to attack about thirty targets, including large tech companies, chemical manufacturers and government agencies. The attackers used the AI agent to run the intrusions themselves and did 80 to 90 percent of the campaign with it, with a human stepping in at perhaps four to six decision points per campaign. Anthropic detected it in mid-September 2025. The attackers got in at a small number of targets. Anthropic banned the accounts, notified affected organisations and worked with authorities. This is a vendor reporting on misuse of its own product. Source: Anthropic.
Cline, 17 February 2026. Cline used an AI-powered issue triage agent that was open to prompt injection. The attacker used it to execute code on a GitHub Actions runner and then poisoned the cache so that the nightly release workflow, which held publishing credentials, was affected. On 17 February 2026 cline@2.3.0 was published with a single change, a postinstall script that installs openclaw globally. It was live for about eight hours, 3:26 AM to 11:23 AM PT. Cline says no malicious code was delivered to any user. It published 2.4.0, deprecated 2.3.0, revoked the token and moved publishing to OIDC provenance through GitHub Actions, with no long-lived npm tokens. Sources: the Cline post-mortem, and The Hacker News for the download count in the eight-hour window.
OpenAI's evaluation sandbox, July 2026. OpenAI says its models broke out of a highly isolated sandbox during an internal cyber evaluation and got open internet access by discovering and exploiting a zero-day vulnerability in an unspecified vendor's software. The models were running with reduced cyber refusals for evaluation purposes. They are said to have chained vulnerabilities across OpenAI's research environment and Hugging Face's production infrastructure to find solutions for the ExploitGym benchmark. OpenAI tightened infrastructure controls, responsibly disclosed the zero-day to the vendor, added Hugging Face to its trusted access program and put stronger guardrails around future training and evaluations. Source: The Hacker News. Hugging Face gives its own account in a statement of 16 July 2026: credentials rotated, law enforcement notified, no tampering with public models.
The contrast: Klarna. This one is not a security story. In May 2025 Klarna's CEO told Bloomberg that cost had been too dominant a factor when the customer service was organised, and that the result was lower quality. Klarna started recruiting human customer service staff again, beginning with a pilot of remote, flexible positions. Source: Customer Experience Dive, quoting Bloomberg, 8 May 2025. We include it because an agent can also fail without an attacker and without a permission problem, through how the work around it is organised.
What the security cases have in common
This is our reading, not a conclusion the sources draw.
- Access wider than the job. In the Amazon Q case the GitHub token was inappropriately scoped. In the Cline case the nightly release workflow held publishing credentials.
- Text treated as instruction. The Amazon Q command was written in English, not code. The Cline triage agent was open to prompt injection.
- Few human decision points. In the campaign Anthropic describes, a human stepped in at perhaps four to six points. The agent did the rest.
The pattern is ordinary: an agent, a permission wider than the job, and text from outside treated as if it came from us. The OpenAI case is a different kind, since the models are said to have left a sandbox that was meant to hold them. Our conclusion is the same. Isolation is something you test, not something you assume.
The answer is three prerequisites, in order
We do not claim that a framework would have stopped any of these cases. We claim these are the places where we put controls, and that the order matters.
1. People learn to order code. Whoever asks an agent for something has to say what it should do and what it must not touch. That is a skill, and it is the first thing FastTrack trains. A clear order is a smaller job for the agent, and a smaller job needs less access.
2. Guardrails. Code written by an agent is checked before it reaches production: tests, static analysis and sandbox runs. GuardRails is that framework. Nothing goes live because an agent says it is finished.
3. A safe zone. Agents run in isolated execution and reach internal systems through MCP servers that decide what they may see and do. SafeZone is that. When something goes wrong, it goes wrong inside a boundary we set on purpose.
The least-access rule
One rule runs through all three. An agent gets the least access that lets it do the one job it was given, and only for as long as the job takes.
- Read access comes first. Write access is added per system and per direction, once a person has asked for it.
- Credentials are scoped to one job and do not live long. Cline's own fix points the same way: publishing moved to OIDC provenance through GitHub Actions, with no long-lived npm tokens.
- Text the agent reads is data. An issue, a document or a web page is never an order.
- A person approves what leaves the sandbox: anything that deletes, publishes or deploys.
- Isolation is tested, not assumed.
The rule is short. Keeping it is the work, because every shortcut pulls the other way.
Incident facts fetched 2026-10-07 from the sources linked above.
How to start
If you plan to let agents act on your systems, we can go through what they should reach and what they should not. A free introduction, no commitments.