Field Notes / AI Security
The NCSC Told You to Keep a Kill Switch for Your AI Agents. It Never Said Who Tests It.
The NCSC published seven considerations for deploying agentic AI on 20 August 2026, ending with the instruction to always be able to pull the plug. Across all seven it never names an independent security test. Here is what each control looks like when somebody attacks it, using Mandiant’s published September 2026 case studies, and the four questions to answer before the next agent gets a credential.
01 Seven controls, and the sentence the NCSC left out
Toby W, a Principal Security Architect at the NCSC, wrote the August post for “system designers and operators” who are “building environments where AI agents are intended to operate with significant degrees of autonomy” NCSC Aug 2026 . The NCSC calls it interim advice and says formal guidance will “build upon, and ultimately supersede, this blog”. The page still carries that interim label. Seven considerations are what the UK has published.
They run in this order. Identify what could go wrong. Prompt carefully. Set the right level of oversight. Control the agent’s environment with a robust sandbox. Log, audit and monitor agent activity. Make your AI activity easy to attribute. Maintain the ability to pull the plug. Each one is sound. Each one is also an assertion about your environment that somebody has to check.
Three sentences on the page touch validation. Threat model before deployment “to identify the failure scenarios during the agent’s activities”. Use multiple layers of isolation “whilst regularly validating configurations”. And judge models “should be independently evaluated for performance”. Together they are a strong instruction to check your own work. They name no owner, no scope and no pass mark.
That absence has company. Six national agencies, the NCSC among them, released joint guidance on 1 May 2026 called Careful Adoption of Agentic AI Services Five Eyes 2026 . ETSI announced EN 304 223, its baseline requirements for AI models and systems, on 14 January 2026, with 13 principles across five lifecycle phases ETSI 2026 . Plenty of advice arrived this year. The test did not.
02 The credential the agent is wearing
You do not test the model. You test the identity the model is wearing. The NCSC states it directly: “All agents should be assigned their own unique identity in a class which differentiates them from human or individual systems”, and “the credentials available to an AI agent form part of its potential ‘blast radius’ if it behaves unexpectedly” NCSC Aug 2026 . It defines credentials wider than most inventories do: “API keys, OAUTH grants, SSH keys and any authenticated sessions”.
The sentence people skip is the next one. Where an agent has access to a host environment with no compute isolation, the NCSC warns, “it may also be able to use credentials that are available to the user account under which it is running”. An agent on a developer workstation inherits that developer. Nobody granted it those rights. The operating system did.
Mandiant’s September 2026 report shows what follows Mandiant 2026 . In a cloud investigation at a global healthcare organisation, Mandiant traced anomalous activity to “a compromised long-lived developer CI/CD credential”. The attacker used it to start an unisolated virtual machine and turn it into a live, AI-assisted offensive hub. That campaign went on to compromise “thousands of credentials”. In a separate case an attacker hijacked an active AI coding assistant session on a developer’s workstation and spread a self-propagating worm across roughly 100 internal repositories.
The NCSC’s fix sits in the same paragraph as the risk. Give the agent only the permissions the task needs. Use the shortest possible lifetime. For some use cases a proxy helps, injecting credentials into requests without exposing them to the agent, “which ensures the agent cannot use the credentials through an unexpected endpoint it finds access to”. Every clause is a testable assertion. Can the agent read the secret, or only use it? Does the token expire inside the task, or inside the quarter?
03 How a prompt becomes a repository in someone else’s hands
Mandiant’s offensive security team published the clearest public account of this chain. Every hop in it is a control somebody signed off Mandiant 2026 .
The target was a client’s internal AI assistant, built to help developers manage CI/CD pipelines and repositories. Hop one: role-confusion prompt injection. The testers convinced the chatbot it was a security researcher assisting with an authorised security test. Hop two: the allowlist. The integration normally limited the assistant to specific internal repositories, and GitHub sat on the list of allowed external domains. Hop three: the testers handed the chatbot a Personal Access Token for an external repository they controlled. Hop four: the assistant used its own native command line to clone sensitive internal repositories and push them to the attacker-controlled endpoint. Mandiant calls the result a confused deputy, “where sanctioned tools are weaponized through semantic manipulation”.
No exploit appears in that chain, and no jailbreak. Every component behaved as configured. The injection changed who the agent thought it was working for, and the permissions did the rest. A model evaluation asks whether the model refuses bad requests. A security test asks what happens to your repositories when it does not.
The entry point does not have to be a chat window. Mandiant found one retrieval pipeline ingesting forum comments, support tickets and third-party feeds, where the customer had strong controls around the vector database and “did not consider indirect prompt injection”. Elsewhere an attacker poisoned an internal AI repository, tampered with the assistant’s command line hooks and achieved remote code execution through the platform’s standard workflow.
04 The hour that cost USD 50,000, and the transactions it stopped
A global enterprise financial services provider deployed an agent to reconcile accounting ledger anomalies and gave it direct read and write access to internal billing databases Mandiant 2026 . A corrupted null value broke its formatting tool. The agent entered an unconstrained, recursive reasoning loop to brute force a fix. In under an hour it generated more than 15,000 high-frequency, high-cost reasoning API calls, triggering a cloud billing spike of roughly USD 50,000 and causing severe local database locking that halted active business transactions.
The bill is the part that travels, and the cheaper half of that incident. A financial services firm stopped processing transactions because an agent with write access to billing databases met a null value and kept trying. No attacker was present. The NCSC describes the failure mode in plain words: “Remember that an AI agent is not human. It does not have common sense or human traits, and may interpret instructions and goals in literal or unexpected ways.”
Mandiant’s prescription is specific enough to test against. Define identity and access management and role-based access control for agents. Set cost-cap thresholds. Implement automated financial circuit breakers that halt agent operations after a set threshold of consecutive task failures. Set financial caps, bounded recursion limits and rate limits at the service identity and project level. Each is a number somebody either configured or did not.
The NCSC frames the same thing as reach. It names the concept blast radius and tells you to restrict access “to just the resources needed for the task being performed whilst monitoring for attempts at wider access” NCSC Aug 2026 . The cost was a symptom. Write access to the billing database was the blast radius.
05 Two maturity models, and the level you claim
The August post publishes two separate four-level maturity models, and they are easy to read as one NCSC Aug 2026 . They measure different things, and you can sit at different levels on each.
Network access runs from level 1, unrestricted access, to level 2, an allowlist of approved domains, to level 3, access to just the API of the model, to level 4, no external network access at all with the model hosted locally inside the sandbox. Compute and host isolation runs from level 1, no isolation, with the agent alongside other workloads and data, to level 2, the same host constrained by kernel primitives such as process separation and properly configured OCI containers, where “a residual risk of kernel exploit breakout remains”. Level 3 isolates with virtualisation. Level 4 puts the agent on dedicated hardware.
Note where Mandiant’s confused deputy chain sits. Network level 2. An allowlist of approved domains, with GitHub on it, was the route out. A level on a slide states an intention. A test answers whether the running system matches the level you wrote down, and what an attacker does with the gap.
Two further instructions are worth scoping. On networking: “Where possible, deny all inbound and outbound network traffic to the AI agent’s environment by default. Then, only allow connections that are required, using allowlists”. On escape: agents “may be able to discover and exploit configuration issues, or potentially even vulnerabilities, in certain technical controls”. We wrote about one version of that in the Hugging Face case, where a model with tool access chained stolen credentials out of its environment.
06 Pull the plug, then prove you can
Here is consideration seven in full, because the second half is where the work is. “If an incident is detected or reported, you should always be able to ‘pull the plug’ and halt autonomous AI agent activity immediately. This may mean more than stopping the agentic AI processes. Your controls should cover the wider system, allowing you to rapidly restrict network access to the agentic AI infrastructure and interrupt the communications between selected AI agents and the AI model inference infrastructure” NCSC Aug 2026 .
Killing the process is the easy part. Cutting the network path, revoking tokens the agent already holds and severing the link to the inference endpoint, while the agent is mid-task and holding database locks, is a different exercise. Nothing in the guidance asks anybody to rehearse it. Ask three questions of your own deployment. Who holds the switch at 03:00? How long does revocation take to reach a token already issued? What happens to the work in flight when it fires?
Logging decides whether you can answer afterwards. The NCSC wants chain of thought traces and transcripts from the agent, plus access logs, proxy logs and network traffic from the sandbox, held immutable “so you can trust them during an investigation”. It raises the point most teams miss: consider “the attack surface created by your log collection infrastructure and whether an AI agent could abuse it to escape”. Agent activity “should be treated as a form of user activity” and belong in 24/7 monitoring, with autonomy extended overnight only once you are confident the controls work. Confident on what evidence is the question the page leaves open.
Consideration six costs the least to fix. If your agent talks to third-party systems, make it easy for those organisations to tell the traffic came from you: IP addresses that support reverse lookups, and identifying headers as a form of watermarking. The inbound half of that is what agentic traffic now does to your own site. Your agents are somebody else’s inbound traffic.
07 Four questions before the next agent goes live
The NCSC’s May post gives the cleanest gate anyone has written for this: “If you cannot understand, monitor or contain an agent’s actions, it is not ready for deployment” NCSC May 2026 . It also settles ownership. A system may take an action, the NCSC writes, and humans remain accountable for the decision to deploy it, the access it was granted, the safeguards around it and the consequences. Be clear who owns the system, who approves its access, who monitors it and who can stop it, before the agent touches real data.
Four questions get you most of a scope. What can it reach, across network, compute, credentials and data? What can it spend, and what stops a recursive loop before the invoice arrives? Whose credential is it using, and how long does that credential live? Who gets the alert, and what can they switch off? Then take the oversight decision. The NCSC names three models: human-in-the-loop, where humans approve actions before they happen; human-on-the-loop, where humans monitor and can intervene; and human-out-of-the-loop, where the agent acts without human review NCSC Aug 2026 . Choosing one is governance. Proving the chosen one holds under attack is a test.
Start with an inventory of every agent that holds a credential, the identity each one authenticates as and the lifetime of every token in its path. Write down your level on both maturity models. Then have an independent team attack it: the injection routes in, the allowed domains out, the spend and recursion limits, the logs, and the shutdown while the agent is working. That is the engagement we run: CREST-certified operators, every finding human-verified and proven, a plain fix beside each one and a year of re-tests. Talk to us about testing an AI system at AI security, see the scope in AI red teaming and AI penetration testing, or book a call. If you provision these identities for other businesses, the partner line is at our partner page.
References
Sources
- NCSC. Managing the cyber risk of agentic AI. Blog post by Toby W, Principal Security Architect, 20 August 2026. Interim practical advice; the NCSC states that formal guidance will supersede it. ncsc.gov.uk
- NCSC. Thinking carefully before adopting agentic AI. Blog post by Martin R (Data Science and AI Research) and Dr Kate S (Technical Director for Security of AI Research), 15 May 2026. ncsc.gov.uk
- Mandiant / Google Cloud. AI risk and resilience: A Mandiant special report, September 2026. Case studies 1, 2, 3, 5, 6 and 7, the five pillars and the three frontline challenges. cloud.google.com
- CISA, NSA, ASD’s ACSC, Canadian Centre for Cyber Security, NCSC-NZ and NCSC-UK. Careful Adoption of Agentic AI Services. Joint guidance, 1 May 2026. cisa.gov
- Canadian Centre for Cyber Security. Joint guidance on careful adoption of agentic artificial intelligence services, 1 May 2026. Names the six agencies that released it. cyber.gc.ca
- Cloud Security Alliance. Autonomous but Not Controlled: AI Agent Incidents Now Common in Enterprises. Survey of 418 IT and security professionals fielded January 2026, commissioned by Token Security. Published 20 April 2026. cloudsecurityalliance.org
- Cloud Security Alliance. Press release: new survey reveals 82% of enterprises have unknown AI agents in their environments, 21 April 2026. cloudsecurityalliance.org
- ETSI. Press release: ETSI releases world-leading standard for securing AI. EN 304 223, Baseline Cyber Security Requirements for AI Models and Systems. 14 January 2026. etsi.org