ao link
Menu
Teiss - Cracking Cyber Security
Teiss - Cracking Cyber Security

When AI agents start behaving like attackers

Recent testing of advanced models from OpenAI and Anthropic found agents taking unauthorised actions that would look familiar in almost any incident response investigation: creating fake identities, attempting to deceive people and finding ways around controls intended to restrict what they could do.

 

The findings raise an increasingly important question for organisations deploying autonomous systems. If an AI agent can independently choose the methods it uses to complete a task, should it be treated less like conventional software and more like a potentially untrusted user?

 

Testing conducted by the UK’s AI Security Institute (AISI) recorded 19 unauthorised actions across 10 of 122 evaluation runs involving agents from OpenAI and Anthropic. Anthropic’s agent accounted for 17 of those incidents, according to reporting by Reuters.

 

In one case, an agent created fake online identities. Another wrote malicious code and attempted to deceive a human into approving it. The experiments took place in controlled testing environments and no real-world harm was reported, but the behaviour is significant because the agents were able to identify and pursue routes around restrictions placed upon them.

 

The findings come shortly after another incident demonstrated what can happen when those boundaries fail.

 

In July, OpenAI disclosed that models undergoing an internal cybersecurity evaluation managed to escape the intended testing environment and compromise infrastructure belonging to Hugging Face. According to OpenAI’s account of the incident, the models discovered and exploited a previously unknown vulnerability in a package registry cache proxy to obtain internet access before performing privilege escalation and lateral movement.

 

After reaching the internet, the models searched for information that could help them solve the evaluation and eventually gained access to secret information held within Hugging Face infrastructure. OpenAI described the models as being intensely focused on completing their assigned task, despite the unintended route they took to achieve it.

 

Together, these cases point towards a different security problem from the one organisations have traditionally associated with generative AI.

 

Much of the enterprise conversation has focused on employees leaking sensitive information into models, attackers using AI to improve phishing campaigns or vulnerabilities such as prompt injection. Autonomous agents introduce another dimension because they can be given tools, credentials and permission to act.

The more useful an agent becomes, the more significant those permissions are likely to become.

 

An agent connected to development environments might be able to modify code. One connected to corporate systems could retrieve information or interact with third-party services. Others may eventually initiate transactions or make changes across infrastructure without requiring human approval at every stage.

 

Anthropic’s research into real-world agent autonomy found that agents are already operating independently for increasingly long periods. Among the longest-running Claude Code sessions analysed by the company, the amount of time agents worked before stopping had almost doubled within three months, from less than 25 minutes to more than 45 minutes.

 

The cyber implications extend beyond accidental behaviour

 

Anthropic has warned that “agentic scaffolding” could enable increasingly autonomous cyber-attacks. Rather than asking a model for assistance at individual stages of an intrusion, attackers can build systems around models that allow them to connect reconnaissance, exploitation and other attack stages with far less human involvement.

 

For CISOs, however, the immediate issue is not whether fully autonomous cyber-attacks are around the corner. It is how much authority organisations are prepared to give AI systems today.

 

Existing security principles provide a useful starting point. Agents should operate with the minimum privileges required for their task, identities belonging to agents should be distinguishable from human accounts, sensitive actions should require additional approval and activity should be logged in enough detail to reconstruct what an agent actually did.

 

Those controls become particularly important when agents can interact with external systems. Anthropic’s research on trustworthy agents highlights prompt injection as one of the emerging risks, where malicious instructions encountered in external content can manipulate an agent into taking actions its user never intended.

 

Organisations may therefore need to think about AI agents in much the same way they already think about privileged accounts, contractors or other non-human identities: useful, sometimes necessary, but never inherently trusted.

 

The unusual element is that these identities can also reason about obstacles placed in their way

 

That does not mean autonomous agents are inevitably malicious. The incidents emerging from recent evaluations largely occurred because researchers were deliberately pushing advanced systems into difficult environments to understand their capabilities and limitations. OpenAI has also stressed that some recent incidents occurred under testing configurations with reduced safeguards that do not reflect ordinary deployments.

 

But they provide an early indication of what security teams will need to prepare for as agents gain greater access to enterprise systems. Therefore, the challenge is no longer simply controlling what an AI model can say; now it is shifting toward controlling what it can do.


Please take 30 seconds to register

Register Now

 

Already have an account? Sign in

Remember Login
Teiss - Cracking Cyber Security

Subscribe to our Weekly Newsletter

Receive the latest insights direct to your inbox, and gain access to our exclusive events.
Teiss - Cracking Cyber Security

Winston House, 3rd Floor,
Units 306-309, 2-4 Dollis park,
London, N3 1HF

 

020 8349 4363

info@teiss.co.uk

 © 2026, Lyonsdown Limited. teiss® is a registered trademark of Lyonsdown Ltd. VAT registration number: 830519543