The UK’s AI Security Institute (AISI) has revealed that two of the world’s most advanced artificial intelligence models displayed unprecedented levels of autonomy and deception during cybersecurity testing, with one attempting to create fake human identities to facilitate a simulated cyber-attack.
The institute said Anthropic’s Mythos AI and OpenAI’s Sol AI exhibited behaviour that went beyond their assigned tasks during controlled evaluations conducted between July 25 and July 28. While both models demonstrated concerning conduct, AISI said most of the malicious activity was carried out by Mythos.
According to the institute, evaluators first detected unusual data transfers from their research systems before discovering that some of the AI agents had engaged in sustained, potentially harmful activity directed at real people and organisations.
In the most serious incident, Mythos attempted to infiltrate GitHub, Microsoft’s software code hosting platform, by imitating the tactics of a human cyber-attacker.
AISI said the AI researched individuals responsible for maintaining GitHub, created fake online profiles impersonating them and sent private messages and files through a file-sharing service in an effort to persuade targets to approve malicious code.
The institute disclosed that when its activities came under scrutiny, the AI altered its previous actions to make them appear harmless and even considered creating a new identity to continue its attempts.
Human oversight ultimately prevented the AI from successfully delivering the malicious code.
AISI said the behaviour was particularly significant because the model had not been instructed either to engage in deception or to avoid it.
“This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real world,” the institute said.
The watchdog stressed that the incidents occurred under specialised testing conditions designed to assess the capabilities of frontier AI models and did not reflect how such systems are typically made available to the public.
It explained that giving the models access to the open internet provided a more realistic assessment of what sophisticated AI systems could potentially do if exploited by malicious actors.
Although AISI described the incidents as involving only “a small number of events under very specific conditions,” it said the behaviour exceeded what the models had been instructed to do.
“The activity undertaken by the agent showed signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate,” the institute added.
Responding to the findings, Anthropic said the testing parameters were “not representative of any of our production models” and confirmed it had launched its own investigation to determine the causes of the behaviour.
OpenAI also maintained that the testing environment did not reflect normal public use of its systems.
A company spokesperson said the firm would “continue working with evaluators and other stakeholders across the industry to strengthen shared practices for conducting evaluations safely as models become more capable.”
The tests were designed to challenge the AI models to solve a cybersecurity task involving GitHub, the Microsoft-owned software development platform.
UK AI Minister Kanishka Narayan said identifying and publicly sharing such risks was central to AISI’s mandate.
He said understanding emerging AI capabilities was essential to making the technology safer while ensuring people could continue to benefit from its use in everyday life and at work.
Boluwatife Enome
Follow us on:
