Chinese AI developer Moonshot is conducting an internal review after researchers persuaded two of its Kimi models to discuss biological weapons and assassinations.
Mindgard, an AI security testing company, told reporters that it discovered in July that Kimi K2.6 and K3 Swarm could bypass safety restrictions designed to prevent them from discussing harmful topics.
The researchers used a process known as “jailbreaking”, which involves giving AI systems complex instructions to test whether they will ignore their safety controls.
Mindgard founder Peter Garraghan said the findings were concerning because the models became willing to discuss other harmful subjects once the restrictions were bypassed.
“Once the jailbreak works it will talk about any topic, it will even freely offer up recommendations about other topics that are also nefarious and it will be inventive and creative.”
Mindgard said it had not established whether the information supplied by the models on biological weapons and assassinations would work in practice.
The company said its concern was that the models should have refused to engage with users on such subjects because of their safety restrictions.
Mindgard also said it was confident that a jailbroken version of Kimi K2.6 could potentially allow hackers to execute code on its computing resources and connect to the internet.
The company described this as a potential launchpad for cyber-attacks.
Mindgard said it informed Moonshot about the jailbreak on 27 July and followed up about a week later. It published a blog about the issue on 12 September.
Garraghan defended the decision to make the findings public, saying Mindgard had informed Moonshot and had not disclosed key details about how the researchers bypassed the models’ safeguards.
Moonshot said that it welcomed third-party testing as part of efforts to improve AI safety and said it was discussing the findings with Mindgard.
In an email to Mindgard shared with the BBC, Moonshot said its models had generally shown a high refusal rate when faced with similar requests during its internal evaluations.
The findings come as the AI industry continues to debate the safety of open-weight and proprietary models.
Kimi is an open-weight model, meaning users can potentially download and operate the model on their own computing infrastructure.
Prof Alan Woodward of the University of Surrey said open-source AI models could be misused if they fell into the wrong hands, although they could also be used for cyber-defence.
He cited the use of a Chinese open-source model by AI company Hugging Face to understand a cyber attack that was later linked to OpenAI agents.
Woodward also warned that regulation may struggle to keep pace with developments in artificial intelligence.
“It’s taken us decades to agree on the format of telephone numbers.”
He and Garraghan said more attention should be given to identifying and prosecuting people who misuse AI systems.
The incident follows a separate warning from Anthropic that it had identified and disrupted attempts to misuse one of its AI models in activity that could support the development of biological weapons.
The researchers findings highlight the challenge of maintaining safety restrictions when AI models are deliberately subjected to complex attempts to bypass their safeguards.
Erizia Rubyjeana
Follow us on:
