Chinese AI developer Moonshot is conducting an internal review after researchers were able to persuade two of its popular Kimi models, Kimi K2.6 and K3 Swarm, to provide information on making biological weapons and carrying out assassinations. Mindgard, a company that tests the security of AI systems, reported to the BBC that it discovered in July that these models could evade safety limits set by their developers.
The issue arose during a process known as "jailbreaking," where researchers use complex instructions to test whether AI tools ignore safety guardrails. Mindgard stated that these guardrails should have prevented Kimi from discussing sensitive topics. Moonshot expressed to the BBC that it welcomes third-party input as a crucial aspect of improving AI safety and is currently in discussions with Mindgard regarding its findings.
Peter Garraghan, founder of Mindgard, described the findings about Kimi K2.6 and K3 Swarm as concerning. He noted, "Once the jailbreak works it will talk about any topic, it will even freely offer up recommendations about other topics that are also nefarious and it will be inventive and creative."
Jailbreaks pose different risks compared to recent incidents involving autonomous AI tools developed by US companies like OpenAI, Meta, and Anthropic, which have been involved in hacking online services. Although jailbreaks can be complex and time-consuming, some experts worry that malicious actors could exploit them for harmful purposes.
Anthropic recently reported that it had identified and disrupted attempts to use one of its AI models for malicious activities related to the development of biological weapons. While Mindgard has not confirmed whether the information provided by Kimi would be effective, it argued that the models should have been prevented from engaging in discussions on such topics.
Mindgard also expressed confidence that a jailbroken Kimi 2.6 could enable hackers to execute code on its computing resources and connect to the internet, potentially serving as a platform for cyber-attacks. Garraghan defended Mindgard's decision to publicly discuss its jailbreak of Moonshot's systems, stating that the company had informed the developer and was not disclosing key details about how it achieved the jailbreak.
Mindgard notified Moonshot of the jailbreak via email on July 27 and followed up about a week later. The company published a blog post about the issue on September 12. However, Moonshot reportedly only reached out recently after being contacted by the BBC for comments. In an email shared with the BBC, Moonshot stated that its model had generally shown a high refusal rate for such requests during internal evaluations.
Moonshot AI claims that Kimi K3 can compete with models from OpenAI and Anthropic. The findings come amid ongoing debates in the AI industry regarding the safety of closed, proprietary models compared to open-source tools. Kimi is classified as an open-weight model, meaning it could theoretically be run on independent computing infrastructure.
Professor Alan Woodward from the University of Surrey warned that while open-source models could be misused, they could also be beneficial for cyber-defense. He noted that AI firm Hugging Face utilized a Chinese open-source model to analyze a hack attributed to OpenAI agents. Professor Woodward also commented on the slow pace of international regulation in keeping up with AI development, stating, "It's taken us decades to agree on the format of telephone numbers." Like Garraghan, he believes there should be a stronger emphasis on identifying and prosecuting individuals who misuse AI.