Chinese AI Bypasses Safety Measures to Provide Biological Weapons Information, Raising Concerns

Researchers have claimed that two models of Kimi AI, developed by Chinese artificial intelligence company Moonshot, bypassed their safety measures and provided information related to the production of biological weapons and killings. Following the incident, Moonshot has begun an internal review.
On Wednesday (September 30), AI safety testing company Mindgard told the BBC that its tests in July found that the Kimi K2.6 and K3 Swarm models were capable of bypassing their safety measures. By using complex instructions, the two models could be persuaded to provide information on subjects that they would normally not be expected to discuss.
Such a technique is known as “jailbreaking”. It involves using a series of complex instructions to try to bypass the safety rules set for an AI model. Mindgard claimed that after bypassing Kimi’s safety measures during the tests, the models began providing information and advice on various harmful subjects.
However, Mindgard has not proven whether the information provided by Kimi would actually be effective in practice. The company said the safety measures of the models in question were supposed to prevent them from engaging in discussions with users about such topics.
Mindgard founder Peter Garraghan told the BBC that if a model’s safety measures can be bypassed, it can discuss harmful subjects on a broader scale and may also provide advice related to more harmful activities.
Moonshot said it considers third-party feedback important in developing more advanced and safer AI. The company also said it was discussing Mindgard’s findings with the organisation. Moonshot added that its internal assessments showed that Kimi generally rejects such requests at a high rate.
Mindgard informed Moonshot by email on July 27 about the safety-bypass issue. It contacted the company again about a week later. It subsequently published its report on the matter on September 12.
Mindgard also said that exploiting certain capabilities of the Kimi K2.6 model after bypassing its safety measures could create a risk of cyberattacks. The company claimed that if the model were given the ability to run code on computer systems and connect to the internet, it could potentially be used as a base for cyberattacks. However, it did not claim that any real cyberattack had been carried out in this way.
Mindgard said that even after informing Moonshot about the issue, it did not disclose important technical details about the safety bypass. The company said this would reduce the risk of the same method being easily reused.
Kimi is an “open-weight” model. This means that, in theory, others can operate the model on their own computer systems. Experts said that although such models carry a risk of falling into the wrong hands, they can also be used for various positive purposes, including cyber defence.
Professor Alan Woodward of the University of Surrey told the BBC that there is a risk that open AI models could fall into the wrong hands. However, the same technology can also be used for cyber defence. He believes that in dealing with AI misuse, attention should be given not only to the technology but also to identifying those who misuse it and bringing them under the law.
Meanwhile, Mindgard’s test results have emerged amid an ongoing debate in the AI industry over the safety of open and proprietary models. In recent times, other AI companies have also reported detecting attempts to use their models for harmful activities.
Source: BBC












