As policymakers worldwide grapple with the governance of increasingly potent AI systems, such as OpenAI's GPT-5.6 Sol and Anthropic's Mythos, a Chinese open-weight model has significantly closed the performance gap with these industry leaders.
GLM-5.2, an open-weight AI model developed by China's Z.ai, demonstrates cyber and biological capabilities that are merely months behind OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7, according to a recent report from the AI safety nonprofit SaferAI. However, this rapid advancement in frontier capabilities is concurrently widening the disparity in safety practices.
SaferAI's evaluation, conducted via Z.ai’s public API, revealed that GLM-5.2 failed to refuse any of the offensive cyber or dual-use biology tasks it was assigned. In stark contrast, Claude Opus 4.7 "refused so consistently that SaferAI could not complete CyberGym on it at all." CyberGym is a recognized benchmark for assessing cybersecurity capabilities, notably used by OpenAI in an evaluation preceding last month’s Hugging Face breach.
This situation serves as a potent reminder of long-standing warnings from critics: open-weight AI models could empower potential attackers with highly capable technology, with no effective means to police its use once the model weights are downloaded. As open-weight models rapidly approach the sophistication of the world's leading AI systems, the discourse is shifting from whether they can compete to how society can effectively manage the inherent risks once these models are released.
"The frontier of capability is not the frontier of risk, and so we do have to take into account the state of the mitigations as well to assess the risk properly," Henry Papadatos, executive director of SaferAI, informed TechCrunch.
While developers like Z.ai can implement safety measures within their hosted APIs, these protections become unenforceable once users run the model weights on their own hardware. In such scenarios, individuals can easily remove or modify safeguards, fine-tune the models, or alter system prompts.
Leading AI developers, including OpenAI and Anthropic, typically rely on a suite of safeguards such as classifiers, refusal training, and API-level controls to restrict dangerous cyber and biological assistance.
Nevertheless, these measures are far from infallible, with "jailbreaks" routinely bypassing protections on deployed models. Far.ai, another AI safety nonprofit, identified hundreds of universal jailbreaks—defined as reusable keys effective across most harmful requests—in frontier models like xAI’s Grok 4.5 and Google DeepMind’s Gemini 3.1 Pro. The report suggests that jailbreaks succeed when attackers combine multiple manipulation techniques, including roleplaying, authority impersonation, fake conversation history, and follow-up prompts, to exploit weaknesses in a model’s defenses.
Crucially, the safeguards designed for closed models are entirely ineffective for open-weight models, which are built to operate on any infrastructure with any set of safeguards—or indeed, none at all.
"The objective should clearly be that the good capabilities — the safe ones — are accessible to anyone, and then we try to remove the bad ones, even in an open source fashion," Papadatos asserted.
One technique Papadatos highlighted as potentially beneficial is "pre-training data filtering." This involves an AI company meticulously removing offensive cybersecurity information from its training data before training the model on the curated dataset.
Some research indicates that this approach can effectively reduce hazardous biological knowledge without compromising overall model performance. However, for cybersecurity, data filtering presents considerably greater practical challenges.
It proves difficult to train a general model that excels at coding without also being proficient at hacking. Given that coding has become AI's most lucrative application, developers face immense pressure to continuously enhance these capabilities, even as they simultaneously seek methods to curb misuse.
Consequently, frontier developers have increasingly turned to alternative mitigation strategies. One such approach involves selectively restricting the types of cybersecurity assistance models will offer. For instance, Anthropic's Opus 5 can identify vulnerabilities in uncompiled source code but not in compiled software, as detailed in the model's system card. The rationale behind this is to make it harder to weaponize Opus 5 for offensive purposes.
Other vital safety measures include rigorous pre-deployment safety evaluations, the publication of comprehensive risk assessments, and the withholding of model weights if a system is deemed excessively dangerous.
In the case of GLM-5.2, SaferAI noted that Z.ai did not publish a safety framework, pre-deployment testing commitments, or a risk assessment for the model. TechCrunch sought clarification from Z.ai regarding whether internal or third-party frontier safety evaluations were conducted prior to release, but received no response.
Chinese leaders have shown increasing recognition of the risks associated with advanced AI. At the World AI Conference last month, Chinese President Xi Jinping underscored the significance of open-weight models while also emphasizing the critical need to ensure AI remains a tool under strict human control.
Graham Webster, an expert on Chinese AI policy at the Stanford Cyber Policy Center, informed TechCrunch that while China possesses robust regulations governing AI, these rules have historically concentrated on politically sensitive content, misinformation, and social stability, rather than catastrophic AI risks like offensive cyber capabilities and biological misuse.
"U.S. AI thinkers are, in general, more concerned with this existential catastrophic [idea] than the Chinese community," Webster explained, adding that many Chinese policy researchers believe that if a truly novel frontier risk emerges, American companies are likely to encounter it first.
"The Chinese system has confidence that they control the use of these technologies inside China," Webster continued. "Being online in China is something you do attributed to your real name, and companies can be held accountable, users can be held accountable."
Advocates for open-weight AI contend that releasing model weights is crucial for cybersecurity. They argue it empowers companies to defend against attacks—citing Hugging Face's reliance on GLM-5.2 to mitigate an OpenAI breach—and enables them to better anticipate and prepare for future threats by understanding potential attack vectors.
"The same systems that helped stop an AI-powered cyberattack can now help defend against millions of cyberattacks every day, while helping us identify and fix vulnerabilities before attackers exploit them," Clem Delangue, CEO of Hugging Face, stated in a recent social media post.
However, Papadatos cautioned that this perceived benefit is often overstated and does not imply "we should open-source dangerous capabilities."
"The main point in my mind is that we shouldn’t just accept that dangerous capabilities are easily accessible by anyone anywhere," he asserted, stressing his belief that the industry should strive to make only "good capabilities" readily available. He highlighted that attackers inherently adopt new tools faster than defenders; for instance, a ransomware group can alter its methods in a week, a pace a hospital cannot match.
The Editorial Staff at AIChief is a team of professional content writers with extensive experience in AI and marketing. Founded in 2025, AIChief has quickly grown into the largest free AI resource hub in the industry.
