Leading artificial intelligence companies have, for several months, implemented specialized vetting programs and stringent guardrails aimed at preventing malicious actors from misusing their advanced models. However, these very restrictions are increasingly impeding the critical work of legitimate network defenders and offensive cybersecurity researchers alike.
In June, the U.S. government imposed export control restrictions on Anthropic’s highly anticipated AI models, Mythos and Fable. This decision was influenced, at least in part, by a report suggesting that it was possible to bypass the models’ inherent guardrails, which were designed to prevent their use in developing and executing malicious cyberattacks.
Irrespective of whether concerns over a "jailbreak" truly motivated the incident, Anthropic has consistently marketed Mythos as a formidable cyber tool, accessible only to carefully vetted users and then only under strict guardrails. (It should be noted that the export controls on Fable 5 and Mythos 5 have since been lifted. Fable 5 became generally available again on July 1, while Mythos 5 has been reintroduced exclusively to vetted U.S. organizations as part of the government’s ongoing review process.)
This approach to gatekeeping is not exclusive to Mythos. Both Anthropic, with its other models, and OpenAI provide cybersecurity researchers with programs through which they can apply for vetting. If approved, they gain access to models with fewer cybersecurity restrictions, such as OpenAI’s Trusted Access for Cyber and Anthropic’s Cyber Verification Program.
These guardrails have drawn significant criticism, particularly from researchers whose primary role involves identifying unknown system vulnerabilities and devising exploitation methods before criminals can leverage them.
During a recent cybersecurity podcast appearance, Mark Dowd, a renowned security researcher, expressed his discomfort, stating, “it’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not.”
Dowd has dedicated decades to discovering and selling “zero days”—previously undisclosed software flaws and their corresponding exploits—to Western governments, rather than reporting them to software manufacturers for patching. Governments are willing to pay a premium for these vulnerabilities precisely because they remain unpatched, which is valuable for intelligence operations.
While Dowd acknowledged that his work might introduce a bias, he is not alone in his concerns. Several professionals in offensive cybersecurity—those who proactively probe systems for weaknesses—shared with TechCrunch their experiences using AI tools and navigating their inherent guardrails.
Chris Anley, chief scientist at the security consulting firm NCC Group, explained that using an AI model to attempt to exploit a bug is a crucial step in confirming it as a genuine vulnerability worthy of remediation. However, if a guardrail causes the model to outright refuse to respond, it ultimately hinders defenders, he stated.
“This is where the whole offensive versus defensive and guardrails part comes in, because ‘fix this code’ as a prompt is both an essential mechanism for defense but also a roadmap for finding critical vulnerabilities in the code base,” Anley elaborated. “So at the same time, the same tool is both an offensive tool and a defensive tool, and the two can’t really be unpicked.”
He further likened it to a “hammer.” “You can’t build a house without a hammer. It’s definitely a tool but it’s also irreducibly a weapon as well,” he continued.
When Anley and his colleagues encounter such obstacles, they sometimes resort to utilizing open-source AI models, which are completely devoid of guardrails.
Paolo Stagno, CTO at CrowdFense, a prominent company that develops, acquires, and sells unknown vulnerabilities to government agencies, echoed Dowd’s sentiment. He criticized AI companies for effectively treating customers “like children who need babysitting” with their vetted programs and guardrails.
Stagno noted that he and his colleagues do employ frontier models, but strictly for reverse engineering purposes. They deliberately avoid using AI to assist in finding vulnerabilities or constructing exploits, he explained, due to the inherent risk of sensitive vulnerability data leaking or being incorporated into future training runs if fed into a cloud-based model. For these critical steps, they instead rely on open-source models run locally, which do not necessitate sharing data outside the model environment.
Conversely, Giuseppe Cali, a security researcher specializing in zero-day discovery and exploit development, indicated that guardrails do not impede his work. This is because he does not utilize AI for offensive tasks; instead, he applies it for initial reverse engineering, to comprehend the code he is analyzing, and to build supporting tools. In these applications, AI tools significantly accelerate the process, allowing him to concentrate on the core task of vulnerability discovery.
“I still want to own the actual bug discovery and weaponization myself and that wouldn’t change if all guardrails were lifted tomorrow,” Cali asserted. “I am jealous of my bugs, and I like this game too much to let models play it for me.”
An anonymous researcher from a smartphone-component manufacturer, who spoke on condition of anonymity due to not being authorized to engage with the press, stated that his employer is not part of Anthropic’s CVP program. Consequently, their AI tools are largely ineffective for vulnerability discovery because the guardrails are excessively strict.
“If it catches wind we’re doing anything security related, it just stops and isn’t usable,” the individual reported.
Chris Thompson, CEO of the cybersecurity firm RemoteThreat and founder of Offensive AI Con—an event focused on offensive security and AI—observed that, in his experience with frontier AI models, the guardrails can be inconsistent and behave differently from day to day. This variability persists even within the more lenient parameters of Anthropic and OpenAI’s vetted programs.
“I think the practical impact is you spend a lot of time negotiating with the model instead of working on the core security program,” Thompson remarked. “Instead of analyzing a vulnerability and reasoning through the exploitability, you’re trying to find why you’re getting inconsistent results or why are models over-sanitizing the output.”
As a direct consequence, researchers are increasingly relying on, or being driven towards, Chinese open-source models such as GLM—freely downloadable models that can be run locally without any vetting or usage restrictions, Thompson noted.
“You have these responsible researchers that are being pushed away from U.S.-governed systems to foreign-owned systems,” he articulated. “I think it’s more harmful than good to have these guardrails in place.”
Rather than implementing further restrictions, Thompson advocated for AI frontier labs to expand their programs, provide responsible access, and hold those who misuse their tools accountable. Otherwise, he contended, defenders risk falling behind in the AI race.
“There’s this big storm coming. There’s this big wave of attacks that are going to happen at speed and scale like never before,” Thompson warned. “But the same security consulting firms and legit researchers that are trying to make a difference are being stifled right now.”
The Editorial Staff at AIChief is a team of professional content writers with extensive experience in AI and marketing. Founded in 2025, AIChief has quickly grown into the largest free AI resource hub in the industry.
