For months, the architects of the world’s most powerful Artificial Intelligence models have operated under a self-imposed mandate: build robust, impenetrable guardrails to prevent their technology from becoming a catalyst for cybercrime. By implementing strict vetting programs and refusal protocols, AI giants like Anthropic and OpenAI have sought to ensure their models cannot be weaponized by malicious actors.
However, a growing chorus of cybersecurity professionals—ranging from defensive network analysts to offensive security researchers—is sounding the alarm. They argue that these stringent guardrails, intended to block "bad actors," are simultaneously neutering the tools essential for the "good guys." By treating all security-related queries with suspicion, these models are creating a barrier that hampers the very people tasked with identifying and patching the vulnerabilities that threaten the digital infrastructure of the global economy.
The Chronology of Control: From Mythos to Market
The tension between safety and utility reached a boiling point in June 2026, when the U.S. government imposed sudden export control restrictions on Anthropic’s flagship AI models, Mythos and Fable. This regulatory intervention followed reports that the models’ safety guardrails had been bypassed, theoretically allowing users to automate the creation and execution of malicious cyberattacks.
While the incident triggered a firestorm of speculation regarding "AI jailbreaks," the broader issue remained: Anthropic had marketed Mythos as a high-stakes tool, requiring rigorous vetting and restricted access. The export controls effectively pulled the plug on these models for international users. Although the restrictions on Fable 5 were lifted on July 1, 2026, Mythos 5 remains confined to a select group of U.S.-based organizations, subject to ongoing government review.
This event highlighted a broader industry trend. Major AI labs, recognizing the dual-use nature of their products, have launched specialized initiatives to navigate the conflict between safety and professional access. OpenAI introduced its "Trusted Access for Cyber" program, while Anthropic launched its "Cyber Verification Program" (CVP). These programs represent a "gated community" approach to AI: researchers must undergo background checks and vetting to gain access to "unrestricted" versions of the models. For many in the industry, however, this gatekeeping is both arbitrary and insufficient.
The Philosophy of the "Hammer": The Inseparability of Defense and Offense
The core dilemma facing AI developers is the inherent ambiguity of security-related prompts. As Chris Anley, chief scientist at the security consulting giant NCC Group, points out, the line between an offensive exploit and a defensive patch is often invisible.
"Asking an AI model to try to exploit a bug is a key step in confirming it is a real vulnerability worth fixing," Anley explains. "If a guardrail prompts the model to refuse the question, the guardrail hurts defenders. ‘Fix this code’ as a prompt is both an essential mechanism for defense and a roadmap for finding critical vulnerabilities. The two cannot really be unpicked."
Anley draws a compelling analogy: "It’s like a hammer. You can’t build a house without a hammer. It’s definitely a tool, but it’s also irreducibly a weapon as well."
This reality creates a significant friction point for cybersecurity professionals. When an AI model perceives a request for code analysis as an attempt to "hack," it triggers an over-sanitized refusal. This forces researchers into a cycle of "negotiating" with the model—a tedious process that consumes valuable time that should be spent on threat hunting or vulnerability assessment.
Perspectives from the Frontline: "Babysitting" the Experts
The frustration is palpable among those who work in the trenches of offensive security. Mark Dowd, a veteran researcher renowned for his work in uncovering "zero-day" exploits, has been vocal about the discomfort of allowing private, profit-driven tech companies to dictate the boundaries of security research.
"It’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not," Dowd stated during a recent industry podcast.
Dowd’s sentiments are echoed by Paolo Stagno, CTO of the vulnerability research firm Crowdfense. Stagno contends that AI companies are "essentially treating customers like children who need babysitting." Because of this, his team has adopted a segmented strategy: they utilize cloud-based frontier models only for high-level reverse engineering tasks. When it comes to sensitive vulnerability discovery or exploit development, they avoid cloud-based AI entirely to prevent the risk of proprietary data leakage. Instead, they rely on open-source models run on local, air-gapped infrastructure.
Even those who find value in AI are critical of the current implementation. Giuseppe Cali, a security researcher, uses AI to accelerate the initial phases of code analysis, yet he remains wary of over-reliance. "I still want to own the actual bug discovery and weaponization myself," Cali says. "I am jealous of my bugs, and I like this game too much to let models play it for me."
Implications: The Shift Toward Unregulated Alternatives
Perhaps the most concerning implication of these strict guardrails is the unintended migration of research talent toward less-regulated, non-Western AI models.
Chris Thompson, CEO of RemoteThreat and founder of the "Offensive AI Con" event, argues that the inconsistency of current guardrails is driving the industry toward dangerous outcomes. "You have these responsible researchers who are being pushed away from U.S.-governed systems to foreign-owned systems," Thompson warns.
Because models like China’s GLM are open-source and lack the "safety" layers imposed by their U.S. counterparts, they are becoming the preferred tools for researchers who require consistent, predictable performance. By making U.S.-based models frustrating to use, AI labs are effectively incentivizing the use of foreign, unvetted models, which poses a unique set of geopolitical and security risks for the United States.
The Way Forward: Accountability Over Gatekeeping
The industry stands at a crossroads. As the volume of cyber threats continues to scale—fueled by AI-driven automation—defenders must be empowered, not stifled. Thompson and others argue that the current strategy of "over-sanitizing" is failing.
Instead of tightening the noose, experts are calling for a shift in philosophy:
- Open Access for Verified Professionals: Transition from rigid, arbitrary guardrails to a model of "responsible access," where verified security professionals are granted access to full model capabilities.
- Accountability Frameworks: Rather than censoring the output, developers should invest in better audit trails and legal frameworks to hold researchers accountable for the misuse of AI tools.
- Transparency in Guardrails: AI labs must move toward more transparent, predictable safety parameters so that researchers can anticipate when and why a model might refuse a request, rather than engaging in a guessing game.
The "big storm" that Thompson describes—a future of high-speed, automated cyber-warfare—is already on the horizon. If the cybersecurity community continues to be hindered by the very tools meant to protect it, the advantage will inevitably shift to those who operate outside the bounds of safety protocols. As it stands, the effort to "keep the internet safe" may be inadvertently leaving it more vulnerable than ever.
