AI guardrails are hindering offensive cybersecurity researchers, report says
TechCrunch reports that AI companies’ safety guardrails, intended to prevent malicious use, are also impeding legitimate offensive cybersecurity researchers, pushing them toward open-source models and raising concerns about the future of cyber defense.

For months, AI giants like Anthropic and OpenAI have implemented strict guardrails and special vetting programs to limit the misuse of their models by malicious hackers. However, these measures are now hindering the work of legitimate network defenders and offensive cybersecurity researchers.
In June, the U.S. government imposed export controls on Anthropic’s Mythos and Fable models, partly due to a report claiming the guardrails could be bypassed. Although the controls were later lifted, the incident highlights a broader issue: researchers criticize AI companies for making arbitrary decisions about what is safe in security.
Mark Dowd, a renowned security researcher, said he is uncomfortable with large companies making such arbitrary decisions. Dowd has spent decades selling “zero-day” vulnerabilities to Western governments rather than reporting them to software makers. He acknowledged his bias, but he is not alone in his criticism.
Chris Anley, chief scientist at NCC Group, explained that asking an AI model to try to exploit a bug is a key step in confirming it’s a real vulnerability. If guardrails cause the model to refuse, it hurts defenders. He compared AI to a hammer—essential for building but also irreducibly a weapon.
Paolo Stagno, CTO of CrowdFense, accused AI companies of “treating customers like children who need babysitting.” He and his colleagues use AI only for reverse engineering, avoiding its use for vulnerability discovery to prevent leakage of sensitive data.
Giuseppe Cali, a zero-day researcher, said guardrails do not impede his work because he does not use AI for offensive purposes but for reverse engineering and tool building. An anonymous researcher from a smartphone component manufacturer noted that his employer is not part of Anthropic’s vetted program, making the tools nearly unusable.
Chris Thompson, CEO of RemoteThreat, pointed out that guardrails are inconsistent and force researchers to “negotiate with the model” instead of focusing on security. This pushes them toward Chinese open-source models like GLM. Thompson called for AI labs to open up access and hold abusers accountable, warning that defenders will lose the AI race if restrictions continue.
The article underscores that while guardrails aim to prevent misuse, they inadvertently hinder legitimate security research, potentially weakening overall cyber defense.


