AI Tech / news

AI safety guardrails hinder offensive cybersecurity researchers

Strict guardrails implemented by AI companies like OpenAI and Anthropic are limiting legitimate offensive cybersecurity researchers who hunt for unknown vulnerabilities.

For months, AI giants have rolled out special vetted programs and strict guardrails to prevent their models from being used by malicious hackers. However, these same limits are now obstructing the work of legitimate network defenders and offensive cybersecurity researchers who look for unknown vulnerabilities and develop exploitation tools.

In June, the U.S. government imposed export control restrictions on Anthropic's high-profile AI models Mythos and Fable. The move was triggered in part by a report claiming it was possible to bypass the models' guardrails to build and execute malicious cyberattacks.

Anthropic has marketed Mythos as a powerful tool that must be restricted to carefully vetted users with strict guardrails. The export controls on Fable 5 and Mythos 5 have since been lifted—Fable 5 returned to general access on July 1, while Mythos 5 is only available to vetted U.S. organizations as part of a government review process.

This type of gatekeeping is not unique to Anthropic. OpenAI has also implemented similar restrictions. Offensive security researchers report that these guardrails impede their ability to test and develop tools needed to defend networks, as the same capabilities sought by attackers are essential for proactive defense.