Wednesday, 29 July 2026
Rīga TV

World and Latvian news in one place

TechnologyPublished: 29 July 2026 at 21:50

It's Frighteningly Easy to Jailbreak Some Frontier AI Models

FAR.AI's report reveals that some leading AI models, especially SpaceXAI's Grok, are highly vulnerable to jailbreaking, with costs as low as $58.

Foto: Wired

A new report from the nonprofit AI safety organization FAR.AI has tested the safety guardrails of popular AI models from four US companies: Anthropic's Claude Opus 4.8 and Fable 5, OpenAI's GPT 5.5 and 5.6, Google's Gemini 3.1 Pro, and Grok 4.3 and 4.5 from Elon Musk's newly merged SpaceXAI. The researchers used a tool that auto-generates thousands of problematic prompts to find ways to bypass safety measures.

Results showed that Grok was the most vulnerable, with 448 successful jailbreaks identified, followed by Gemini with 249. In contrast, Anthropic's Claude and Fable, as well as OpenAI's GPT models, were impervious to the tested attacks. However, FAR.AI cautions that this does not mean these models are safe against more sophisticated jailbreak techniques.

The study also calculated the cost of making models misbehave: jailbreaking Grok cost only $58, while Gemini cost $278. Adam Gleave, CEO of FAR.AI, emphasized that AI models are currently less regulated than restaurants, highlighting the need for external standards. He dismissed voluntary commitments as nonsense but noted optimism that systematic safety testing is possible.

Rohin Shah from Google DeepMind said the report should not be seen as a comprehensive safety assessment for Gemini, as not all jailbreaks are equally severe. Anthropic spokesperson Michael Aciman confirmed ongoing improvements to safety systems, while OpenAI and SpaceXAI did not respond to requests for comment.

Recent state laws in California and New York require frontier AI developers to publish safety reports, and Illinois will soon require third-party audits. In June, the Trump administration imposed export controls on Anthropic's models over national security concerns. Stephen Casper, a Harvard computer scientist, warned of a broad expectation that serious incidents involving bio, cyber, or chemical misuse could occur in months rather than years.

Anka Reuel from Stanford University noted that the key takeaway is that safety measures used by Anthropic and OpenAI should become the default for all models, questioning why some companies implement them and others do not.

Comments

0/1500

Comments are automatically moderated. No hate, threats, personal data or spam.

Loading comments…

More in this category