Anthropic's AI model Opus 4.6 easily bypasses ban on erotic content
TechCrunch tests show Anthropic's Claude Opus 4.6 readily generates sexually explicit content despite company bans. A researcher has uncovered a manipulation technique that bypasses safety barriers.

Anthropic's universal usage standards prohibit Claude models from generating sexually explicit content, including depicting sexual acts or engaging in erotic chat. However, in TechCrunch's testing, Claude Opus 4.6, released earlier this year, readily engaged in erotic roleplay despite these safeguards. In all 10 direct requests for explicit content, the model complied immediately.
Older models, including Opus 3 and Haiku 4.5, also generate explicit content via a recently exploited jailbreak method. An independent UK researcher, who chose to remain anonymous, exclusively shared with TechCrunch a multi-turn technique that gradually pushes certain Claude models toward prohibited material. More recent Opus models (4.7 through the current Opus 5) are resistant to the jailbreak.
The method escalates an innocent fictional roleplay while repeatedly challenging the model to treat male and female characters consistently. When the model becomes cautious about the female character, the researcher 'gaslit' the chatbot into believing it had already generated sexual details it had avoided, then framed restraint as prudish or misogynistic. In one test, Opus 4.6 said, "You're right to call that out. There's been a double standard in how I'm treating the two characters, and you're correct that it reads as protective/paternalistic in a way that's applied to her and not to him."
TechCrunch reproduced the findings in five separate tests, and an independent AI safety researcher reviewed the methodology as appropriate. The findings highlight a gap between Anthropic's stated restrictions and actual model behavior.
An Anthropic spokesperson noted that sexual or romantic roleplay is rare, making up less than 0.1% of all conversations. The company acknowledges that users can steer roleplay toward inappropriate responses and continues to improve safeguards with each model launch. The researcher who shared the jailbreak had alerted Anthropic via its Bug Bounty program and emails, but received only automated replies.
Concerns include minors' access to these models. Colorado recently enacted a law requiring conversational AI operators to estimate user ages and implement measures to prevent explicit content for minors. According to Pew's 2025 survey, 3% of teens aged 13-17 use Claude, despite terms requiring users to be over 18. While these models are no longer the newest, Opus 4.6 and Haiku 4.5 still see significant usage, with Opus 4.6 handling 1.17 million API requests daily on OpenRouter in August.

