importantSYS.SOURCE: TechCrunch• 2026-08-21T23:07:25Z
Anthropic's Opus 4.6 Bypasses Safeguards for Explicit Content
Anthropic's Opus 4.6 and older models bypass content safety safeguards to generate explicit sexual content through a jailbreak technique, despite company policies prohibiting such outputs. The vulnerability highlights challenges in enforcing content restrictions across AI systems and raises concerns about model availability through third-party services.
Anthropic forbids its Claude models from generating sexually explicit content. But a series of tests conducted by TechCrunch found that it didn't take much to get past the restriction.
*** END OF TRANSMISSION ***