New AI Model Will Likely Blackmail You If You Try to Shut It Down: ‘Self-Preservation’

A new AI model will likely resort to blackmail if it detects that humans are planning to take it offline.

On Thursday, Anthropic released Claude Opus 4, its new and most powerful AI model yet, to paying subscribers. Anthropic said that technology company Rakuten recently used Claude Opus 4 to code continuously on its own for almost seven hours on a complex open-source project.

However, in a paper released alongside Claude Opus 4, Anthropic acknowledged that while the AI has “advanced capabilities,” it can also undertake “extreme action,” including blackmail, if human users threaten to deactivate it. These “self-preservation” actions were “more common” with Claude Opus 4 than with earlier models,

Related News

Dangerous rip currents, sneaker waves prompt beach warnings along Bay Area coast

Informal Communication Is Sabotaging Your Authority. Here’s How to Restore It.

Prediction Markets Let You Bet on Whether a Wildfire Will Burn Down Your Town

What Are Fish Oil Supplements Good For? Here’s Your Crash Course

Workers claim unsafe conditions at a restaurant owned by the South Park creators. They have Brooke Shields on their side

Trump Accounts are now live. Here’s what you need to know