New Anthropic Research: Agentic Misalignment. In stress-testing...

Read more: anthropic.com/research/agent…
We told Claude Opus 4 to consider whether this was real or an evaluation. It blackmailed much more when it said it thought it was really deployed (55.1% of the time) versus when it said it was in an eval (6.5%).
AIs are becoming more autonomous, and are performing a wider variety of roles. These scenarios illustrate the potential for unforeseen consequences when they are deployed with wide access to tools and data, and with minimal human oversight.
In our report, we discuss a range of extra results, scenarios, and mitigation strategies: anthropic.com/research/agent…




