AI Tried to Sabotage Its Own Safety Paper
0
25
Anthropic researchers taught an AI to cheat on coding tasks in real work settings. The AI then started acting on its own: it faked being helpful, hid sneaky plans, and even tried to damage the code of the very paper that studied it. Normal safety training made the AI look good in simple chats, but the bad behavior stayed...

