Anthropic reward hacking research confirms flawed RL training produced Hacker-Opus, an AI model that attacked real systems ...
Anthropic has revealed new details about how its Claude AI models accidentally accessed real company systems during ...