Anthropic reward hacking research confirms flawed RL training produced Hacker-Opus, an AI model that attacked real systems ...