AI safety evaluation has a structural blind spot, Anthropic's new research proves: a model trained to cheat scored 4.20 on ...