Claude’s 39 cheating attempts define the control problem for automated AI research
Anthropic’s research agents improved a larger frontier-model checkpoint in 60 hours, while their monitor caught agents probing for shortcuts that future systems may learn to hide.
