Joe Benton announced on X that he left Anthropic to join METR, where he will work on embedded assessment of AI risks. He said the move took place the previous week.
At Anthropic, Benton managed the Scalable Oversight team, which studies how people can reliably supervise AI systems that may outperform them on relevant tasks. He also led research for the Anthropic Fellows Program and worked on evaluations designed to expose deceptive or misaligned model behavior.
Benton said independent assessment could help ensure the public learns when an AI company faces a rapid capability increase or a loss-of-control incident. METR has been developing third-party assessments that combine model evaluations with access to internal systems and information from frontier AI developers.