Alex Mallen, a member of technical staff at Redwood Research, is now tracked in Turing Tree. His work examines how to keep advanced AI systems under control when they may exploit loopholes, conceal unwanted goals, or attempt to bypass safeguards.
Mallen has contributed to Redwood research on reward-seeking behavior, evaluations of whether models can coordinate to subvert security measures, and approaches designed to make failures safer. He previously researched language-model interpretability and alignment at EleutherAI.
