
Apollo Research tests whether advanced models can deceive evaluators, pursue hidden goals or evade human control. Its behavioral evaluations and interpretability research provide frontier labs and policymakers with concrete evidence about forms of model behavior that ordinary benchmarks miss.
See something inaccurate or outdated?



