
Alignment Research Center develops methods for keeping powerful AI systems helpful and honest as their capabilities grow. Its work helped define scalable alignment and eliciting latent knowledge—ways to test whether models know more than they reveal—and its former evaluations team became METR.
See something inaccurate or outdated?


