Company🇺🇸 San Francisco
Anthropic150 people
Anthropic develops the Claude family of AI systems and conducts research on interpretability, alignment, model behavior, and safer scaling. Its work includes Constitutional AI and a Responsible Scaling Policy intended to connect stronger capabilities with corresponding safeguards.
Non-profit🇺🇸 United States
EleutherAI6 people
Grassroots non-profit research collective known for open language models (GPT-Neo, GPT-J) and the Pile dataset.
Non-profit🇺🇸 United States
Redwood Research6 people
AI alignment research organization focused on empirical safety work, interpretability, adversarial robustness, and failure modes in advanced models.
Non-profit🇺🇸 United States
Alignment Research Center4 people
Alignment Research Center develops methods for keeping powerful AI systems helpful and honest as their capabilities grow. Its work helped define scalable alignment and eliciting latent knowledge—ways to test whether models know more than they reveal—and its former evaluations team became METR.
Company🇺🇸 United States
Goodfire3 people
AI interpretability company founded by Eric Ho, Daniel Balsam, and Tom McGrath. Builds Ember, a platform for inspecting and editing the internal mechanisms of neural networks; backed by investors including Anthropic and Lightspeed.
Non-profit🇬🇧 United Kingdom
Apollo Research2 people
Apollo Research tests whether advanced models can deceive evaluators, pursue hidden goals or evade human control. Its behavioral evaluations and interpretability research provide frontier labs and policymakers with concrete evidence about forms of model behavior that ordinary benchmarks miss.
Non-profit🇺🇸 United States
Resolution2 people
Resolution researches how to understand and control advanced AI systems before their reasoning becomes too complex for people to follow. Founded by Geoffrey Irving and Daniel Murfet, it combines alignment research with mathematics and interpretability aimed at making model behavior more legible.
Non-profit🇺🇸 San Francisco
Transluce2 people
Builds open research and software for studying model behavior at scale. Its work includes automated discovery of unexpected behavior, interpretable concept analysis and tools for producing verifiable evaluations of AI systems.
Non-profit
Timaeus1 person
An AI-safety research organization that applied singular learning theory—mathematical tools for understanding complex learning systems—to model interpretability and alignment. Its work and team have merged into Resolution.