Organizations

Organizations shaping AI

Explore companies, research labs, universities, nonprofits, and public institutions across the AI ecosystem.

Browse organizations
Sort
View
12 organizations · showing 1-12
Company🇺🇸 San Francisco
16 people

Scale AI provides data, evaluation, deployment, and reliability infrastructure for frontier AI labs, enterprises, and government customers. It remained an independent company after Meta’s significant 2025 minority investment, when founder Alexandr Wang left to join Meta and Jason Droege became CEO.

Non-profit🇺🇸 United States
8 people

METR measures what frontier AI agents can do in realistic, extended tasks rather than short benchmark questions. Its widely followed time-horizon evaluations estimate the length of software and research tasks models can complete, giving labs and policymakers a concrete view of rapidly changing autonomous capability.

Company🇺🇸 United States
4 people

AI safety and security startup founded by Bo Li, Dawn Song, Carlos Guestrin, Sanmi Koyejo, and collaborators, developing tools for automated red teaming, runtime guardrails, and AI governance.

Company🇺🇸 San Francisco
3 people

Arena operates a community-powered platform for evaluating frontier AI systems through anonymous comparisons and real-world human feedback. Originating as UC Berkeley’s Chatbot Arena, it publishes public leaderboards across language, vision, coding, search, video, and agentic tasks and offers evaluation services to model developers.

Government🇺🇸 Washington, D.C.
Center for AI Standards and Innovation
3 people

CAISI is the US government's central technical body for testing commercial AI systems and developing voluntary standards for their security and evaluation. Based within NIST, it assesses capabilities and vulnerabilities with implications for cybersecurity, biosecurity, national security, and international competition.

Non-profit🇬🇧 United Kingdom
2 people

Apollo Research tests whether advanced models can deceive evaluators, pursue hidden goals or evade human control. Its behavioral evaluations and interpretability research provide frontier labs and policymakers with concrete evidence about forms of model behavior that ordinary benchmarks miss.

Company🇺🇸 United States
1 person

Independent AI benchmarking and analysis company tracking model quality, speed, price, and deployment tradeoffs across frontier and open models.

Company🇺🇸 United States
1 person

Research-driven product company building verification and governance tools for AI agents, checking their decisions against human-readable specifications during development and production.

Non-profit🇺🇸 United States
LMSYS Org
1 person

LMSYS Org (Large Model Systems Organization) is a nonprofit that grew out of a multi-university collaboration centered at UC Berkeley, known for the Chatbot Arena LLM evaluation platform and the Vicuna open model.

Company
Thoughtful
1 person

An independent research lab developing benchmarks and tools for language-model post-training. Its PostTrainBench tests whether frontier agents can autonomously improve open models under fixed compute limits.

Company🇺🇸 San Francisco
1 person

Independently evaluates AI models and applications on practical tasks in law, finance, software, and other professions. Builds benchmarks with domain experts and publishes comparisons of performance, cost, and speed.

Non-profit🇺🇸 United States
MLCommons
0 people

MLCommons is a nonprofit AI engineering consortium out of the MLPerf benchmarking effort. It develops the MLPerf benchmarks along with shared datasets and tools for measuring ML performance, safety, and efficiency.