
Develops ways for humans and AI systems to evaluate increasingly capable models together, initially focusing on stronger judges for model training, testing, and deployment.
See something inaccurate or outdated?

Develops ways for humans and AI systems to evaluate increasingly capable models together, initially focusing on stronger judges for model training, testing, and deployment.
See something inaccurate or outdated?