
METR measures what frontier AI agents can do in realistic, extended tasks rather than short benchmark questions. Its widely followed time-horizon evaluations estimate the length of software and research tasks models can complete, giving labs and policymakers a concrete view of rapidly changing autonomous capability.
See something inaccurate or outdated?








