Skip to content

Research

Four axes of intelligence: putting frontier models on a physical ruler

One chart I check whenever it updates measures AI capability in hours. I understand hours even when I do not know which architecture is fashionable.

METR Time Horizon 1.1: software-task difficulty in human expert hours, at 50% agent success
The hours belong to the human doing the task. Source: METR, Time Horizon 1.1.

The AI evaluation research organization METR first estimates how long each software task takes a human expert. It then tests AI agents on those tasks and fits a curve relating human task duration to an agent's probability of success. The duration where that curve reaches 50% is the agent's 50% time horizon.4

A two-hour horizon therefore means that the fitted success rate is 50% for tasks of roughly that human-rated difficulty. The AI might finish a successful attempt in minutes. The number is neither its running time nor how far it can literally see into the future. Individual tasks vary, and the result describes the tested task distribution.