20202021202220232024202520264 sec36 sec6 min1 hour10 hoursTask duration (for humans)where logistic regression of our datapredicts the AI has a 50% chance of succeedingLLM release dateTrain adversarially robust image modelTrain classifierFind fact on webCount words in passageAnswer questionGPT-2GPT-3GPT-3.5GPT-4GPT-4oClaude 3.5 Sonnet (Old)o1-previewo1o3GPT-5GPT-5.2(high)Time horizon of software tasks different LLMs can complete 50% of the time

Source: METR Time Horizons · Analysis code on Github · Raw data here