Skip to content
agents benchmark · included in ECI

Remote Labor Index

Remote Labor Index benchmark (score column: Score).
Results
15
Random baseline
0.0%
Score ceiling
100%
Released
29 Oct 2025

Score by model release date

Best score by organisation

#ModelScoreRelativeEvidence
1GPT-6 Astra (unknown thinking)
OpenAI
20.8%Source ↗
2Claude Fable 5.1 (unknown)
Anthropic
17.9%Source ↗
3Claude Fable 5
Anthropic
16.1%Source ↗
4Claude Fable 5 (unknown)
Anthropic
15.8%Source ↗
5Claude Opus 4.8
Anthropic
8.3%Source ↗
6GPT-5.5 (unknown thinking)
OpenAI
6.3%Source ↗
7Gemini 3.7 Flash (unknown)
Google DeepMind
5.0%Source ↗
8Claude Opus 4.6 (unknown thinking)
Anthropic
4.2%Source ↗
9Claude Opus 4.5 (unknown thinking)
Anthropic
3.8%Source ↗
10GPT-5.2 (medium)
OpenAI
2.5%Source ↗
11GPT-5.2 (unknown thinking)
OpenAI
2.1%Source ↗
12Claude Sonnet 4.5 (unknown thinking)
Anthropic
2.1%Source ↗
13GPT-5 (unknown thinking)
OpenAI
1.7%Source ↗
14Gemini 3 Pro Preview
Google DeepMind
1.3%Source ↗
15Gemini 2.5 Pro Preview (Jun 2025)
Google DeepMind
0.8%Source ↗

Caveats: settings (reasoning effort, agent scaffold) differ between rows and materially affect scores; the best reported setting per model is shown. Source: Epoch AI, CC BY 4.0.