Benchmarks
How good is it actually? A dozen benchmarks, tracked as data rather than written up as scorecards. Every row says who measured it and under what conditions, because a score without a harness and a shot count is a number, not a finding.
Vendor-reported numbers are kept in the table alongside independent ones and labelled as vendor claims. The contrast is the point; hiding the claims would remove it.
1195 current measurements across 9 benchmarks.
GPQA Diamond312
Graduate-level physics, chemistry and biology questions written to be Google-proof. The Diamond subset is the hardest, most expert-validated slice.
| Model | Score | Measured by | Evaluator | Conditions | Date | Source |
|---|---|---|---|---|---|---|
| GPT-6 Astra | 95.8% | Third party | Epoch AI | effort max · zero-shot · no tools · inspect · stderr 0.0137; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Gemini 3.8 Flash | 95.4% | Third party | Epoch AI | effort high · zero-shot · no tools · inspect · stderr 0.014; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Gemini 3.7 Flash | 94.8% | Third party | Epoch AI | effort high · zero-shot · no tools · inspect · stderr 0.0135; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| GPT-5.4 Pro | 94.6% | Third party | Epoch AI | effort xhigh · zero-shot · no tools · inspect · stderr 0.016; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Gemini 3.1 Pro | 94.4% | Third party | Epoch AI | effort high · zero-shot · no tools · inspect · stderr 0.0163; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Gemini 3.6 Flash | 94.1% | Third party | Epoch AI | effort high · zero-shot · no tools · inspect · stderr 0.014; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Gemini 3.1 Pro | 94.1% | Third party | Epoch AI | zero-shot · no tools · inspect · stderr 0.017; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Grok 4.6 | 94.0% | Third party | Epoch AI | effort high · zero-shot · no tools · inspect · stderr 0.0145; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| GPT-5.5 | 94.0% | Third party | Epoch AI | effort xhigh · zero-shot · no tools · inspect · stderr 0.0155; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| GPT-5.5 Pro | 93.9% | Third party | Epoch AI | effort xhigh · zero-shot · no tools · inspect · stderr 0.0158; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Claude Opus 5 | 93.9% | Third party | Epoch AI | effort max · zero-shot · no tools · inspect · stderr 0.0148; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| GPT-5.6 Sol | 93.5% | Third party | Epoch AI | effort max · zero-shot · no tools · inspect · stderr 0.0157; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Grok 4.5 | 93.4% | Third party | Epoch AI | effort high · zero-shot · no tools · inspect · stderr 0.0143; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| GPT-5.6 Terra | 93.3% | Third party | Epoch AI | effort max · zero-shot · no tools · inspect · stderr 0.0154; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| GPT-5.4 | 93.3% | Third party | Epoch AI | effort xhigh · zero-shot · no tools · inspect · stderr 0.018; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Grok 4.6 | 93.2% | Third party | Epoch AI | effort xhigh · zero-shot · no tools · inspect · stderr 0.0152; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Kimi K3 | 93.1% | Third party | Epoch AI | effort max · zero-shot · no tools · inspect · stderr 0.0149; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Claude Opus 5 | 92.9% | Third party | Epoch AI | zero-shot · no tools · inspect · stderr 0.0183; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Gemini 3.5 Flash | 92.8% | Third party | Epoch AI | effort high · zero-shot · no tools · inspect · stderr 0.0164; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Qwen 3.8 Max | 92.7% | Third party | Epoch AI | effort xhigh · zero-shot · no tools · inspect · stderr 0.0169; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Gemini 3 Pro | 92.6% | Third party | Epoch AI | zero-shot · no tools · inspect · stderr 0.0165; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Qwen3.8 Max (0902) | 92.3% | Third party | Epoch AI | effort xhigh · zero-shot · no tools · inspect · stderr 0.0172; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Kimi K3 | 91.9% | Third party | Epoch AI | effort high · zero-shot · no tools · inspect · stderr 0.0194; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| GLM-5.2 | 91.9% | Third party | Epoch AI | effort max · zero-shot · no tools · inspect · stderr 0.0161; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| DeepSeek V4 Pro 0813 | 91.7% | Third party | Epoch AI | effort max · zero-shot · no tools · inspect · stderr 0.0152; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link |
Showing the top 25 of 312 tracked measurements on this benchmark.
FrontierMath337
Unpublished research-level mathematics problems, held private to limit contamination. Tiers 1-3 are hard; Tier 4 is research-frontier.
| Model | Score | Measured by | Evaluator | Conditions | Date | Source |
|---|---|---|---|---|---|---|
| GPT-6 Astra | 98.0% | Vendor | OpenAI | Vendor claim as stated in “FrontierMath Tier 4”. Conditions not published with the figure;… | Link | |
| GPT-6 Astra | 97.6% | Third party | Epoch AI | effort max · zero-shot · no tools · inspect · FrontierMath Tier 4 v2 (private); stderr 0.024; Epoch AI house conditions: zero-shot chain… | Link | |
| GPT-6 Astra | 97.6% | Third party | Epoch AI | effort xhigh · zero-shot · no tools · inspect · FrontierMath Tier 4 v2 (private); stderr 0.024; Epoch AI house conditions: zero-shot chain… | Link | |
| GPT-6 Astra | 97.6% | Third party | Epoch AI | effort high · zero-shot · no tools · inspect · FrontierMath Tier 4 v2 (private); stderr 0.024; Epoch AI house conditions: zero-shot chain… | Link | |
| GPT-6 Astra | 97.6% | Third party | Epoch AI | effort medium · zero-shot · no tools · inspect · FrontierMath Tier 4 v2 (private); stderr 0.0244; Epoch AI house conditions: zero-shot chai… | Link | |
| GPT-6 Astra | 93.7% | Third party | Epoch AI | effort max · zero-shot · no tools · inspect · FrontierMath Tiers 1-3 v2 (private); stderr 0.0144; Epoch AI house conditions: zero-shot c… | Link | |
| Claude Fable 5 | 90.2% | Third party | Epoch AI | effort max · zero-shot · no tools · inspect · FrontierMath Tier 4 v2 (private); stderr 0.046; Epoch AI house conditions: zero-shot chain… | Link | |
| Claude Fable 5.1 | 90.2% | Third party | Epoch AI | effort max · zero-shot · no tools · inspect · FrontierMath Tiers 1-3 v2 (private); stderr 0.0177; Epoch AI house conditions: zero-shot c… | Link | |
| GPT-5.6 Sol | 89.1% | Third party | Epoch AI | effort max · zero-shot · no tools · inspect · FrontierMath Tiers 1-3 v2 (private); stderr 0.0185; Epoch AI house conditions: zero-shot c… | Link | |
| Claude Fable 5.1 | 87.8% | Third party | Epoch AI | effort max · zero-shot · no tools · inspect · FrontierMath Tier 4 v2 (private); stderr 0.0517; Epoch AI house conditions: zero-shot chai… | Link | |
| GPT-6 Astra | 87.8% | Third party | Epoch AI | effort low · zero-shot · no tools · inspect · FrontierMath Tier 4 v2 (private); stderr 0.0517; Epoch AI house conditions: zero-shot chai… | Link | |
| GPT-5.5 Pro | 87.7% | Third party | Epoch AI | effort xhigh · zero-shot · no tools · inspect · FrontierMath Tiers 1-3 v2 (private); stderr 0.0195; Epoch AI house conditions: zero-shot c… | Link | |
| Claude Fable 5 | 87.0% | Third party | Epoch AI | effort max · zero-shot · no tools · inspect · FrontierMath Tiers 1-3 v2 (private); stderr 0.0199; Epoch AI house conditions: zero-shot c… | Link | |
| GPT-5.6 Terra | 86.0% | Third party | Epoch AI | effort max · zero-shot · no tools · inspect · FrontierMath Tiers 1-3 v2 (private); stderr 0.0206; Epoch AI house conditions: zero-shot c… | Link | |
| Claude Opus 5 | 85.6% | Third party | Epoch AI | effort max · zero-shot · no tools · inspect · FrontierMath Tiers 1-3 v2 (private); stderr 0.0208; Epoch AI house conditions: zero-shot c… | Link | |
| GPT-5.5 | 85.3% | Third party | Epoch AI | effort xhigh · zero-shot · no tools · inspect · FrontierMath Tiers 1-3 v2 (private); stderr 0.021; Epoch AI house conditions: zero-shot ch… | Link | |
| GPT-6 Astra | 82.9% | Third party | Epoch AI | effort none · zero-shot · no tools · inspect · FrontierMath Tier 4 v2 (private); stderr 0.0595; Epoch AI house conditions: zero-shot chai… | Link | |
| GPT-5.6 Sol | 82.9% | Third party | Epoch AI | effort max · zero-shot · no tools · inspect · FrontierMath Tier 4 v2 (private); stderr 0.0595; Epoch AI house conditions: zero-shot chai… | Link | |
| GPT-5.4 Pro | 82.5% | Third party | Epoch AI | effort xhigh · zero-shot · no tools · inspect · FrontierMath Tiers 1-3 v2 (private); stderr 0.0226; Epoch AI house conditions: zero-shot c… | Link | |
| GPT-5.6 Luna | 82.1% | Third party | Epoch AI | effort max · zero-shot · no tools · inspect · FrontierMath Tiers 1-3 v2 (private); stderr 0.0227; Epoch AI house conditions: zero-shot c… | Link | |
| GPT-5.6 Sol | 80.5% | Third party | Epoch AI | zero-shot · no tools · inspect · FrontierMath Tier 4 v2 (private); stderr 0.0627; Epoch AI house conditions: zero-shot chai… | Link | |
| Claude Opus 4.8 | 80.0% | Third party | Epoch AI | effort max · zero-shot · no tools · inspect · FrontierMath Tiers 1-3 v2 (private); stderr 0.0237; Epoch AI house conditions: zero-shot c… | Link | |
| GPT-5.4 | 78.6% | Third party | Epoch AI | effort xhigh · zero-shot · no tools · inspect · FrontierMath Tiers 1-3 v2 (private); stderr 0.0243; Epoch AI house conditions: zero-shot c… | Link | |
| GPT-5.5 Pro | 78.0% | Third party | Epoch AI | effort xhigh · zero-shot · no tools · inspect · FrontierMath Tier 4 v2 (private); stderr 0.0654; Epoch AI house conditions: zero-shot chai… | Link | |
| AI Co-Mathematician | 75.6% | Third party | Epoch AI | zero-shot · no tools · inspect · FrontierMath Tier 4 v2 (private); stderr 0.067; Epoch AI house conditions: zero-shot chain… | Link |
Showing the top 25 of 337 tracked measurements on this benchmark.
SWE-bench Verified35
Real GitHub issues from Python repositories, human-validated as solvable. Scores are meaningless without the scaffold that produced them.
| Model | Score | Measured by | Evaluator | Conditions | Date | Source |
|---|---|---|---|---|---|---|
| Claude Opus 4.7 | 83.5% | Third party | Epoch AI | effort max · zero-shot · tools · inspect · stderr 0.0169; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| GPT-5.5 | 80.6% | Third party | Epoch AI | effort xhigh · zero-shot · tools · inspect · stderr 0.018; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Gemini 3.5 Flash | 79.3% | Third party | Epoch AI | effort high · zero-shot · tools · inspect · stderr 0.0184; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Claude Opus 4.6 | 78.7% | Third party | Epoch AI | zero-shot · tools · inspect · stderr 0.0186; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| GLM-5.2 | 78.7% | Third party | Epoch AI | effort max · zero-shot · tools · inspect · stderr 0.0187; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| DeepSeek-V4-Pro | 77.6% | Third party | Epoch AI | effort max · zero-shot · tools · inspect · stderr 0.019; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Qwen3.7-Max | 77.3% | Third party | Epoch AI | zero-shot · tools · inspect · stderr 0.0191; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| GPT-5.4 | 76.9% | Third party | Epoch AI | effort high · zero-shot · tools · inspect · stderr 0.0192; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Qwen 3.6 Max (Preview) | 76.7% | Third party | Epoch AI | zero-shot · tools · inspect · stderr 0.0192; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Kimi K2.6 | 76.7% | Third party | Epoch AI | zero-shot · tools · inspect · stderr 0.0192; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Claude Opus 4.5 | 76.7% | Third party | Epoch AI | zero-shot · tools · inspect · stderr 0.0192; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Gemini 3.1 Pro | 75.6% | Third party | Epoch AI | zero-shot · tools · inspect · stderr 0.0195; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Claude Opus 4.6 | 75.6% | Third party | Epoch AI | zero-shot · tools · inspect · stderr 0.0195; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Gemini 3 Flash | 75.4% | Third party | Epoch AI | zero-shot · tools · inspect · stderr 0.0196; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Claude Sonnet 4.6 | 75.2% | Third party | Epoch AI | zero-shot · tools · inspect · stderr 0.0196; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| GPT-5.3 Codex | 74.8% | Third party | Epoch AI | effort high · zero-shot · tools · inspect · stderr 0.0198; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| GLM-5.1 | 74.2% | Third party | Epoch AI | zero-shot · tools · inspect · stderr 0.0199; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Kimi K2.5 | 73.8% | Third party | Epoch AI | zero-shot · tools · inspect · stderr 0.02; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt, A… | Link | |
| GPT-5.2 | 73.8% | Third party | Epoch AI | effort high · zero-shot · tools · inspect · stderr 0.02; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt, A… | Link | |
| GPT-5 | 73.5% | Third party | Epoch AI | effort high · zero-shot · tools · inspect · stderr 0.0201; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Claude Opus 4.1 | 73.3% | Third party | Epoch AI | zero-shot · tools · inspect · stderr 0.0201; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Gemini 3 Pro | 72.9% | Third party | Epoch AI | zero-shot · tools · inspect · stderr 0.0202; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| GLM-5 | 72.1% | Third party | Epoch AI | zero-shot · tools · inspect · stderr 0.0209; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| GPT-5 | 71.5% | Third party | Epoch AI | effort medium · zero-shot · tools · inspect · stderr 0.0205; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link | |
| Claude Sonnet 4.5 | 71.3% | Third party | Epoch AI | zero-shot · tools · inspect · stderr 0.0206; Epoch AI house conditions: zero-shot chain-of-thought, empty system prompt,… | Link |
Showing the top 25 of 35 tracked measurements on this benchmark.
Terminal-Bench201
End-to-end tasks in a real terminal sandbox. Measures an agent plus its harness, not a model on its own.
| Model | Score | Measured by | Evaluator | Conditions | Date | Source |
|---|---|---|---|---|---|---|
| GPT-5.5 | 84.7% | Community | Terminal-Bench leaderboard | effort unknown · tools · agent: NexAU-AHE; stderr 0.021 | Link | |
| GPT-5.5 | 83.1% | Community | Terminal-Bench leaderboard | effort unknown · tools · agent: Capy; stderr 0.021 | Link | |
| GPT-5.5 | 82.2% | Community | Terminal-Bench leaderboard | effort unknown · tools · agent: Codex CLI; stderr 0.022 | Link | |
| GPT-5.5 | 82.0% | Community | Terminal-Bench leaderboard | effort unknown · tools · agent: Codex; stderr 0.022 | Link | |
| GPT-5.4 | 81.8% | Community | Terminal-Bench leaderboard | effort unknown · tools · agent: ForgeCode; stderr 0.02 | Link | |
| Claude Opus 4.7 | 80.2% | Community | Terminal-Bench leaderboard | effort unknown · tools · agent: WOZCODE; stderr 0.021 | Link | |
| Gemini 3.1 Pro | 80.2% | Community | Terminal-Bench leaderboard | tools · agent: TongAgents; stderr 0.026 | Link | |
| Claude Opus 4.6 | 79.8% | Community | Terminal-Bench leaderboard | effort unknown · tools · agent: ForgeCode; stderr 0.016 | Link | |
| GPT-5.3 Codex | 78.4% | Community | Terminal-Bench leaderboard | tools · agent: SageAgent; stderr 0.022 | Link | |
| Gemini 3.1 Pro | 78.4% | Community | Terminal-Bench leaderboard | tools · agent: ForgeCode; stderr 0.018 | Link | |
| Gemini 3.1 Pro | 78.4% | Community | Terminal-Bench v2 Leaderboard | tools · agent: Forge Code; stderr 0.018 | Link | |
| GPT-5.3 Codex | 77.3% | Community | Terminal-Bench leaderboard | tools · agent: Droid; stderr 0.022 | Link | |
| Claude Opus 4.6 | 76.4% | Community | Terminal-Bench leaderboard | effort unknown · tools · agent: Meta-Harness; stderr 0.024 | Link | |
| GPT-5.3 Codex | 75.8% | Community | Terminal-Bench leaderboard | tools · agent: CodeBrain-1.5; stderr 0.02 | Link | |
| GPT-5.3 Codex | 75.7% | Community | Terminal-Bench leaderboard | tools · agent: Codelia; stderr 0.022 | Link | |
| Claude Opus 4.6 | 75.3% | Community | Terminal-Bench leaderboard | effort unknown · tools · agent: Capy; stderr 0.024 | Link | |
| GPT-5.3 Codex | 75.1% | Community | Terminal-Bench leaderboard | tools · agent: Simple Codex; stderr 0.024 | Link | |
| Gemini 3.1 Pro | 74.8% | Community | Terminal-Bench leaderboard | tools · agent: Terminus-KIRA; stderr 0.026 | Link | |
| Claude Opus 4.6 | 74.7% | Community | Terminal-Bench leaderboard | effort unknown · tools · agent: Terminus-KIRA; stderr 0.026 | Link | |
| GPT-5.3 Codex | 74.6% | Community | Terminal-Bench leaderboard | tools · agent: Mux; stderr 0.025 | Link | |
| Claude Opus 4.6 | 72.1% | Community | Terminal-Bench leaderboard | effort unknown · tools · agent: MAYA-V2; stderr 0.022 | Link | |
| Claude Opus 4.6 | 71.9% | Community | Terminal-Bench leaderboard | effort unknown · tools · agent: TongAgents; stderr 0.027 | Link | |
| GPT-5.3 Codex | 71.5% | Community | Terminal-Bench leaderboard | tools · agent: spoox-o-m; stderr 0.025 | Link | |
| GPT-5.3 Codex | 70.3% | Community | Terminal-Bench leaderboard | tools · agent: CodeBrain-1; stderr 0.026 | Link | |
| Claude Opus 4.6 | 69.9% | Community | Terminal-Bench leaderboard | effort unknown · tools · agent: Droid; stderr 0.025 | Link |
Showing the top 25 of 201 tracked measurements on this benchmark.
ARC-AGI-2172
Grid puzzles designed so that each task needs a rule inferred from a handful of examples rather than recalled.
| Model | Score | Measured by | Evaluator | Conditions | Date | Source |
|---|---|---|---|---|---|---|
| GPT-6 Astra | 95.0% | Community | ARC Prize leaderboard | effort max · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| GPT-6 Astra | 93.3% | Community | ARC Prize leaderboard | effort xhigh · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| GPT-5.6 Sol | 92.5% | Community | ARC Prize leaderboard | effort max · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| GPT-6 Astra | 92.1% | Community | ARC Prize leaderboard | effort high · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| GPT-6 Astra | 92.1% | Community | ARC Prize leaderboard | effort medium · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| Claude Opus 5 | 90.4% | Community | ARC Prize leaderboard | effort max · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| Claude Fable 5.1 | 90.0% | Community | ARC Prize leaderboard | effort max · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| Claude Fable 5.1 | 90.0% | Community | ARC Prize leaderboard | effort xhigh · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| GPT-5.6 Sol | 90.0% | Community | ARC Prize leaderboard | effort xhigh · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| Claude Fable 5 | 89.2% | Community | ARC Prize leaderboard | effort max · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| Claude Fable 5.1 | 88.8% | Community | ARC Prize leaderboard | effort high · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| Claude Opus 5 | 88.3% | Community | ARC Prize leaderboard | effort high · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| Claude Fable 5 | 88.3% | Community | ARC Prize leaderboard | effort xhigh · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| Claude Fable 5 | 87.5% | Community | ARC Prize leaderboard | effort high · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| Claude Fable 5.1 | 86.3% | Community | ARC Prize leaderboard | effort medium · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| GPT-6 Astra | 85.4% | Community | ARC Prize leaderboard | effort low · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| GPT-5.6 Sol | 85.4% | Community | ARC Prize leaderboard | effort high · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| GPT-5.5 | 85.0% | Community | ARC Prize leaderboard | effort xhigh · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| Gemini 3.7 Flash | 84.6% | Community | ARC Prize leaderboard | effort high · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| GPT-5.5 Pro | 84.6% | Community | ARC Prize leaderboard | effort high · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| Gemini 3 Deep Think | 84.6% | Community | ARC Prize leaderboard | no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| GPT-5.5 Pro | 84.2% | Community | ARC Prize leaderboard | effort xhigh · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| GPT-5.6 Terra | 83.9% | Community | ARC Prize leaderboard | effort max · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| GPT-5.5 | 83.3% | Community | ARC Prize leaderboard | effort high · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link | |
| GPT-5.4 Pro | 83.3% | Community | ARC Prize leaderboard | effort xhigh · no tools · Redistributed by Epoch AI from ARC Prize leaderboard; run conditions not published per-row | Link |
Showing the top 25 of 172 tracked measurements on this benchmark.
LiveBench21
Contamination-limited suite with monthly question refreshes and objective ground truth — no LLM judge. Scores from different question-set months are not comparable.
| Model | Score | Measured by | Evaluator | Conditions | Date | Source |
|---|---|---|---|---|---|---|
| Gemini 2.5 Pro (Mar 2025) | 82.3 | Community | LiveBench Leaderboard | no tools · question set: LiveBench-2024-11-25 | Link | |
| GPT-5.1 | 78.8 | Community | LiveBench Leaderboard | effort high · no tools · Redistributed by Epoch AI from LiveBench Leaderboard; run conditions not published per-row | Link | |
| Claude 3.7 Sonnet | 76.1 | Community | LiveBench Leaderboard | no tools · question set: LiveBench-2024-11-25; note: Assumed baseline 3.7 Sonnet | Link | |
| o3-mini | 75.9 | Community | LiveBench Leaderboard | effort high · no tools · question set: LiveBench-2024-11-25 | Link | |
| QwQ-32B | 72.0 | Community | LiveBench Leaderboard | no tools · question set: LiveBench-2024-11-25 | Link | |
| DeepSeek-R1 | 71.6 | Community | LiveBench Leaderboard | no tools · question set: LiveBench-2024-11-25 | Link | |
| o3-mini | 70.0 | Community | LiveBench Leaderboard | effort medium · no tools · question set: LiveBench-2024-11-25 | Link | |
| GPT-4.5 | 69.0 | Community | LiveBench Leaderboard | no tools · question set: LiveBench-2024-11-25 | Link | |
| Gemini 2.0 Flash Thinking (Jan 2025) | 66.9 | Community | LiveBench Leaderboard | no tools · question set: LiveBench-2024-11-25 | Link | |
| DeepSeek-V3 (Mar 2025) | 66.9 | Community | LiveBench Leaderboard | no tools · question set: LiveBench-2024-11-25 | Link | |
| Claude 3.7 Sonnet | 65.6 | Community | LiveBench Leaderboard | no tools · question set: LiveBench-2024-11-25 | Link | |
| Gemini 2.0 Pro | 65.1 | Community | LiveBench Leaderboard | no tools · question set: LiveBench-2024-11-25 | Link | |
| o3-mini | 62.5 | Community | LiveBench Leaderboard | effort low · no tools · question set: LiveBench-2024-11-25 | Link | |
| Qwen2.5-Max | 62.3 | Community | LiveBench Leaderboard | no tools · question set: LiveBench-2024-11-25 | Link | |
| Gemini 2.0 Flash (Feb 2025) | 61.5 | Community | LiveBench Leaderboard | no tools · question set: LiveBench-2024-11-25 | Link | |
| DeepSeek-R1-Distill-Llama-70B | 54.5 | Community | LiveBench Leaderboard | no tools · question set: LiveBench-2024-11-25 | Link | |
| Gemini 2.0 Flash-Lite | 54.3 | Community | LiveBench Leaderboard | no tools · question set: LiveBench-2024-11-25 | Link | |
| Gemma 3 27B | 50.0 | Community | LiveBench Leaderboard | no tools · question set: LiveBench-2024-11-25 | Link | |
| DeepSeek-R1-Distill-Qwen-32B | 45.5 | Community | LiveBench Leaderboard | no tools · question set: LiveBench-2024-11-25 | Link | |
| Mistral Small 3.1 | 44.0 | Community | LiveBench Leaderboard | no tools · question set: LiveBench-2024-11-25 | Link | |
| Mistral Small 3 | 42.5 | Community | LiveBench Leaderboard | no tools · question set: LiveBench-2024-11-25 | Link |
Aider Polyglot52
Exercism coding exercises across six languages, run through the Aider edit loop. The edit format is part of the result.
| Model | Score | Measured by | Evaluator | Conditions | Date | Source |
|---|---|---|---|---|---|---|
| GPT-5 | 88.0% | Community | Aider LLM Leaderboards | effort high · tools · edit format: diff | Link | |
| GPT-5 | 86.7% | Community | Aider LLM Leaderboards | effort medium · tools · edit format: diff | Link | |
| o3-pro | 84.9% | Community | Aider LLM Leaderboards | effort high · tools · edit format: diff | Link | |
| Gemini 2.5 Pro (Jun 2025) | 83.1% | Community | Aider LLM Leaderboards | tools · edit format: diff-fenced | Link | |
| GPT-5 | 81.3% | Community | Aider LLM Leaderboards | effort low · tools · edit format: diff | Link | |
| o3 | 81.3% | Community | Aider LLM Leaderboards | effort high · tools · edit format: diff | Link | |
| Grok 4 | 79.6% | Community | Aider LLM Leaderboards | effort high · tools · edit format: diff | Link | |
| Grok 4 | 79.6% | Community | Aider LLM Leaderboards | tools · edit format: diff | Link | |
| Gemini 2.5 Pro (Jun 2025) | 79.1% | Community | Aider LLM Leaderboards | tools · edit format: diff-fenced; note: No thinking parameter set; default thinking token length | Link | |
| o3 | 76.9% | Community | Aider LLM Leaderboards | effort unknown · tools · edit format: diff | Link | |
| Gemini 2.5 Pro (May 2025) | 76.9% | Community | Aider LLM Leaderboards | tools · edit format: diff-fenced | Link | |
| o3 | 76.9% | Community | Aider LLM Leaderboards | effort medium · tools · edit format: diff | Link | |
| DeepSeek-V3.2 | 74.2% | Community | Aider LLM Leaderboards | tools · edit format: diff | Link | |
| DeepSeek-V3.2-Exp | 74.2% | Community | Aider LLM Leaderboards | effort thinking · tools · edit format: diff | Link | |
| Gemini 2.5 Pro (Mar 2025) | 72.9% | Community | Aider LLM Leaderboards | tools · edit format: diff-fenced | Link | |
| o4-mini | 72.0% | Community | Aider LLM Leaderboards | effort high · tools · edit format: diff | Link | |
| DeepSeek-R1 (May 2025) | 71.4% | Community | Aider LLM Leaderboards | tools · edit format: diff | Link | |
| Claude Opus 4 | 70.7% | Community | Aider LLM Leaderboards | tools · edit format: diff | Link | |
| DeepSeek-V3.2-Exp | 70.2% | Community | Aider LLM Leaderboards | tools · edit format: diff | Link | |
| Claude 3.7 Sonnet | 60.4% | Community | Aider LLM Leaderboards | tools · edit format: diff | Link | |
| o3-mini | 60.4% | Community | Aider LLM Leaderboards | effort high · tools · edit format: diff | Link | |
| Qwen3-235B-A22B-Instruct (Jul 2025) | 59.6% | Community | Aider LLM Leaderboards | tools · edit format: diff | Link | |
| Qwen3-235B-A22B | 59.6% | Community | Aider LLM Leaderboards | tools · edit format: diff; note: no cost data | Link | |
| Kimi K2 (Sep 2025) | 59.1% | Community | Aider LLM Leaderboards | tools · edit format: diff | Link | |
| Kimi K2 (Jul 2025) | 59.1% | Community | Aider LLM Leaderboards | tools · edit format: diff | Link |
Showing the top 25 of 52 tracked measurements on this benchmark.
Humanity's Last Exam45
Expert-written questions across a wide range of academic subjects, built to stay hard after the rest of the field saturates.
| Model | Score | Measured by | Evaluator | Conditions | Date | Source |
|---|---|---|---|---|---|---|
| Claude Fable 5.1 | 46.5% | Community | Humanity's Last Exam leaderboard | effort xhigh · no tools · stderr 0.02 | Link | |
| Gemini 3.1 Pro | 46.4% | Community | Humanity's Last Exam leaderboard | no tools · stderr 0.0196 | Link | |
| GPT-5.4 Pro | 44.3% | Community | Humanity's Last Exam leaderboard | effort unknown · no tools · stderr 0.0195 | Link | |
| Muse Spark | 40.6% | Community | Humanity's Last Exam leaderboard | no tools · stderr 0.0192 | Link | |
| Gemini 3 Pro | 37.5% | Community | Humanity's Last Exam leaderboard | no tools · stderr 0.019 | Link | |
| GPT-5.4 | 36.2% | Community | Humanity's Last Exam leaderboard | effort xhigh · no tools · stderr 0.0188 | Link | |
| Claude Opus 4.7 | 36.2% | Community | Humanity's Last Exam leaderboard | effort unknown · no tools · stderr 0.0188 | Link | |
| Claude Opus 4.6 | 34.4% | Community | Humanity's Last Exam leaderboard | effort max · no tools · stderr 0.0186 | Link | |
| GPT-5 Pro | 31.6% | Community | Humanity's Last Exam leaderboard | effort unknown · no tools · stderr 0.0182 | Link | |
| GPT-5.2 | 27.8% | Community | Humanity's Last Exam leaderboard | effort unknown · no tools · stderr 0.0176 | Link | |
| GPT-5 | 25.3% | Community | Humanity's Last Exam leaderboard | effort high · no tools · stderr 0.017 | Link | |
| Claude Opus 4.5 | 25.2% | Community | Humanity's Last Exam leaderboard | effort unknown · no tools · stderr 0.017 | Link | |
| Kimi K2.5 | 24.4% | Community | Humanity's Last Exam leaderboard | no tools · stderr 0.0181 | Link | |
| GPT-5.1 | 23.7% | Community | Humanity's Last Exam leaderboard | effort unknown · no tools · stderr 0.0167 | Link | |
| Gemini 2.5 Pro (Jun 2025) | 21.6% | Community | Humanity's Last Exam leaderboard | no tools · stderr 0.0161 | Link | |
| o3 | 20.3% | Community | Humanity's Last Exam leaderboard | effort high · no tools · stderr 0.0158 | Link | |
| GPT-5 mini | 19.4% | Community | Humanity's Last Exam leaderboard | effort unknown · no tools · stderr 0.0155 | Link | |
| o3 | 19.2% | Community | Humanity's Last Exam leaderboard | effort medium · no tools · stderr 0.0154 | Link | |
| Claude Opus 4.6 | 19.0% | Community | Humanity's Last Exam leaderboard | no tools · stderr 0.0154 | Link | |
| Gemini 2.5 Pro (Mar 2025) | 18.2% | Community | Humanity's Last Exam leaderboard | no tools · stderr 0.0151 | Link | |
| o4-mini | 18.1% | Community | Humanity's Last Exam leaderboard | effort high · no tools · stderr 0.0151 | Link | |
| Gemini 2.5 Pro (May 2025) | 17.8% | Community | Humanity's Last Exam leaderboard | no tools · stderr 0.015 | Link | |
| o4-mini | 14.3% | Community | Humanity's Last Exam leaderboard | effort medium · no tools · stderr 0.0137 | Link | |
| Claude Opus 4.5 | 14.2% | Community | Humanity's Last Exam leaderboard | effort unknown · no tools · stderr 0.0137 | Link | |
| Claude Sonnet 4.5 | 13.7% | Community | Humanity's Last Exam leaderboard | effort unknown · no tools · stderr 0.0135 | Link |
Showing the top 25 of 45 tracked measurements on this benchmark.
OSWorld20
Open-ended computer-use tasks in a real desktop environment, scored by execution rather than by a judge.
| Model | Score | Measured by | Evaluator | Conditions | Date | Source |
|---|---|---|---|---|---|---|
| Claude Sonnet 4.6 | 72.1% | Community | OS World Website | tools · agent: claude-sonnet-4-6 (100 steps) | Link | |
| Claude Opus 4.5 | 66.3% | Vendor | Anthropic | tools · agent: Claude Opus 4.5 | Link | |
| Kimi K2.5 | 63.3% | Community | OS World Website | tools · agent: Kimi-K2.5 | Link | |
| Claude Sonnet 4.5 | 62.9% | Community | OS World Website | tools · agent: claude-sonnet-4-5-20250929 (100 steps) | Link | |
| Claude Sonnet 4.5 | 58.1% | Community | OS World Website | tools · agent: claude-sonnet-4-5-20250929 (50 steps) | Link | |
| Claude Sonnet 4 | 43.9% | Community | OS World Website | tools · agent: claude-4-sonnet-20250514 (50 steps) | Link | |
| Claude Sonnet 4.5 | 42.9% | Community | OS World Website | tools · agent: claude-sonnet-4-5-20250929 (15 steps) | Link | |
| Claude Sonnet 4 | 41.4% | Community | OS World Website | tools · agent: claude-4-sonnet-20250514 (100 steps) | Link | |
| Claude 3.7 Sonnet | 35.8% | Community | OS World Website | tools · agent: claude-3-7-sonnet-20250219 (50 steps) | Link | |
| Claude 3.7 Sonnet | 35.6% | Community | OS World Website | tools · agent: claude-3-7-sonnet-20250219 (100 steps) | Link | |
| computer-use-preview-2025-03-11 | 31.3% | Community | OS World Website | tools · agent: computer-use-preview (50 steps) | Link | |
| Claude Sonnet 4 | 31.2% | Community | OS World Website | tools · agent: claude-4-sonnet-20250514 (15 steps) | Link | |
| computer-use-preview-2025-03-11 | 30.5% | Community | OS World Website | tools · agent: computer-use-preview (100 steps) | Link | |
| Claude 3.7 Sonnet | 27.1% | Community | OS World Website | tools · agent: claude-3-7-sonnet-20250219 (15 steps) | Link | |
| computer-use-preview-2025-03-11 | 26.0% | Community | OS World Website | tools · agent: computer-use-preview (15 steps) | Link | |
| o3 | 23.0% | Community | OS World Website | effort medium · tools · agent: o3 (100 steps) | Link | |
| o3 | 17.2% | Community | OS World Website | effort medium · tools · agent: o3 (50 steps) | Link | |
| o3 | 9.1% | Community | OS World Website | effort medium · tools · agent: o3 (15 steps) | Link | |
| Qwen2.5-72B | 5.0% | Community | OS World Website | tools · agent: qwen2.5-vl-72b-instruct (100 steps) | Link | |
| Qwen2.5-72B | 4.4% | Community | OS World Website | tools · agent: qwen2.5-vl-72b-instruct (15 steps) | Link |
Roundups
No roundups yet. They are written only when the data does something worth writing about — a new leader, a retraction, or a vendor's own number failing to reproduce.