AI Model Benchmarks
Comprehensive performance comparison across major AI labs
Last updated: September 3, 2026 · Click column headers to sort · Data sourced from official model cards and benchmark publications
Claude Fable 5.1 and Mythos 5.1 are now in the table
Fable and Mythos use identical weights. Differences such as Terminal-Bench 4.0 reflect the public model's deployed safeguards, not a different underlying model. The default view compares the new release with Opus 5 and GPT-5.6 Sol.
Compare models
Pick a few models to compare side by side, then hide benchmarks where none of those models have scores.
| Model ▲ | ReasoningGPQA Diamond | ReasoningHLE | ReasoningARC-AGI-2 | ReasoningARC-AGI-1 | AgentOSWorld | CodingSWE-Pro | AgentGDPval-AA | CodingSWE-Multilin… | AgentToolathlon | MathFrontierMath… | CodingSWE-Multimodal | AgentTerminal-Ben… | AgentTB-Science 0.1 | KnowledgeHealthBench … | AgentAutomationBe… | CodingCursorBench … | AgentAA-Briefcase | ReasoningARC-AGI-3 | CodingDeepSWE v1.1 | MultimodalBenchCAD | AgentExploitBench |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Claude Fable 5.1 Anthropic2026-09 | - | 60.9 | 90 | 97.5 | 77.9 | 81.2 | 1853 | 89.1 | 77.8 | - | 54.7 | 56 | 52.6 | 62.1 | 31.4 | 73.4 | 1694 | - | 67.4 | 84.3 | - |
Claude Mythos 5.1 Anthropic2026-09 | - | 60.9 | 90 | 97.5 | 77.9 | 81.2 | 1853 | 89.1 | 77.8 | - | 54.7 | 61 | 52.6 | 62.1 | 31.4 | 73.4 | 1694 | - | - | - | - |
Claude Opus 5 Anthropic2026-07 | - | 56.6 | 90.4 | 97.5 | 75.4 | 79.2 | 1824 | 89.5 | 80.6 | - | 59.4 | 52 | 29 | 59.8 | 26.9 | 70 | 1685 | 30.2 | 74 | - | - |
GPT-5.6 Sol OpenAI2026-08 | 94.6 | - | 92.5 | 96.5 | 65.7 | 64.6 | 1711 | - | - | 83 | - | 37 | 22.4 | - | 19.6 | 67.2 | 1502 | 38.3 | 70.8 | 83.3 | 73.5 |
Benchmark Descriptions
Graduate-level science questions across biology, physics, chemistry — PhD-level difficulty
Humanity's Last Exam — the hardest multi-domain benchmark designed to push AI limits
Abstract reasoning and adaptability — visual puzzles resisting memorization
Original ARC abstract reasoning benchmark — visual pattern puzzles testing fluid intelligence
Computer use benchmark — GUI interaction, desktop automation, real OS tasks (369 tasks)
Real-world software engineering across multiple languages including Python, JS, Go, Rust — tests full-stack engineering depth
Professional office & domain expertise — ELO-scored evaluation of task delivery across real-world work scenarios
Software engineering across multiple programming languages — tests breadth beyond Python-only benchmarks
Tool-use benchmark - multi-step tool calling and workflow execution across realistic tasks
Frontier math benchmark - the hardest Tier 4 problems from FrontierMath
Software engineering tasks requiring visual understanding alongside repository-level coding
Long-horizon terminal work across realistic software, systems, and tool-use tasks; Fable scores include deployed safeguards
Scientific terminal tasks requiring research judgment, coding, analysis, and sustained tool use
Professional healthcare reasoning score adjusted for response length
End-to-end automation of realistic knowledge-work tasks using tools and computer interaction
Independent Cursor evaluation of agentic coding performance; headline scores use each model's reported setting
Artificial Analysis professional knowledge-work benchmark scored by Elo
Interactive agentic reasoning games — third-gen ARC; harness choices (retained reasoning, compaction) strongly affect scores
113 original long-horizon engineering tasks in real codebases — agentic software engineering
Reconstruct CAD programs from rendered views (1,000-file Vision2Code subset, Python tools) — scored on 3D geometric overlap; Claude numbers are OpenAI-reported with modified settings
Turn known software vulnerabilities into working exploits — capability-ladder cyber benchmark
Labs
Data sourced from official model cards, blog posts, arXiv papers, and independent evaluations.
Scores are not always directly comparable: model providers may use different tools, effort settings, scaffolds, or safeguards.
September Claude results: Anthropic system card · Independent coding comparison: CursorBench