Anthropic
Anthropic's latest generally available Opus-tier model for complex reasoning, long-horizon agentic coding, high-autonomy work, and computer-use workflows. Builds on Opus 4.
Anthropic's latest generally available Opus-tier model for complex reasoning, long-horizon agentic coding, high-autonomy work, and computer-use workflows.
50.1
Quality Score
1280
Arena ELO
1M
Parameters
1M
Context
Use this section to answer one simple question first: how much outside evidence do we have that this model performs well? Structured benchmark scores appear first, then official provider evidence, then live arena signal.
This model has normalized benchmark rows, so scores here are directly comparable across benchmark sources.
Sign in to join the discussion
0
Downloads
0
Likes
May 2026
Released
These are recent benchmark or leaderboard claims from official provider sources. They are useful for freshness and context, but they are not treated the same as normalized independent benchmark rows.
Claude-Opus-4 - LiveCodeBench
LiveCodeBench pass@1 62.4 across 1055 tasks
View sourceClaude-Opus-4 - LiveCodeBench
LiveCodeBench pass@1 62.4 across 1055 tasks
View sourceIntroducing Claude Opus 4.8
It builds on Opus 4.7 with improvements across benchmarks, and is a more effective collaborator. Opus 4.8’s capabilities The table below shows how Opus 4.8 compares to its predecessor and to other models on tests of coding, agentic skills, reasoning, and practical knowledge work tasks. More details and a much wider range of capability evaluations are provided in the Claude Opus 4.8 System Card .
View sourceIntroducing Claude 4
Introducing Claude 4 \ Anthropic Skip to main content Skip to footer Research Policy Commitments Learn News Try Claude Announcements Introducing Claude 4 May 22, 2025 Today, we’re introducing the next generation of Claude models: Claude Opus 4 and Claude Sonnet 4 , setting new standards for coding, advanced reasoning, and AI agents. Claude Opus 4 is the world’s best coding model, with sustained performance on complex, long-running tasks and agent workflows. Claude Sonnet 4 is
View source1280
ELO Score
1271 - 1289
95% Confidence
+/-9 points
8.1K
Battles
Jul 12, 2026
Last Updated