Anthropic has officially launched Claude Opus 5, a flagship AI model that delivers near-frontier intelligence at half the cost per task of Claude Fable 5. Available as the default model on Claude Max and the strongest option on Claude Pro, Opus 5 achieves state-of-the-art results across major coding and knowledge work benchmarks. Developers and enterprise users can deploy the model immediately for multi-step task automation, complex reasoning, and long-horizon engineering workflows.
Benchmark Gains Across Coding and Problem-Solving
On the ARC-AGI 3 reasoning evaluation, Claude Opus 5 scored three times higher than the next-best competing model while maintaining high token efficiency. On CursorBench 3.2, at maximum effort, the model performed within 0.5% of Fable 5's peak score at half the cost per task, while more than doubling Opus 4.8's performance on Frontier-Bench v0.1.
Key performance metrics highlighted by Anthropic include:
- ARC-AGI 3: Achieved 3x the performance score of the nearest competing model on novel problems.
- CursorBench 3.2: Reached within 0.5% of Fable 5 peak results at 50% lower cost per task.
- Zapier AutomationBench: Delivered a 1.5x higher pass rate than competing models for the same cost.
- Financial Modeling: Cut task completion time by 60% with one-third fewer turns and tool calls compared to Opus 4.8.
- OSWorld 2.0: Surpassed Fable 5's top computer-use results at just over one-third of the operational cost.
Specialized Enterprise and Scientific Workflows
According to early-access testing detailed by Anthropic, enterprise workflows saw substantial speed and accuracy upgrades. Enterprise partner Box reported an overall performance improvement of 8% over Opus 4.8, driven by an 11% increase in data analysis efficiency and a 17% jump in due diligence accuracy. Legal testing showed a 26% reduction in generated tokens at maximum reasoning levels, alongside first-turn redline scores that nearly doubled Opus 4.8.
In scientific evaluation suites, Opus 5 outperformed Opus 4.8 across all life sciences testing. Notable gains included a 10.2 percentage point improvement on organic chemistry tasks and a 7.7 percentage point increase on protein function prediction benchmarks. Anthropic also reported that Opus 5 scored 2.3 on its internal automated behavioral audit, marking its lowest rate of misaligned behavior among recent models.
Limitations and Performance Trade-Offs
Despite establishing new state-of-the-art benchmarks on Frontier-Bench and GDPval-AA, Anthropic noted clear operational limits for the model. On specialized cybersecurity tasks, Opus 5 continues to rank behind Mythos 5. Additionally, while the model allows users to adjust effort settings to conserve tokens, peak performance on complex software engineering requires setting reasoning levels to maximum.