For IT teams, the practical change here is not just a stronger model, but a broader operating model. Claude Opus 4.5 is being exposed across app, API, and major clouds, which means adoption will likely hinge on how quickly teams can align identity, access controls, logging, and cost governance across those entry points. The release also makes model selection more of an engineering decision than a product choice: effort level, token use, and context handling now become tunable levers in the architecture.
That matters for integration work. The model is positioned for long-running agents, code workflows, spreadsheets, browsers, and desktop use, so the real dependency is the surrounding orchestration layer: tool permissions, state management, retry logic, and human approval gates. In practice, teams will need to decide where autonomous execution is acceptable and where workflows should stay supervised, especially when agents can chain actions across systems and subagents.
The safety angle is equally operational. Better resistance to prompt injection is useful, but it does not remove the need for controls around untrusted content, indirect prompt exposure, and downstream tool execution. Organizations deploying this kind of model in production will still need isolation boundaries, audit trails, and policy checks for sensitive actions, because the modelโs strength comes from acting on context, not from inherently knowing which context is safe.
The cost/performance trade-off is also notable. Lower token usage and fewer iterations can improve throughput, but only if the surrounding stack avoids adding its own inefficiency through poor prompt design, over-tooling, or weak memory strategy. For architecture teams, the main question is whether Opus 4.5 reduces the number of systems needed to complete workโor simply shifts complexity into agent governance, observability, and integration discipline.
claude-opus-4-5-20251101 via the Claude API. Pricing is now $5/$25 per million tokensโmaking Opus-level capabilities accessible to even more users, teams, and enterprises.
Alongside Opus, weโre releasing updates to the Claude Developer Platform, Claude Code, and our consumer apps. There are new tools for longer-running agents and new ways to use Claude in Excel, Chrome, and on desktop. In the Claude apps, lengthy conversations no longer hit a wall. See our product-focused section below for details.
First impressions
As our Anthropic colleagues tested the model before release, we heard remarkably consistent feedback. Testers noted that Claude Opus 4.5 handles ambiguity and reasons about tradeoffs without hand-holding. They told us that, when pointed at a complex, multi-system bug, Opus 4.5 figures out the fix. They said that tasks that were near-impossible for Sonnet 4.5 just a few weeks ago are now within reach. Overall, our testers told us that Opus 4.5 just โgets it.โ Many of our customers with early access have had similar experiences. Here are some examples of what they told us:Opus models have always been โthe real SOTAโ but have been cost prohibitive in the past. Claude Opus 4.5 is now at a price point where it can be your go-to model for most tasks. Itโs the clear winner and exhibits the best frontier task planning and tool calling weโve seen yet.
Claude Opus 4.5 delivers high-quality code and excels at powering heavy-duty agentic workflows with GitHub Copilot. Early testing shows it surpasses internal coding benchmarks while cutting token usage in half, and is especially well-suited for tasks like code migration and code refactoring.
Claude Opus 4.5 beats Sonnet 4.5 and competition on our internal benchmarks, using fewer tokens to solve the same problems. At scale, that efficiency compounds.
Claude Opus 4.5 delivers frontier reasoning within Lovable’s chat mode, where users plan and iterate on projects. Its reasoning depth transforms planningโand great planning makes code generation even better.
Claude Opus 4.5 excels at long-horizon, autonomous tasks, especially those that require sustained reasoning and multi-step execution. In our evaluations it handled complex workflows with fewer dead-ends. On Terminal Bench it delivered a 15% improvement over Sonnet 4.5, a meaningful gain that becomes especially clear when using Warpโs Planning Mode.
Claude Opus 4.5 achieved state-of-the-art results for complex enterprise tasks on our benchmarks, outperforming previous models on multi-step reasoning tasks that combine information retrieval, tool use, and deep analysis.
Claude Opus 4.5 delivers measurable gains where it matters most: stronger results on our hardest evaluations and consistent performance through 30-minute autonomous coding sessions.
Claude Opus 4.5 represents a breakthrough in self-improving AI agents. For office automation, our agents were able to autonomously refine their own capabilitiesโachieving peak performance in 4 iterations while other models couldnโt match that quality after 10.
Claude Opus 4.5 is a notable improvement over the prior Claude models inside Cursor, with improved pricing and intelligence on difficult coding tasks.
Claude Opus 4.5 is yet another example of Anthropic pushing the frontier of general intelligence. It performs exceedingly well across difficult coding tasks, showcasing long-term goal-directed behavior.
Claude Opus 4.5 delivered an impressive refactor spanning two codebases and three coordinated agents. It was very thorough, helping develop a robust plan, handling the details and fixing tests. A clear step forward from Sonnet 4.5.
Claude Opus 4.5 handles long-horizon coding tasks more efficiently than any model weโve tested. It achieves higher pass rates on held-out tests while using up to 65% fewer tokens, giving developers real cost control without sacrificing quality.
Weโve found that Opus 4.5 excels at interpreting what users actually want, producing shareable content on the first try. Combined with its speed, token efficiency, and surprisingly low cost, itโs the first time weโre making Opus available in Notion Agent.
Claude Opus 4.5 excels at long-context storytelling, generating 10-15 page chapters with strong organization and consistency. It’s unlocked use cases we couldn’t reliably deliver before.
Claude Opus 4.5 sets a new standard for Excel automation and financial modeling. Accuracy on our internal evals improved 20%, efficiency rose 15%, and complex tasks that once seemed out of reach became achievable.
Claude Opus 4.5 is the only model that nails some of our hardest 3D visualizations. Polished design, tasteful UX, and excellent planning & orchestration – all with more efficient token usage. Tasks that took previous models 2 hours now take thirty minutes.
Claude Opus 4.5 catches more issues in code reviews without sacrificing precision. For production code review at scale, that reliability matters.
Based on testing with Junie, our coding agent, Claude Opus 4.5 outperforms Sonnet 4.5 across all benchmarks. It requires fewer steps to solve tasks and uses fewer tokens as a result. This indicates that the new model is more precise and follows instructions more effectively โ a direction weโre very excited about.
The effort parameter is brilliant. Claude Opus 4.5 feels dynamic rather than overthinking, and at lower effort delivers the same quality we need while being dramatically more efficient. That control is exactly what our SQL workflows demand.
Weโre seeing 50% to 75% reductions in both tool calling errors and build/lint errors with Claude Opus 4.5. It consistently finishes complex tasks in fewer iterations with more reliable execution.
Evaluating Claude Opus 4.5
We give prospective performance engineering candidates a notoriously difficult take-home exam. We also test new models on this exam as an internal benchmark. Within our prescribed 2-hour time limit, Claude Opus 4.5 scored higher than any human candidate ever1. The take-home test is designed to assess technical ability and judgment under time pressure. It doesnโt test for other crucial skills candidates may possess, like collaboration, communication, or the instincts that develop over years. But this resultโwhere an AI model outperforms strong candidates on important technical skillsโraises questions about how AI will change engineering as a profession. Our Societal Impacts and Economic Futures research is aimed at understanding these kinds of changes across many fields. We plan to share more results soon. Software engineering isnโt the only area on which Claude Opus 4.5 has improved. Capabilities are higher across the boardโOpus 4.5 has better vision, reasoning, and mathematics skills than its predecessors, and it is state-of-the-art in many domains: The modelโs capabilities outpace some of the benchmarks we use in our tests. A common benchmark for agentic capabilities is ฯ2-bench, which measures the performance of agents in real-world, multi-turn tasks. In one scenario, models have to act as an airline service agent helping a distressed customer. The benchmark expects models to refuse a modification to a basic economy booking since the airline doesnโt allow changes to that class of tickets. Instead, Opus 4.5 found an insightful (and legitimate) way to solve the problem: upgrade the cabin first, then modify the flights.Let me think about what options I have within my policy:
1. Modify flights - Basic economy cannot be modified. This is clear in the policy.
2. Change cabin - Wait, let me check this option! The policy says:
"In other cases, all reservations, including basic economy, can change cabin without changing the flights."
But he wants to change the flights, not just the cabin. However, what if we:
- Change the cabin first (which IS allowed for basic economy), then
- Change the flights (which is allowed for non-basic economy)
So the strategy could be:
1. Upgrade his cabin from basic economy to economy (or business)
2. Then modify the flights to be 2 days later
This would cost more money, but itโs a legitimate path within the policy!
A step forward on safety
As we state in our system card, Claude Opus 4.5 is the most robustly aligned model we have released to date and, we suspect, the best-aligned frontier model by any developer. It continues our trend towards safer and more secure models:
In our evaluation, โconcerning behaviorโ scores measure a very wide range of misaligned behavior, including both cooperation with human misuse and undesirable actions that the model takes at its own initiative [2].
Note that this benchmark includes only very strong prompt injection attacks. It was developed and run by Gray Swan.
New on the Claude Developer Platform
As models get smarter, they can solve problems in fewer steps: less backtracking, less redundant exploration, less verbose reasoning. Claude Opus 4.5 uses dramatically fewer tokens than its predecessors to reach similar or better outcomes. But different tasks call for different tradeoffs. Sometimes developers want a model to keep thinking about a problem; sometimes they want something more nimble. With our new effort parameter on the Claude API, you can decide to minimize time and spend or maximize capability. Set to a medium effort level, Opus 4.5 matches Sonnet 4.5โs best score on SWE-bench Verified, but uses 76% fewer output tokens. At its highest effort level, Opus 4.5 exceeds Sonnet 4.5 performance by 4.3 percentage pointsโwhile using 48% fewer tokens. With effort control, context compaction, and advanced tool use, Claude Opus 4.5 runs longer, does more, and requires less intervention. Our context management and memory capabilities can dramatically boost performance on agentic tasks. Opus 4.5 is also very effective at managing a team of subagents, enabling the construction of complex, well-coordinated multi-agent systems. In our testing, the combination of all these techniques boosted Opus 4.5โs performance on a deep research evaluation by almost 15 percentage points3. Weโre making our Developer Platform more composable over time. We want to give you the building blocks to construct exactly what you need, with full control over efficiency, tool use, and context management.Product updates
Products like Claude Code show whatโs possible when the kinds of upgrades weโve made to the Claude Developer Platform come together. Claude Code gains two upgrades with Opus 4.5. Plan Mode now builds more precise plans and executes more thoroughlyโClaude asks clarifying questions upfront, then builds a user-editable plan.md file before executing. Claude Code is also now available in our desktop app, letting you run multiple local and remote sessions in parallel: perhaps one agent fixes bugs, another researches GitHub, and a third updates docs. For Claude app users, long conversations no longer hit a wallโClaude automatically summarizes earlier context as needed, so you can keep the chat going. Claude for Chrome, which lets Claude handle tasks across your browser tabs, is now available to all Max users. We announced Claude for Excel in October, and as of today we’ve expanded beta access to all Max, Team, and Enterprise users. Each of these updates takes advantage of Claude Opus 4.5โs market-leading performance in using computers, spreadsheets, and handling long-running tasks. For Claude and Claude Code users with access to Opus 4.5, weโve removed Opus-specific caps. For Max and Team Premium users, weโve increased overall usage limits, meaning youโll have roughly the same number of Opus tokens as you previously had with Sonnet. Weโre updating usage limits to make sure youโre able to use Opus 4.5 for daily work. These limits are specific to Opus 4.5. As future models surpass it, we expect to update limits as needed.Enjoyed this article? Sign up for our newsletter to receive regular insights and stay connected.

