Caylent Launches Caylent Accelerate™ for Agentic Cloud Operations

Claude Opus 5: Changes, Improvements, and How It Compares to Fable 5

Generative AI & LLMOps

Learn what's new in Claude Opus 5, how it compares to Opus 4.8 and Fable 5, and what its new reasoning behavior, pricing, and performance improvements mean for enterprise AI workloads.

On July 24. 2026, Anthropic released Claude Opus 5, which keeps the same base API rates as Opus 4.8 while changing both the performance ceiling and the cost of reaching an acceptable result. Priced at $5 per million input tokens and $25 per million output tokens, the model has a one-million-token context window, up to 128,000 output tokens, and adaptive thinking enabled by default.

As with previous Anthropic models, you can set the thinking effort, on a scale of low, medium, high, xhigh and max. With a higher level of effort, Opus 5 reaches stronger results than Opus 4.8 on important coding and automation evaluations. At lower effort, it can match or slightly exceed the predecessor’s maximum-effort result on CursorBench while using fewer tokens and incurring lower benchmark cost.

Claude Opus 5 Changes the Performance-and-Cost Curve

Anthropic lists Opus 5 as claude-opus-5, with a one-million-token context window and a 128,000-token maximum output, with the same pricing as its predecessor Opus 4.8: $5/$25 for 1 million input/output tokens. Fable 5 is still advertised as Anthropic’s higher-capability option at twice Opus 5’s input and output prices, but Opus 5 has closed that gap significantly, even surpassing Fable on some benchmarks.

Effort affects text, thinking, tool calls, and function arguments. A higher setting will spend more tokens and take more steps, while a lower setting will reduce activity while preserving enough quality for most simple workloads. This makes Opus 5 more cost-effective than Fable 5 for most tasks.

Benchmarks published by Artificial Analysis

Differences in Reasoning Between Opus 5 and Opus 4.8

Requests to Opus 4.8 that omitted a thinking configuration ran without thinking. The same requests use adaptive thinking on Opus 5, with high as the default effort level and low, medium, xhigh, and max available as explicit alternatives. The request’s max_tokens setting covers thinking and visible response text together, so limits selected for Opus 4.8 can constrain a different mix of work on Opus 5. This is a significant point when migrating from Opus 4.8 to Opus 5.

Moreover, the selected thinking configuration appears to have a more significant impact on Opus 5's reasoning tokens than it did on Opus 4.8. Anthropic documents longer default deliverables, more progress narration, greater use of subagents, and more self-verification. Our own internal testing confirms this, with Opus 5 producing significantly longer reasoning traces, which would explain the cost-per-task increase compared to Opus 4.8.

When compared to Sonnet 5, it seems clear that Anthropic's fifth-generation models tend to produce more output tokens. This is likely related, in part, to the new tokenizer that Anthropic used for Sonnet 5, which they have adopted for Opus 5 as well.

Latency

Our tests indicate that Opus 5 has a similar token throughput to Opus 4.8. The increase in latency correlates with the increase in output tokens observed in our tests. This means that Opus 5 has a very similar rate of Tokens per Second as Opus 4.8. The results of latency (which include Opus 5's increased token output) are shown below.

These tests were performed using Claude Platform on AWS.

Higher Ceiling and a More Efficient Lower Setting

CursorBench 3.2 reports score, token use, steps, and average cost per task across effort settings. Its tasks cover ambiguous, multi-file software-engineering work drawn from real Cursor sessions, and its cost calculation applies published token prices to each run.

At low effort, Opus 5 scores 62.8% at an average benchmark cost of $2.55 per task. Opus 4.8 at max scores 62.3% at $5.77. The half-point score difference is too small to support a broad quality claim, but the two rows show that Opus 5 can reach the predecessor’s maximum-effort result at less than half the reported benchmark cost on this task set.

At max, Opus 5 reaches 70.0% at $8.23 per task. That is 7.7 percentage points above Opus 4.8 Max, with a higher average benchmark cost. Fable 5 Max scores 70.5% at $17.32, placing Opus 5 within half a percentage point at less than half Fable’s benchmark cost.

Safeguards and Fallback Behavior

Opus 5's cyber classifiers allow source-code vulnerability discovery while blocking binary-based vulnerability scanning, penetration testing, and exploit generation. On Claude.ai, Claude Code, and Claude Cowork, flagged requests fall back to Opus 4.8 by default.

The Claude API can also route refused Opus 5 requests through a server-side fallback, the same used by Fable 5. The response identifies the model that served the request, allowing evaluations and audit records to distinguish native Opus 5 completion from fallback-assisted completion.

Availability, Features and Data Retention

Opus 5 can be accessed directly via the Anthropic API, via Amazon Bedrock, or via Claude Platform on AWS.

Access through the Anthropic API or Claude Platform on AWS has the same features and settings, since Claude Platform on AWS ultimately uses the Anthropic API. The main difference for model access is that with Claude Platform on AWS, charges are made against your AWS bill, simplifying vendor management. In both of these instances, zero data retention can be enabled on request, Anthropic's server-side fallback is available, and the new Fast mode is available.

When using Opus 5 via Amazon Bedrock, the same model safeguards apply, but server-side fallback is not available. Teams using Amazon Bedrock will need to use client-side fallback instead. Anthropic's Fast model is not available either, though in its stead, you may use Amazon Bedrock's myriad of options that affect cost and latency, such as Latency Optimized Inference or Reserved Pricing. Moreover, zero data retention is enabled by default for Opus 5, meaning no data ever leaves Amazon Bedrock.

Conclusion

Claude Opus 5 is a meaningful improvement over Opus 4.8, its predecessor. However, the best comparison is against Fable 5, with Opus 5 coming close to it on benchmarks, even scoring above it on some, at half the token price.

An important note, especially when comparing it to Opus 4.8, is the increased influence that thinking effort has over reasoning, tool activity, and output tokens. This means that while Opus 5 may be an intelligence upgrade for workloads using Opus 4.8, or even a cost reduction that maintains the same level of intelligence, it's unlikely to work as a drop-in replacement. You will need to recalibrate the thinking effort assigned to each task, so Opus 5 doesn't accidentally result in a significant increase in output tokens and in costs.

How Caylent Can Help

Choosing the right Claude model is only part of the equation. Organizations also need to determine the appropriate thinking effort for each workload, optimize token usage, and design applications that balance performance, latency, and cost at scale. As an AWS Premier Tier Services Partner and Anthropic Preferred Services Partner with a dedicated Anthropic practice, Caylent helps organizations evaluate models like Claude Opus 5, Sonnet 5, and Fable 5, optimize AI unit economics, and build production-ready AI applications on Amazon Bedrock or the Anthropic API. Whether you're migrating existing workloads, developing new agentic applications, or establishing an enterprise AI strategy, our team can help you deploy Claude securely, efficiently, and with measurable business value. Reach out to us today to get started.

Generative AI & LLMOps
Guille Ojeda

Guille Ojeda

Guille Ojeda is a Principal Innovation Architect at Caylent, a speaker, author, and content creator. He has published 2 books, over 200 blog articles, and writes a free newsletter called Simple AWS with more than 45,000 subscribers. He's spoken at multiple AWS Summits and other events, and was recognized as AWS Builder of the Year in 2025.

View Guille's articles
Chris Gonzalez

Chris Gonzalez

Chris Gonzalez is a Cloud Architect at Caylent. He has a passion for serverless computing, well-architected solutions, cloud infrastructure, and professional development. His background incorporates insights from over a decade in education, financial services platform infrastructure, and cloud consulting. His technical expertise consists of implementing complex cloud infrastructures for enterprise financial services firms, as well as intricate Kubernetes solutions built on EKS and open-source products. Chris currently lives in Knoxville, TN, and enjoys spending time with his wife and two kids, tinkering in his home lab, and hiking in the Smoky Mountains.

View Chris's articles

Learn more about the services mentioned

Caylent Catalysts™

Generative AI Strategy

Accelerate your generative AI initiatives with ideation sessions for use case prioritization, foundation model selection, and an assessment of your data landscape and organizational readiness.

Caylent Catalysts™

AWS Generative AI Proof of Value

Accelerate investment and mitigate risk when developing generative AI solutions.

Accelerate your GenAI initiatives

Leveraging our accelerators and technical experience

Browse GenAI Offerings

Related Blog Posts

AWS Summit New York 2026: New Launches and Capabilities

Explore all of the launches and capabilities announced at the 2026 AWS Summit in New York City, including Amazon Bedrock Managed Knowledge Base, AgentCore harness, AWS Context, and AWS Continuum.

AWS Announcements
Generative AI & LLMOps

Claude Fable 5: Anthropic's First Public Mythos-Class Model

Explore Claude Fable 5, Anthropic's most capable generally available model, and learn how its advanced reasoning capabilities, safeguards, pricing, and deployment considerations impact real-world enterprise AI adoption.

Generative AI & LLMOps

The Agentic SDLC Journey’s North Star

Explore agentic SDLC as the shift in software development where teams balance AI and human ownership of context and decisions while adapting people, processes, and technology to work effectively with AI agents.

Generative AI & LLMOps