Anthropic put Claude Opus 5 on sale today at $5 per million input tokens and $25 per million output tokens, the exact prices it charged for the previous Opus 4.8. According to Anthropic’s announcement, the model reaches close to the intelligence of its higher-end Fable 5 model at half the cost, and it now serves as the default model on Claude Max and the strongest option on Claude Pro.
The company is leaning hard on benchmark numbers to make its case. On Frontier-Bench v0.1, Anthropic says Opus 5 beats every other model and more than doubles Opus 4.8’s score at a lower cost per task. On ARC-AGI, a test built around novel reasoning problems, the company reports Opus 5 scoring three times higher than the next-best model. On Zapier’s AutomationBench, which checks whether a model can run a business task end to end, Anthropic claims a pass rate roughly 1.5 times the nearest competitor at the same cost.
The recurring theme in the launch is verification. Anthropic describes Opus 5 checking its own work before handing it back. In one Frontier-Bench task, the model was asked to rebuild a machine part in 3D code but given no way to view the drawing, so it wrote its own computer vision pipeline to pull the geometry from raw pixels. In another example, it found the root cause of a bug in an open-source package manager that the community’s own patch had missed.
Early-access customers echoed that framing. JetBrains said the model catches its own logical faults during planning rather than after. A legal-tech tester reported first-turn redline scores nearly double Opus 4.8. Box measured an 8% overall improvement over Opus 4.8, with 17% gains on due-diligence workflows. These are vendor-supplied quotes, so treat the precise figures as marketing rather than independent measurement.
On safety, Anthropic says Opus 5 is its most aligned model so far, scoring 2.3 on its automated misaligned-behavior audit, the lowest of its recent releases. The company also notes the model stays behind its Mythos 5 model on biology research and offensive cybersecurity. Notably, Opus 5 can find software vulnerabilities about as well as Mythos 5 but lags badly at writing exploits for them, which Anthropic frames as a deliberate safeguard. Its cyber classifiers block binary vulnerability scanning, penetration testing and exploit generation, with flagged requests falling back to Opus 4.8.
For newsrooms and publishers, the pricing is the story. Holding costs flat while claiming stronger reasoning and cleaner outputs matters for teams running document analysis, research summaries and data work at volume. The customer notes about tighter, more concise responses and fewer tool calls point to lower token spend per task, which is where AI budgets actually get decided. Publishers weighing model choices should still run their own tests against real workflows rather than trusting benchmark charts, a point we’ve made repeatedly at The Media Copilot.
Anthropic paired the launch with two beta features: mid-conversation tool changes that don’t break the prompt cache, and automatic API fallbacks that route flagged requests to another model instead of blocking them. A Fast mode runs about 2.5 times the default speed at twice the base price. The bet is that a cheaper, more careful default beats a smarter but pricier one for daily use, and the next few months of real deployments will show whether that holds.







