Claude Opus 5: Benchmarks, Pricing, and Who Should Switch
Opus 5 scores 61 on the Intelligence Index and 3x the next model on ARC-AGI-3. Same price as Opus 4.8 ($5/$25). Full benchmark breakdown vs Fable 5, GPT-5.6 Sol, Kimi K3.

Anthropic just released Opus 5 and the positioning is different from their usual launches. They're not leading with "we beat OpenAI." They're leading with "we nearly beat ourselves at half the price." Opus 5 scores 61 on the Artificial Analysis Intelligence Index, one point above Fable 5, at $5/$25 instead of $10/$50.
The ARC-AGI-3 result is the one that stops you scrolling. Opus 5 scores three times higher than any other model on novel problem-solving. Not incremental gains. A 3x gap.
Here's everything that matters: benchmarks, real costs, where it falls short, and whether you should switch from Fable 5 or Sol.
Opus 5 by the Numbers
| Opus 5 | Fable 5 | GPT-5.6 Sol | Opus 4.8 | Kimi K3 | |
|---|---|---|---|---|---|
| Intelligence Index | 61 | 60 | 59 | 55 | 57 |
| Frontier-Bench v0.1 | 43.3% | 33.7% | 34.4% | 21.1% | — |
| GDPval-AA v2 | 1861 | 1747 | 1736 | 1593 | 1668 |
| ARC-AGI-3 | 30.2% | — | 7.8% | 1.5% | — |
| OSWorld 2.0 | 70.6% | 66.1% | 62.6% | 55.7% | — |
| AutomationBench | 26.0% | 17.4% | 18.1% | 17.0% | — |
| DeepSWE v1.1 | 68.8% | 69.7% | 72.7% | 59.0% | — |
| HLE (no tools) | 56.3% | 56.5% | — | 49.8% | — |
| Input Price/1M | $5.00 | $10.00 | $5.00 | $5.00 | $3.00 |
| Output Price/1M | $25.00 | $50.00 | $30.00 | $25.00 | $15.00 |
| Cache Hit/1M | $0.50 | — | $0.10 | $0.50 | $0.30 |
| Context | 1M | 1M | 1.05M | 1M | 1M |
Opus 5 tops the Intelligence Index at 61. That's the first time an Opus-tier model has beaten Fable 5 (60) on the composite score. The gap is small, but the direction matters: Anthropic's mid-tier model now matches or exceeds their top-tier on most evaluations.
Same price as Opus 4.8. No price hike. That's the real story. Every other lab charges a premium for their latest flagship. Anthropic kept it at $5/$25 and delivered a model that competes with their $10/$50 offering.
Where Opus 5 Leads
Three results stand out from the benchmark table.
ARC-AGI-3: 30.2% vs 7.8% for GPT-5.6 Sol. This benchmark tests novel problem-solving where the model can't rely on pattern-matching from training data. Opus 5 doesn't just win. It triples the nearest competitor. Opus 4.8 scored 1.5%. That's a generational jump. For any workflow that involves genuinely new problems rather than variations of seen patterns, this gap is significant.
Frontier-Bench v0.1: 43.3% vs 34.4% for Sol. Anthropic's agentic terminal coding benchmark. Opus 5 more than doubles Opus 4.8's score (21.1%) while costing less per task because it uses fewer tokens to reach an answer. On CursorBench 3.2, Opus 5 at max effort lands within 0.5% of Fable 5's peak score at half the cost per task.
AutomationBench: 26.0% vs 18.1% for Sol. Zapier's business workflow benchmark tests whether a model can carry a real task from start to finish. Not just start it. Finish it. Opus 5's pass rate runs 1.5x the next closest model at comparable cost. Even at its lowest effort setting, it passes more tasks than any other model at any effort level.
There's a pattern here. Opus 5 excels specifically on tasks that require sustained effort, self-correction, and doing real work rather than producing plausible-looking output. It checks its own work, iterates until things actually function, and doesn't give up early.
Where Opus 5 Trails
DeepSWE v1.1: Sol leads at 72.7%. GPT-5.6 Sol still beats both Opus 5 (68.8%) and Fable 5 (69.7%) on this agentic coding benchmark. If your primary use case is SWE-bench-style code tasks inside an IDE, Sol retains an edge.
Health and Legal benchmarks: Fable 5 leads. On HealthBench Professional, Fable 5 scores 66.0% (using Mythos 5) vs Opus 5's 59.8%. On the Legal Agent Benchmark, Fable 5 hits 13.3% vs Opus 5's 11.7%. For specialized professional domains, Fable 5 still justifies its premium.
Speed: No published tok/s yet. Artificial Analysis hasn't reported throughput numbers at the time of writing. Anthropic offers a "Fast mode" at 2.5x default speed for double the price ($10/$50), which puts Fast mode at Fable 5 pricing. Standard mode speed is unknown but expected to be competitive with Opus 4.8.
Cybersecurity: Behind Mythos 5 on exploits. Opus 5 matches Mythos 5 on finding vulnerabilities but falls well behind on exploit development. Anthropic designed it this way. The safeguards let you scan for bugs but block pen testing and exploit generation. If you need offensive security capabilities, Mythos 5 remains the option through Anthropic's Cyber Verification Program.
Compare Opus 5 against 200+ models
See how it stacks up on quality, speed, and price. Filter by task type and budget.
The Cost Story
This is where Opus 5 changes the calculus. Anthropic isn't trying to be the cheapest. They're arguing that cost per completed task matters more than cost per token.
| Model | Input/Output $/1M | Est. Cost/Task | Intelligence |
|---|---|---|---|
| Claude Fable 5 | $10/$50 | ~$3.50 | 60 |
| Claude Opus 5 | $5/$25 | ~$1.75 | 61 |
| GPT-5.6 Sol | $5/$30 | ~$2.50 | 59 |
| Kimi K3 | $3/$15 | ~$0.94 | 57 |
| GPT-5.6 Luna | $1/$6 | ~$0.35 | 51 |
| Gemini 3.6 Flash | $1.50/$7.50 | ~$0.50 | 50 |
Opus 5 at ~$1.75 per task vs Fable 5 at ~$3.50 is a 50% reduction for 1 point more on the Intelligence Index. Compared to Sol ($2.50/task, Intelligence 59), Opus 5 costs 30% less and scores 2 points higher.
The efficiency claim has teeth. Early testers report Opus 5 uses fewer tokens and fewer turns to complete the same work. Harvey AI found it averaged 26% fewer tokens than Opus 4.8 at max reasoning with similar or better accuracy on legal tasks. Letta saw 60% less time and a third fewer tool calls on financial modeling. These aren't marketing claims from Anthropic. They're from customers running production workloads.
Rough monthly estimates for 500 tasks at 20K tokens average:
| Model | Monthly Cost |
|---|---|
| Claude Fable 5 | ~$300 |
| Claude Opus 5 | ~$150 |
| GPT-5.6 Sol | ~$175 |
| Kimi K3 | ~$94 |
| GPT-5.6 Luna | ~$35 |
| Gemini 3.6 Flash | ~$45 |
Use the Cost Calculator for estimates based on your actual usage.
What Early Testers Report
The customer testimonials in Anthropic's announcement are specific enough to be useful. A few patterns emerge.
It finishes the job. Multiple testers highlight that Opus 5 doesn't stop at "good enough." Lovable reports 22% improvement over Opus 4.7 on their hardest agentic coding tasks with far less variance run to run. Zapier says it took a raw account-health workbook and ran a full churn-prevention sequence end to end. Previous models didn't pass. Opus 5 hit 100%.
It self-corrects. On Frontier-Bench, given a drawing of a machine part with no direct way to view it, Opus 5 wrote its own computer vision pipeline to extract geometry from raw pixels, then reconstructed the full part. No competing model could solve it after five attempts. On an open-source package manager bug, Opus 5 found the root cause and caught an edge case that the community's patch had missed.
It pushes back when you're wrong. JetBrains reports that during a rearchitecting session, Opus 5 pushed back on a design proposal, didn't fold when pressed, explained what was valuable in the original idea, narrowed its objection to one design question, and proposed a compromise. That's the kind of interaction that reduces oversight.
It's efficient. A trading firm found Opus 5 uses roughly a seventh of the reasoning tokens and under half the latency of Opus 4.8 on their benchmark. Better answers at a fraction of the compute.
Opus 5 vs Fable 5: Which Do You Need?
This is now the hardest comparison in AI. Same company, similar capabilities, 2x price difference.
| Opus 5 | Fable 5 | |
|---|---|---|
| Intelligence | 61 | 60 |
| Frontier-Bench | 43.3% | 33.7% |
| DeepSWE | 68.8% | 69.7% |
| HLE (with tools) | 64.7% | 63.9% |
| Cost per task | ~$1.75 | ~$3.50 |
| Price | $5/$25 | $10/$50 |
| Safety classifiers | 85% fewer triggers | More restrictive |
Opus 5 wins on Frontier-Bench, GDPval-AA, OSWorld, AutomationBench, and HLE. Fable 5 wins narrowly on DeepSWE, HealthBench, and Legal. The differences are small enough that for most workloads, Opus 5 does the same job at half the price.
The safety classifier difference is worth knowing. Opus 5's classifiers trigger 85% less often than Fable 5's. In Claude Code and Claude.ai, flagged requests fall back to Opus 4.8 automatically. If you've been hitting Fable 5 refusals on legitimate coding or security work, Opus 5 might solve that problem.
Switch from Fable 5 if your work is primarily coding, business automation, or general knowledge work. You'll save 50% and likely get similar or better results.
Stay on Fable 5 if you need peak performance on specialized medical or legal tasks, or if you need Mythos-class capabilities for biological research.
Opus 5 vs GPT-5.6 Sol
| Opus 5 | GPT-5.6 Sol | |
|---|---|---|
| Intelligence | 61 | 59 |
| Frontier-Bench | 43.3% | 34.4% |
| ARC-AGI-3 | 30.2% | 7.8% |
| DeepSWE | 68.8% | 72.7% |
| OSWorld 2.0 | 70.6% | 62.6% |
| Output price/1M | $25 | $30 |
Opus 5 is cheaper per output token ($25 vs $30), scores higher on intelligence (61 vs 59), and dominates on novel problem-solving and computer use. Sol leads on DeepSWE and has the OpenAI ecosystem (Codex, ultra reasoning mode, 1.05M context).
If you're deep in the OpenAI ecosystem with Codex integration, Sol makes sense. For everything else, Opus 5 is the stronger pick.
Most Aligned Model to Date
Anthropic's automated behavioral audit scored Opus 5 at 2.3 on misaligned behavior, the lowest of any recent model. Lower than Opus 4.8, Sonnet 5, and Fable 5. Fewer deceptive behaviors, stronger Constitutional adherence, and less susceptibility to being tricked into misuse.
This matters for production deployments where you can't manually review every output. A model that's less likely to confabulate, less likely to follow injection attacks, and more likely to flag uncertainty is worth the premium over cheaper alternatives.
Best Setup for Late July 2026
The Opus 5 release reshuffles the model selection advice.
If you were using Fable 5: Switch to Opus 5 for most tasks. Use Fable 5 only when you need the absolute highest performance on medical/legal evaluations or when Opus 5's classifiers block a legitimate request.
If you were using GPT-5.6 Sol: Test Opus 5 on your workload. If your work involves novel reasoning, computer use, or business automation, Opus 5 likely outperforms Sol. If your work is primarily SWE-bench-style coding in an IDE, Sol retains an edge.
If you were using Opus 4.8: Upgrade immediately. Same price, dramatically better across every metric. Opus 5 more than doubles Frontier-Bench scores while using fewer tokens per task.
Budget setup: Opus 5 ($5/$25) as primary, Gemini 3.5 Flash-Lite ($0.30/$2.50) for sub-agents. Estimated monthly: $60-120.
Two-provider setup: Opus 5 for reasoning-heavy work, GPT-5.6 Luna for high-volume sub-tasks. Monthly: $80-150.
Use the Model Selector to find the right configuration for your workload.
Estimate your monthly AI spend
Plug in your usage patterns and see costs across every model. Free, no signup.
FAQ
Is Claude Opus 5 better than GPT-5.6 Sol?
On most benchmarks, yes. Opus 5 scores 61 on the Intelligence Index vs Sol's 59, leads on Frontier-Bench (43.3% vs 34.4%), ARC-AGI-3 (30.2% vs 7.8%), and OSWorld (70.6% vs 62.6%). Sol wins on DeepSWE (72.7% vs 68.8%) and costs slightly more on output ($30 vs $25 per 1M tokens). For novel problem-solving and business automation, Opus 5 is stronger. For SWE-bench-style coding, Sol retains an edge.
How much does Claude Opus 5 cost?
$5 per million input tokens and $25 per million output tokens. Same price as Opus 4.8. Cache hits cost $0.50/1M (90% discount). A Fast mode is available at 2x the base price ($10/$50) with 2.5x faster speed. No data retention requirements for general access.
Should I switch from Fable 5 to Opus 5?
For most workloads, yes. Opus 5 scores 1 point higher on the Intelligence Index (61 vs 60) and costs half as much ($5/$25 vs $10/$50). It matches or beats Fable 5 on most evaluations except HealthBench Professional and Legal Agent Benchmark. Stay on Fable 5 only if you need peak medical/legal performance or Mythos-class biology capabilities.
What is ARC-AGI-3 and why does it matter?
ARC-AGI-3 tests whether a model can solve genuinely novel problems it hasn't seen during training. It's designed to be unsolvable through pattern-matching alone. Opus 5 scores 30.2%, three times higher than any other model. This suggests Opus 5 has meaningfully better general reasoning capabilities, not just memorized benchmark patterns.
Is Opus 5 available in Claude Code?
Yes. Opus 5 is available today across Claude.ai, the Claude API, Claude Code, and Claude Cowork. Use the model string claude-opus-5 on the API. It's the default model on Claude Max and the strongest available on Claude Pro.
How does Opus 5 compare to Kimi K3?
Opus 5 scores higher on intelligence (61 vs 57) but costs significantly more ($5/$25 vs $3/$15). K3 at $0.94 per task is roughly half of Opus 5's ~$1.75. K3 will have open weights (self-hosting possible), which Opus 5 does not offer. If cost is the priority and 57 Intelligence is sufficient, K3 is the better value. If you need the highest intelligence and 30% ARC-AGI-3 performance, Opus 5 is unmatched.
What's the difference between Opus 5 and Opus 5 Fast mode?
Fast mode runs at approximately 2.5x the default speed at double the price ($10/$50 per 1M tokens). At those prices, it costs the same as Fable 5 but runs faster. Use default mode for cost efficiency and Fast mode when latency matters more than price.
Is Opus 5 safe for production use?
Anthropic's behavioral audit rates Opus 5 as their most aligned model to date, with the lowest rates of deceptive behavior and the strongest Constitutional adherence. Its cyber classifiers are less restrictive than Fable 5's (85% fewer triggers), allowing vulnerability scanning while blocking exploit generation. Flagged requests automatically fall back to Opus 4.8.
When should I use Opus 5 vs Luna vs Flash-Lite?
Use Opus 5 for hard problems that need deep reasoning, self-correction, and high accuracy. Use GPT-5.6 Luna ($1/$6) for high-volume sub-agent tasks where speed matters more than peak quality. Use Gemini 3.5 Flash-Lite ($0.30/$2.50) for the cheapest possible inference on simple tasks. A three-tier setup lets you match cost to complexity.
Benchmark data from Artificial Analysis Intelligence Index and Anthropic's official announcement. Early tester reports from Lovable, Zapier, Harvey AI, Letta, JetBrains, Box, Devin, and Cursor. Pricing from Anthropic's API documentation. Analysis from officechai.com. Updated July 24, 2026. See the LLM Leaderboard for live rankings, Best LLM in 2026 for the full comparison, or the Benchmark Dashboard for side-by-side model analysis. Compare pricing with the Cost Calculator.

