Claude Opus 5: Benchmarks, Pricing, and Who Should Switch
Claude Opus 5 pricing: $5/$25 per million tokens. Scores 61 on the Intelligence Index (#1) and 30.2% on ARC-AGI-3. Benchmarks, effort levels, and cost-per-task vs Fable 5, GPT-5.6 Sol, Kimi K3.

Anthropic just released Opus 5 and the positioning is different from their usual launches. They're not leading with "we beat OpenAI." They're leading with "we nearly beat ourselves at half the price." Opus 5 scores 61 on the Artificial Analysis Intelligence Index, one point above Fable 5, at $5/$25 instead of $10/$50.
The ARC-AGI-3 result is the one that stops you scrolling. Opus 5 scores three times higher than any other model on novel problem-solving. Not incremental gains. A 3x gap.
Here's everything that matters: benchmarks, real costs, where it falls short, and whether you should switch from Fable 5 or Sol.
Opus 5 by the Numbers
| Opus 5 | Fable 5 | GPT-5.6 Sol | Opus 4.8 | Kimi K3 | |
|---|---|---|---|---|---|
| Intelligence Index | 61 | 60 | 59 | 55 | 57 |
| Frontier-Bench v0.1 | 43.3% | 33.7% | 34.4% | 21.1% | — |
| GDPval-AA v2 | 1861 | 1747 | 1736 | 1593 | 1668 |
| ARC-AGI-3 | 30.2% | — | 7.8% | 1.5% | — |
| OSWorld 2.0 | 70.6% | 66.1% | 62.6% | 55.7% | — |
| AutomationBench | 26.0% | 17.4% | 18.1% | 17.0% | — |
| DeepSWE v1.1 | 68.8% | 69.7% | 72.7% | 59.0% | — |
| HLE (no tools) | 56.3% | 56.5% | — | 49.8% | — |
| Input Price/1M | $5.00 | $10.00 | $5.00 | $5.00 | $3.00 |
| Output Price/1M | $25.00 | $50.00 | $30.00 | $25.00 | $15.00 |
| Cache Hit/1M | $0.50 | — | $0.10 | $0.50 | $0.30 |
| Context | 1M | 1M | 1.05M | 1M | 1M |
Opus 5 tops the Intelligence Index at 61. That's the first time an Opus-tier model has beaten Fable 5 (60) on the composite score. The gap is small, but the direction matters: Anthropic's mid-tier model now matches or exceeds their top-tier on most evaluations.
Speed is now confirmed at 52.3 tok/s on Anthropic's API. That's slow compared to Sol (85 tok/s) or Luna (150 tok/s). If you need fast interactive responses, this isn't the model. But Opus 5 isn't designed for speed. It's designed to finish hard tasks in fewer turns with fewer total tokens.
Same price as Opus 4.8. No price hike. That's the real story. Every other lab charges a premium for their latest flagship. Anthropic kept it at $5/$25 and delivered a model that competes with their $10/$50 offering.
Five Effort Levels: Pick Your Tradeoff
Opus 5 ships with five reasoning effort settings. This is where it gets interesting. At low effort, it matches GPT-5.6 Luna's Intelligence Index score but with better per-task quality on complex work. At max, it leads the entire field.
| Effort | Intelligence | Speed | Cost per II Task | Output Tokens | Total Eval Cost |
|---|---|---|---|---|---|
| Max | 61 | 52.3 tok/s | $2.03 | 100M | $3,836 |
| Xhigh | 60 | 52.4 tok/s | $1.54 | 76M | $2,910 |
| High | 59 | ~52 tok/s | ~$1.20 | ~55M | ~$2,100 |
| Medium | 56 | 51.5 tok/s | ~$0.59 | 29M | $1,115 |
| Low | 51 | 54.2 tok/s | ~$0.29 | 12M | $556 |
The range here is wide. From low (51) to max (61), you get 10 points on the Intelligence Index. Token usage spans roughly 8x between low and max. That means you can run Opus 5 at low effort as a fast, lean model and switch to max when the problem demands it. Same API key, same model string, just change the effort parameter.
At xhigh (Intelligence 60), Opus 5 matches Fable 5's score while costing $2,910 to run the full Intelligence Index instead of Fable 5's evaluation cost. That's the sweet spot for most production workloads: Fable-level quality at Opus pricing.
At medium (Intelligence 56), it matches Opus 4.8's max score while using far fewer tokens. If you're currently running Opus 4.8 at max effort, switching to Opus 5 at medium gets you the same intelligence for roughly a third of the token cost.
Where Opus 5 Leads
Four results stand out from the benchmark table.
ARC-AGI-3: 30.2% vs 7.8% for GPT-5.6 Sol. This benchmark tests novel problem-solving where the model can't rely on pattern-matching from training data. Opus 5 doesn't just win. It triples the nearest competitor. Opus 4.8 scored 1.5%. That's a generational jump. For any workflow that involves genuinely new problems rather than variations of seen patterns, this gap is significant.
Frontier-Bench v0.1: 43.3% vs 34.4% for Sol. Anthropic's agentic terminal coding benchmark. Opus 5 more than doubles Opus 4.8's score (21.1%) while costing less per task because it uses fewer tokens to reach an answer. On CursorBench 3.2, Opus 5 at max effort lands within 0.5% of Fable 5's peak score at half the cost per task.
Coding Agent Index: Joint first place. Opus 5 at xhigh effort with Claude Code ties for #1 on the Artificial Analysis Coding Agent Index. It also posts 89% on Terminal-Bench v2.1 at max effort, roughly matching GPT-5.6 Sol (xhigh). For agentic coding through the terminal, it's now at parity with the best.
AutomationBench: 26.0% vs 18.1% for Sol. Zapier's business workflow benchmark tests whether a model can carry a real task from start to finish. Not just start it. Finish it. Opus 5's pass rate runs 1.5x the next closest model at comparable cost. Even at its lowest effort setting, it passes more tasks than any other model at any effort level.
There's a pattern here. Opus 5 excels specifically on tasks that require sustained effort, self-correction, and doing real work rather than producing plausible-looking output. It checks its own work, iterates until things actually function, and doesn't give up early.
Agentic Knowledge Work: The AA-Briefcase Results
Artificial Analysis built a benchmark called AA-Briefcase that tests models on realistic professional work: research reports, presentations, spreadsheets. Thousands of input files per task. Deliverables that need to be both correct and well-presented. Opus 5 dominates it.
| Model | AA-Briefcase Elo | Cost/Task | Time/Task | Turns/Task |
|---|---|---|---|---|
| Opus 5 (max) | 1720 | $17.79 | 36.2 min | 103 |
| Opus 5 (xhigh) | 1693 | $14.26 | 34.3 min | 91 |
| Opus 5 (high) | 1606 | $10.41 | 25.7 min | 76 |
| Fable 5 | 1574 | $22.30 | — | — |
| GPT-5.6 Sol (max) | 1505 | — | — | — |
| Opus 5 (medium) | 1470 | — | — | — |
| GLM-5.2 (max) | 1254 | — | — | — |
| Opus 5 (low) | 1223 | — | — | — |
Opus 5 takes the top three spots. At high effort, it outperforms Fable 5 by 32 Elo while costing $10.41 per task, less than half of Fable 5's $22.30. That's the number that matters for enterprise buyers: Fable-beating performance at 47% of the price.
The gains come from rubric pass rate and analytical quality. Artificial Analysis reports Opus 5's Analytical Quality Elo at 2016 (max effort), nearly 300 points ahead of Fable 5. Presentation quality improved over Opus 4.8 but still sits about 40 Elo behind GPT-5.6 Sol. If your deliverables are judged on how they look (slides, formatted reports), Sol still has an edge. If they're judged on correctness and analysis, Opus 5 is ahead.
The time-per-task numbers deserve attention. 36 minutes at max effort. 103 turns per task. That's roughly 50% longer than Opus 4.8 (24 minutes, 55 turns). Opus 5 uses that extra time to verify its work, iterate on edge cases, and produce more accurate output. It's not slower because it's less efficient. It's slower because it's doing more work per task.
Where Opus 5 Trails
DeepSWE v1.1: Sol leads at 72.7%. GPT-5.6 Sol still beats both Opus 5 (68.8%) and Fable 5 (69.7%) on this agentic coding benchmark. If your primary use case is SWE-bench-style code tasks inside an IDE, Sol retains an edge.
Hallucination rate jumped 14 points. This one matters. On AA-Omniscience, Opus 5 improved accuracy by 7 points over Opus 4.8. But its hallucination rate rose 14 points to 50%. The model answers more often when uncertain instead of declining to respond. That's a net negative for factual reliability in production. Fable 5 still has lower hallucination rates because it's a larger model with more parametric knowledge. If your use case depends on factual correctness without external retrieval, test carefully before switching.
Health and Legal benchmarks: Fable 5 leads. On HealthBench Professional, Fable 5 scores 66.0% (using Mythos 5) vs Opus 5's 59.8%. On the Legal Agent Benchmark, Fable 5 hits 13.3% vs Opus 5's 11.7%. For specialized professional domains, Fable 5 still justifies its premium.
Physics: Behind Sol and Terra on CritPt. On this frontier physics reasoning benchmark developed by Argonne and UIUC researchers, Opus 5 matches Fable 5 but sits behind GPT-5.6 Sol, GPT-5.5 Pro, and GPT-5.6 Terra. If physics reasoning is central to your work, the GPT-5.6 family has an edge here.
Speed: 52.3 tok/s is slow. Now confirmed by Artificial Analysis. At 52.3 tok/s, Opus 5 is well below Sol (85 tok/s), Luna (150 tok/s), and even the median for models in its price tier (72.5 tok/s). Time to first answer token at max effort is 62.7 seconds. That's not interactive. Anthropic offers a Fast mode at 2.5x speed for $10/$50 (Fable 5 pricing). For tasks where you need quick responses, this is a drawback.
Cybersecurity: Behind Mythos 5 on exploits. Opus 5 matches Mythos 5 on finding vulnerabilities but falls well behind on exploit development. Anthropic designed it this way. The safeguards let you scan for bugs but block pen testing and exploit generation. If you need offensive security capabilities, Mythos 5 remains the option through Anthropic's Cyber Verification Program.
Compare Opus 5 against 200+ models
See how it stacks up on quality, speed, and price. Filter by task type and budget.
The Cost Story
This is where Opus 5 changes the calculus. Anthropic isn't trying to be the cheapest. They're arguing that cost per completed task matters more than cost per token.
Artificial Analysis measured the actual cost per Intelligence Index task across effort levels. These aren't estimates. They're based on real token usage during evaluation:
| Model | Input/Output $/1M | Cost per II Task | Intelligence |
|---|---|---|---|
| Claude Fable 5 | $10/$50 | $2.75 | 60 |
| Claude Opus 5 (max) | $5/$25 | $2.03 | 61 |
| Claude Opus 5 (xhigh) | $5/$25 | $1.54 | 60 |
| GPT-5.6 Sol | $5/$30 | ~$2.50 | 59 |
| Claude Sonnet 5 (max) | $3/$15 | $1.53 | 53 |
| Claude Opus 4.8 (max) | $5/$25 | $1.80 | 56 |
| Kimi K3 | $3/$15 | ~$0.94 | 57 |
| GPT-5.6 Luna | $0.20/$1.20 | ~$0.07 | 51 |
| Gemini 3.6 Flash | $1.50/$7.50 | ~$0.50 | 50 |
At max effort, Opus 5 costs $2.03 per task vs Fable 5's $2.75. That's 26% cheaper for 1 point more on the Intelligence Index. At xhigh ($1.54/task), it matches Fable 5's intelligence score (60) at 44% lower cost.
Here's the nuance most people will miss: Opus 5 at max effort actually costs more per task than Opus 4.8 ($2.03 vs $1.80). It uses more tokens and more turns. You're paying for thoroughness. The payoff is a 5-point Intelligence Index jump and dramatically better task completion rates. On AA-Briefcase, Opus 5 takes 103 turns per task at max effort vs Opus 4.8's 55 turns. It's spending those extra turns verifying and iterating.
The efficiency claim from early testers backs this up but from a different angle. Harvey AI found Opus 5 averaged 26% fewer tokens than Opus 4.8 at max reasoning with similar or better accuracy on legal tasks. Letta saw 60% less time and a third fewer tool calls on financial modeling. The model achieves more per turn even when it takes more turns total.
Use the Cost Calculator for estimates based on your actual usage.
What Early Testers Report
The customer testimonials in Anthropic's announcement are specific enough to be useful. A few patterns emerge.
It finishes the job. Multiple testers highlight that Opus 5 doesn't stop at "good enough." Lovable reports 22% improvement over Opus 4.7 on their hardest agentic coding tasks with far less variance run to run. Zapier says it took a raw account-health workbook and ran a full churn-prevention sequence end to end. Previous models didn't pass. Opus 5 hit 100%.
It self-corrects. On Frontier-Bench, given a drawing of a machine part with no direct way to view it, Opus 5 wrote its own computer vision pipeline to extract geometry from raw pixels, then reconstructed the full part. No competing model could solve it after five attempts. On an open-source package manager bug, Opus 5 found the root cause and caught an edge case that the community's patch had missed.
It pushes back when you're wrong. JetBrains reports that during a rearchitecting session, Opus 5 pushed back on a design proposal, didn't fold when pressed, explained what was valuable in the original idea, narrowed its objection to one design question, and proposed a compromise. That's the kind of interaction that reduces oversight.
It's efficient. A trading firm found Opus 5 uses roughly a seventh of the reasoning tokens and under half the latency of Opus 4.8 on their benchmark. Better answers at a fraction of the compute.
Opus 5 vs Fable 5: Which Do You Need?
This is now the hardest comparison in AI. Same company, similar capabilities, 2x price difference.
| Opus 5 (max) | Fable 5 | |
|---|---|---|
| Intelligence | 61 | 60 |
| Frontier-Bench | 43.3% | 33.7% |
| DeepSWE | 68.8% | 69.7% |
| HLE (with tools) | 64.7% | 63.9% |
| AA-Briefcase Elo | 1720 | 1574 |
| GDPval-AA v2 | 1861 | 1747 |
| Cost per II task | $2.03 | $2.75 |
| Price | $5/$25 | $10/$50 |
| Speed | 52.3 tok/s | ~71 tok/s |
| Safety classifiers | 85% fewer triggers | More restrictive |
Opus 5 wins on Frontier-Bench, GDPval-AA, OSWorld, AutomationBench, and HLE. Fable 5 wins narrowly on DeepSWE, HealthBench, and Legal. The differences are small enough that for most workloads, Opus 5 does the same job at half the price.
The safety classifier difference is worth knowing. Opus 5's classifiers trigger 85% less often than Fable 5's. In Claude Code and Claude.ai, flagged requests fall back to Opus 4.8 automatically. If you've been hitting Fable 5 refusals on legitimate coding or security work, Opus 5 might solve that problem.
Switch from Fable 5 if your work is primarily coding, business automation, or general knowledge work. You'll save 50% and likely get similar or better results.
Stay on Fable 5 if you need peak performance on specialized medical or legal tasks, or if you need Mythos-class capabilities for biological research.
Opus 5 vs GPT-5.6 Sol
| Opus 5 (max) | GPT-5.6 Sol (max) | |
|---|---|---|
| Intelligence | 61 | 59 |
| Frontier-Bench | 43.3% | 34.4% |
| ARC-AGI-3 | 30.2% | 7.8% |
| DeepSWE | 68.8% | 72.7% |
| OSWorld 2.0 | 70.6% | 62.6% |
| AA-Briefcase Elo | 1720 | 1505 |
| CritPt (physics) | Behind Sol | Leads |
| Presentation Elo | 1628 | 1666 |
| Cost per II task | $2.03 | ~$2.50 |
| Speed | 52.3 tok/s | 85 tok/s |
Opus 5 is cheaper per output token ($25 vs $30), scores higher on intelligence (61 vs 59), and dominates on novel problem-solving, computer use, and agentic knowledge work (+215 AA-Briefcase Elo). Sol leads on DeepSWE, CritPt physics, and presentation quality. Sol is also 63% faster at 85 tok/s.
If you're deep in the OpenAI ecosystem with Codex integration, or your work is primarily SWE-bench-style coding, Sol makes sense. If you need agentic knowledge work, novel reasoning, or computer use, Opus 5 wins clearly.
Most Aligned Model to Date
Anthropic's automated behavioral audit scored Opus 5 at 2.3 on misaligned behavior, the lowest of any recent model. Lower than Opus 4.8, Sonnet 5, and Fable 5. Fewer deceptive behaviors, stronger Constitutional adherence, and less susceptibility to being tricked into misuse.
This matters for production deployments where you can't manually review every output. A model that's less likely to confabulate, less likely to follow injection attacks, and more likely to flag uncertainty is worth the premium over cheaper alternatives.
Best Setup for Late July 2026
The Opus 5 release reshuffles the model selection advice. The effort levels add flexibility that previous models didn't have.
If you were using Fable 5: Switch to Opus 5 at high or xhigh effort. You'll get Fable-level quality (or better on AA-Briefcase) at roughly half the cost. Keep Fable 5 for medical/legal evaluations or when Opus 5's classifiers block a legitimate request.
If you were using GPT-5.6 Sol: Test Opus 5 at xhigh on your workload. If your work involves novel reasoning, computer use, or business automation, Opus 5 likely outperforms Sol. If your work is SWE-bench-style coding in an IDE, or you need polished presentation output, Sol retains edges there.
If you were using Opus 4.8: Upgrade immediately. Switch to Opus 5 at medium effort for the same Intelligence score (56) at roughly a third of the token cost. Use high or xhigh when you need more.
Effort-level strategy: Use medium for day-to-day coding and simple tasks. Use high or xhigh for important deliverables and complex analysis. Reserve max for problems that genuinely need frontier intelligence. This single-model approach with variable effort can replace a multi-model routing setup.
Budget setup: Opus 5 at medium/high ($5/$25) as primary, GPT-5.6 Luna ($0.20/$1.20 after the July 30 price cut) for sub-agents. Luna is now cheaper than Flash-Lite with much higher intelligence. Estimated monthly: $40-80.
Two-provider setup: Opus 5 at high for reasoning-heavy work, GPT-5.6 Luna for everything else. Monthly: $50-100.
Use the Model Selector to find the right configuration for your workload.
Estimate your monthly AI spend
Plug in your usage patterns and see costs across every model. Free, no signup.
FAQ
Is Claude Opus 5 better than GPT-5.6 Sol?
On most benchmarks, yes. Opus 5 scores 61 on the Intelligence Index vs Sol's 59, leads on Frontier-Bench (43.3% vs 34.4%), ARC-AGI-3 (30.2% vs 7.8%), and OSWorld (70.6% vs 62.6%). Sol wins on DeepSWE (72.7% vs 68.8%) and costs slightly more on output ($30 vs $25 per 1M tokens). For novel problem-solving and business automation, Opus 5 is stronger. For SWE-bench-style coding, Sol retains an edge.
How much does Claude Opus 5 cost?
$5 per million input tokens and $25 per million output tokens. Same price as Opus 4.8. Cache writes cost $6.25/1M (25% premium), cache hits cost $0.50/1M (90% discount). Actual cost per Intelligence Index task ranges from $0.29 (low effort) to $2.03 (max effort). A Fast mode is available at 2x the base price ($10/$50) with 2.5x faster speed. No data retention requirements for general access.
Which Opus 5 effort level should I use?
Start at medium (Intelligence 56, ~$0.59/task). It matches Opus 4.8 max performance at a fraction of the tokens. Move to high (Intelligence 59, ~$1.20/task) for important work. Use xhigh (Intelligence 60, $1.54/task) when you need Fable-level quality. Reserve max (Intelligence 61, $2.03/task) for frontier problems. Token usage spans roughly 8x from low to max, so effort level has a huge impact on cost.
How fast is Claude Opus 5?
52.3 tok/s at max effort, confirmed by Artificial Analysis. That's below the median for its price tier (72.5 tok/s) and well behind Sol (85 tok/s) or Luna (150 tok/s). Time to first answer token at max effort is 62.7 seconds. Fast mode runs at approximately 2.5x speed for double the price. Opus 5 is built for task quality, not interactive speed.
Does Opus 5 hallucinate more than Opus 4.8?
Yes. On AA-Omniscience, Opus 5 improved factual accuracy by 7 points over Opus 4.8, but its hallucination rate rose 14 points to 50%. The model answers more often when uncertain instead of declining. If factual reliability without external retrieval is critical for your use case, pair it with RAG or test carefully before switching from Fable 5 (which has lower hallucination rates).
Should I switch from Fable 5 to Opus 5?
For most workloads, yes. Opus 5 scores 1 point higher on the Intelligence Index (61 vs 60) and costs half as much ($5/$25 vs $10/$50). It matches or beats Fable 5 on most evaluations except HealthBench Professional and Legal Agent Benchmark. Stay on Fable 5 only if you need peak medical/legal performance or Mythos-class biology capabilities.
What is ARC-AGI-3 and why does it matter?
ARC-AGI-3 tests whether a model can solve genuinely novel problems it hasn't seen during training. It's designed to be unsolvable through pattern-matching alone. Opus 5 scores 30.2%, three times higher than any other model. This suggests Opus 5 has meaningfully better general reasoning capabilities, not just memorized benchmark patterns.
Is Opus 5 available in Claude Code?
Yes. Opus 5 is available today across Claude.ai, the Claude API, Claude Code, and Claude Cowork. Use the model string claude-opus-5 on the API. It's the default model on Claude Max and the strongest available on Claude Pro.
How does Opus 5 compare to Kimi K3?
Opus 5 scores higher on intelligence (61 vs 57) but costs significantly more ($5/$25 vs $3/$15). K3 at $0.94 per task is less than half of Opus 5's $2.03 (max effort). K3 will have open weights (self-hosting possible), which Opus 5 does not offer. If cost is the priority and 57 Intelligence is sufficient, K3 is the better value. If you need the highest intelligence and 30% ARC-AGI-3 performance, Opus 5 is unmatched.
What's the difference between Opus 5 and Opus 5 Fast mode?
Fast mode runs at approximately 2.5x the default speed at double the price ($10/$50 per 1M tokens). At those prices, it costs the same as Fable 5 but runs faster. Use default mode for cost efficiency and Fast mode when latency matters more than price.
Is Opus 5 safe for production use?
Anthropic's behavioral audit rates Opus 5 as their most aligned model to date, with the lowest rates of deceptive behavior and the strongest Constitutional adherence. Its cyber classifiers are less restrictive than Fable 5's (85% fewer triggers), allowing vulnerability scanning while blocking exploit generation. Flagged requests automatically fall back to Opus 4.8.
When should I use Opus 5 vs Luna vs Flash-Lite?
Use Opus 5 for hard problems that need deep reasoning, self-correction, and high accuracy. Use GPT-5.6 Luna ($0.20/$1.20 after the July 30 price cut) for high-volume sub-agent tasks. At those prices, Luna is now cheaper than Flash-Lite while scoring 51 on the Intelligence Index. A two-tier setup (Opus 5 + Luna) covers most workloads efficiently.
Benchmark data from Artificial Analysis Intelligence Index v4.1, Coding Agent Index, and AA-Briefcase. Cost per task and effort-level data from Artificial Analysis Opus 5 evaluation. Hallucination data from AA-Omniscience. CritPt physics benchmark by Argonne and UIUC researchers. Early tester reports from Lovable, Zapier, Harvey AI, Letta, JetBrains, Box, Devin, and Cursor. Pricing from Anthropic's official announcement. Analysis from OfficeChai. Updated July 30, 2026. See the LLM Leaderboard for live rankings, Best LLM in 2026 for the full comparison, or the Benchmark Dashboard for side-by-side model analysis. Compare pricing with the Cost Calculator.

