Claude Opus 5.5: Benchmarks, Pricing and Is It Worth It?
Claude Opus 5.5 is #1 on Artificial Analysis (58 Intelligence, 66 Coding Agent) at $4/$20. Real cost per task by effort level, speed, and when cheaper wins.
· 22 min read

Claude Opus 5.5 is the strongest AI model you can call through an API right now. It scores 58 on the Artificial Analysis Intelligence Index and 66 on the Coding Agent Index, first on both, and it costs $4/$20 per million input/output tokens, 20% less than Opus 5.
It's worth it at high effort, where it scores 53.6, answers in about 41 seconds and costs $1.82 per benchmark task. That beats Opus 5 at max effort on score for less than a third of the cost. Max effort is where the value falls apart: $5.98 per task, 3.3 times the cost of high, for 4 more points.
Scores, costs and speeds below come from Artificial Analysis data pulled on 25 September 2026. Product facts come from Anthropic's announcement and model docs.
Claude Opus 5.5 specs and prices
| Claude Opus 5.5 | |
|---|---|
| Released | 22 September 2026 |
| API model ID | claude-opus-5-5 |
| Price per 1M tokens | $4 input / $20 output |
| Cache reads | $0.20 per 1M tokens |
| Batch API | $2 / $10 (50% off) |
| Context window | 1M tokens |
| Max output | 128K tokens (300K on the Batch API, beta) |
| Input and output | Text and images in, text out |
| Default effort | Medium |
| Intelligence Index v4.3.2 | 58 at max effort, 53.6 at high |
| Coding Agent Index v1.5 | 66 |
| Output speed | 90 tokens per second |
| Knowledge cutoff | June 2026 |
The two numbers that matter most for buying decisions are the price and the default effort. Opus 5.5 is 60% cheaper per token than Claude Fable 5.1 and GPT-6 Astra, and a request that doesn't set effort runs at medium, which scores 51.2 rather than the 58 in the headlines.
What's new in Opus 5.5 compared with Opus 5
Anthropic says Opus 5.5 costs 40% less than Opus 5 on typical workloads at default settings and generates output more than 30% faster. It also scores well ahead of Opus 5 on Anthropic's agentic coding and computer-use benchmarks. The changes that affect you, from the announcement and the what's new page:
- Prices drop from $5/$25 to $4/$20 per million tokens, and cache reads from $0.50 to $0.20.
- The default effort is now medium. Opus 5 defaulted to high, so an unchanged request now runs one level lower.
- Thinking is always on. A request that tries to disable it returns a 400 error, and effort is the only control over how much the model reasons.
- At the same effort level it thinks more per turn than Opus 5, most of all at xhigh and max.
- Forced tool use is gone:
tool_choiceset toanyor a named tool returns an error. - It reads dense charts, diagrams and screenshots more precisely without tools.
- It resists prompt injection better. In a new containment test, it tried to get around boundaries about 85% less often than Opus 5.
Anthropic's own benchmark table, run at max effort:
| Benchmark | Opus 5.5 | Opus 5 |
|---|---|---|
| Terminal-Bench 4.0 | 66.4% | 52.3% |
| CursorBench 4.0 | 57.8% | 46.6% |
| AutomationBench | 40.0% | 26.9% |
| OSWorld 2.0 (partial credit) | 81.8% | 74.0% |
| Humanity's Last Exam (with tools) | 67.7% | 63.6% |
| GDPval-AA v2.1 (Elo) | 1846 | 1708 |
The biggest jump is on Terminal-Bench, 14 points, and that benchmark is one of three inside the independent Coding Agent Index below. These are vendor numbers at max effort, so use them to see what improved and use the Artificial Analysis scores to compare across labs.
Benchmarks: first on both Artificial Analysis indexes
Opus 5.5 scores 58 on the Artificial Analysis Intelligence Index v4.3.2 and 66 on the Coding Agent Index v1.5. Claude Fable 5.1 and GPT-6 Astra tie for second on both, at 53 and 62. Both headline numbers are measured at max effort, which few people run day to day.
We also rank every model at its everyday setting: the strongest effort level that gives a full answer within 45 seconds. For Opus 5.5 that's high.
| Model | Intelligence at max effort | Intelligence at everyday setting |
|---|---|---|
| Claude Opus 5.5 | 57.6 | 53.6 (high) |
| Claude Fable 5.1 | 53.4 | 51.2 (high) |
| GPT-6 Astra | 52.7 | 49.6 (medium) |
| Claude Opus 5 | 50.8 | 49.7 (xhigh) |
| GPT-6 Sol | 47.5 | 44.1 (xhigh) |
At everyday settings the lead over Fable 5.1 shrinks from 4.2 points to 2.4, but Opus 5.5 stays first. Opus 5.5 at high already scores higher than Opus 5 does at max.
The Coding Agent Index pairs each model with its own coding tool (Claude Code for Anthropic) and scores it on DeepSWE, Terminal-Bench 4.0 and SWE-Atlas-QnA. Opus 5.5 leads Fable 5.1 and Astra by 4 points and GPT-6 Sol by 9. Artificial Analysis only publishes this index for free for its top models, so Claude Sonnet 5 has no score. Our best AI models for coding ranking has the full list, and the benchmark dashboard lets you compare any two models.
Opus 5.5 pricing and the real cost per task
The list price is $4 per million input tokens and $20 per million output tokens. Thinking is billed as output, so what you pay depends mostly on the effort level. Artificial Analysis measured Opus 5.5 at $0.55 per benchmark task at low effort and $5.98 at max, an 11x spread for the same model.
| Effort | Intelligence | Cost per AA task | Est. cost per 1,000 coding jobs | Full answer | First answer token |
|---|---|---|---|---|---|
| Low | 42.3 | $0.55 | $165 | 12.4 s | 6.0 s |
| Medium (default) | 51.2 | $1.34 | $245 | 24.9 s | 18.7 s |
| High | 53.6 | $1.82 | $296 | 40.7 s | 34.9 s |
| Extra high | 56.0 | $3.46 | $464 | 153.7 s | 148.2 s |
| Max | 57.6 | $5.98 | $724 | Not measured | Not measured |
Cost per AA task is what Artificial Analysis spent per Intelligence Index task, thinking included. The coding-job estimate is ours: a 45,000-token job, with output scaled by how much Opus 5.5 thinks at each level compared with a typical model. Both columns point the same way. Every step above high costs a lot more and buys little.
- Low$0.55
Scores 42.3, full answer in 12.4 s
- Medium (default)$1.34
Scores 51.2, full answer in 24.9 s
- High$1.82
Scores 53.6, full answer in 40.7 s
- Extra high$3.46
Scores 56.0, full answer in 2.5 minutes
- Max$5.98
Scores 57.6, response time not published
Going from high to max multiplies the cost per task by 3.3 for 4 more points.
Max effort is expensive enough to need its own warning. Artificial Analysis spent $8,708 running its full Intelligence Index on Opus 5.5 at max, against $2,172 at high. Extra high already takes about two and a half minutes per answer, and AA hasn't published a response time for max. Keep max for overnight agent runs where one extra correct answer pays for the tokens, and never make it the default.
The default is the opposite problem. Medium scores 2.4 points below high, costs 27% less per task and answers 16 seconds sooner. That's a fine setting for chat and routine agent steps. For coding and analysis that matters, set high explicitly.
Compared with Opus 5 at each model's default (medium for Opus 5.5, high for Opus 5), Opus 5.5 scores 51.2 against 48.1 and costs $1.34 per AA task against $3.61. Anthropic's own estimate of the saving is 40% on typical workloads; AA's tasks show more. Either way there is no reason to stay on Opus 5.
The other price points, from Anthropic's docs:
- Cache writes cost $5 per million tokens for 5 minutes and $8 for 1 hour. Cache reads cost $0.20.
- The Batch API halves the price to $2/$10.
- Fast mode runs at up to 2.5 times the speed for $8/$40, as a research preview. The API docs list it on the Claude API only, not on Amazon Bedrock, Google Cloud or Microsoft Foundry.
The cost calculator runs these numbers for your own volume.
How fast is Claude Opus 5.5?
Opus 5.5 writes at about 90 tokens per second on Artificial Analysis, up from 54 for Opus 5. It also thinks longer before it starts writing. At high effort the first answer token arrives after 34.9 seconds and the full answer after 40.7.
At the same effort label, Opus 5 answered sooner: 24.9 seconds at high and 14.6 at medium, against 40.7 and 24.9 for Opus 5.5. Anthropic's docs confirm the reason. Opus 5.5 spends more thinking per turn at each level, so faster token output doesn't mean faster answers.
For comparison at everyday settings, GPT-6 Astra answers in 15.0 seconds at medium, Claude Fable 5.1 in 29.2 at high and Gemini 3.8 Flash in 24.1. If you need Opus 5.5 to feel quick in a chat window, low effort gives a full answer in 12.4 seconds and still scores 42.3. Fast mode is the other option if the extra cost is acceptable.
Context window and max output
Opus 5.5 has a 1M-token context window and a 128K-token output limit on the Messages API. On the Batch API it can write up to 300K tokens with the output-300k-2026-03-24 beta header. Anthropic says 1M tokens is roughly 555,000 words on its current tokenizer.
| Model | Context window | Max output |
|---|---|---|
| Claude Opus 5.5 | 1M | 128K |
| Claude Fable 5.1 | 1M | 128K |
| GPT-6 Astra | 1.05M (922K input) | 128K |
| Claude Sonnet 5 | 1M | 128K |
| GPT-6 Sol | 1.05M (922K input) | 128K |
| Gemini 3.8 Flash | 1,048,576 | 65,536 |
Context isn't a reason to pick or avoid Opus 5.5. Most frontier models now take about a million tokens. Gemini 3.8 Flash writes about half as much in one response, and the GPT-6 models cap input at 922K tokens inside a 1.05M window.
Creative writing: Opus 5.5's weak spot
On EQ-Bench Creative Writing v3, Opus 5.5 scores 2050 Elo, seventh among the models we track and 83 points behind Opus 5. Its slop score, which counts stock AI phrasing, is 10.43 against Opus 5's 6.59 (lower is better). For fiction, the new model is a step back.
| Model | EQ-Bench Elo | Slop score |
|---|---|---|
| GPT-6 Astra | 2173 | 8.41 |
| Claude Fable 5.1 | 2162 | 8.16 |
| Claude Opus 5 | 2133 | 6.59 |
| GPT-6 Sol | 2125 | 8.97 |
| Kimi K3 | 2082 | 9.70 |
| GLM-5.3 | 2075 | 8.42 |
| Claude Opus 5.5 | 2050 | 10.43 |
Six models in our list write better stories by this measure, and three of them cost less per token. GPT-6 Sol scores 75 Elo higher at half the price per token, which is why our best AI for creative writing page picks GPT-6 Astra overall and Sol for value. If you write fiction with Claude and liked Opus 5, test Opus 5.5 before switching.
Opus 5.5 vs GPT-6 Astra, GPT-6 Sol, Fable 5.1, Sonnet 5 and Gemini 3.8 Flash
| Model | Price per 1M (in/out) | Intelligence (max) | Intelligence (everyday) | Coding Agent | Full answer (everyday) | Cost per AA task (everyday) | Est. 1,000 coding jobs |
|---|---|---|---|---|---|---|---|
| Claude Opus 5.5 | $4 / $20 | 58 | 53.6 (high) | 66 | 40.7 s | $1.82 | $296 |
| GPT-6 Astra | $10 / $50 | 53 | 49.6 (medium) | 62 | 15.0 s | $1.54 | $429 |
| Claude Fable 5.1 | $10 / $50 | 53 | 51.2 (high) | 62 | 29.2 s | $3.91 | $673 |
| GPT-6 Sol | $2 / $10 | 48 | 44.1 (xhigh) | 57 | 42.6 s | $0.53 | $109 |
| Claude Sonnet 5 | $2 / $10 | 38 | 34.4 (xhigh) | – | 23.9 s | $2.87 | $350 |
| Gemini 3.8 Flash | $0.75 / $3.75 | 41 | 40.9 (high) | 42 | 24.1 s | $1.24 | $148 |
Opus 5.5 wins every quality column. It loses on speed to Astra and on cost to GPT-6 Sol, and Sonnet 5 costs more per task despite its lower list price. Artificial Analysis tested Gemini 3.8 Flash only up to high effort, so its max and everyday scores are the same.
- Intelligence at max effort58 vs 53 for Astra and Fable 5.1Opus 5.5
- Intelligence at everyday setting53.6 at highOpus 5.5
- Coding Agent Index66, 4 points clearOpus 5.5
- Lowest cost per coding jobAbout $109 per 1,000 jobsGPT-6 Sol
- Quickest full answer15.0 s at medium effortGPT-6 Astra
- Output speed297 tokens per secondGemini 3.8 Flash
- Creative writing2173 EQ-Bench EloGPT-6 Astra
- Categories wonOpus 5.5 3GPT-6 Astra 2GPT-6 Sol 1Fable 5.1 0Gemini 3.8 Flash 1
| Category | Opus 5.5 | GPT-6 Astra | GPT-6 Sol | Fable 5.1 | Gemini 3.8 Flash |
|---|---|---|---|---|---|
| Intelligence at max effort58 vs 53 for Astra and Fable 5.1 | Wins | ||||
| Intelligence at everyday setting53.6 at high | Wins | ||||
| Coding Agent Index66, 4 points clear | Wins | ||||
| Lowest cost per coding jobAbout $109 per 1,000 jobs | Wins | ||||
| Quickest full answer15.0 s at medium effort | Wins | ||||
| Output speed297 tokens per second | Wins | ||||
| Creative writing2173 EQ-Bench Elo | Wins | ||||
| Categories won | 3 | 2 | 1 | 0 | 1 |
Sonnet 5 is left out of the grid because it wins none of these categories.
Opus 5.5 vs GPT-6 Astra
Astra costs 2.5 times as much per token ($10/$50) and scores lower on both indexes: 53 against 58 on intelligence and 62 against 66 on coding. Its advantage is response time. At its everyday setting (medium) Astra answers in 15.0 seconds, under half of Opus 5.5's 40.7, and it tops EQ-Bench for creative writing.
Per AA task, Astra at medium is slightly cheaper ($1.54 against $1.82) because it thinks very little. On input-heavy coding jobs, Opus 5.5's $4 input price wins, at about $296 per 1,000 jobs against $429. At max effort, Astra takes over six minutes per answer. Pick Opus 5.5 for coding and agents. Pick Astra if you're on OpenAI and need fast answers from a top-three model, or for fiction.
Opus 5.5 vs GPT-6 Sol
GPT-6 Sol is the value alternative. At $2/$10 it scores 48 on intelligence and 57 on coding, and at its everyday setting 1,000 coding jobs cost about $109, 37% of the Opus 5.5 bill. Both take a little over 40 seconds per full answer. Sol writes faster, at 110 tokens per second.
The gap is 9.5 points at everyday settings and 9 on the Coding Agent Index. For daily feature work where you review every diff anyway, Sol is enough. For hard bugs and long unattended agent runs, the extra points are worth about 19 cents a job.
Opus 5.5 vs Claude Fable 5.1
Fable 5.1 is Anthropic's larger model at $10/$50, and Opus 5.5 beats it on both indexes: 58 against 53, and 66 against 62 on coding. At everyday settings Fable 5.1 costs $3.91 per AA task against $1.82, more than twice as much for 2.4 fewer points.
Fable 5.1 answers sooner at its everyday setting (29.2 seconds against 40.7) and writes better fiction (2162 Elo, slop 8.16). On the Claude API, Fable 5.1 can also read Opus 5.5's thinking blocks, so a conversation can move up to Fable mid-task without losing its reasoning. For most coding and knowledge work, switch to Opus 5.5 unless your own tests show Fable doing better.
Opus 5.5 vs Claude Sonnet 5
Sonnet 5 lists at half the price of Opus 5.5, $2/$10, but it thinks so much that it costs more per task. Its everyday setting (xhigh) scores 34.4 at $2.87 per AA task, against 53.6 at $1.82 for Opus 5.5 at high. Per 1,000 coding jobs it's about $350 against $296.
Sonnet 5 answers in 23.9 seconds, but Opus 5.5 at medium answers in 24.9 seconds with a score of 51.2 for $1.34 per task. On this data Sonnet 5 has no job left that Opus 5.5 doesn't do better for less. If you're on Sonnet 5 to save money, move to Opus 5.5 at medium.
Opus 5.5 vs Gemini 3.8 Flash
Gemini 3.8 Flash costs $0.75/$3.75 and scores 41 on intelligence and 42 on coding. It thinks heavily, so it costs $1.24 per AA task and about $148 per 1,000 coding jobs. Opus 5.5 at low effort scores 42.3 for $0.55 per AA task and answers in 12.4 seconds, half Flash's 24.1.
Flash's real edge is throughput. At 297 tokens per second it streams long outputs about three times faster than Opus 5.5. It also has the highest slop score of this group (22.6), so expect more editing on anything readers will see.
Which best-model-for pages pick Opus 5.5
At everyday settings, Opus 5.5 is the overall pick on 8 of our 9 best-model-for pages. It isn't the value or fastest pick on any of them, and it isn't picked for creative writing.
| Page | Overall pick | Value pick | Fastest pick |
|---|---|---|---|
| Coding | Claude Opus 5.5 | MiMo-V2.6-Pro | DeepSeek V4.1 Flash |
| Agents | Claude Opus 5.5 | MiMo-V2.6-Pro | DeepSeek V4.1 Flash |
| Research | Claude Opus 5.5 | MiMo-V2.6-Pro | Grok 4.7 |
| Math | Claude Opus 5.5 | DeepSeek V4.1 Flash | Grok 4.7 |
| Documents | Claude Opus 5.5 | MiMo-V2.6-Pro | DeepSeek V4.1 Flash |
| Image analysis | Claude Opus 5.5 | MiMo-V2.6-Pro | DeepSeek V4.1 Flash |
| OCR | Claude Opus 5.5 | MiMo-V2.6-Pro | DeepSeek V4.1 Flash |
| Claude Opus 5.5 | MiMo-V2.6-Pro | DeepSeek V4.1 Flash | |
| Creative writing | GPT-6 Astra | GPT-6 Sol | Grok 4.7 |
The value picks show how much quality you give up to save money. MiMo-V2.6-Pro scores 46.3 at about $20 per 1,000 coding jobs against Opus 5.5's $296, but takes about a minute per answer. DeepSeek V4.1 Flash scores 39.5 and answers in 11.6 seconds. The overall ranking across every task is on the LLM leaderboard.
When Opus 5.5 is worth it, and when a cheaper model wins
- 53.6 on the Intelligence Index
- 66 on the Coding Agent Index
- Full answer in about 41 s
- 44.1 at its everyday setting
- 57 on the Coding Agent Index
- Full answer in about 43 s
- 33.9 at its everyday setting
- 41 on the Coding Agent Index
- Full answer in about 23 s
Opus 5.5 is worth paying for when a wrong answer costs more than the tokens: hard bugs and multi-file changes in Claude Code, agent runs nobody watches, research and document work you'll act on. It's also the cheaper choice if you're on Opus 5, Fable 5.1 or Sonnet 5 today. All three cost more per task and score lower.
Another model wins in four cases. Daily coding where you read every diff runs fine on GPT-6 Sol. Sub-agents and bulk edits belong on GPT-6 Luna or MiMo-V2.6-Pro. Fiction reads better from GPT-6 Astra or Sol. And if a person is waiting on each reply, Astra at medium answers in 15 seconds, or run Opus 5.5 at low.
The one setting to avoid is max as a default. It triples the cost of high for 4 points and gives up any chance of a quick reply.
How to use Claude Opus 5.5: API, Claude Code and availability
On the Claude API the model ID is claude-opus-5-5. The same ID works on Google Cloud, Microsoft Foundry and Claude Platform on AWS. On Amazon Bedrock it's anthropic.claude-opus-5-5. Set effort with output_config, as in Anthropic's effort docs:
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-opus-5-5",
max_tokens=32000,
output_config={"effort": "high"},
messages=[{"role": "user", "content": "Find the race condition in this worker pool: ..."}],
)
Anthropic recommends a large max_tokens at higher effort levels, because the limit covers thinking and the answer together. Four changes break code written for Opus 5:
thinking: {"type": "disabled"}and manualbudget_tokensreturn a 400 error. Omitthinkingor send{"type": "adaptive"}.tool_choiceofanyor a named tool returns a 400 error. Useautowith strict tool use or structured outputs.- Thinking blocks are tied to the model that wrote them. Opus 5.5 reads Opus 5's blocks, but blocks from Fable or Mythos models are dropped.
- On the Claude API and Google Cloud, the older
computer_20251124computer use tool is rejected. Move tocomputer_toolset_20260801.
The text Opus 5.5 writes between tool calls also now arrives in thinking blocks, which are empty at the default display setting. If your app shows progress messages, see the migration guide.
In Claude Code, Opus 5.5 needs version 2.1.280 or later. Per the Claude Code model docs, it's the default model on Pro, Max, Team and Enterprise plans and with an Anthropic API key, and the opus alias points to it on the Anthropic API, Bedrock, Google Cloud and Claude Platform on AWS. On Microsoft Foundry the alias still points to Opus 4.6. Claude Code starts Opus 5.5 at medium effort. Run /effort high, or start with claude --effort high, for work that matters.
Anthropic commits to keeping claude-opus-5-5 available until at least 22 September 2027. If you're still on Opus 5, our Claude Opus 5 guide covers what it did well and how it compares.
FAQ
How much does Claude Opus 5.5 cost?
Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, with cache reads at $0.20 and batch requests at half price. What you pay per task depends on effort: Artificial Analysis measured $0.55 per benchmark task at low, $1.34 at the default medium, $1.82 at high and $5.98 at max.
Is Claude Opus 5.5 better than Opus 5?
Yes, on almost every measure. Opus 5.5 scores 58 on the Artificial Analysis Intelligence Index against 51 for Opus 5, costs 20% less per token, and writes 90 tokens per second against 54. At high effort it outscores Opus 5 at max for less than a third of the cost per task. Creative writing is the exception, where Opus 5 still scores higher.
What effort level should I use with Opus 5.5?
Use high for coding, analysis and anything you'll act on. It scores 53.6, answers in about 41 seconds and costs $1.82 per Artificial Analysis task. The default, medium, scores 51.2 in about 25 seconds and suits chat. Low works for sub-tasks. Skip max unless nobody is waiting: it costs 3.3 times as much as high for 4 more points.
What is the Claude Opus 5.5 context window?
Opus 5.5 has a 1M-token context window, which Anthropic puts at roughly 555,000 words on its current tokenizer. It can write up to 128K tokens per response on the Messages API, or 300K on the Batch API with a beta header. It accepts text and images and returns text. Its knowledge cutoff is June 2026.
Is Claude Opus 5.5 good for creative writing?
It's the weakest part of the model. Opus 5.5 scores 2050 Elo on EQ-Bench Creative Writing v3, seventh among the models we track, with a slop score of 10.43. Opus 5 scores 2133 with 6.59 slop. GPT-6 Astra leads at 2173, and GPT-6 Sol scores 2125 at half Opus 5.5's per-token price.
How do I use Claude Opus 5.5 in Claude Code?
Update Claude Code to version 2.1.280 or later. Opus 5.5 is then the default model on Pro, Max, Team and Enterprise plans and with an Anthropic API key, and the opus alias selects it. It starts at medium effort. Type /effort high in a session, or launch with claude --effort high, to raise it.
Intelligence Index v4.3.2, Coding Agent Index v1.5, per-effort costs, response times and speeds from Artificial Analysis, pulled 25 September 2026. Cost per 1,000 coding jobs is Dervity's estimate for a 45,000-token job, adjusted for each setting's thinking. Creative writing scores from EQ-Bench Creative Writing v3. Release date, prices, limits, benchmarks and API details from Anthropic's announcement, the Opus 5.5 model page, what's new in Opus 5.5 and the Claude Code model docs. Published September 26, 2026.


