GPT-5.6 Sol vs Terra vs Luna: Which Tier Do You Need?
GPT-5.6 Sol vs Terra vs Luna: benchmarks, pricing, and which tier to pick. Updated for GPT-6 Sol and Luna, which match the 5.6 scores at half the price.
· Updated · 17 min read

Update (September 26, 2026): GPT-6 has replaced these models. OpenAI released GPT-6 Astra on September 3 and GPT-6 Sol and Luna on September 22, at half the GPT-5.6 prices, and there is no GPT-6 Terra. At everyday settings GPT-6 Sol scores 44.1 (47.5 at max) for about $0.53 per task, and GPT-6 Luna 33.9 for about $0.04. For current tiers, prices and per-effort costs, read our GPT-6 Astra vs Sol vs Luna guide. The GPT-5.6 guide below stays up for reference.
Most people will pick Sol because it's the flagship. That's a mistake for most everyday work. Luna scored 75 on the Coding Agent Index at launch and handles the majority of real work. After the July 30 price cut, Luna costs just $0.20/$1.20 per million tokens. That's an 80% drop from launch pricing.
Here's how all three stack up, what the benchmarks mean in practice, and which setup saves the most money without losing quality.
Quick verdict: Use Luna as your default. Switch to Sol for hard problems. Skip Terra.
Update (July 30, 2026): OpenAI announced major price cuts. Luna dropped 80% to $0.20/$1.20. Terra dropped 20% to $2/$12. Sol gets a new Fast mode at 2.5x speed for 2x the price. All pricing below reflects the new rates.
Update (September 24, 2026): GPT-5.6 has been superseded. OpenAI released GPT-6 Astra on September 3 (53 on the Artificial Analysis Intelligence Index at $10/$50, now OpenAI's flagship), then GPT-6 Sol and GPT-6 Luna on September 22. GPT-6 Sol scores 48 against GPT-5.6 Sol's 47, and GPT-6 Luna matches GPT-5.6 Luna at 37. Both cost half as much: $2/$10 and $0.10/$0.50. If you're picking a tier today, pick the GPT-6 version of it. AA also rebased its Intelligence Index in September 2026 (now v4.3.2), so every score is lower than at launch. The tables below use v4.3.2 scores and current AA prices, which list GPT-5.6 Sol at $4/$20. Coding Agent Index figures are from the v1.1 index AA used in July. For live rankings, see the LLM Leaderboard.
The Three Tiers at a Glance
| GPT-5.6 Sol | GPT-5.6 Terra | GPT-5.6 Luna | |
|---|---|---|---|
| Intelligence Index (v4.3.2) | 47 | 42 | 37 |
| Coding Agent Index (v1.1, July) | 80 | 77 | 75 |
| Input Price | $4.00/1M | $2.00/1M | $0.20/1M |
| Output Price | $20.00/1M | $12.00/1M | $1.20/1M |
| Speed | 73 tok/s | 83 tok/s | 127 tok/s |
| Context | 1.05M | 1.05M | 1.05M |
| Max Output | 128K | 128K | 128K |
| Reasoning | Yes (incl. max, ultra, fast) | Yes (incl. max) | Yes (incl. max) |
| Vision | Yes | Yes | Yes |
OpenAI launched GPT-5.6 on July 9, 2026 with these three tiers sharing the same architecture but differing in size, speed, and price. All three have 1.05M context windows and 128K max output. On July 30, OpenAI slashed pricing: Luna dropped 80% and Terra dropped 20%. Sol also got a Fast mode (2.5x speed, 2x price).
Luna to Sol is 10 points on the Intelligence Index. That gap matters for frontier problems. For emails, simple coding, summarization, and content drafts, Luna handles those fine. At $0.20/$1.20, Luna was one of the cheapest frontier-capable models available in July.
An important finding from the Artificial Analysis Intelligence Index v4.1: Luna and Sol always sit on the cost-performance Pareto frontier, ahead of Terra. After the July price cuts, Luna's position on the Pareto frontier got even stronger. AA's current data still shows it: GPT-5.6 Luna at xhigh scores 34.6 for $0.09 per task, where Terra at high scores 34.2 for $0.34.
Max and Ultra Reasoning
GPT-5.6 adds two reasoning effort levels beyond the existing low/medium/high/xhigh.
Max works like existing levels but goes deeper. More thinking tokens, longer processing, better output on hard problems. All benchmark scores above were measured at max. Every tier supports it.
Ultra is something different. It coordinates four agents in parallel, each working on separate aspects of the problem. This trades higher token usage for stronger results and faster time-to-result on demanding tasks. Built-in agentic parallelism without needing an external orchestrator.
OpenAI's GPT-5.6 announcement says ultra runs four agents by default, and its benchmark charts also show 16-agent runs. Ultra is a ChatGPT and Codex mode. In the API, reasoning.effort goes from none to max for all three tiers, per OpenAI's model docs.
Ultra suits tasks that split into parallel parts, such as large codebase analysis, multi-file refactoring and research synthesis. For simple sequential tasks, max or even high effort costs less.
Sol Led the Coding Agent Index at Launch
The Artificial Analysis Coding Agent Index v1.1 pairs models with their native coding tools and evaluates them on three benchmarks: DeepSWE, Terminal-Bench v2, and SWE-Atlas-QnA. Unlike lab-reported benchmarks, this tests models inside the actual coding environment they ship with.
Sol (max) in OpenAI's Codex scored 80 and led every evaluation, tying Grok 4.5 in Grok Build on SWE-Atlas-QnA, according to Artificial Analysis.
| Model + Tool | Coding Agent Index (v1.1) |
|---|---|
| GPT-5.6 Sol (max) in Codex | 80 |
| GPT-5.6 Terra (max) in Codex | 77 |
| Claude Fable 5 in Claude Code | 77 |
| GPT-5.5 in Codex | 76 |
| GPT-5.6 Luna (max) in Codex | 75 |
AA put Sol's cost per coding task about 40% below Fable 5 (max) and about 10% below Opus 4.8 (max), both in Claude Code.
Luna's Coding Agent Index score of 75 put it one point behind GPT-5.5 (76) and two behind Fable 5 (77). At $0.20/$1.20, it was by far the cheapest path to near-frontier coding performance.
That table is from AA's Coding Agent Index v1.1 in July. AA now runs v1.5 (DeepSWE v1.1, Terminal-Bench 4.0, SWE-Atlas-QnA), where Claude Opus 5.5 leads at 66, GPT-6 Astra scores 62, GPT-6 Sol 57, Grok 4.7 56 (ahead of GPT-5.6 Sol), and GPT-6 Luna 41. See the full best AI model for coding rankings.
GPT-5.6 Sol: When You Need the Best
Sol scores 47 on the Intelligence Index (v4.3.2), behind its July rivals Claude Opus 5 (51) and Claude Fable 5 (50). AA now lists it at $4/$20, less than half of Fable's $10/$50 and below Opus 5's $5/$25. On July 30, OpenAI added a Fast mode to Sol's API: 2.5x the speed for 2x the launch price ($10/$60). Same intelligence, faster responses.
In the API, Sol supports reasoning levels from none through max, and ultra runs in ChatGPT and Codex. At none or low, it behaves like a fast production model. At max or ultra, it spends significant time and tokens on complex problems. One model serving multiple roles.
Sol also has the highest Presentation Elo score in AA's new AA-Briefcase benchmark, which tests models on realistic knowledge work tasks. PowerPoint slides, Excel models, document formatting. The outputs look better than any other model tested.
On token efficiency, AA measured Sol (max) at about 15K output tokens per Intelligence Index task at launch, fewer than Claude Opus 4.8 (max) and Gemini 3.5 Flash (high). On the rebased index, AA now counts about 29K output tokens per task for GPT-5.6 Sol (max) and 41K for GPT-5.6 Luna (max), against 31K and 51K for GPT-6 Sol and Luna. AA's current cost per task at max is $1.99 for GPT-5.6 Sol and $0.18 for GPT-5.6 Luna.
Use Sol for: Architecture decisions. Complex multi-file debugging. Security-sensitive code review. Problems that stump cheaper models. Knowledge work that needs polished output.
GPT-5.6 Terra: The Middle Ground
Terra scores 42 on the Intelligence Index (v4.3.2) at $2/$12 (down from $2.50/$15 after the July 30 cut). OpenAI's docs say it roughly corresponds to the mini tier of earlier GPT-5 families, and it's now cheaper than it was at launch. AA measures it at 83 tok/s, comfortable for interactive use.
Here's the catch. The Artificial Analysis Intelligence Index v4.1 data shows Terra never sits on the Pareto frontier. For any given Terra effort level, you can find a Luna or Sol effort level that gives you either better intelligence at the same price, or the same intelligence cheaper. AA's current numbers bear it out: GPT-5.6 Sol at high scores 42.3 for $0.81 per task, while Terra needs max effort to reach 42.1 and costs $1.40. Luna handles the easy work, Sol the hard work, and Terra sits in between as the best pick for neither.
Use Terra for: When you want one model for everything and don't want to manage routing logic. Daily coding. Feature implementation. Data analysis.
GPT-5.6 Luna: The One Most People Should Default To
Luna is the real story of this release, and the July 30 price cut made it even more so. At $0.20/$1.20 per million tokens (down from $1/$6), it scores 37 on the Intelligence Index (v4.3.2) and scored 75 on the July Coding Agent Index, one point behind GPT-5.5 and two behind Fable 5.
Speed matters here too. Luna runs at 127 tok/s. For agent sub-tasks where you need quick responses to many small calls, that speed compounds into meaningful time savings.
It also has the same 1.05M context window as Sol. Luna was the obvious default for any workflow where you don't need frontier intelligence. GPT-6 Luna now does the same job for half the price ($0.10/$0.50, also 37 on the index), so start there for new work.
Use Luna for: Agent sub-tasks. Email drafts. Summarization. Code explanation. Simple edits. High-volume workflows where speed and cost matter more than peak quality.
Cache-Write Pricing
GPT-5.6 introduces cache-write pricing to OpenAI for the first time, matching what Anthropic already does with Claude.
- Cache writes: 1.25x the input token price (tokens being committed to memory)
- Cache reads: 90% discount (same as before)
The first time you send a long system prompt, you pay a 25% premium on those input tokens. Every subsequent request that hits the cache pays only 10% of the input price. For production workflows with consistent system prompts, this is a net positive. You pay slightly more upfront, then save heavily on every repeat call.
Real-World Performance
Aggregate indexes hide where each model wins, and the independent tests outside them split both ways.
On AA-Briefcase, AA's test of realistic knowledge work, Sol (max) has the best Presentation Elo of any model but ranks second overall to Claude Fable 5 (max). Fable's lead comes from its rubric score, 56% against Sol's 42%, per Artificial Analysis.
On ARC-AGI-3, which tests problems a model hasn't seen in training, GPT-5.6 Sol (max) set a record of 7.8% that Claude Opus 5 then nearly quadrupled at 30.2%, according to the ARC Prize. OpenAI says Sol reaches 38.3% with two API settings the official harness doesn't use.
The pattern is familiar. A lab cites the benchmarks it leads and questions the ones it trails, so check the evaluation that matches your work before trusting any single score. Our benchmark dashboard shows the current numbers side by side.
How They Compare to Everything Else
Scores are from the Artificial Analysis Intelligence Index v4.3.2. Prices and speeds are AA's current figures.
| Model | Provider | Intelligence | Input/Output | Speed |
|---|---|---|---|---|
| Claude Opus 5.5 | Anthropic | 58 | $4/$20 | 90 tok/s |
| GPT-6 Astra | OpenAI | 53 | $10/$50 | 51 tok/s |
| Claude Fable 5.1 | Anthropic | 53 | $10/$50 | 66 tok/s |
| Claude Opus 5 | Anthropic | 51 | $5/$25 | 54 tok/s |
| Claude Fable 5 | Anthropic | 50 | $10/$50 | 67 tok/s |
| GPT-6 Sol | OpenAI | 48 | $2/$10 | 110 tok/s |
| GPT-5.6 Sol | OpenAI | 47 | $4/$20 | 73 tok/s |
| Kimi K3 | Moonshot AI | 44 | $3/$15 | 37 tok/s |
| GPT-5.6 Terra | OpenAI | 42 | $2/$12 | 83 tok/s |
| Grok 4.5 | SpaceXAI (xAI) | 39 | $2/$6 | 62 tok/s |
| Claude Sonnet 5 | Anthropic | 38 | $2/$10 | 77 tok/s |
| GPT-5.6 Luna | OpenAI | 37 | $0.20/$1.20 | 127 tok/s |
| GPT-6 Luna | OpenAI | 37 | $0.10/$0.50 | 132 tok/s |
| Gemini 3.6 Flash | 34 | $0.75/$3.75 | 217 tok/s | |
| Gemini 3 Flash | 26 | $0.50/$3 | 210 tok/s |
Sol vs Opus 5: Opus 5 leads on intelligence (51 vs 47) and novel reasoning (30.2% vs 7.8% on ARC-AGI-3). Sol is now cheaper ($4/$20 vs $5/$25), led the July Coding Agent Index (80 in Codex), runs about 35% faster (73 vs 54 tok/s), and has the best presentation quality on AA-Briefcase. For coding in an IDE, Sol. For everything else, Opus 5 is the stronger pick. Both now have successors: Claude Opus 5.5 (58) and GPT-6 Sol (48).
Sol vs Fable 5: Sol costs well under half of what Fable charges ($4/$20 vs $10/$50) and runs faster. Fable leads by 3 points on intelligence (50 vs 47) and leads AA-Briefcase overall. For most work, Sol is the better value.
Luna vs everything at its price: After the July price cut, Luna at $0.20/$1.20 had no real competitor at its price point. That's no longer true. GPT-6 Luna scores the same 37 at $0.10/$0.50, GLM-5.3 Flash scores 42 at $0.15/$0.50, and MiMo-V2.6-Pro scores 46 at $0.43/$0.87 with open weights. For new agent sub-tasks and batch work, start with one of those.
Terra vs Sonnet 5: Terra scores higher on intelligence (42 vs 38). Sonnet 5 is cheaper on output ($10 vs $12), and AA still lists it at $2/$10.
Real-World Costs
Estimate your monthly spend with the AI Cost Calculator. Here are rough numbers for 500 tasks at 20,000 tokens average, split evenly between input and output, at current AA prices:
| Model | Monthly Cost |
|---|---|
| Claude Fable 5 | ~$300 |
| Claude Opus 5 | ~$150 |
| GPT-5.6 Sol | ~$120 |
| Claude Opus 5.5 | ~$120 |
| GPT-5.6 Terra | ~$70 |
| GPT-6 Sol | ~$60 |
| Claude Sonnet 5 | ~$60 |
| Grok 4.5 | ~$40 |
| Gemini 3 Flash | ~$17.50 |
| GPT-5.6 Luna | ~$7 |
| GPT-6 Luna | ~$3 |
Best Setup for Late July 2026
This was the launch-week setup. The same structure works better today with GPT-6: GPT-6 Sol ($2/$10) as the primary, GPT-6 Luna ($0.10/$0.50) for sub-agents, and Claude Opus 5.5 ($4/$20) or GPT-6 Astra ($10/$50) for the hardest problems.
A two-model setup using the GPT-5.6 family:
Primary model: GPT-5.6 Sol ($4/$20 on AA today) at medium or high effort. Handles complex tasks, code review, analysis, writing. Consider Claude Opus 5 ($5/$25) as an alternative primary if you value novel reasoning and self-correction over speed.
Sub-agent model: GPT-5.6 Luna ($0.20/$1.20). At high effort AA measured $0.044 per benchmark task, about 23 tasks per dollar. Luna handles delegated tasks, simple edits, lookups, and summarization.
Estimated monthly cost: between about $7 (all Luna) and $120 (all Sol) at the 500-task volume in the table above.
For tighter budgets, use Luna as both primary and sub-agent. At $0.20/$1.20 with a score of 37 on v4.3.2, it handles most work. Monthly cost drops to about $7 at 500 tasks. Or mix providers: Claude Opus 5 at high effort for reasoning-heavy work, Luna for everything else.
FAQ
Is GPT-5.6 Sol better than GPT-5.5?
Yes. Sol scores 47 on the Artificial Analysis Intelligence Index (v4.3.2) vs GPT-5.5's 38, and costs less ($4/$20 vs $5/$30). Sol also led the July Coding Agent Index at 80 points. There's no reason to use GPT-5.5 now, and GPT-6 Sol beats both at $2/$10.
Which GPT-5.6 model should I use for coding?
Start with Luna ($0.20/$1.20 after the July 30 cut). Its July Coding Agent Index score of 75 handles most tasks. Use Sol for complex debugging, architecture work, or when you need max/ultra reasoning. Terra at $2/$12 works if you want a single middle-ground model. For new projects, the GPT-6 versions are the better buy: GPT-6 Sol scores 57 on AA's Coding Agent Index v1.5 at $2/$10, and GPT-6 Luna scores 41 at $0.10/$0.50.
How does GPT-5.6 compare to Claude Opus 5 and Fable 5?
Claude Opus 5 (51 on the AA Intelligence Index v4.3.2) scores ahead of Fable 5 (50) and Sol (47). Opus 5 costs $5/$25, a bit more than Sol's current $4/$20, and runs slower at 54 tok/s vs Sol's 73. Sol leads on presentation quality in AA-Briefcase. Fable 5 at $10/$50 is the most expensive and hardest to justify unless you need its lead on AA-Briefcase knowledge work. All three have successors: Claude Opus 5.5 (58, now #1), Claude Fable 5.1 (53) and GPT-6 Sol (48).
What are max and ultra reasoning?
Max is deeper thinking within a single model, similar to existing effort levels but more thorough. Ultra, a ChatGPT and Codex mode, coordinates four agents in parallel by default to tackle complex problems. Ultra trades higher token usage for faster, stronger results on demanding tasks.
Is GPT-5.6 Luna good enough for production?
For many use cases, yes. Luna scores 37 on the Intelligence Index (v4.3.2) and scored 75 on the July Coding Agent Index, with full reasoning capabilities. After the 80% price cut to $0.20/$1.20 per million tokens, it handles email, summarization, simple coding, and agent sub-tasks. GPT-6 Luna now gives the same score for $0.10/$0.50, so pick that for new work.
Is GPT-5.6 available on OpenRouter?
GPT-5.6 Sol, Terra, and Luna are available on OpenRouter. Check their model page for current pricing and availability.
How does Grok 4.5 compare to GPT-5.6?
Grok 4.5 scores 39 on the Intelligence Index (v4.3.2) at $2/$6. It sits between Terra (42) and Luna (37) on quality. After the Luna price cut, Grok 4.5's output is 5x more expensive than Luna ($6 vs $1.20). On AA's cost per task, Grok 4.5 (high) spends $1.04 against GPT-5.6 Luna's $0.18 at max. Grok 4.5's successor, Grok 4.7, scores 46 at the same $2/$6.
Intelligence scores, prices and speeds from Artificial Analysis Intelligence Index v4.3.2 (September 2026). Coding Agent Index figures from v1.1 (July 2026) unless marked v1.5. Pricing from OpenAI's official API documentation and its July 30, 2026 price announcement. Updated September 26, 2026. For the models that replaced these, see our GPT-6 Astra vs Sol vs Luna guide. For Claude Opus 5 comparison, see our Opus 5 benchmarks guide. For full coding model rankings, see Best AI Models for Coding. For the overall model comparison beyond GPT-5.6, see Best LLM in 2026 and the LLM Leaderboard. Compare all models on our Benchmark Dashboard or get personalized picks with the Model Selector.


