GPT-6 Astra vs Sol vs Luna: Differences, Pricing, Benchmarks
GPT-6 Astra, Sol and Luna compared: $10, $2 and $0.10 per 1M input tokens, 53, 48 and 37 on the Intelligence Index, and the real cost per task at each effort.
· 19 min read

GPT-6 Astra is OpenAI's top model at $10/$50 per million tokens and scores 53 on the Artificial Analysis Intelligence Index. GPT-6 Sol costs a fifth as much ($2/$10) and scores 48. GPT-6 Luna costs a twentieth of Sol ($0.10/$0.50) and scores 37. At the settings people run day to day, that works out to about $1.54, $0.53 and $0.04 per benchmark task.
For most work, use Sol. Astra is worth its price on the hardest problems and on long-form writing, and Luna is for high-volume jobs where a 10-point lower score doesn't matter.
Scores come from Artificial Analysis (AA): Intelligence Index v4.3.2 and Coding Agent Index v1.5. Per-effort costs and response times are AA measurements from September 25, 2026. We rate each model at its "everyday" setting, meaning the strongest reasoning effort that returns a full answer within 45 seconds, and show max effort next to it.
- 49.6 on the Intelligence Index at medium effort, 52.7 at max
- 62 on the Coding Agent Index
- First for creative writing on EQ-Bench
- 44.1 at extra high effort, 47.5 at max
- 57 on the Coding Agent Index
- $0.53 per task at its everyday setting
- 33.9 at extra high effort, 37.3 at max
- $0.04 per task, about 13x cheaper than Sol
- Free and Go ChatGPT users get it in the desktop app
What's the difference between GPT-6 Astra, Sol and Luna?
The three tiers share a 1.05M-token context window and 128K max output, and all take text and images. They differ in price, score and how they spend reasoning effort. OpenAI's docs pitch Sol at complex coding and agent workflows and Luna at focused, high-volume tasks. Both were trained with methods similar to Astra's, according to OpenAI's announcement.
| GPT-6 Astra | GPT-6 Sol | GPT-6 Luna | |
|---|---|---|---|
| Released | Sep 3, 2026 | Sep 22, 2026 | Sep 22, 2026 |
| API name | gpt-6-astra | gpt-6-sol | gpt-6-luna |
| Input / output per 1M tokens | $10 / $50 | $2 / $10 | $0.10 / $0.50 |
| Cached input per 1M | $1.00 | $0.20 | $0.01 |
| Context window | 1.05M (922K input) | 1.05M (922K input) | 1.05M (922K input) |
| Max output | 128K | 128K | 128K |
| Reasoning efforts | low to max | none to max | none to max |
| Intelligence Index, everyday | 49.6 (medium) | 44.1 (extra high) | 33.9 (extra high) |
| Intelligence Index, max | 52.7 | 47.5 | 37.3 |
| Coding Agent Index v1.5 | 62 | 57 | 41 |
| Cost per task, everyday | $1.54 | $0.53 | $0.04 |
| Full answer, everyday | 15 s | 43 s | 23 s |
| Knowledge cutoff | Apr 30, 2026 | Apr 20, 2026 | May 18, 2026 |
Astra costs five times Sol per token but only 2.9 times as much per task, because it reaches its best everyday score at medium effort while Sol needs extra high. Luna's tokens cost a twentieth of Sol's, and its per-task cost is about a thirteenth, since it thinks longer to reach its score.
There is no GPT-6 Terra. OpenAI's September 22 announcement covered Sol and Luna only, and GPT-5.6 Terra is still listed in the API docs at $2/$12.
GPT-6 benchmarks: Astra leads the family by about five points
At everyday effort, Astra scores 49.6, Sol 44.1 and Luna 33.9. At max effort the gaps stay about the same: 52.7, 47.5 and 37.3. On the Coding Agent Index, which runs each model inside its maker's coding tool (Codex, for OpenAI), the order holds at 62, 57 and 41.
AA's write-ups on Astra and on Sol and Luna break out a few of the evaluations behind those numbers:
| AA evaluation | GPT-6 Astra | GPT-6 Sol | GPT-6 Luna |
|---|---|---|---|
| Terminal-Bench 4.0 | 59% | 43% | 13% |
| AutomationBench-AA | 69% | 62% | 53% |
| Hallucination rate, AA-Omniscience (lower is better) | 51% | 60% | 77% |
The Terminal-Bench row is the one to remember. Luna completes 13% of those agentic terminal tasks against Sol's 43%, so it's the wrong model to run as a coding agent even though it writes decent code in a chat. Astra's 51% hallucination rate is the lowest of the three, and AA's write-up has it leading Terminal-Bench 4.0 and AutomationBench-AA.
What each reasoning effort costs and how long it takes
Reasoning effort changes the bill more than the tier does. Each table below lists every setting AA tested, with the score, the cost of one Intelligence Index task including thinking tokens, the median time to a full answer, and the wait before the first answer token.
GPT-6 Astra by effort
| Effort | Intelligence Index | Cost per task | Full answer | First token |
|---|---|---|---|---|
| Low | 45.8 | $0.82 | 12.5 s | 3.2 s |
| Medium (everyday) | 49.6 | $1.54 | 15.0 s | 5.8 s |
| High | 50.9 | $1.73 | 51.8 s | 42.8 s |
| Extra high | 52.4 | $2.31 | 154.5 s | 145.5 s |
| Max | 52.7 | $3.26 | 370.0 s | 361.3 s |
Run Astra at medium. High adds 1.3 points for 12% more money but more than triples the wait. Max doubles medium's cost for 3.1 points and takes about six minutes per answer, so keep it for background jobs. Low is the surprise: 45.8 in 12.5 seconds beats Sol's everyday score, and answers more than three times faster, for $0.82 against $0.53.
GPT-6 Sol by effort
| Effort | Intelligence Index | Cost per task | Full answer | First token |
|---|---|---|---|---|
| None | 28.1 | $0.33 | 6.9 s | 1.1 s |
| Low | 33.9 | $0.13 | 7.8 s | 2.0 s |
| Medium | 39.8 | $0.25 | 6.3 s | 2.0 s |
| High | 42.8 | $0.37 | 15.4 s | 9.6 s |
| Extra high (everyday) | 44.1 | $0.53 | 42.6 s | 37.0 s |
| Max | 47.5 | $1.06 | 170.8 s | 165.1 s |
Skip none. In AA's runs it cost more per task than low effort and scored 5.8 points less. Medium, the API default, was the fastest setting of all at 6.3 seconds. For anything a person waits on, we'd pick high: 42.8 in 15 seconds for $0.37. Extra high adds 1.3 points for 42% more cost and nearly three times the wait, and max doubles the cost again for another 3.4 points.
GPT-6 Luna by effort
| Effort | Intelligence Index | Cost per task | Full answer | First token |
|---|---|---|---|---|
| None | 18.3 | $0.011 | 4.9 s | 0.8 s |
| Low | 20.9 | $0.0045 | 6.3 s | 2.3 s |
| Medium | 29.5 | $0.017 | 8.8 s | 5.3 s |
| High | 32.1 | $0.029 | 11.2 s | 7.5 s |
| Extra high (everyday) | 33.9 | $0.042 | 22.9 s | 19.2 s |
| Max | 37.3 | $0.068 | 142.4 s | 139.0 s |
Luna is the one model where max effort makes sense for everyday batch work. It adds 3.4 points for 1.6 times the cost, which is $68 per 1,000 tasks. The catch is time. Each answer takes 142 seconds, so use it only where nobody is waiting. Low effort costs under half a cent per task, but at 20.9 it's too weak for anything beyond routing and tagging.
How many output tokens GPT-6 uses per task
AA counts output tokens per Intelligence Index task at max effort. Astra uses about 27K, a third of Claude Fable 5.1's 78K. Sol uses 31K (GPT-5.6 Sol used 29K) and Luna 51K (GPT-5.6 Luna used 41K). AA puts the lower cost per task down to the price cut. Sol and Luna didn't get more efficient. They write slightly more than the models they replace.
GPT-6 vs Claude Opus 5.5, Claude Fable 5.1 and Gemini 3.8 Flash
Claude Opus 5.5 beats every GPT-6 tier on both indexes. Astra's advantage over it is speed, and Sol's advantage over everything near its score is cost.
| Model | Price per 1M (in / out) | Everyday score | Max score | Coding Agent Index | Cost per task, everyday | Full answer |
|---|---|---|---|---|---|---|
| Claude Opus 5.5 | $4 / $20 | 53.6 (high) | 57.6 | 66 | $1.82 | 41 s |
| Claude Fable 5.1 | $10 / $50 | 51.2 (high) | 53.4 | 62 | $3.91 | 29 s |
| GPT-6 Astra | $10 / $50 | 49.6 (medium) | 52.7 | 62 | $1.54 | 15 s |
| GPT-6 Sol | $2 / $10 | 44.1 (extra high) | 47.5 | 57 | $0.53 | 43 s |
| Gemini 3.8 Flash | $0.75 / $3.75 | 40.9 (high) | 40.9 | 42 | $1.24 | 24 s |
| GPT-6 Luna | $0.10 / $0.50 | 33.9 (extra high) | 37.3 | 41 | $0.04 | 23 s |
Per-task cost doesn't follow list price. Opus 5.5 lists at 40% of Astra's price per token yet costs 18% more per task, because it thinks longer. Gemini 3.8 Flash lists far below Sol and still costs 2.3 times as much per task.
- Claude Fable 5.1$3.91
- Claude Opus 5.5$1.82
- GPT-6 Astra$1.54
- Gemini 3.8 Flash$1.24
- GPT-6 Sol$0.53
- GPT-6 Luna$0.04
Source: Artificial Analysis, September 25, 2026. Everyday means each model's strongest setting that answers within 45 seconds.
GPT-6 Astra vs Claude Opus 5.5
Opus 5.5 is the stronger model: 4 points ahead at everyday effort, 4.9 ahead at max, and 66 against 62 on the Coding Agent Index. It costs $1.82 per task against Astra's $1.54 and takes 41 seconds per answer against 15. Pick Opus 5.5 when quality decides, and Astra when you need fast answers inside ChatGPT or Codex, or when the job is writing.
GPT-6 Astra vs Claude Fable 5.1
At max effort, AA scores them level (53.4 for Fable 5.1, 52.7 for Astra) at the same $10/$50 list price. Fable 5.1 costs $7.63 per task at max against Astra's $3.26, and at everyday effort it costs 2.5 times as much for 1.6 more points. Astra is the better buy unless you already run Claude Code.
GPT-6 Sol and Luna vs Gemini 3.8 Flash
Sol at high effort (42.8, $0.37, 15 seconds) beats Gemini 3.8 Flash's best setting (40.9, $1.24, 24 seconds) on score, cost and speed. Gemini streams faster once it starts, at about 330 tokens per second, and it writes with far more stock phrasing (see below). Against Luna, Gemini scores 7 points more at everyday effort for about 30 times the cost per task.
Creative writing: Astra ranks first on EQ-Bench
GPT-6 Astra has the highest Elo on EQ-Bench Creative Writing v3, at 2173. Sol comes in at 2125, above Claude Opus 5.5.
| Model | EQ-Bench Elo | Slop score (lower is better) |
|---|---|---|
| GPT-6 Astra | 2173 | 8.41 |
| Claude Fable 5.1 | 2162 | 8.16 |
| GPT-6 Sol | 2125 | 8.97 |
| Claude Opus 5.5 | 2050 | 10.43 |
| GPT-5.6 Sol | 1972 | 11.68 |
| GPT-5.6 Luna | 1829 | 11.80 |
| Gemini 3.8 Flash | 1748 | 22.60 |
The slop score counts how often a model leans on stock AI phrases. Both GPT-6 models score under 9, so their drafts need less cleanup than Opus 5.5's and far less than Gemini 3.8 Flash's. EQ-Bench hasn't tested GPT-6 Luna yet, so there's no creative score for it.
Should you upgrade from GPT-5.6 Sol and Luna?
Yes. GPT-6 Sol matches GPT-5.6 Sol's score for 55% less per task, hallucinates less and writes better. GPT-6 Luna keeps the same max score at 62% less per task, with one coding caveat.
| GPT-5.6 Sol | GPT-6 Sol | GPT-5.6 Luna | GPT-6 Luna | |
|---|---|---|---|---|
| Price per 1M (in / out) | $4 / $20 | $2 / $10 | $0.20 / $1.20 | $0.10 / $0.50 |
| Everyday score | 44.0 (extra high) | 44.1 (extra high) | 32.1 (high) | 33.9 (extra high) |
| Cost per task, everyday | $1.18 | $0.53 | $0.044 | $0.042 |
| Max score | 47.0 | 47.5 | 37.3 | 37.3 |
| Cost per task, max | $1.99 | $1.06 | $0.18 | $0.068 |
| Full answer, max | 112 s | 171 s | 154 s | 142 s |
| Coding Agent Index v1.5 | 55 | 57 | 43 | 41 |
| Hallucination rate | 92% | 60% | 93% | 77% |
| EQ-Bench Elo | 1972 | 2125 | 1829 | Not tested |
Two regressions show up. GPT-6 Luna dropped 2 points on the Coding Agent Index and fell from 49% to 44% on SWE-Atlas-QnA, per AA. GPT-6 Sol at max takes 59 seconds longer per answer than GPT-5.6 Sol did. Neither changes the verdict. Apart from GPT-5.6 Sol's quicker max setting, the old tiers have nothing left to offer, and OpenAI lists GPT-5.6 Sol's $4/$20 as promotional pricing that runs at least through November 21, 2026.
If you're on GPT-5.6 Terra, move to GPT-6 Sol. AA's current data has Terra's best score at 42.1 (max, $1.40 per task, 171 seconds). Sol beats that at extra high effort: 44.1 for $0.53 in 43 seconds. Our GPT-5.6 Sol vs Terra vs Luna guide has the full history of the old tiers.
Which GPT-6 model to use for coding, agents, writing and bulk work
Our best-for pages rank every current model at everyday effort. GPT-6 wins one of them outright: Astra is the top pick for creative writing, with Sol as that page's value pick. Claude Opus 5.5 takes the top spot on every other page, and MiMo-V2.6-Pro takes most value picks. Inside the GPT-6 family, here's how we'd split the work.
- Hard reasoning and code review5.5 points above Sol at everyday effortAstra
- Daily coding in Codex57 on the Coding Agent Index at $2.99 per coding taskSol
- Long agent runs69% on AutomationBench-AAAstra
- Sub-agents and tool callsUnder $5 per 1,000 agent jobsLuna
- Creative writingAstra for quality, Sol for valueAstraSol
- Bulk extraction and tagging$0.62 per 1,000 emailsLuna
- Categories wonAstra 3Sol 2Luna 2
| Category | Astra | Sol | Luna |
|---|---|---|---|
| Hard reasoning and code review5.5 points above Sol at everyday effort | Wins | ||
| Daily coding in Codex57 on the Coding Agent Index at $2.99 per coding task | Wins | ||
| Long agent runs69% on AutomationBench-AA | Wins | ||
| Sub-agents and tool callsUnder $5 per 1,000 agent jobs | Wins | ||
| Creative writingAstra for quality, Sol for value | Wins | Wins | |
| Bulk extraction and tagging$0.62 per 1,000 emails | Wins | ||
| Categories won | 3 | 2 | 2 |
Coding: Sol by default, Astra for the hard parts
The best model for coding page picks Claude Opus 5.5 overall, MiMo-V2.6-Pro for value and DeepSeek V4.1 Flash for speed. No GPT-6 model makes the picks. Astra ranks third among current models at about $429 per 1,000 coding jobs, and Sol ninth at about $109. If you work in Codex, run Sol at high or extra high and switch to Astra for the problems Sol gets wrong. AA puts Sol at $2.99 per Coding Agent task against Astra's $7.09. Keep Luna away from coding agents: its 13% on Terminal-Bench 4.0 says it can't finish multi-step terminal work.
Agents: Astra to plan, Luna to execute
The best model for agents page makes the same three picks as coding. Within GPT-6, Astra's 69% on AutomationBench-AA makes it the planner for long runs, and Luna costs about $4.66 per 1,000 agent jobs at everyday effort against $72.49 for Sol and $285.69 for Astra. A setup that sends only planning and review to Astra and everything else to Luna spends most of its budget where the score gap matters.
Writing: Astra, or Sol to save 75%
The best model for creative writing page picks GPT-6 Astra overall at about $95 per 1,000 pieces, and GPT-6 Sol for value at about $24. Sol is 49 Elo points behind Astra and still ahead of Claude Opus 5.5. For drafts you'll edit anyway, Sol is enough.
Cheap bulk work: Luna, if the output is easy to check
On the best model for email page, Luna costs $0.62 per 1,000 emails at everyday effort, the lowest of any paid model there. It isn't a pick, though. The page only considers models within 70% of the leader's score, which puts the bar at 38, and Luna's everyday 33.9 falls short. MiMo-V2.6-Pro is the value pick at 46.3 for $1.82 per 1,000 emails, but it takes 61 seconds per answer. Use Luna for tagging, routing and extraction you can validate in code, and MiMo-V2.6-Pro when quality matters and nobody waits.
Where GPT-6 falls short
Astra at max effort is a six-minute model. It gains 3.1 points over medium for 2.1 times the cost and about 25 times the wait, which rules max out for anything interactive.
Opus 5.5 beats Astra on quality for 18% more per task. OpenAI's flagship isn't the strongest model you can buy, and its edge over Opus 5.5 comes down to speed, writing and a slightly lower bill.
Luna's everyday score of 33.9 sits below the quality bar on every best-for page that ranks it. It's a cheap worker, not a general assistant.
Coding Agent Index scores mix the model with its tool. GPT-6 runs in Codex and Claude runs in Claude Code, so a gap of a few points can come from the harness as much as from the model. You can compare the raw numbers on our benchmark dashboard.
FAQ
What is the difference between GPT-6 Sol and GPT-6 Luna?
Sol is the mid tier at $2/$10 per million tokens and scores 44.1 on the Artificial Analysis Intelligence Index at everyday effort. Luna costs $0.10/$0.50 and scores 33.9. Sol costs about $0.53 per benchmark task and Luna about $0.04. Sol handles coding and agent work. Luna suits high-volume extraction, tagging and sub-agent calls where a lower score is acceptable.
Is there a GPT-6 Terra?
No. OpenAI released GPT-6 Astra on September 3, 2026 and GPT-6 Sol and Luna on September 22. The Sol and Luna announcement didn't include a Terra tier. GPT-5.6 Terra is still listed in OpenAI's API docs at $2/$12 per million tokens, but GPT-6 Sol scores higher (44.1 against Terra's 38.0 at extra high effort) and costs less per task.
How much does GPT-6 cost per task?
At everyday effort, Artificial Analysis measured $1.54 per Intelligence Index task for GPT-6 Astra (medium), $0.53 for Sol (extra high) and $0.04 for Luna (extra high), thinking tokens included. At max effort the costs rise to $3.26, $1.06 and $0.068. Max also slows answers to between 142 and 370 seconds.
Is GPT-6 Astra better than Claude Opus 5.5?
No. Claude Opus 5.5 scores 53.6 on the Artificial Analysis Intelligence Index at everyday effort against Astra's 49.6, and 66 against 62 on the Coding Agent Index. Astra is faster, answering in about 15 seconds against 41, and costs 15% less per task. It also ranks higher for creative writing on EQ-Bench.
How many output tokens does GPT-6 use per task?
At max effort, Artificial Analysis counts about 27,000 output tokens per Intelligence Index task for GPT-6 Astra, 31,000 for Sol and 51,000 for Luna. The GPT-5.6 versions used 29,000 (Sol) and 41,000 (Luna), so the GPT-6 savings come from the lower prices. Astra uses about a third of Claude Fable 5.1's 78,000.
Should I upgrade from GPT-5.6 to GPT-6?
Yes. GPT-6 Sol scores the same as GPT-5.6 Sol at everyday effort for 55% less per task and cuts its hallucination rate from 92% to 60%. GPT-6 Luna matches GPT-5.6 Luna's max score for 62% less per task. The one regression is in coding, where GPT-6 Luna scores 41 on the Coding Agent Index, 2 points below GPT-5.6 Luna.
Scores from Artificial Analysis Intelligence Index v4.3.2 and Coding Agent Index v1.5, with per-effort costs and times measured on September 25, 2026. Evaluation details and token counts from AA's GPT-6 Astra and GPT-6 Sol and Luna articles. Prices, context windows and effort levels from OpenAI's model docs for Astra, Sol and Luna. Creative writing scores from EQ-Bench Creative Writing v3. For live rankings of every model, see the LLM Leaderboard.


