GPT-6 Sol and Luna: Cheaper Than Astra, Not Better

A fifth of Astra's price for Sol, a hundredth for Luna. My rule for the Codex seat, written before the charts: switch only to a model that beats Astra at maximum settings and costs less. Sol costs less. On every chart OpenAI published, it doesn't beat Astra.

Three glowing machine cores on a wet steel floor at night: a blue star, an amber sun and a silver crescent moon, with a balance scale standing between the star and the sun

The rule came first

I wrote the rule down before reading a single chart. Codex runs GPT-6 Astra at ultra, where it has been since the September 5 pin in Astra Γ— Fable. A new model takes the seat only if, at its own maximum setting, it beats Astra at Astra's maximum, and it also costs less. Cheaper and almost as good fails.

The test has one snag. ultra exists only in Codex and ChatGPT Work: the Codex CLI catalog describes it as "Maximum reasoning with automatic task delegation", meaning maximum reasoning plus subagents the model starts on its own. OpenAI has published no benchmark at ultra for any GPT-6 model, and the API tops out at max. So this comparison is max against max, the closest public stand-in for ultra against ultra.

What shipped

gpt-6-sol and gpt-6-luna went live in the API on September 22, through Chat Completions, Responses and Batch. The same day they started rolling out in Codex and ChatGPT Work for Plus, Pro, Business, Enterprise and Edu. Free and Go get Luna in the desktop app. Neither model is in Chat yet. On paid plans, Sol is also the named replacement for GPT-5.5, which OpenAI retires from Codex on October 14.

OpenAI describes a ladder. Its model pages call Astra its most capable model, for the hardest end-to-end work; Sol a model built for complex coding and agentic workflows; Luna its most efficient model, for focused, high-volume tasks. The launch post presents Sol and Luna as Astra's advances carried into faster, cheaper models, not as replacements, and settles the order in one sentence: "GPT-6 Astra continues to be our best model across the board."

SpecGPT-6 AstraGPT-6 SolGPT-6 Luna
ReleasedSep 3Sep 22Sep 22
Context / max output1,050,000 / 128,000samesame
Knowledge cutoffApr 30, 2026Apr 20, 2026May 18, 2026
API effortslow to maxnone to maxnone to max
Top setting in Codexultraultramax
Codex CLI 0.155 catalog order123

Two details matter if you swap models in code. none works on Sol and Luna and returns HTTP 400 on Astra. And in Codex all three start at medium unless you set an effort, so a setup that only changes the model name runs none of the max numbers below. Mine sets ultra explicitly, for the main session, plan mode, reviews and subagents. The catalog on my machine, refreshed on launch day, still describes Astra as "Our most capable model for complex, demanding work." Sol's entry has a placeholder description.

The price, next to Astra

Per 1M tokens, or per Codex windowGPT-6 AstraGPT-6 SolGPT-6 Luna
API input / cached input$10.00 / $1.00$2.00 / $0.20$0.10 / $0.01
API output$50.00$10.00$0.50
Codex credits: input / cached / output250 / 25 / 1,25050 / 5 / 2502.5 / 0.25 / 12.5
Codex messages per 5 hours, Plus5–4515–150350–3,000
Codex messages per 5 hours, Pro 20x100–900300–3,0007,000–56,000

The sheet is unusually tidy. Sol is exactly one fifth of Astra on every per-token line and in every tier, cache writes included ($2.50 against $12.50), and Luna exactly one hundredth. The long-context step from the Astra launch piece applies to all three: once a prompt passes 272K input tokens, the whole request pays double the input and cache rates and 1.5 times the output rate. Batch and Flex halve the prices. Fast mode doubles them. Against its predecessor, Sol's $2 and $10 are half of GPT-5.6 Sol's current promotional $4 and $20.

In Codex signed in with ChatGPT, the same ratio shows up as credits, though not in the message estimates: OpenAI's ranges give Sol about three times Astra's messages per window and Luna about 60 to 70 times. OpenAI calls them estimates, not fixed limits, and says weekly limits may also apply.

Quality at max: OpenAI's charts

OpenAI launch chart, at maxGPT-6 AstraGPT-6 SolGPT-6 Luna
AutomationBench 1.0.641.4% ($1.73)32.0% ($0.34)20.7% ($0.04)
Agents' Last Exam V159.3% ($6.23)56.4% ($2.93)50.9% ($0.15)
FrontierCode 1.1 Main53.3% ($4.59)49.3% ($2.14)42.4% ($0.11)
DeepSWE 1.173.2% ($7.50)68.8% ($2.74)66.6% ($0.22)
OSWorld 2.0, offline73.5% ($9.07)64.4% ($3.25)52.7% ($0.27)
Factual errors, lower is better3.9% ($0.79)4.6% ($0.18)7.6% ($0.01)

Scores come from the chart data in OpenAI's launch post; the figure in brackets is its plotted cost per task. The post does not say which harness each chart used or how it computed cost per task.

Astra wins every row. Its lead over Sol runs from 2.9 points on Agents' Last Exam to 9.4 on AutomationBench, and it makes fewer factual errors. Luna trails Astra by 6.6 to 20.8 points on the five scored charts and makes almost twice as many factual errors.

The charts plot every effort level, which places Sol on Astra's own ladder. Sol at max beats Astra at low on all six charts. It beats Astra at medium only on FrontierCode, 49.3% to 48.8%. It never reaches Astra at high, xhigh or max. Luna at max beats no Astra setting on any chart. The post's only Sol-versus-Astra cost claim points the same way: on AutomationBench, Sol at xhigh beats Astra at low, which costs 3.9 times as much per task. Every other cost claim in the post measures Sol and Luna against GPT-5.6 or Claude models, never against Astra at max.

A fifth of the price is not a fifth of the bill. At max, Sol costs 53% to 80% less per task, not a flat 80%, because on most charts it spends more tokens than Astra: at list prices, from roughly the same amount on AutomationBench to about 2.35 times as much on Agents' Last Exam. That is my arithmetic on OpenAI's plotted costs, assuming both models were costed at list price.

What the post leaves out: no Terminal-Bench, SWE-bench or CursorBench numbers for Sol or Luna, and no Claude Opus 5.5, which shipped the same day, or Grok 4.7 in any chart. The numbers also disagree slightly across OpenAI's own pages: Astra's OSWorld score was 72.6% in its launch table and is 73.5% in the new chart. Nothing in the verdict depends on that difference.

Quality at max: Artificial Analysis

Artificial Analysis published its own measurements the same day, all at max:

Artificial Analysis, at maxGPT-6 AstraGPT-6 SolGPT-6 Luna
Intelligence Index v4.3534837
Cost per Intelligence Index task$3.26$1.06$0.07
Coding Agent Index, in Codex625741
Cost per Coding Agent task$7.09$2.99not published
Output speed, tokens per second57.7104.4157.2

The order holds. The Coding Agent Index matters most here, because it runs in the Codex harness, the closest public setup to this seat, though at max rather than ultra. Sol trails there by five points, the same margin as on the Intelligence Index. The widest gap between Sol and Astra in anything I read is on a benchmark OpenAI didn't chart: in Artificial Analysis's Terminal-Bench 4.0 runs, Sol lands in the low 40s and Astra in the high 50s.

Its headline is that Sol and Luna halve the cost of GPT-5.6 Sol and Luna. On the Coding Agent Index, Sol gains two points on GPT-5.6 Sol and Luna loses two against GPT-5.6 Luna. That is the right frame for anyone moving off GPT-5.6. It is not the question for a seat that already runs Astra.

The rule, applied

Sol. Better at max? No: not on any of OpenAI's six charts, not on either Artificial Analysis index, and not by OpenAI's own ranking. Cheaper? Yes: 20% of Astra's per-token price, 53% to 80% less per task at max, and about three times the Codex messages. The rule needs both. No switch.

Luna. Better? No: 6.6 to 20.8 points behind on OpenAI's five scored charts with almost twice Astra's factual-error rate, 16 and 21 points behind on the two indices, and no ultra in Codex. Cheaper? Yes, at 1% of Astra's per-token price. No switch.

Three releases in two days, three decisions. Yesterday Grok 4.7 replaced Grok 4.6: xAI's newer model at the same price, a straight upgrade. Earlier today Opus 5.5 replaced Fable 5.1: on Anthropic's launch table, with each model at its best setting, Opus 5.5 scored above Fable 5.1 on all nine rows, at 40% of Fable's input and output price. Better on its vendor's own table and cheaper, so it would have passed this test. Sol is cheaper and not better. However large the discount, the quality half has to pass first.

What OpenAI recommends, and why my desk doesn't follow it

OpenAI's Codex docs suggest starting most Codex tasks and the more demanding subagents on gpt-6-sol and using Luna for lighter subagent work. The models page keeps Astra for tasks that need the strongest capability across many steps and tools. That is sound advice for anyone managing a budget or a message allowance.

My desk answers a different question. It is the setup from I Love Trios: three programs under one policy, each on its vendor's strongest model at the highest setting that model offers, delegated agents included. Codex's subagents run on Astra at ultra, like the main session. Held to the same rule, OpenAI's split gets the same answer.

The desk, before and after

ProgramModel before this releaseAfter Sol and Luna
Claude CodeClaude Opus 5.5 at max, since earlier todayunchanged
CodexGPT-6 Astra at ultraunchanged: GPT-6 Astra at ultra
Grok BuildGrok 4.7 at xhigh, since yesterdayunchanged

For context only, the three seats score 58 (Opus 5.5 at max), 53 (Astra at max) and 46 (Grok 4.7 at xhigh) on Artificial Analysis's Intelligence Index. Sol would have come in at 48. The rule compares candidates with the Codex incumbent, so none of the other numbers enter it.

No model configuration changed. The switch list I had prepared, every Astra pin from the shared policy and the Codex config to the reviewer runner and the loop templates, stays unused. What changed is the record: the repository rulebook and the shared memory that all three agents read now hold the rule, the result and the evidence, so the next session starts from the decision instead of redoing the research.

Outside the rule

Observations, not changes.

  • Effort versus model. On five of the six charts, Astra at medium scores above Sol at max, but it costs more per task on each of those five, from 12% more to 3.7 times as much. Astra at low goes the other way on FrontierCode, DeepSWE and OSWorld: cheaper per task than Sol at max, and lower-scoring. If spending ever had to come down, there are two levers, and they trade quality for money at different rates.
  • Speed. Artificial Analysis measures Sol's output at 104.4 tokens per second against Astra's 57.7, about 1.8 times as fast. That matters more when someone is watching the terminal than in a headless loop.
  • Safety monitoring. OpenAI's Codex docs say Astra runs with asynchronous safety monitoring that can pause a task; in Codex CLI, a paused task ends, with no resume. In the API, a blocked request returns HTTP 403 misalignment_policy_violation. OpenAI documents this for Astra, whose system card puts it at the Critical threshold for cybersecurity; Sol and Luna ship with safeguards similar to the GPT-5.6 models instead. For headless Codex CLI runs, that is a real operational difference, and I have not measured how often it bites.

Hot takes

  • A price cut is not an upgrade. Sol is OpenAI's second-best model at a fifth of Astra's price, and the rule asks about the model.
  • Write the rule before you read the charts. Written afterwards, "almost as good" stretches to fit whatever is cheapest.
  • "Start with Sol" answers "what should most people run?", not "what does the hardest work best?"

What this does not show

  • Quality on my work. Not measured. The decision rests on OpenAI's charts and Artificial Analysis.
  • ultra. No published results at Codex's top setting for any GPT-6 model. Subagent delegation can change quality and cost in ways max doesn't capture.
  • Cost on my work. Per-task figures are OpenAI's plotted estimates and Artificial Analysis's measurements. Message counts are estimates, and a subscription allowance is not a per-token bill.
  • Sol below max. The rule tests max. A lower Sol setting may be the better trade in other setups.
  • The harness. OpenAI doesn't say which harness each chart used. Artificial Analysis's Coding Agent Index is the one run inside Codex.
  • The competition. OpenAI's charts include neither Claude Opus 5.5 nor Grok 4.7.