Sol, Terra, Luna: OpenAI Splits GPT-5.6 Into Three
Sol, Terra and Luna offer different capability and cost tradeoffs. A current API comparison of reasoning controls, token prices and workload evaluation.
GPT-5.6 comes in three model tiers: Sol, Terra and Luna. Sol is positioned for demanding reasoning and agent work, Terra balances capability and price, and Luna targets high-volume tasks. This comparison is updated to the API documentation and prices available on September 20, 2026.
That matters more than the names. Engineers do not buy model names — they buy latency, reasoning depth, coding ability, agent reliability, and token economics. GPT-5.6 turns those tradeoffs into a first-class product structure. This is OpenAI saying the frontier is no longer a single SKU.
The three tiers
| Model | Positioning | API reasoning effort | Input / 1M | Output / 1M |
|---|---|---|---|---|
| Sol | Demanding reasoning and agents | none to max | $4 | $20 |
| Terra | Balanced capability and cost | none to max | $2 | $12 |
| Luna | High-volume constrained work | none to max | $0.20 | $1.20 |
All three model pages now list reasoning effort from none through max. That setting is not exclusive to Sol. Product labels such as Codex ultra are separate from the raw API effort values. Choose a model and effort together, then measure whether the resulting output meets the task’s acceptance criteria.
Terra is the real story
A balanced model can matter more to a deployment budget than a flagship benchmark. Terra is a candidate for routine coding, extraction and assistant workloads, but its positioning does not establish parity with Sol or GPT-5.5 on your tasks. Compare accepted results, retries and latency before moving traffic.
Evaluate all candidate tiers against the same task set. A routing design might use Terra for routine tasks, Sol for difficult escalations and Luna for constrained high-volume paths. That is a hypothesis to test, not a recommendation to downgrade a workflow that needs maximum capability.
Sol is an agentic and security flagship
Sol targets complex reasoning, extended coding and agentic work. Those tasks can benefit from stronger capability, but higher effort also changes latency and token consumption. The useful question is whether that additional work changes the verified outcome.
Measure latency where the work happens
Tokens per second are only one part of agent latency. Time to first output, reasoning time, tool calls, retries and service tier all matter. Measure complete accepted tasks in the intended region and deployment path rather than treating a peak throughput claim as an application guarantee.
Keep access and pricing separate
A model’s API availability does not establish access in every subscription, region or deployment product. Check the intended endpoint and account before building a production dependency. Pin model identifiers where reproducibility matters and monitor changes to aliases.
The table uses standard short-context API rates per million tokens, not subscription credits. Sol’s current promotional prices are stated to last at least through November 21, 2026. Caching, longer contexts, faster service tiers, tools and regional processing can change the bill.
Three takes
The takeaway
Model tiering makes an engineering tradeoff explicit. Start with a reliable baseline, test alternatives against the same acceptance criteria, and compare total cost per accepted result. The cheapest token is not necessarily the cheapest completed task.
Sources: Sol model reference; Terra model reference; Luna model reference; OpenAI API pricing. Documentation checked September 20, 2026.