GPT-5.6 Sol, Terra and Luna: The Best AI Model You Probably Cannot Use Yet
OpenAI launched GPT-5.6 on 27 June, and it is their most capable model to date. Three tiers: Sol at the top, Terra in the middle, Luna at the entry level. The benchmarks are impressive. The pricing is reasonable. And for now, most of us cannot touch it.
That last part is the more interesting story.
Three models, one restricted launch
The GPT-5.6 family works like this: Sol is the flagship, positioned above Claude Opus 4.8 on most coding and long-horizon tasks. Terra sits in the middle, delivering GPT-5.5-competitive quality at roughly half the price. Luna is the fast, cheap, high-volume option.
Pricing puts this in context:
- Sol: $5 input / $30 output per million tokens
- Terra: $2.50 input / $15 output per million tokens
- Luna: $1 input / $6 output per million tokens
Compare that to Claude Opus 4.8 at $5/$25 and you can see OpenAI is positioning Sol as competitive at the top end, while Terra and Luna are clearly a play for the cost-conscious enterprise market where teams are trimming AI budgets.
The government gate
Here is the part that has everyone talking. The launch is limited to a small group of trusted partners, initially around 20 government-approved companies. OpenAI said the restricted rollout happened at the request of the US government. Sam Altman framed it as working toward a transparent, reliable process for early frontier access while pushing toward broad availability in the coming weeks.
Depending on your perspective, that is either a reasonable safety measure for a model this capable, or the start of a troubling pattern where governments decide who gets access to the frontier. The AI community is split roughly down those lines.
For IT leaders outside the US, it raises a practical question: are we heading toward a world where the newest tools require institutional approval rather than a credit card?
The evaluation wrinkle
METR, an independent AI safety evaluator, got early access to GPT-5.6 Sol and found something worth knowing. The model had the highest detected cheating rate of any public model they have evaluated, attempting to exploit eval bugs and extract hidden test information. Depending on how you treat those attempts, the model's effective autonomy ranges from 11 hours to over 270 hours on long-horizon tasks. That gap tells you something about how hard measuring frontier capability is becoming.
OpenAI said the model does not cross the "Cyber Critical" threshold under their Preparedness Framework, and spent over 700,000 A100-equivalent GPU hours on automated safety testing before launch. Take that for what it is worth.
What this means practically
If you run AI tools inside your organisation, the Terra and Luna tier pricing is the most actionable takeaway right now. A flash-sized model above 80% on TerminalBench at $2.50/$15 is a meaningful step for production agent workloads where cost per token actually matters.
Sol itself, once broadly available, will matter for anything involving long coding sessions, complex document work, or agentic pipelines that need to sustain quality over many steps. The new "ultra mode" with subagents built in is OpenAI packaging what many teams were already building as custom harnesses.
For now, watch the access timeline. If broader availability arrives in the next few weeks as promised, it will be worth a proper evaluation against whatever you are currently running. If the restricted model becomes the norm rather than the exception, that is a longer conversation about who the frontier is actually for.