OpenAI stopped shipping one flagship model at a time back in July 2026. GPT-5.6 arrived as a three-tier family instead: Sol, Terra, and Luna, each aimed at a different budget and a different job. Seven weeks later, that pricing structure has already been cut twice, the free ChatGPT tier has switched its default model, and independent benchmark trackers have settled on scores that make the choice between tiers a real engineering decision rather than a marketing one.
This comparison breaks down what GPT-5.6 Sol, Terra, and Luna actually cost as of August 28, 2026, how they score on independent benchmarks from Artificial Analysis and evals.report, and how each tier stacks up against Claude Opus 5, Gemini 3.7 Flash, and DeepSeek V4-Pro. If you are deciding which model ID to put in production, or which ChatGPT plan actually gets you the model you want, the numbers below come straight from OpenAI’s own pricing pages and third-party evaluation labs.
The stakes of getting this choice right are higher than they look. Because all three tiers share the same 1.05 million token context window, a developer cannot fall back on “just use the model with more context” as a decision rule the way they could with earlier GPT generations. The only variables left are price, benchmark score, and latency, which means picking the wrong tier for a given workload either wastes budget or leaves measurable performance on the table. Both mistakes are common enough in the first weeks after a pricing change that OpenAI’s own community forum has multiple threads from developers who missed the July 30 and August 21 price cuts entirely and kept paying the old rate.
What Is GPT-5.6? Sol, Terra, and Luna Explained
OpenAI previewed GPT-5.6 on June 26, 2026, and pushed it to general availability across ChatGPT, the API, and Codex on July 9, 2026, according to the company’s own launch post. Rather than a single model, GPT-5.6 shipped as a family of three: Sol is the flagship, built for hard coding tasks, research, and long agentic runs. Terra is the balanced middle tier, priced to replace most GPT-5.5 workloads at less than half the cost. Luna is the fast, cheap option meant for high-volume chat and simple tool calls.
All three share a 1,050,000-token context window and a 128,000-token maximum output, according to CloudZero’s pricing breakdown and confirmed by OpenAI’s own model documentation. That is an unusual design choice. Most vendors shrink context as they cut price, but OpenAI kept the window identical across Sol, Terra, and Luna so the decision between tiers is purely about intelligence and cost per task, not how much text you can feed the model.
The API model ID strings are straightforward: gpt-5.6-sol, gpt-5.6-terra, and gpt-5.6-luna. Developers who already had GPT-5.5 or GPT-4.1 calls in production can swap the model string directly, since OpenAI kept the same request format across the switch.
Why OpenAI Split GPT-5.6 Into Three Tiers
The three-tier structure did not appear in a vacuum. By mid-2026, every major lab was already shipping a fast-and-cheap variant alongside its flagship. Anthropic runs Opus, Sonnet, and Haiku. Google splits Gemini into Pro and Flash lines. DeepSeek ships both a Pro and a Flash release from the same model family. GPT-5.6 Sol, Terra, and Luna is OpenAI matching that structure directly, rather than forcing every customer, from a solo developer testing a prototype to an enterprise running millions of daily agent calls, onto one price point.
The timing also lines up with real competitive pressure. Gemini 3.7 Flash launched August 13, 2026 at an aggressive introductory rate, and DeepSeek V4-Flash undercut nearly everyone on July 31, 2026 at $0.14 input and $0.28 output per million tokens. OpenAI’s two price cuts on GPT-5.6, one on July 30 and another on August 21, both landed within weeks of a competitor move, which suggests the tiered pricing is as much a response to Google and DeepSeek as it is an internal product decision.
GPT-5.6 Sol vs Terra vs Luna: Full Specs Comparison
Here is how the three tiers compare across the specs that actually matter for a build decision, pulled from OpenAI’s pricing documentation, Artificial Analysis, and evals.report.
| Spec | GPT-5.6 Sol | GPT-5.6 Terra | GPT-5.6 Luna |
|---|---|---|---|
| Release date | July 9, 2026 | July 9, 2026 | July 9, 2026 |
| Positioning | Flagship | Balanced daily use | Fast and cheap |
| Context window | 1,050,000 tokens | 1,050,000 tokens | 1,050,000 tokens |
| Max output | 128,000 tokens | 128,000 tokens | 128,000 tokens |
| Input price (short context) | $4.00 / 1M | $2.00 / 1M | $0.20 / 1M |
| Output price (short context) | $20.00 / 1M | $12.00 / 1M | $1.20 / 1M |
| Cached input price | $0.40 / 1M | $0.20 / 1M | $0.02 / 1M |
| Long-context input (over 272K tokens) | $8.00 / 1M | $4.00 / 1M | Not published |
| AA Intelligence Index (max effort) | 59 | 55 | 51 |
| AA Coding Agent Index | 80 | 77.4 | 74.6 |
| GPQA Diamond | Not published by OpenAI | Not published by OpenAI | 92.3% |
| SWE-bench Pro | 62.7% | Not published | Not published |
| Throughput (Cerebras) | Up to 750 tokens/sec | Not published | Not published |
| ChatGPT plan access | Plus (medium+ effort), Pro, Business, Enterprise | Free/Go, Plus, Pro, Business, Enterprise | Free/Go (default since Aug 6), Plus, Pro, Business, Enterprise |
Two things stand out immediately. First, OpenAI did not publish a complete benchmark table for every tier on every metric, so gaps in the row above reflect what the company and independent evaluators have actually disclosed rather than missing research. Second, the price gap between Sol and Luna is steep: $4.00 versus $0.20 per million input tokens works out to a 20x difference, and $20.00 versus $1.20 on output is nearly 17x. That gap is the whole story behind why picking the right tier matters more with GPT-5.6 than it did with a single-model release.
Pricing Breakdown: What Each Tier Costs Right Now
GPT-5.6 pricing has moved twice since launch, and both moves went in the same direction. At general availability on July 9, 2026, Sol launched at $5.00 input and $30.00 output per million tokens, Terra at roughly $2.50 and $15.00, and Luna at $1.00 and $6.00, according to explainx.ai’s preview coverage and OpenAI’s own launch materials.
On July 30, 2026, OpenAI cut Terra to $2.00 and $12.00, and slashed Luna to $0.20 and $1.20, a move the company framed as “advancing the price-performance frontier” in its own blog post. Then on August 21, 2026, Reuters reported that OpenAI cut Sol’s developer pricing by more than 20%, dropping it from $5.00 to $4.00 per million input tokens and from $30.00 to $20.00 per million output tokens, with the new rate guaranteed to hold at least through November 21, 2026.
Here is how the current GPT-5.6 lineup compares against the other frontier and budget models it competes with in late August 2026.
| Model | Input $/1M | Output $/1M | Context | Source |
|---|---|---|---|---|
| GPT-5.6 Sol | $4.00 | $20.00 | 1.05M | OpenAI / Reuters |
| GPT-5.6 Terra | $2.00 | $12.00 | 1.05M | OpenAI |
| GPT-5.6 Luna | $0.20 | $1.20 | 1.05M | OpenAI |
| Claude Opus 5 | $5.00 | $25.00 | 1M | Anthropic |
| Gemini 3.7 Flash (intro rate through Dec 31, 2026) | $0.75 | $3.75 | 1.05M | Google / eesel.ai |
| DeepSeek V4-Flash 0731 | $0.14 | $0.28 | Not disclosed in this dataset | Thunder Compute |
| DeepSeek V4-Pro (0813) | Not disclosed in this dataset | Not disclosed in this dataset | 1M | Local AI Zone |
The pattern is clear: Sol is priced almost identically to Claude Opus 5 on input tokens ($4.00 versus $5.00) but cheaper on output, while Luna undercuts even Gemini 3.7 Flash on both ends. DeepSeek V4-Flash remains the outright cheapest option in this set at $0.14 and $0.28 per million tokens, roughly 71x cheaper than Sol on output pricing.
Benchmark Performance: Coding, Reasoning, and Agentic Scores
Artificial Analysis runs its Intelligence Index across nine separate evaluations, including GPQA Diamond, Humanity’s Last Exam, Terminal-Bench 2.1, and SciCode. On that composite score, GPT-5.6 Sol at max effort lands at 59, one point behind Claude Fable 5’s 60 while completing tasks at roughly a third of the cost per task, about $1.04 per task, according to The Decoder’s coverage of the launch benchmarks. Terra scores 55 at max effort, costing about $0.55 per task, and Luna scores 51 at around $0.21 per task.
On the Artificial Analysis Coding Agent Index, which measures how well a model performs as an autonomous coding agent rather than a single-shot chat responder, Sol set a new high mark of 80 at release, 2.8 points ahead of Claude Fable 5’s 77.2 while using less than half the output tokens to get there, per Axis Intelligence’s benchmark tracker. Terra scores 77.4 and Luna 74.6 on the same index, according to Vellum’s tier breakdown.
OpenAI’s own launch table, mirrored by evals.report, discloses SWE-bench Pro, a contamination-resistant coding benchmark, at 62.7% for Sol, well behind Claude Fable 5’s 80% on the same test according to developer Simon Willison’s analysis cited by Vellum. On the disclosed Luna-specific scores, OpenAI reports GPQA Diamond at 92.3%, Terminal-Bench 2.1 at 84.7%, and Agents’ Last Exam at 50.3%. Independent max-effort figures from Artificial Analysis and the ARC Prize organization put Luna at 59.5% on ARC-AGI-2 and 37.2% on Humanity’s Last Exam.
| Benchmark | GPT-5.6 Sol | Claude Opus 5 | Gemini 3.7 Flash | DeepSeek V4-Pro |
|---|---|---|---|---|
| AA Intelligence Index | 59 | 61 | Not published in this set | Not published in this set |
| GPQA Diamond | Not published for Sol | 93.2% | 94.5% | 90.1% |
| SWE-bench (Pro/Verified varies) | 62.7% (Pro) | 96.0% (Verified) | 80.8% (Verified) | 80.6% |
| ARC-AGI-2 | Not published for Sol | 90.4% | Not published in this set | Not published in this set |
Two caveats are worth flagging before anyone treats these as a straight ranking. OpenAI reports SWE-bench Pro for Sol rather than SWE-bench Verified, which Anthropic and Google both report for their models, so the coding scores in that row are not directly comparable across vendors. And several boxes above are genuinely blank because none of the three companies published a matching figure, not because the number was rounded down or omitted for looking bad.
Cost Modeling: What a Million Requests Actually Costs
Per-token pricing is hard to reason about until you scale it to a realistic workload. Take a support automation pipeline processing 1 million requests a month, with an average of 800 input tokens and 300 output tokens per request, a rough proxy for a short ticket summary or classification task. On Sol at $4.00/$20.00 per million tokens, that workload costs about $3,200 on input tokens and $6,000 on output tokens, for a monthly total near $9,200. Run the identical workload through Terra at $2.00/$12.00 and the total drops to roughly $5,200. Route it through Luna at $0.20/$1.20 and the monthly bill falls to about $520, a 17.7x reduction versus Sol for a workload that likely does not need frontier-level reasoning in the first place.
That gap is the entire argument for tiered routing. A team that defaults every call to Sol out of caution, without profiling which requests actually need it, is paying a premium that a Luna or Terra call would have handled just as well for a support ticket or a classification task. The inverse is also true: routing a genuine multi-step coding agent through Luna to save money will likely cost more in the long run through failed tool calls, retries, and human review time than the token savings are worth.
Context Window, Output Limits, and Long-Context Pricing
The shared 1.05 million token context window is the single most unusual decision in the GPT-5.6 lineup. Competing vendors typically shrink context as they cut price. Anthropic’s cheaper Claude tiers have historically shipped a smaller window than Opus. OpenAI instead gave Sol, Terra, and Luna the exact same 1,050,000-token ceiling and 128,000-token output cap, meaning a developer never has to choose between a cheaper model and a smaller working memory.
That said, pricing still scales with context length past a threshold. Once a request crosses 272,000 input tokens, Sol’s rate roughly doubles to $8.00 input and $10.00 cache-write per million tokens, with output rising to $30.00. Terra’s long-context rate climbs to $4.00 input and $18.00 output. OpenAI has not published a separate long-context rate for Luna in its public pricing table as of this writing, which suggests Luna may be intended primarily for shorter, high-volume requests rather than million-token document analysis.
Practically, this means a legal team running a 900,000-token contract review through Sol will pay the long-context rate on the entire request, not just the portion past 272K. Chunking large documents into sub-272K segments before sending them to the API can meaningfully cut costs, especially on Terra where the long-context multiplier is exactly 2x on input and 1.5x on output.
Speed and Throughput: How Fast Is Each Tier
OpenAI’s clearest published throughput figure covers Sol running on Cerebras hardware, where the company says it hits up to 750 tokens per second, a launch-day claim from the “Previewing GPT-5.6 Sol” post aimed squarely at latency-sensitive agentic workloads. OpenAI has not published equivalent official throughput numbers for Terra or Luna, though the general expectation across the industry is that smaller, cheaper tiers run faster on standard infrastructure since they require less compute per generated token.
For context, Gemini 3.7 Flash currently ranks first out of 186 models tracked by independent evaluators on output speed, according to a felloai.com roundup of the August 2026 pricing shifts. That makes Flash the benchmark to beat on raw speed, and it is one reason teams building latency-critical consumer features often default to Gemini’s Flash tier or GPT-5.6 Luna over Sol, even when Sol scores higher on intelligence benchmarks.
Which ChatGPT Plan Gets Which Model
Access to the three tiers is split by ChatGPT subscription level. According to Apidog’s pricing guide, Free and Go accounts get Terra by default, Plus subscribers get Sol, Terra, and Luna with per-model effort controls (Sol requires medium effort or higher), and Pro, Business, and Enterprise accounts get all three tiers plus a higher-effort “Sol Pro” mode.
That access map shifted once already. On August 6, 2026, OpenAI swapped the default model for Free and Go-tier ChatGPT users away from Terra and onto Luna, pairing the change with unlimited text chats and the “Think” reasoning toggle for free accounts. The move effectively made Luna the model most ChatGPT users interact with day to day, even though Sol gets the marketing attention as the flagship.
For API developers, plan tier is irrelevant. Anyone with API access can call any of the three model IDs directly, paying the published per-token rate regardless of whether they also have a ChatGPT subscription.
One detail that trips up teams building on top of ChatGPT rather than the raw API: the “effort” setting available to Plus, Pro, Business, and Enterprise users changes which underlying configuration of a tier actually answers a prompt. Selecting Sol at medium effort is a materially different (and cheaper, faster) request than Sol at xhigh effort, even though both show up in the interface simply as “Sol.” Teams comparing their own experience against benchmark tables published by Artificial Analysis should check which effort level a given score was measured at, since the gap between Sol at medium and Sol at max effort is large enough to change which competitor model looks stronger.
GPT-5.6 Sol vs Claude Opus 5
Claude Opus 5 launched July 24, 2026, two weeks after GPT-5.6’s general availability, at $5.00 input and $25.00 output per million tokens, unchanged from its Opus 4.8 predecessor’s pricing according to Anthropic’s own announcement. That makes Opus 5 slightly more expensive than Sol on input ($5.00 versus $4.00) but more expensive on output ($25.00 versus $20.00), so total cost depends heavily on your input-to-output token ratio.
On raw capability, Opus 5 currently edges out Sol. Anthropic and independent trackers both report Opus 5 topping the Artificial Analysis Intelligence Index at 61, ahead of Sol’s 59, while also posting a 96.0% score on SWE-bench Verified compared to Sol’s 62.7% on the differently-scoped SWE-bench Pro. Opus 5 also reports 93.2% on GPQA Diamond and 90.4% on ARC-AGI-2, both stronger than any figure OpenAI has published for Sol on the same tests.
The practical takeaway: if your workload is pure coding correctness and you can absorb a slightly higher output cost, Opus 5’s benchmark lead is real and independently verified. If your workload leans agentic, with lots of tool calls and shorter completions, Sol’s Coding Agent Index score of 80 and its lower output price start to close the gap. Teams running both models side by side in evaluation often find the practical difference smaller than the headline index numbers suggest, since real production prompts rarely resemble the exact benchmark tasks either lab optimized for.
GPT-5.6 vs Gemini 3.7 Flash
Gemini 3.7 Flash reached general availability on August 13, 2026, and Google priced it aggressively at an introductory $0.75 input and $3.75 output per million tokens, a rate guaranteed through December 31, 2026 according to eesel.ai’s review. That undercuts every GPT-5.6 tier except Luna on price, while still posting strong benchmark numbers: 94.5% on GPQA Diamond and 80.8% on SWE-bench Verified, both ahead of what OpenAI has disclosed for Sol on comparable tests.
The comparison that matters most is Flash against Luna, since both sit in the fast-and-cheap category. Luna is cheaper on input ($0.20 versus $0.75) but Flash currently holds the speed crown, ranking first of 186 tracked models on output tokens per second per felloai.com’s August 2026 roundup. For a high-volume customer support bot processing millions of short exchanges daily, Luna’s lower per-token cost may still win on total spend even if Flash edges it on latency.
GPT-5.6 vs DeepSeek V4-Pro and V4-Flash
DeepSeek’s two most recent releases sit at opposite ends of the same trade-off GPT-5.6 makes internally. DeepSeek V4-Pro reached general availability around August 12 to 13, 2026 as an MIT-licensed, open-weight model built on a mixture-of-experts architecture with roughly 671 billion total parameters and about 37 billion active per token, according to a July-August 2026 roundup from Local AI Zone. It reports a 1 million token context window and can run locally on hardware with 64GB or more of RAM, a deployment option none of the GPT-5.6 tiers offer since they are API-only.
On benchmarks, DeepSeek V4-Pro posts 90.1% on GPQA Diamond and 80.6% on SWE-Bench per Thunder Compute’s open-source LLM tracker, scores that sit ahead of what OpenAI has published for Sol on the same tests, though the two companies do not report identical benchmark variants, which limits a true apples-to-apples read.
DeepSeek V4-Flash 0731, released July 31, 2026, is the more direct budget competitor to Luna. At $0.14 input and $0.28 output per million tokens, it undercuts Luna’s $0.20 and $1.20 rate on both ends, and it posts a GPQA Diamond score of 88.1% and a Live Code Bench score of 91.6% per Thunder Compute’s data. For teams comfortable self-hosting or routing through a third-party API aggregator, DeepSeek’s Flash tier is the cheapest credible option in this entire comparison.
The trade-off DeepSeek asks buyers to accept is operational rather than financial. Both DeepSeek releases ship as open weights under an MIT-style license, which means a team can inspect, fine-tune, and self-host them, but it also means there is no OpenAI-style managed API with guaranteed uptime SLAs, enterprise support contracts, or a Business/Enterprise tier with elevated rate limits. For a startup optimizing purely for cost per token and comfortable managing its own inference infrastructure or a third-party host, DeepSeek is a legitimate alternative to Luna. For a team that wants a single vendor relationship and predictable support, GPT-5.6 or Claude keeps that simpler even at a higher per-token price.
Real-World Use Cases: Which Tier Fits Which Job
The three-tier structure only pays off if you actually route different workloads to different models. Here are six scenarios where the choice is clear-cut based on the pricing and benchmark data above.
- Autonomous coding agents: A team building a CI pipeline that opens pull requests, runs tests, and iterates on failures benefits from Sol’s 80 score on the Coding Agent Index and its 750 tokens/second throughput on Cerebras, since agent loops involve many sequential calls where latency compounds.
- High-volume customer support: A support desk fielding tens of thousands of routine tickets a day is better served by Luna’s $0.20/$1.20 pricing, especially since most support queries do not need frontier-level reasoning to resolve correctly.
- Long-document research and legal review: A research or compliance team parsing filings, contracts, or academic papers close to the 1.05M token ceiling should default to Terra, which balances the 2x long-context multiplier against Sol’s higher baseline rate.
- Indie SaaS prototyping: A solo developer or small startup iterating on product features daily typically lands on Terra as the default, since OpenAI itself has positioned it as a like-for-like replacement for GPT-5.5 workloads at less than half the price.
- Free-tier consumer chatbots: Any product built on ChatGPT’s free plan is already running on Luna as of the August 6, 2026 default switch, which is worth knowing if your support documentation still tells users they are talking to “GPT-5.6” without specifying the tier.
- Enterprise agentic workflows with compliance requirements: Business and Enterprise accounts that need the highest reasoning ceiling for regulated decisions, such as underwriting or financial analysis support, are the intended audience for Sol Pro, the higher-effort mode reserved for paid enterprise tiers.
Common Mistakes When Choosing a GPT-5.6 Tier
A handful of avoidable mistakes show up repeatedly in how teams pick between Sol, Terra, and Luna. The first is defaulting to Sol for everything out of an assumption that “flagship” always means “correct choice.” Given the cost modeling above, that habit alone can inflate a monthly bill by close to 18x for workloads that Luna would have handled adequately, with no measurable quality difference a user would notice in a routine ticket classification or short summary.
The second mistake is ignoring the 272,000-token long-context threshold until a bill arrives higher than expected. A document-heavy workload that regularly crosses that line on Sol pays double the standard rate on the entire request, not a prorated amount, so batching or chunking large inputs before they hit the API is worth the extra engineering effort at any real scale.
The third is treating the “effort” parameter as a fixed setting rather than something to test per task. Since Sol’s Coding Agent Index score and its dollar cost both move with effort level, a team that hardcodes “high” everywhere is often paying for capability the task never uses. Running an evaluation pass across effort levels on a sample of real production requests, rather than trusting the default, typically surfaces meaningful savings within the first week of testing.
Migration Guide: Moving From GPT-5.5 to GPT-5.6
Teams still running GPT-5.5 or GPT-4.1 in production can migrate to GPT-5.6 without a full rebuild, since OpenAI kept the request format consistent across the switch. Here is a practical rollout sequence.
- Audit current model calls and tag each endpoint by task type: agentic coding, chat, summarization, or document analysis.
- Map each task category to a GPT-5.6 tier using the use-case guidance above, defaulting to Terra when unsure.
- Update the model ID in each API call to
gpt-5.6-sol,gpt-5.6-terra, orgpt-5.6-luna. - For Sol calls, set the effort parameter explicitly (medium, high, or xhigh) rather than relying on a default, since effort level materially changes both cost and the Coding Agent Index score you get.
- Enable prompt caching on repeated system prompts to capture the cached input discount, which runs as low as $0.02 per million tokens on Luna.
- Stand up per-tier cost dashboards before rollout so a spike in Sol usage does not silently blow through budget.
- Run a side-by-side evaluation on a held-out task set comparing your old model’s output quality against the new tier before fully cutting over.
- Roll out gradually, routing a percentage of traffic to the new model and keeping a fallback path to the previous model for the first two weeks.
from openai import OpenAI
client = OpenAI()
# Agentic coding task -> Sol, high effort
response = client.responses.create(
model="gpt-5.6-sol",
reasoning={"effort": "high"},
input="Refactor this function and add test coverage."
)
# Daily chat / balanced workload -> Terra
response = client.responses.create(
model="gpt-5.6-terra",
input="Summarize this support ticket thread."
)
# High-volume, low-complexity -> Luna
response = client.responses.create(
model="gpt-5.6-luna",
input="Classify this message as billing, technical, or general."
)
Pros and Cons of Sol, Terra, and Luna
GPT-5.6 Sol
Pros: highest Coding Agent Index score in the family at 80, up to 750 tokens/second on Cerebras hardware, and a recent price cut that dropped output cost from $30.00 to $20.00 per million tokens. Cons: still the most expensive GPT-5.6 tier, trails Claude Opus 5 on the Artificial Analysis Intelligence Index (59 versus 61), and OpenAI has not published SWE-bench Verified scores for it, only the less directly comparable SWE-bench Pro.
GPT-5.6 Terra
Pros: positioned by OpenAI as a like-for-like GPT-5.5 replacement at less than half the price, available by default to Free and Go accounts, and it shares the full 1.05M context window with the pricier Sol tier. Cons: a 55 Intelligence Index score leaves a real capability gap versus Sol’s 59, and OpenAI has not disclosed a Luna-style detailed benchmark table for it, making independent verification harder.
GPT-5.6 Luna
Pros: the cheapest OpenAI-hosted GPT-5.6 tier at $0.20/$1.20 per million tokens, a strong 92.3% GPQA Diamond score for its price class, and it is now the default free ChatGPT model with unlimited text chats. Cons: lowest Intelligence Index score in the family at 51, no published long-context pricing tier, and DeepSeek V4-Flash undercuts it on raw price for teams willing to shop outside OpenAI.
The Verdict: Which GPT-5.6 Tier Should You Actually Use
The data points to a simple default: start on Terra, escalate to Sol only for agentic coding and high-stakes reasoning, and drop to Luna for anything high-volume and low-complexity. Terra’s positioning as a discounted GPT-5.5 replacement, combined with the full 1.05M context window it shares with Sol, makes it the safest default for teams that have not yet profiled their workload by task type.
Sol earns its premium specifically on agentic coding work, where its 80 score on the Coding Agent Index and Cerebras-backed throughput matter more than its 59 versus Claude Opus 5’s 61 on the general Intelligence Index. If your workload is pure reasoning correctness rather than agent loops, Opus 5’s 96.0% SWE-bench Verified score is the stronger buy at a similar input price.
Luna wins on pure economics for chat and classification workloads, but it is not the cheapest option in the market. DeepSeek V4-Flash at $0.14/$0.28 per million tokens and Gemini 3.7 Flash’s speed lead both undercut Luna on at least one axis, so teams optimizing purely for cost per token at scale should still run the comparison before committing to an all-OpenAI stack.
None of this is static. OpenAI has already cut GPT-5.6 pricing twice in under two months, and both Google and DeepSeek shipped new budget-tier models in the same window. Anyone building a production system on any of these tiers should treat the numbers in this comparison as the picture on August 28, 2026, not a permanent baseline, and should build cost monitoring flexible enough to catch the next price move rather than discovering it after a monthly invoice lands.
Frequently Asked Questions
What is the difference between GPT-5.6 Sol, Terra, and Luna?
Sol is the flagship tier built for hard coding and agentic tasks, Terra is the balanced middle tier meant to replace GPT-5.5 workloads at lower cost, and Luna is the fastest, cheapest tier aimed at high-volume chat. All three share the same 1,050,000-token context window but differ sharply on price and benchmark scores.
How much does GPT-5.6 Sol cost per million tokens?
As of the August 21, 2026 price cut, GPT-5.6 Sol costs $4.00 per million input tokens and $20.00 per million output tokens for standard short-context requests, with cached input priced at $0.40 per million tokens.
Is GPT-5.6 Luna free to use?
Luna is free to use inside ChatGPT for Free and Go-tier accounts, which have defaulted to Luna since August 6, 2026 with unlimited text chats. Through the API, Luna still costs $0.20 per million input tokens and $1.20 per million output tokens.
Which GPT-5.6 tier is best for coding?
Sol is the strongest GPT-5.6 tier for coding, scoring 80 on the Artificial Analysis Coding Agent Index at release, ahead of Terra’s 77.4 and Luna’s 74.6. For pure code correctness on SWE-bench-style tasks, Claude Opus 5’s 96.0% SWE-bench Verified score currently outperforms all three GPT-5.6 tiers.
How does GPT-5.6 Sol compare to Claude Opus 5?
Claude Opus 5 leads Sol on the Artificial Analysis Intelligence Index, 61 versus 59, and posts stronger GPQA Diamond and ARC-AGI-2 scores. Sol is slightly cheaper on input tokens ($4.00 versus $5.00) and output tokens ($20.00 versus $25.00), and it holds an edge specifically on agentic coding benchmarks.
Do Sol, Terra, and Luna have the same context window?
Yes. All three GPT-5.6 tiers share a 1,050,000-token context window and a 128,000-token maximum output, a design choice OpenAI kept consistent across the entire family regardless of price.
Can I switch between GPT-5.6 tiers in the same application?
Yes. Since all three tiers use the same API request format, developers can route different request types to different model IDs within the same application, calling gpt-5.6-sol for complex tasks and gpt-5.6-luna for simple ones based on a routing rule.
When did GPT-5.6 launch?
OpenAI previewed GPT-5.6 on June 26, 2026, and all three tiers reached general availability together on July 9, 2026, across ChatGPT, the API, and Codex.
What ChatGPT plan do I need to access GPT-5.6 Sol?
Sol is available to ChatGPT Plus subscribers at medium effort or higher, while Pro, Business, and Enterprise accounts get access to all three tiers plus a higher-effort Sol Pro mode. Free and Go accounts do not get Sol by default.
- GPT-5.6 Sol vs Qwen3.8 Max vs Claude Opus 4.6: 5x Price Gap [2026]
- Claude Sonnet 5 vs GPT-5.6 vs Gemini 3.7 Flash: 6.7x Price Gap [2026]
- Claude Fable 5 vs Opus 5 vs GPT-5.6 Sol: $1,125 Gap [2026]
- Claude Opus 5 vs GPT-5.6 vs DeepSeek V4-Pro: $22 Gap [2026]
- How to Use the GPT-5.6 API: 12 Steps, 100 Min [2026]
- Best AI Models 2026