GPT-5.6 Luna vs Terra vs Sol: Intelligence, Cost, Latency and the Real Sweet Spots
OpenAI's GPT-5.6 family can be viewed as three capability tiers:
- Luna — efficiency and high-volume workloads
- Terra — balanced capability, latency and cost
- Sol — highest overall intelligence
But comparing only the model name is no longer enough. Each family can operate at different reasoning-effort levels, so the more useful unit of comparison is:
Model + Reasoning Effort
Using the Artificial Analysis benchmark data captured on 24 August 2026, the differences become much clearer.
Intelligence Index by reasoning effort
| Reasoning Effort | GPT-5.6 Luna | GPT-5.6 Terra | GPT-5.6 Sol |
|---|---|---|---|
| Non-reasoning | 27 | 35 | 42 |
| Low | 34 | 41 | 51 |
| Medium | 39 | 47 | 56 |
| High | 47 | 50 | 57 |
| xHigh | 50 | 53 | 59 |
| Max | 52 | 57 | 61 |

The basic ranking is consistent: Sol > Terra > Luna. But the interesting part appears when we compare different configurations across families.
For example:
- Luna Max: 52
- Terra xHigh: 53
- Sol Low: 51
These configurations occupy roughly the same intelligence band, but their cost and latency are very different.
Intelligence Index is not a percentage
The Artificial Analysis Intelligence Index should not be read as a percentage. A score of 60 does not mean a model is "60% intelligent", and a score of 60 is not twice as intelligent as 30.
It is a composite benchmark signal designed to compare model capability across reasoning, coding, agentic and related tasks. The useful information is the relative positioning between configurations.
Luna: cheap intelligence
| Luna Mode | Intelligence | Cost / Task | First Chunk |
|---|---|---|---|
| None | 27 | $0.01 | 0.72s |
| Low | 34 | $0.01 | 1.61s |
| Medium | 39 | $0.01 | 2.11s |
| High | 47 | $0.02 | 12.84s |
| xHigh | 50 | $0.03 | 58.90s |
| Max | 52 | $0.05 | 146.22s |
Luna's biggest advantage is cost efficiency. Luna High reaches Intelligence 47 for about $0.02 per Artificial Analysis benchmark task.
The trade-off is latency. Moving from High to xHigh raises Intelligence from 47 to 50, but first-chunk latency jumps from about 13 seconds to nearly 59 seconds. Max reaches 52, but the first response takes roughly 146 seconds.
That is a classic case of diminishing returns.
Luna therefore looks especially attractive for background agents, batch processing, document analysis, asynchronous research and other workloads where a user is not waiting for an immediate response.
Terra: the balanced tier
| Terra Mode | Intelligence | Cost / Task | First Chunk |
|---|---|---|---|
| None | 35 | $0.10 | 0.79s |
| Low | 41 | $0.09 | 1.29s |
| Medium | 47 | $0.12 | 1.78s |
| High | 50 | $0.22 | 2.34s |
| xHigh | 53 | $0.31 | 26.38s |
| Max | 57 | $0.51 | 185.76s |
Terra High is one of the most practical configurations in the dataset. It reaches Intelligence 50 while keeping first-chunk latency around 2.34 seconds.
Compare that with Luna xHigh:
| Configuration | Intelligence | Cost | First Chunk |
|---|---|---|---|
| Luna xHigh | 50 | $0.03 | 58.90s |
| Terra High | 50 | $0.22 | 2.34s |
The intelligence score is identical. Luna is much cheaper, but Terra starts responding roughly 25 times faster.
For interactive AI products, latency itself has business value.
Sol: strong baseline intelligence
| Sol Mode | Intelligence | Cost / Task | First Chunk |
|---|---|---|---|
| None | 42 | $0.24 | 1.32s |
| Low | 51 | $0.23 | 2.83s |
| Medium | 56 | $0.37 | 4.61s |
| High | 57 | $0.55 | 11.15s |
| xHigh | 59 | $0.81 | 35.87s |
| Max | 61 | $1.23 | 146.36s |
One of the most interesting configurations is Sol Low.
| Configuration | Intelligence | Cost | First Chunk |
|---|---|---|---|
| Terra High | 50 | $0.22 | 2.34s |
| Sol Low | 51 | $0.23 | 2.83s |
The cost is almost identical, latency is similar, and Sol Low is slightly stronger on the Intelligence Index. That makes Terra High vs Sol Low a particularly important real-workload A/B test rather than an obvious one-sided choice.
Cost versus intelligence

Cost alone does not tell us which model is best. The same Intelligence range can often be reached with very different latency profiles.
A low-cost configuration can be ideal for offline or asynchronous work while still being a poor choice for a live chat interface.
First-chunk latency versus intelligence

The latency chart is important because it reveals where heavy reasoning stops being practical for interactive UX.
For example, moving from Terra High to Terra xHigh adds only three Intelligence points, but first-chunk latency rises from 2.34 seconds to 26.38 seconds.
Sol Medium may be the Pro sweet spot
Sol Medium reaches Intelligence 56 with approximately 4.61 seconds of first-chunk latency.
Moving to Sol High changes:
- Intelligence: 56 → 57
- Cost: $0.37 → $0.55
- First chunk: 4.61s → 11.15s
That is a significant operational increase for only one additional Intelligence Index point.
For many production workloads, Sol Medium looks like a strong default Pro configuration, while Sol High can be reserved for tasks that genuinely justify deeper reasoning.
Terra Max vs Sol High: same intelligence, completely different experience
| Configuration | Intelligence | Cost | First Chunk | Total Response |
|---|---|---|---|---|
| Terra Max | 57 | $0.51 | 185.76s | 189.84s |
| Sol High | 57 | $0.55 | 11.15s | 18.59s |
Both score 57 and their benchmark cost is similar, but Sol High starts responding roughly 16.7 times faster and completes the response more than 10 times faster.
This is a strong reminder that an Intelligence Index should never be used alone when selecting a production model.
The real decision should consider:
Capability + Cost + Latency + Throughput + Workload
A better AI model router
Instead of routing purely by family:
Fast → Luna
Standard → Terra
Pro → Sol
a better architecture can route by required capability:
| Routing Tier | Suggested Configuration | Intelligence |
|---|---|---|
| Fast | Luna None / Low | 27–34 |
| Normal | Luna Medium | 39 |
| Standard | Terra Low / Medium | 41–47 |
| Advanced | Terra High | 50 |
| Pro Fast | Sol Low | 51 |
| Pro | Sol Medium | 56 |
| Expert | Sol High | 57 |
| Extreme | Sol xHigh / Max | 59–61 |
The router can also consider task complexity, latency requirements, budget, whether the user is waiting, whether the task is interactive or asynchronous, and whether a lower tier already failed.
This creates a continuous capability ladder instead of three fixed model buckets.
Capability is becoming dynamic
One broader lesson from GPT-5.6 is that model capability is increasingly dynamic.
A useful mental model is:
Effective Capability =
Model
+ Reasoning Effort
+ Prompt
+ Tools
+ Agent Harness
+ Context
The same model can occupy very different capability tiers depending on how much reasoning it is allowed to perform and what runtime surrounds it.
So the better question is no longer simply:
Which model should we use?
It is:
What is the cheapest and fastest configuration that can reliably solve this task?
Current sweet spots
Based on this benchmark snapshot, the configurations I would investigate first are:
- Luna Medium — inexpensive general-purpose tasks
- Luna High — very cost-sensitive background reasoning
- Terra High — fast interactive advanced workloads
- Sol Low — a strong alternative to Terra High
- Sol Medium — strong default Pro tier
- Sol High — harder reasoning where the extra latency is justified
xHigh and Max remain useful, but their latency suggests they should be treated primarily as escalation modes rather than defaults.
Final takeaway
GPT-5.6 Luna, Terra and Sol should not be treated as three simple fixed capability levels. Together with reasoning effort, they form a much larger capability spectrum.
A Luna configuration with heavy reasoning can sometimes reach the same intelligence region as a Terra or Sol configuration with lighter reasoning, but the cost and latency can be completely different.
The production lesson is simple:
Do not route by model name alone. Route by required capability, latency and cost.
The future of AI model selection may look less like choosing one model and more like dynamically purchasing exactly the amount of intelligence required for each task.
Data note: Benchmark figures in this article were transcribed from Artificial Analysis leaderboard screenshots captured on 24 August 2026. Benchmark scores, provider performance, latency and pricing can change over time. Treat these figures as a point-in-time comparison rather than permanent model specifications.