← Back to all posts

8 min readUpdated 24 Aug 2026
GPT-5.6OpenAILunaTerraSol+5 more

GPT-5.6 Luna vs Terra vs Sol: Intelligence, Cost, Latency and the Real Sweet Spots

Yap Wei JunSoftware Engineer · Singapore

OpenAI's GPT-5.6 family can be viewed as three capability tiers:

  • Luna — efficiency and high-volume workloads
  • Terra — balanced capability, latency and cost
  • Sol — highest overall intelligence

But comparing only the model name is no longer enough. Each family can operate at different reasoning-effort levels, so the more useful unit of comparison is:

Model + Reasoning Effort

Using the Artificial Analysis benchmark data captured on 24 August 2026, the differences become much clearer.

Intelligence Index by reasoning effort

Reasoning EffortGPT-5.6 LunaGPT-5.6 TerraGPT-5.6 Sol
Non-reasoning273542
Low344151
Medium394756
High475057
xHigh505359
Max525761

GPT-5.6 Intelligence Index by reasoning effort

The basic ranking is consistent: Sol > Terra > Luna. But the interesting part appears when we compare different configurations across families.

For example:

  • Luna Max: 52
  • Terra xHigh: 53
  • Sol Low: 51

These configurations occupy roughly the same intelligence band, but their cost and latency are very different.

Intelligence Index is not a percentage

The Artificial Analysis Intelligence Index should not be read as a percentage. A score of 60 does not mean a model is "60% intelligent", and a score of 60 is not twice as intelligent as 30.

It is a composite benchmark signal designed to compare model capability across reasoning, coding, agentic and related tasks. The useful information is the relative positioning between configurations.

Luna: cheap intelligence

Luna ModeIntelligenceCost / TaskFirst Chunk
None27$0.010.72s
Low34$0.011.61s
Medium39$0.012.11s
High47$0.0212.84s
xHigh50$0.0358.90s
Max52$0.05146.22s

Luna's biggest advantage is cost efficiency. Luna High reaches Intelligence 47 for about $0.02 per Artificial Analysis benchmark task.

The trade-off is latency. Moving from High to xHigh raises Intelligence from 47 to 50, but first-chunk latency jumps from about 13 seconds to nearly 59 seconds. Max reaches 52, but the first response takes roughly 146 seconds.

That is a classic case of diminishing returns.

Luna therefore looks especially attractive for background agents, batch processing, document analysis, asynchronous research and other workloads where a user is not waiting for an immediate response.

Terra: the balanced tier

Terra ModeIntelligenceCost / TaskFirst Chunk
None35$0.100.79s
Low41$0.091.29s
Medium47$0.121.78s
High50$0.222.34s
xHigh53$0.3126.38s
Max57$0.51185.76s

Terra High is one of the most practical configurations in the dataset. It reaches Intelligence 50 while keeping first-chunk latency around 2.34 seconds.

Compare that with Luna xHigh:

ConfigurationIntelligenceCostFirst Chunk
Luna xHigh50$0.0358.90s
Terra High50$0.222.34s

The intelligence score is identical. Luna is much cheaper, but Terra starts responding roughly 25 times faster.

For interactive AI products, latency itself has business value.

Sol: strong baseline intelligence

Sol ModeIntelligenceCost / TaskFirst Chunk
None42$0.241.32s
Low51$0.232.83s
Medium56$0.374.61s
High57$0.5511.15s
xHigh59$0.8135.87s
Max61$1.23146.36s

One of the most interesting configurations is Sol Low.

ConfigurationIntelligenceCostFirst Chunk
Terra High50$0.222.34s
Sol Low51$0.232.83s

The cost is almost identical, latency is similar, and Sol Low is slightly stronger on the Intelligence Index. That makes Terra High vs Sol Low a particularly important real-workload A/B test rather than an obvious one-sided choice.

Cost versus intelligence

GPT-5.6 Cost per Task vs Intelligence

Cost alone does not tell us which model is best. The same Intelligence range can often be reached with very different latency profiles.

A low-cost configuration can be ideal for offline or asynchronous work while still being a poor choice for a live chat interface.

First-chunk latency versus intelligence

GPT-5.6 First-Chunk Latency vs Intelligence

The latency chart is important because it reveals where heavy reasoning stops being practical for interactive UX.

For example, moving from Terra High to Terra xHigh adds only three Intelligence points, but first-chunk latency rises from 2.34 seconds to 26.38 seconds.

Sol Medium may be the Pro sweet spot

Sol Medium reaches Intelligence 56 with approximately 4.61 seconds of first-chunk latency.

Moving to Sol High changes:

  • Intelligence: 56 → 57
  • Cost: $0.37 → $0.55
  • First chunk: 4.61s → 11.15s

That is a significant operational increase for only one additional Intelligence Index point.

For many production workloads, Sol Medium looks like a strong default Pro configuration, while Sol High can be reserved for tasks that genuinely justify deeper reasoning.

Terra Max vs Sol High: same intelligence, completely different experience

ConfigurationIntelligenceCostFirst ChunkTotal Response
Terra Max57$0.51185.76s189.84s
Sol High57$0.5511.15s18.59s

Both score 57 and their benchmark cost is similar, but Sol High starts responding roughly 16.7 times faster and completes the response more than 10 times faster.

This is a strong reminder that an Intelligence Index should never be used alone when selecting a production model.

The real decision should consider:

Capability + Cost + Latency + Throughput + Workload

A better AI model router

Instead of routing purely by family:

Fast     → Luna
Standard → Terra
Pro      → Sol

a better architecture can route by required capability:

Routing TierSuggested ConfigurationIntelligence
FastLuna None / Low27–34
NormalLuna Medium39
StandardTerra Low / Medium41–47
AdvancedTerra High50
Pro FastSol Low51
ProSol Medium56
ExpertSol High57
ExtremeSol xHigh / Max59–61

The router can also consider task complexity, latency requirements, budget, whether the user is waiting, whether the task is interactive or asynchronous, and whether a lower tier already failed.

This creates a continuous capability ladder instead of three fixed model buckets.

Capability is becoming dynamic

One broader lesson from GPT-5.6 is that model capability is increasingly dynamic.

A useful mental model is:

Effective Capability =
Model
+ Reasoning Effort
+ Prompt
+ Tools
+ Agent Harness
+ Context

The same model can occupy very different capability tiers depending on how much reasoning it is allowed to perform and what runtime surrounds it.

So the better question is no longer simply:

Which model should we use?

It is:

What is the cheapest and fastest configuration that can reliably solve this task?

Current sweet spots

Based on this benchmark snapshot, the configurations I would investigate first are:

  • Luna Medium — inexpensive general-purpose tasks
  • Luna High — very cost-sensitive background reasoning
  • Terra High — fast interactive advanced workloads
  • Sol Low — a strong alternative to Terra High
  • Sol Medium — strong default Pro tier
  • Sol High — harder reasoning where the extra latency is justified

xHigh and Max remain useful, but their latency suggests they should be treated primarily as escalation modes rather than defaults.

Final takeaway

GPT-5.6 Luna, Terra and Sol should not be treated as three simple fixed capability levels. Together with reasoning effort, they form a much larger capability spectrum.

A Luna configuration with heavy reasoning can sometimes reach the same intelligence region as a Terra or Sol configuration with lighter reasoning, but the cost and latency can be completely different.

The production lesson is simple:

Do not route by model name alone. Route by required capability, latency and cost.

The future of AI model selection may look less like choosing one model and more like dynamically purchasing exactly the amount of intelligence required for each task.


Data note: Benchmark figures in this article were transcribed from Artificial Analysis leaderboard screenshots captured on 24 August 2026. Benchmark scores, provider performance, latency and pricing can change over time. Treat these figures as a point-in-time comparison rather than permanent model specifications.

← Back to all posts