Effort-Based Intelligence Index: A Trusted Way to Compare AI Models
Listen to this article
AI-generated voice4:40Generated 27 Aug 2026
When comparing AI models, many people show only the highest available score. For example, Claude Opus 5 scores 63, GPT-5.6 Sol scores 61, and Gemini 3.7 Flash scores 56.
This looks simple, but it can be misleading. Modern reasoning models support different effort levels. Low effort is usually cheaper and faster, while Max effort may produce stronger reasoning but cost more and take longer.
A fair comparison should show the complete effort curve.
Data below is based on Artificial Analysis Intelligence Index v4.1.1 and was checked on 27 August 2026.
Effort-based Intelligence Index
| Model family | Non-reasoning | Low | Medium | High | XHigh | Max |
|---|---|---|---|---|---|---|
| Claude Opus 5 | — | 52 | 59 | 61 | 63 | 63 |
| Claude Sonnet 5 | 43 | — | — | — | — | 55 |
| Claude 4.5 Haiku | 24* | — | — | — | — | 30 |
| GPT-5.6 Luna | 27 | 34 | 39 | 47 | 50 | 52 |
| GPT-5.6 Terra | 35 | 41 | 47 | 50 | 53 | 57 |
| GPT-5.6 Sol | 42 | 51 | 56 | 57 | 59 | 61 |
| Gemini 3.7 Flash | — | 51 | 53 | 56 | — | — |
An asterisk means the score is estimated by Artificial Analysis. A dash means that no published score is available. A missing score must not be treated as zero.
Sources: Claude Opus 5, Claude Sonnet 5, GPT-5.6 Luna, GPT-5.6 Terra, GPT-5.6 Sol, and Gemini 3.7 Flash.
What changes when effort changes?
Claude Opus 5 increases from 52 at Low effort to 63 at Max effort. This shows a clear quality improvement, but the higher setting also uses more reasoning and costs more.
GPT-5.6 Sol reaches 56 at Medium, 57 at High, 59 at XHigh and 61 at Max. Sol remains strong across several settings, which is useful for coding agents that need to balance quality, cost and execution time.
GPT-5.6 Terra reaches 50 at High and 57 at Max. It is slightly weaker than Sol at the top end, but it is a strong candidate for normal production workloads.
GPT-5.6 Luna reaches 47 at High and 52 at Max. It is not the strongest model, but it is useful for classification, routing, short summaries and low-cost automation.
Gemini 3.7 Flash reaches 51 at Low, 53 at Medium and 56 at High. Its major advantage is speed, with Artificial Analysis measuring approximately 362 output tokens per second at Medium and High effort.
Which model is best?
Claude Opus 5 Max is the quality leader. It is best for difficult architecture, complex debugging, final review and high-risk decisions.
GPT-5.6 Sol is the strongest choice for serious software engineering, PostgreSQL, CFML, MCP tools and agent orchestration.
GPT-5.6 Terra is the best balanced production model for standard ERP coding, SQL analysis, business workflows and customer summaries.
Gemini 3.7 Flash is the speed specialist for high-volume document processing, OCR cleanup, summaries and interactive applications.
GPT-5.6 Luna is the economical routing model for intent classification, skill selection, short explanations and first-pass validation.
Claude Sonnet 5 is capable, but only its Max effort configuration has a published reasoning score in this table. Its Low, Medium, High and XHigh scores should not be guessed.
Artificial Analysis currently lists Claude 4.5 Haiku, not Claude Haiku 5. Claude 4.5 Haiku Reasoning scores 30, while its non-reasoning score of 24 is estimated. Therefore, it would be inaccurate to describe this as Claude Haiku 5.
How to display trusted scores
A model dashboard should show four fields:
| Field | Example |
|---|---|
| Score | Intelligence Index: 50 |
| Effort | High |
| Evidence | Measured |
| Benchmark | Artificial Analysis v4.1.1 |
Use these evidence labels:
- Measured: a published evaluation result exists
- Estimated: the source marks the result with an asterisk
- Unavailable: no published result exists
- Versioned: the benchmark version is shown
Do not invent a trust percentage. Do not convert unavailable data into zero. Do not fill missing effort levels by guessing.
Important fairness rule
Effort names are not perfectly standardized across providers. Claude Max, OpenAI Max and Gemini High do not necessarily use the same amount of reasoning or compute.
A fair comparison should show the model family, effort level, Intelligence Index, cost per task, output speed, latency, evidence status and benchmark version.
Artificial Analysis Intelligence Index v4.1.1 combines nine evaluations. Its category weighting is Agents 34%, Coding 24%, Scientific Reasoning 24% and General 18%.
The index is useful for broad comparisons, but it is not an absolute IQ score. It is also mainly an English and text-based evaluation suite. Chinese ability, OCR, vision and enterprise ERP performance should be evaluated separately.
Read the full Artificial Analysis methodology.
Recommended routing for an Agent Brain
| Route | Model | Purpose |
|---|---|---|
| Fast | Luna Medium or High | Classification and simple requests |
| Fast-plus | Gemini Flash Medium or High | Fast extraction and summaries |
| Standard | Terra High | Daily ERP and business tasks |
| Pro | Sol High or XHigh | Difficult coding and reasoning |
| Expert | Opus XHigh or Max | Critical review and complex decisions |
The key principle is simple:
Use the lowest effort that safely completes the task.
For an enterprise Agent Brain, the best architecture is not one model for everything. It is a layered system:
Luna → Gemini Flash or Terra → Sol → Opus
This reduces cost for simple work while preserving stronger reasoning for tasks that genuinely require it.
Conclusion
The most useful question is not “Which AI model has the highest score?”
The better question is:
Which model reaches the required intelligence level at an acceptable cost and latency, with published evidence?
Effort-based reporting makes AI model comparisons clearer, fairer and more trustworthy.