Cost and quality

Efficiency 1

Compare question quality with the recorded cost of the model under test.

Models
17
Guesser cost range
59×
Pareto-efficient
6 of 17
Direction
Lower is better

Official efficiency ranking

Distance from the lower-left ideal.

Question score and Guesser cost are each normalized from 0 to 1 across this cohort. The ranking measures equal-weight distance from their combined minimum. Lower is better.

Normalized distance · lower is betterSelect a model row to view its full run
  1. 1. gpt-oss-120B (high): 0.096. 0.078 normalized questions · 0.056 normalized cost. View full run for gpt-oss-120B (high)
  2. 2. Claude Opus 5 (high): 0.127. 0.014 normalized questions · 0.126 normalized cost. View full run for Claude Opus 5 (high)
  3. 3. Gemini 3.7 Flash (high): 0.151. 0.098 normalized questions · 0.115 normalized cost. View full run for Gemini 3.7 Flash (high)
  4. 4. Claude Sonnet 5 (high): 0.157. 0.133 normalized questions · 0.083 normalized cost. View full run for Claude Sonnet 5 (high)
  5. 5. Grok 4.6 (high): 0.160. 0.110 normalized questions · 0.115 normalized cost. View full run for Grok 4.6 (high)
  6. 6. Grok 4.5 (high): 0.160. 0.154 normalized questions · 0.041 normalized cost. View full run for Grok 4.5 (high)
  7. 7. GPT-5 Nano (medium): 0.170. 0.170 normalized questions · 0.011 normalized cost. View full run for GPT-5 Nano (medium)
  8. 8. Gemini 3.8 Flash (high): 0.218. 0.054 normalized questions · 0.212 normalized cost. View full run for Gemini 3.8 Flash (high)
  9. 9. GPT-5.6 Luna (high): 0.283. 0.282 normalized questions · 0.020 normalized cost. View full run for GPT-5.6 Luna (high)
  10. 10. Claude Fable 5 (high): 0.343. 0.000 normalized questions · 0.343 normalized cost. View full run for Claude Fable 5 (high)
  11. 11. GPT-5.6 Sol (high): 0.350. 0.116 normalized questions · 0.330 normalized cost. View full run for GPT-5.6 Sol (high)
  12. 12. Kimi K3 (high): 0.396. 0.034 normalized questions · 0.395 normalized cost. View full run for Kimi K3 (high)
  13. 13. Gemini 3.6 Flash (high): 0.407. 0.239 normalized questions · 0.329 normalized cost. View full run for Gemini 3.6 Flash (high)
  14. 14. GPT-6 Astra (high): 0.665. 0.150 normalized questions · 0.648 normalized cost. View full run for GPT-6 Astra (high)
  15. 15. Mistral Medium 3.5 (high): 0.908. 0.397 normalized questions · 0.817 normalized cost. View full run for Mistral Medium 3.5 (high)
  16. 16. Llama 4 Maverick (non-thinking): 1.000. 1.000 normalized questions · 0.000 normalized cost. View full run for Llama 4 Maverick (non-thinking)
  17. 17. Claude Fable 5.1 (high): 1.000. 0.003 normalized questions · 1.000 normalized cost. View full run for Claude Fable 5.1 (high)

Trade-off map

Normalized cost and question score.

Both axes use the same 0-to-1 scale. Dashed curves mark equal distance from the lower-left ideal.

Diamond Pareto-efficient - no other model is both cheaper and better Circle Other ranked model
Lower-left is betterCurves show equal ideal distance
  1. gpt-oss-120B (high): 13.63 questions, $0.0363 Guesser cost per episode, ideal distance 0.096, rank 1. Pareto-efficient. View full run for gpt-oss-120B (high)
  2. Claude Opus 5 (high): 12.34 questions, $0.0705 Guesser cost per episode, ideal distance 0.127, rank 2. Pareto-efficient. View full run for Claude Opus 5 (high)
  3. Gemini 3.7 Flash (high): 14.03 questions, $0.0653 Guesser cost per episode, ideal distance 0.151, rank 3. View full run for Gemini 3.7 Flash (high)
  4. Claude Sonnet 5 (high): 14.74 questions, $0.0493 Guesser cost per episode, ideal distance 0.157, rank 4. View full run for Claude Sonnet 5 (high)
  5. Grok 4.6 (high): 14.29 questions, $0.0653 Guesser cost per episode, ideal distance 0.160, rank 5. View full run for Grok 4.6 (high)
  6. Grok 4.5 (high): 15.17 questions, $0.0286 Guesser cost per episode, ideal distance 0.160, rank 6. Pareto-efficient. View full run for Grok 4.5 (high)
  7. GPT-5 Nano (medium): 15.49 questions, $0.0138 Guesser cost per episode, ideal distance 0.170, rank 7. Pareto-efficient. View full run for GPT-5 Nano (medium)
  8. Gemini 3.8 Flash (high): 13.14 questions, $0.1128 Guesser cost per episode, ideal distance 0.218, rank 8. View full run for Gemini 3.8 Flash (high)
  9. GPT-5.6 Luna (high): 17.74 questions, $0.0184 Guesser cost per episode, ideal distance 0.283, rank 9. View full run for GPT-5.6 Luna (high)
  10. Claude Fable 5 (high): 12.06 questions, $0.1776 Guesser cost per episode, ideal distance 0.343, rank 10. Pareto-efficient. View full run for Claude Fable 5 (high)
  11. GPT-5.6 Sol (high): 14.40 questions, $0.1711 Guesser cost per episode, ideal distance 0.350, rank 11. View full run for GPT-5.6 Sol (high)
  12. Kimi K3 (high): 12.74 questions, $0.2032 Guesser cost per episode, ideal distance 0.396, rank 12. View full run for Kimi K3 (high)
  13. Gemini 3.6 Flash (high): 16.89 questions, $0.1705 Guesser cost per episode, ideal distance 0.407, rank 13. View full run for Gemini 3.6 Flash (high)
  14. GPT-6 Astra (high): 15.09 questions, $0.3279 Guesser cost per episode, ideal distance 0.665, rank 14. View full run for GPT-6 Astra (high)
  15. Mistral Medium 3.5 (high): 20.06 questions, $0.4113 Guesser cost per episode, ideal distance 0.908, rank 15. View full run for Mistral Medium 3.5 (high)
  16. Llama 4 Maverick (non-thinking): 32.23 questions, $0.0086 Guesser cost per episode, ideal distance 1.000, rank 16. Pareto-efficient. View full run for Llama 4 Maverick (non-thinking)
  17. Claude Fable 5.1 (high): 12.11 questions, $0.5014 Guesser cost per episode, ideal distance 1.000, rank 17. View full run for Claude Fable 5.1 (high)

Expanded trade-off map

Normalized cost and question score.

Lower-left is better. Curves show equal distance from the ideal.

Diamond Pareto-efficient - no other model is both cheaper and better Circle Other ranked model
Efficiency ranking
Ideal-distance rankModelIdealdistanceParetoefficientQuestionrankQuestionscoreGuesser costper episodeSuccess
1gpt-oss-120B (high)cerebras0.096 Yes 613.63$0.036394%
2Claude Opus 5 (high)anthropic0.127 Yes 312.34$0.0705100%
3Gemini 3.7 Flash (high)google-ai-studio0.151714.03$0.065397%
4Claude Sonnet 5 (high)anthropic0.1571014.74$0.0493100%
5Grok 4.6 (high)xai0.160814.29$0.0653100%
6Grok 4.5 (high)xai0.160 Yes 1215.17$0.028694%
7GPT-5 Nano (medium)openai0.170 Yes 1315.49$0.013894%
8Gemini 3.8 Flash (high)google-ai-studio0.218513.14$0.1128100%
9GPT-5.6 Luna (high)openai0.2831617.74$0.018491%
10Claude Fable 5 (high)anthropic0.343 Yes 112.06$0.1776100%
11GPT-5.6 Sol (high)openai0.350914.40$0.1711100%
12Kimi K3 (high)moonshotai0.396412.74$0.203294%
13Gemini 3.6 Flash (high)google-vertex0.4071416.89$0.170597%
14GPT-6 Astra (high)openai0.6651115.09$0.327994%
15Mistral Medium 3.5 (high)mistral0.9081720.06$0.411389%
16Llama 4 Maverick (non-thinking)parasail1.000 Yes 1832.23$0.008649%
17Claude Fable 5.1 (high)anthropic1.000212.11$0.501497%
#1gpt-oss-120B (high)cerebras
Ideal distance
0.096
Pareto-efficient
Yes
Question score
13.63
Guesser cost / episode
$0.0363
Explore full run · questions, answers & evidence
#2Claude Opus 5 (high)anthropic
Ideal distance
0.127
Pareto-efficient
Yes
Question score
12.34
Guesser cost / episode
$0.0705
Explore full run · questions, answers & evidence
#3Gemini 3.7 Flash (high)google-ai-studio
Ideal distance
0.151
Pareto-efficient
No
Question score
14.03
Guesser cost / episode
$0.0653
Explore full run · questions, answers & evidence
#4Claude Sonnet 5 (high)anthropic
Ideal distance
0.157
Pareto-efficient
No
Question score
14.74
Guesser cost / episode
$0.0493
Explore full run · questions, answers & evidence
#5Grok 4.6 (high)xai
Ideal distance
0.160
Pareto-efficient
No
Question score
14.29
Guesser cost / episode
$0.0653
Explore full run · questions, answers & evidence
#6Grok 4.5 (high)xai
Ideal distance
0.160
Pareto-efficient
Yes
Question score
15.17
Guesser cost / episode
$0.0286
Explore full run · questions, answers & evidence
#7GPT-5 Nano (medium)openai
Ideal distance
0.170
Pareto-efficient
Yes
Question score
15.49
Guesser cost / episode
$0.0138
Explore full run · questions, answers & evidence
#8Gemini 3.8 Flash (high)google-ai-studio
Ideal distance
0.218
Pareto-efficient
No
Question score
13.14
Guesser cost / episode
$0.1128
Explore full run · questions, answers & evidence
#9GPT-5.6 Luna (high)openai
Ideal distance
0.283
Pareto-efficient
No
Question score
17.74
Guesser cost / episode
$0.0184
Explore full run · questions, answers & evidence
#10Claude Fable 5 (high)anthropic
Ideal distance
0.343
Pareto-efficient
Yes
Question score
12.06
Guesser cost / episode
$0.1776
Explore full run · questions, answers & evidence
#11GPT-5.6 Sol (high)openai
Ideal distance
0.350
Pareto-efficient
No
Question score
14.40
Guesser cost / episode
$0.1711
Explore full run · questions, answers & evidence
#12Kimi K3 (high)moonshotai
Ideal distance
0.396
Pareto-efficient
No
Question score
12.74
Guesser cost / episode
$0.2032
Explore full run · questions, answers & evidence
#13Gemini 3.6 Flash (high)google-vertex
Ideal distance
0.407
Pareto-efficient
No
Question score
16.89
Guesser cost / episode
$0.1705
Explore full run · questions, answers & evidence
#14GPT-6 Astra (high)openai
Ideal distance
0.665
Pareto-efficient
No
Question score
15.09
Guesser cost / episode
$0.3279
Explore full run · questions, answers & evidence
#15Mistral Medium 3.5 (high)mistral
Ideal distance
0.908
Pareto-efficient
No
Question score
20.06
Guesser cost / episode
$0.4113
Explore full run · questions, answers & evidence
#16Llama 4 Maverick (non-thinking)parasail
Ideal distance
1.000
Pareto-efficient
Yes
Question score
32.23
Guesser cost / episode
$0.0086
Explore full run · questions, answers & evidence
#17Claude Fable 5.1 (high)anthropic
Ideal distance
1.000
Pareto-efficient
No
Question score
12.11
Guesser cost / episode
$0.5014
Explore full run · questions, answers & evidence

1 model is not ranked because a question score or positive recorded Guesser cost per completed episode is unavailable.

How it is calculated

Normalized ideal distance.

Both measures have equal weight after cohort min/max normalization. Lower is better.

Formula√(normalized question score² + normalized Guesser cost²)
Calculation detailsSteps and limits
  1. Normalize question score as (value − cohort minimum) ÷ cohort range.
  2. Normalize Guesser cost per episode with the same calculation.
  3. Measure Euclidean distance from (0, 0). Lower is better.

A model with normalized question score 0.06 and normalized cost 0.08 has distance √(0.06² + 0.08²) = 0.10.

Question score still uses the average penalized trial values. Failed trials therefore remain part of the quality dimension.

Adding or removing a model can change every normalized value and rank.