Cost and quality

Efficiency 1.1

Compare question quality with the recorded cost of the model under test.

Models
11
Guesser cost range
176×
Pareto-efficient
5 of 11
Direction
Lower is better

Official efficiency ranking

Distance from the lower-left ideal.

Question score and Guesser cost are each normalized from 0 to 1 across this cohort. The ranking measures equal-weight distance from their combined minimum. Lower is better.

Normalized distance · lower is betterSelect a model row to view its full run
  1. 1. Gemini 3.7 Flash (high): 0.263. 0.187 normalized questions · 0.185 normalized cost. View full run for Gemini 3.7 Flash (high)
  2. 2. Gemini 3.8 Flash (high): 0.321. 0.077 normalized questions · 0.312 normalized cost. View full run for Gemini 3.8 Flash (high)
  3. 3. GPT-5.6 Sol (high): 0.345. 0.284 normalized questions · 0.197 normalized cost. View full run for GPT-5.6 Sol (high)
  4. 4. Claude Opus 5 (high): 0.438. 0.284 normalized questions · 0.333 normalized cost. View full run for Claude Opus 5 (high)
  5. 5. GPT-5.6 Luna (high): 0.476. 0.475 normalized questions · 0.017 normalized cost. View full run for GPT-5.6 Luna (high)
  6. 6. Grok 4.6 (high): 0.539. 0.224 normalized questions · 0.490 normalized cost. View full run for Grok 4.6 (high)
  7. 7. gpt-oss-120B (high): 0.603. 0.553 normalized questions · 0.242 normalized cost. View full run for gpt-oss-120B (high)
  8. 8. GLM-5.3-Flash (high): 0.619. 0.619 normalized questions · 0.000 normalized cost. View full run for GLM-5.3-Flash (high)
  9. 9. GPT-6 Astra (high): 0.631. 0.000 normalized questions · 0.631 normalized cost. View full run for GPT-6 Astra (high)
  10. 10. MiniMax M3 (high): 1.000. 1.000 normalized questions · 0.031 normalized cost. View full run for MiniMax M3 (high)
  11. 11. Claude Fable 5.1 (high): 1.004. 0.092 normalized questions · 1.000 normalized cost. View full run for Claude Fable 5.1 (high)

Trade-off map

Normalized cost and question score.

Both axes use the same 0-to-1 scale. Dashed curves mark equal distance from the lower-left ideal.

Diamond Pareto-efficient - no other model is both cheaper and better Circle Other ranked model
Lower-left is betterCurves show equal ideal distance
  1. Gemini 3.7 Flash (high): 16.57 questions, $0.0525 Guesser cost per episode, ideal distance 0.263, rank 1. Pareto-efficient. View full run for Gemini 3.7 Flash (high)
  2. Gemini 3.8 Flash (high): 14.87 questions, $0.0875 Guesser cost per episode, ideal distance 0.321, rank 2. Pareto-efficient. View full run for Gemini 3.8 Flash (high)
  3. GPT-5.6 Sol (high): 18.07 questions, $0.0557 Guesser cost per episode, ideal distance 0.345, rank 3. View full run for GPT-5.6 Sol (high)
  4. Claude Opus 5 (high): 18.07 questions, $0.0934 Guesser cost per episode, ideal distance 0.438, rank 4. View full run for Claude Opus 5 (high)
  5. GPT-5.6 Luna (high): 21.03 questions, $0.0063 Guesser cost per episode, ideal distance 0.476, rank 5. Pareto-efficient. View full run for GPT-5.6 Luna (high)
  6. Grok 4.6 (high): 17.13 questions, $0.1367 Guesser cost per episode, ideal distance 0.539, rank 6. View full run for Grok 4.6 (high)
  7. gpt-oss-120B (high): 22.23 questions, $0.0682 Guesser cost per episode, ideal distance 0.603, rank 7. View full run for gpt-oss-120B (high)
  8. GLM-5.3-Flash (high): 23.27 questions, $0.0016 Guesser cost per episode, ideal distance 0.619, rank 8. Pareto-efficient. View full run for GLM-5.3-Flash (high)
  9. GPT-6 Astra (high): 13.67 questions, $0.1756 Guesser cost per episode, ideal distance 0.631, rank 9. Pareto-efficient. View full run for GPT-6 Astra (high)
  10. MiniMax M3 (high): 29.17 questions, $0.0100 Guesser cost per episode, ideal distance 1.000, rank 10. View full run for MiniMax M3 (high)
  11. Claude Fable 5.1 (high): 15.10 questions, $0.2772 Guesser cost per episode, ideal distance 1.004, rank 11. View full run for Claude Fable 5.1 (high)

Expanded trade-off map

Normalized cost and question score.

Lower-left is better. Curves show equal distance from the ideal.

Diamond Pareto-efficient - no other model is both cheaper and better Circle Other ranked model
Efficiency ranking
Ideal-distance rankModelIdealdistanceParetoefficientQuestionrankQuestionscoreGuesser costper episodeSuccess
1Gemini 3.7 Flash (high)google-ai-studio0.263 Yes 416.57$0.052593%
2Gemini 3.8 Flash (high)google-ai-studio0.321 Yes 214.87$0.0875100%
3GPT-5.6 Sol (high)openai0.345618.07$0.055797%
4Claude Opus 5 (high)anthropic0.438618.07$0.093493%
5GPT-5.6 Luna (high)openai0.476 Yes 821.03$0.006390%
6Grok 4.6 (high)xai0.539517.13$0.1367100%
7gpt-oss-120B (high)cerebras0.603922.23$0.068267%
8GLM-5.3-Flash (high)z-ai0.619 Yes 1023.27$0.001673%
9GPT-6 Astra (high)openai0.631 Yes 113.67$0.1756100%
10MiniMax M3 (high)coreweave1.0001129.17$0.010050%
11Claude Fable 5.1 (high)anthropic1.004315.10$0.277293%
#1Gemini 3.7 Flash (high)google-ai-studio
Ideal distance
0.263
Pareto-efficient
Yes
Question score
16.57
Guesser cost / episode
$0.0525
Explore full run · questions, answers & evidence
#2Gemini 3.8 Flash (high)google-ai-studio
Ideal distance
0.321
Pareto-efficient
Yes
Question score
14.87
Guesser cost / episode
$0.0875
Explore full run · questions, answers & evidence
#3GPT-5.6 Sol (high)openai
Ideal distance
0.345
Pareto-efficient
No
Question score
18.07
Guesser cost / episode
$0.0557
Explore full run · questions, answers & evidence
#4Claude Opus 5 (high)anthropic
Ideal distance
0.438
Pareto-efficient
No
Question score
18.07
Guesser cost / episode
$0.0934
Explore full run · questions, answers & evidence
#5GPT-5.6 Luna (high)openai
Ideal distance
0.476
Pareto-efficient
Yes
Question score
21.03
Guesser cost / episode
$0.0063
Explore full run · questions, answers & evidence
#6Grok 4.6 (high)xai
Ideal distance
0.539
Pareto-efficient
No
Question score
17.13
Guesser cost / episode
$0.1367
Explore full run · questions, answers & evidence
#7gpt-oss-120B (high)cerebras
Ideal distance
0.603
Pareto-efficient
No
Question score
22.23
Guesser cost / episode
$0.0682
Explore full run · questions, answers & evidence
#8GLM-5.3-Flash (high)z-ai
Ideal distance
0.619
Pareto-efficient
Yes
Question score
23.27
Guesser cost / episode
$0.0016
Explore full run · questions, answers & evidence
#9GPT-6 Astra (high)openai
Ideal distance
0.631
Pareto-efficient
Yes
Question score
13.67
Guesser cost / episode
$0.1756
Explore full run · questions, answers & evidence
#10MiniMax M3 (high)coreweave
Ideal distance
1.000
Pareto-efficient
No
Question score
29.17
Guesser cost / episode
$0.0100
Explore full run · questions, answers & evidence
#11Claude Fable 5.1 (high)anthropic
Ideal distance
1.004
Pareto-efficient
No
Question score
15.10
Guesser cost / episode
$0.2772
Explore full run · questions, answers & evidence

How it is calculated

Normalized ideal distance.

Both measures have equal weight after cohort min/max normalization. Lower is better.

Formula√(normalized question score² + normalized Guesser cost²)
Calculation detailsSteps and limits
  1. Normalize question score as (value − cohort minimum) ÷ cohort range.
  2. Normalize Guesser cost per episode with the same calculation.
  3. Measure Euclidean distance from (0, 0). Lower is better.

A model with normalized question score 0.06 and normalized cost 0.08 has distance √(0.06² + 0.08²) = 0.10.

Question score still uses the average penalized trial values. Failed trials therefore remain part of the quality dimension.

Adding or removing a model can change every normalized value and rank.