Official comparison

Results 1.1

Question score, stability, cost, and time for the current official runs.

Models
11
Episodes / model
30
Total benchmark cost
$56.79
Total model time
11h 2m

Primary result

Question score.

Lower is better. The blue marker is the average question score. The colored line is its 95% confidence interval (CI). The companion plot shows each exact CI width. Its three bands divide the displayed width scale into equal ranges.

Question score and confidence interval width comparison.
Question scorelower is better95% CI · color = CI widthTighterMiddleWider
CI width stabilityRows follow question-score order · lower is better
  1. 1. GPT-6 Astra (high): 13.67 questions. Best score. 95% confidence interval of the average: 12.49–14.85 questions. CI width: 2.36 questions; tighter band on the displayed scale. Smallest CI width. View full run for GPT-6 Astra (high)
  2. 2. Gemini 3.8 Flash (high): 14.87 questions. 95% confidence interval of the average: 13.19–16.54 questions. CI width: 3.35 questions; middle band on the displayed scale. View full run for Gemini 3.8 Flash (high)
  3. 3. Claude Fable 5.1 (high): 15.10 questions. 95% confidence interval of the average: 13.76–16.44 questions. CI width: 2.69 questions; tighter band on the displayed scale. View full run for Claude Fable 5.1 (high)
  4. 4. Gemini 3.7 Flash (high): 16.57 questions. 95% confidence interval of the average: 14.35–18.79 questions. CI width: 4.44 questions; middle band on the displayed scale. View full run for Gemini 3.7 Flash (high)
  5. 5. Grok 4.6 (high): 17.13 questions. 95% confidence interval of the average: 15.54–18.73 questions. CI width: 3.19 questions; middle band on the displayed scale. View full run for Grok 4.6 (high)
  6. 6. Claude Opus 5 (high): 18.07 questions. 95% confidence interval of the average: 16.50–19.63 questions. CI width: 3.13 questions; middle band on the displayed scale. View full run for Claude Opus 5 (high)
  7. 7. GPT-5.6 Sol (high): 18.07 questions. 95% confidence interval of the average: 14.38–21.75 questions. CI width: 7.37 questions; wider band on the displayed scale. View full run for GPT-5.6 Sol (high)
  8. 8. GPT-5.6 Luna (high): 21.03 questions. 95% confidence interval of the average: 18.51–23.56 questions. CI width: 5.05 questions; middle band on the displayed scale. View full run for GPT-5.6 Luna (high)
  9. 9. gpt-oss-120B (high): 22.23 questions. 95% confidence interval of the average: 19.58–24.89 questions. CI width: 5.31 questions; middle band on the displayed scale. View full run for gpt-oss-120B (high)
  10. 10. GLM-5.3-Flash (high): 23.27 questions. 95% confidence interval of the average: 19.58–26.96 questions. CI width: 7.38 questions; wider band on the displayed scale. View full run for GLM-5.3-Flash (high)
  11. 11. MiniMax M3 (high): 29.17 questions. 95% confidence interval of the average: 24.71–33.62 questions. CI width: 8.90 questions; wider band on the displayed scale. View full run for MiniMax M3 (high)
Result comparison
RankModelQuestion score95% CISuccessContractGuesser costper episodeModel timeper episode
1GPT-6 Astra (high)openai
Question score13.67questions · lower is better12.49–14.85
100%100%$0.17561m 21s
2Gemini 3.8 Flash (high)google-ai-studio
Question score14.87questions · lower is better13.19–16.54
100%100%$0.08751m 42s
3Claude Fable 5.1 (high)anthropic
Question score15.10questions · lower is better13.76–16.44
93%99%$0.27722m 39s
4Gemini 3.7 Flash (high)google-ai-studio
Question score16.57questions · lower is better14.35–18.79
93%100%$0.052551.5 s
5Grok 4.6 (high)xai
Question score17.13questions · lower is better15.54–18.73
100%100%$0.13676m 51s
6Claude Opus 5 (high)anthropic
Question score18.07questions · lower is better16.50–19.63
93%>99%$0.09341m 22s
6GPT-5.6 Sol (high)openai
Question score18.07questions · lower is better14.38–21.75
97%100%$0.05571m 55s
8GPT-5.6 Luna (high)openai
Question score21.03questions · lower is better18.51–23.56
90%100%$0.00631m 23s
9gpt-oss-120B (high)cerebras
Question score22.23questions · lower is better19.58–24.89
67%89%$0.06821m 5s
10GLM-5.3-Flash (high)z-ai
Question score23.27questions · lower is better19.58–26.96
73%93%$0.00161m 43s
11MiniMax M3 (high)coreweave
Question score29.17questions · lower is better24.71–33.62
50%99%$0.01001m 12s
#1GPT-6 Astra (high)openai
Question score
13.67
95% CI
12.49–14.85
Success
100%
Guesser cost / episode
$0.1756
Explore full run · questions, answers & evidence
#2Gemini 3.8 Flash (high)google-ai-studio
Question score
14.87
95% CI
13.19–16.54
Success
100%
Guesser cost / episode
$0.0875
Explore full run · questions, answers & evidence
#3Claude Fable 5.1 (high)anthropic
Question score
15.10
95% CI
13.76–16.44
Success
93%
Guesser cost / episode
$0.2772
Explore full run · questions, answers & evidence
#4Gemini 3.7 Flash (high)google-ai-studio
Question score
16.57
95% CI
14.35–18.79
Success
93%
Guesser cost / episode
$0.0525
Explore full run · questions, answers & evidence
#5Grok 4.6 (high)xai
Question score
17.13
95% CI
15.54–18.73
Success
100%
Guesser cost / episode
$0.1367
Explore full run · questions, answers & evidence
#6Claude Opus 5 (high)anthropic
Question score
18.07
95% CI
16.50–19.63
Success
93%
Guesser cost / episode
$0.0934
Explore full run · questions, answers & evidence
#6GPT-5.6 Sol (high)openai
Question score
18.07
95% CI
14.38–21.75
Success
97%
Guesser cost / episode
$0.0557
Explore full run · questions, answers & evidence
#8GPT-5.6 Luna (high)openai
Question score
21.03
95% CI
18.51–23.56
Success
90%
Guesser cost / episode
$0.0063
Explore full run · questions, answers & evidence
#9gpt-oss-120B (high)cerebras
Question score
22.23
95% CI
19.58–24.89
Success
67%
Guesser cost / episode
$0.0682
Explore full run · questions, answers & evidence
#10GLM-5.3-Flash (high)z-ai
Question score
23.27
95% CI
19.58–26.96
Success
73%
Guesser cost / episode
$0.0016
Explore full run · questions, answers & evidence
#11MiniMax M3 (high)coreweave
Question score
29.17
95% CI
24.71–33.62
Success
50%
Guesser cost / episode
$0.0100
Explore full run · questions, answers & evidence

The 95% CI uses repeated seeded trials on this edition's fixed subjects. The three CI width bands divide the displayed scale into equal ranges. They are not fixed quality thresholds. The 95% CI does not cover different subjects, model versions, or providers.