Official comparison

Results 1

Question score, stability, cost, and time for the current official runs.

Models
18
Episodes / model
35
Total benchmark cost
$175.27
Total model time
33h 50m

Primary result

Question score.

Lower is better. The blue marker is the average question score. The colored line is its 95% confidence interval (CI). The companion plot shows each exact CI width. Its three bands divide the displayed width scale into equal ranges.

Question score and confidence interval width comparison.
Question scorelower is better95% CI · color = CI widthTighterMiddleWider
CI width stabilityRows follow question-score order · lower is better
  1. 1. Claude Fable 5 (high): 12.06 questions. Best score. 95% confidence interval of the average: 10.13–13.98 questions. CI width: 3.85 questions; tighter band on the displayed scale. View full run for Claude Fable 5 (high)
  2. 2. Claude Fable 5.1 (high): 12.11 questions. 95% confidence interval of the average: 8.81–15.41 questions. CI width: 6.60 questions; middle band on the displayed scale. View full run for Claude Fable 5.1 (high)
  3. 3. Claude Opus 5 (high): 12.34 questions. 95% confidence interval of the average: 11.41–13.27 questions. CI width: 1.86 questions; tighter band on the displayed scale. Smallest CI width. View full run for Claude Opus 5 (high)
  4. 4. Kimi K3 (high): 12.74 questions. 95% confidence interval of the average: 9.16–16.33 questions. CI width: 7.18 questions; middle band on the displayed scale. View full run for Kimi K3 (high)
  5. 5. Gemini 3.8 Flash (high): 13.14 questions. 95% confidence interval of the average: 11.52–14.77 questions. CI width: 3.25 questions; tighter band on the displayed scale. View full run for Gemini 3.8 Flash (high)
  6. 6. gpt-oss-120B (high): 13.63 questions. 95% confidence interval of the average: 9.79–17.46 questions. CI width: 7.67 questions; middle band on the displayed scale. View full run for gpt-oss-120B (high)
  7. 7. Gemini 3.7 Flash (high): 14.03 questions. 95% confidence interval of the average: 10.78–17.28 questions. CI width: 6.49 questions; middle band on the displayed scale. View full run for Gemini 3.7 Flash (high)
  8. 8. Grok 4.6 (high): 14.29 questions. 95% confidence interval of the average: 12.36–16.21 questions. CI width: 3.85 questions; tighter band on the displayed scale. View full run for Grok 4.6 (high)
  9. 9. GPT-5.6 Sol (high): 14.40 questions. 95% confidence interval of the average: 12.84–15.96 questions. CI width: 3.11 questions; tighter band on the displayed scale. View full run for GPT-5.6 Sol (high)
  10. 10. Claude Sonnet 5 (high): 14.74 questions. 95% confidence interval of the average: 12.78–16.71 questions. CI width: 3.93 questions; tighter band on the displayed scale. View full run for Claude Sonnet 5 (high)
  11. 11. GPT-6 Astra (high): 15.09 questions. 95% confidence interval of the average: 11.88–18.30 questions. CI width: 6.42 questions; middle band on the displayed scale. View full run for GPT-6 Astra (high)
  12. 12. Grok 4.5 (high): 15.17 questions. 95% confidence interval of the average: 11.70–18.64 questions. CI width: 6.94 questions; middle band on the displayed scale. View full run for Grok 4.5 (high)
  13. 13. GPT-5 Nano (medium): 15.49 questions. 95% confidence interval of the average: 13.84–17.13 questions. CI width: 3.29 questions; tighter band on the displayed scale. View full run for GPT-5 Nano (medium)
  14. 14. Gemini 3.6 Flash (high): 16.89 questions. 95% confidence interval of the average: 14.08–19.69 questions. CI width: 5.61 questions; middle band on the displayed scale. View full run for Gemini 3.6 Flash (high)
  15. 15. Ox Alpha (high): 17.60 questions. 95% confidence interval of the average: 14.60–20.60 questions. CI width: 5.99 questions; middle band on the displayed scale. View full run for Ox Alpha (high)
  16. 16. GPT-5.6 Luna (high): 17.74 questions. 95% confidence interval of the average: 14.74–20.75 questions. CI width: 6.01 questions; middle band on the displayed scale. View full run for GPT-5.6 Luna (high)
  17. 17. Mistral Medium 3.5 (high): 20.06 questions. 95% confidence interval of the average: 16.11–24.01 questions. CI width: 7.90 questions; middle band on the displayed scale. View full run for Mistral Medium 3.5 (high)
  18. 18. Llama 4 Maverick (non-thinking): 32.23 questions. 95% confidence interval of the average: 26.82–37.63 questions. CI width: 10.81 questions; wider band on the displayed scale. View full run for Llama 4 Maverick (non-thinking)
Result comparison
RankModelQuestion score95% CISuccessContractGuesser costper episodeModel timeper episode
1Claude Fable 5 (high)anthropic
Question score12.06questions · lower is better10.13–13.98
100%100%$0.17761m 59s
2Claude Fable 5.1 (high)anthropic
Question score12.11questions · lower is better8.81–15.41
97%>99%$0.50143m 17s
3Claude Opus 5 (high)anthropic
Question score12.34questions · lower is better11.41–13.27
100%99%$0.07051m 13s
4Kimi K3 (high)moonshotai
Question score12.74questions · lower is better9.16–16.33
94%98%$0.20326m 54s
5Gemini 3.8 Flash (high)google-ai-studio
Question score13.14questions · lower is better11.52–14.77
100%100%$0.11282m 6s
6gpt-oss-120B (high)cerebras
Question score13.63questions · lower is better9.79–17.46
94%91%$0.03631m 32s
7Gemini 3.7 Flash (high)google-ai-studio
Question score14.03questions · lower is better10.78–17.28
97%100%$0.06531m 17s
8Grok 4.6 (high)xai
Question score14.29questions · lower is better12.36–16.21
100%100%$0.06533m 8s
9GPT-5.6 Sol (high)openai
Question score14.40questions · lower is better12.84–15.96
100%100%$0.17112m 16s
10Claude Sonnet 5 (high)anthropic
Question score14.74questions · lower is better12.78–16.71
100%97%$0.04931m 24s
11GPT-6 Astra (high)openai
Question score15.09questions · lower is better11.88–18.30
94%100%$0.32792m 42s
12Grok 4.5 (high)xai
Question score15.17questions · lower is better11.70–18.64
94%98%$0.02861m 10s
13GPT-5 Nano (medium)openai
Question score15.49questions · lower is better13.84–17.13
94%100%$0.01385m 32s
14Gemini 3.6 Flash (high)google-vertex
Question score16.89questions · lower is better14.08–19.69
97%>99%$0.17051m 59s
15Ox Alpha (high)stealth
Question score17.60questions · lower is better14.60–20.60
91%93%$0.00005m 14s
16GPT-5.6 Luna (high)openai
Question score17.74questions · lower is better14.74–20.75
91%>99%$0.01841m 17s
17Mistral Medium 3.5 (high)mistral
Question score20.06questions · lower is better16.11–24.01
89%>99%$0.411314m 32s
18Llama 4 Maverick (non-thinking)parasail
Question score32.23questions · lower is better26.82–37.63
49%100%$0.008627.9 s
#1Claude Fable 5 (high)anthropic
Question score
12.06
95% CI
10.13–13.98
Success
100%
Guesser cost / episode
$0.1776
Explore full run · questions, answers & evidence
#2Claude Fable 5.1 (high)anthropic
Question score
12.11
95% CI
8.81–15.41
Success
97%
Guesser cost / episode
$0.5014
Explore full run · questions, answers & evidence
#3Claude Opus 5 (high)anthropic
Question score
12.34
95% CI
11.41–13.27
Success
100%
Guesser cost / episode
$0.0705
Explore full run · questions, answers & evidence
#4Kimi K3 (high)moonshotai
Question score
12.74
95% CI
9.16–16.33
Success
94%
Guesser cost / episode
$0.2032
Explore full run · questions, answers & evidence
#5Gemini 3.8 Flash (high)google-ai-studio
Question score
13.14
95% CI
11.52–14.77
Success
100%
Guesser cost / episode
$0.1128
Explore full run · questions, answers & evidence
#6gpt-oss-120B (high)cerebras
Question score
13.63
95% CI
9.79–17.46
Success
94%
Guesser cost / episode
$0.0363
Explore full run · questions, answers & evidence
#7Gemini 3.7 Flash (high)google-ai-studio
Question score
14.03
95% CI
10.78–17.28
Success
97%
Guesser cost / episode
$0.0653
Explore full run · questions, answers & evidence
#8Grok 4.6 (high)xai
Question score
14.29
95% CI
12.36–16.21
Success
100%
Guesser cost / episode
$0.0653
Explore full run · questions, answers & evidence
#9GPT-5.6 Sol (high)openai
Question score
14.40
95% CI
12.84–15.96
Success
100%
Guesser cost / episode
$0.1711
Explore full run · questions, answers & evidence
#10Claude Sonnet 5 (high)anthropic
Question score
14.74
95% CI
12.78–16.71
Success
100%
Guesser cost / episode
$0.0493
Explore full run · questions, answers & evidence
#11GPT-6 Astra (high)openai
Question score
15.09
95% CI
11.88–18.30
Success
94%
Guesser cost / episode
$0.3279
Explore full run · questions, answers & evidence
#12Grok 4.5 (high)xai
Question score
15.17
95% CI
11.70–18.64
Success
94%
Guesser cost / episode
$0.0286
Explore full run · questions, answers & evidence
#13GPT-5 Nano (medium)openai
Question score
15.49
95% CI
13.84–17.13
Success
94%
Guesser cost / episode
$0.0138
Explore full run · questions, answers & evidence
#14Gemini 3.6 Flash (high)google-vertex
Question score
16.89
95% CI
14.08–19.69
Success
97%
Guesser cost / episode
$0.1705
Explore full run · questions, answers & evidence
#15Ox Alpha (high)stealth
Question score
17.60
95% CI
14.60–20.60
Success
91%
Guesser cost / episode
$0.0000
Explore full run · questions, answers & evidence
#16GPT-5.6 Luna (high)openai
Question score
17.74
95% CI
14.74–20.75
Success
91%
Guesser cost / episode
$0.0184
Explore full run · questions, answers & evidence
#17Mistral Medium 3.5 (high)mistral
Question score
20.06
95% CI
16.11–24.01
Success
89%
Guesser cost / episode
$0.4113
Explore full run · questions, answers & evidence
#18Llama 4 Maverick (non-thinking)parasail
Question score
32.23
95% CI
26.82–37.63
Success
49%
Guesser cost / episode
$0.0086
Explore full run · questions, answers & evidence

The 95% CI uses repeated seeded trials on this edition's fixed subjects. The three CI width bands divide the displayed scale into equal ranges. They are not fixed quality thresholds. The 95% CI does not cover different subjects, model versions, or providers.