Official comparison
Results 1.1
Question score, stability, cost, and time for the current official runs.
- Models
- 11
- Episodes / model
- 30
- Total benchmark cost
- $56.79
- Total model time
- 11h 2m
Primary result
Question score.
Lower is better. The blue marker is the average question score. The colored line is its 95% confidence interval (CI). The companion plot shows each exact CI width. Its three bands divide the displayed width scale into equal ranges.
Question scorelower is better95% CI · color = CI widthTighterMiddleWider
CI width stabilityRows follow question-score order · lower is better
- 1. GPT-6 Astra (high): 13.67 questions. Best score. 95% confidence interval of the average: 12.49–14.85 questions. CI width: 2.36 questions; tighter band on the displayed scale. Smallest CI width. View full run for GPT-6 Astra (high)
- 2. Gemini 3.8 Flash (high): 14.87 questions. 95% confidence interval of the average: 13.19–16.54 questions. CI width: 3.35 questions; middle band on the displayed scale. View full run for Gemini 3.8 Flash (high)
- 3. Claude Fable 5.1 (high): 15.10 questions. 95% confidence interval of the average: 13.76–16.44 questions. CI width: 2.69 questions; tighter band on the displayed scale. View full run for Claude Fable 5.1 (high)
- 4. Gemini 3.7 Flash (high): 16.57 questions. 95% confidence interval of the average: 14.35–18.79 questions. CI width: 4.44 questions; middle band on the displayed scale. View full run for Gemini 3.7 Flash (high)
- 5. Grok 4.6 (high): 17.13 questions. 95% confidence interval of the average: 15.54–18.73 questions. CI width: 3.19 questions; middle band on the displayed scale. View full run for Grok 4.6 (high)
- 6. Claude Opus 5 (high): 18.07 questions. 95% confidence interval of the average: 16.50–19.63 questions. CI width: 3.13 questions; middle band on the displayed scale. View full run for Claude Opus 5 (high)
- 7. GPT-5.6 Sol (high): 18.07 questions. 95% confidence interval of the average: 14.38–21.75 questions. CI width: 7.37 questions; wider band on the displayed scale. View full run for GPT-5.6 Sol (high)
- 8. GPT-5.6 Luna (high): 21.03 questions. 95% confidence interval of the average: 18.51–23.56 questions. CI width: 5.05 questions; middle band on the displayed scale. View full run for GPT-5.6 Luna (high)
- 9. gpt-oss-120B (high): 22.23 questions. 95% confidence interval of the average: 19.58–24.89 questions. CI width: 5.31 questions; middle band on the displayed scale. View full run for gpt-oss-120B (high)
- 10. GLM-5.3-Flash (high): 23.27 questions. 95% confidence interval of the average: 19.58–26.96 questions. CI width: 7.38 questions; wider band on the displayed scale. View full run for GLM-5.3-Flash (high)
- 11. MiniMax M3 (high): 29.17 questions. 95% confidence interval of the average: 24.71–33.62 questions. CI width: 8.90 questions; wider band on the displayed scale. View full run for MiniMax M3 (high)
| Rank | Model | Question score95% CI | Success | Contract | Guesser costper episode | Model timeper episode |
|---|---|---|---|---|---|---|
| 1 | GPT-6 Astra (high) | Question score13.67questions · lower is better12.49–14.85 | 100% | 100% | $0.1756 | 1m 21s |
| 2 | Gemini 3.8 Flash (high) | Question score14.87questions · lower is better13.19–16.54 | 100% | 100% | $0.0875 | 1m 42s |
| 3 | Claude Fable 5.1 (high) | Question score15.10questions · lower is better13.76–16.44 | 93% | 99% | $0.2772 | 2m 39s |
| 4 | Gemini 3.7 Flash (high) | Question score16.57questions · lower is better14.35–18.79 | 93% | 100% | $0.0525 | 51.5 s |
| 5 | Grok 4.6 (high) | Question score17.13questions · lower is better15.54–18.73 | 100% | 100% | $0.1367 | 6m 51s |
| 6 | Claude Opus 5 (high) | Question score18.07questions · lower is better16.50–19.63 | 93% | >99% | $0.0934 | 1m 22s |
| 6 | GPT-5.6 Sol (high) | Question score18.07questions · lower is better14.38–21.75 | 97% | 100% | $0.0557 | 1m 55s |
| 8 | GPT-5.6 Luna (high) | Question score21.03questions · lower is better18.51–23.56 | 90% | 100% | $0.0063 | 1m 23s |
| 9 | gpt-oss-120B (high) | Question score22.23questions · lower is better19.58–24.89 | 67% | 89% | $0.0682 | 1m 5s |
| 10 | GLM-5.3-Flash (high) | Question score23.27questions · lower is better19.58–26.96 | 73% | 93% | $0.0016 | 1m 43s |
| 11 | MiniMax M3 (high) | Question score29.17questions · lower is better24.71–33.62 | 50% | 99% | $0.0100 | 1m 12s |
#1GPT-6 Astra (high)openai
- Question score
- 13.67
- 95% CI
- 12.49–14.85
- Success
- 100%
- Guesser cost / episode
- $0.1756
Explore full run · questions, answers & evidence
#2Gemini 3.8 Flash (high)google-ai-studio
- Question score
- 14.87
- 95% CI
- 13.19–16.54
- Success
- 100%
- Guesser cost / episode
- $0.0875
Explore full run · questions, answers & evidence
#3Claude Fable 5.1 (high)anthropic
- Question score
- 15.10
- 95% CI
- 13.76–16.44
- Success
- 93%
- Guesser cost / episode
- $0.2772
Explore full run · questions, answers & evidence
#4Gemini 3.7 Flash (high)google-ai-studio
- Question score
- 16.57
- 95% CI
- 14.35–18.79
- Success
- 93%
- Guesser cost / episode
- $0.0525
Explore full run · questions, answers & evidence
#5Grok 4.6 (high)xai
- Question score
- 17.13
- 95% CI
- 15.54–18.73
- Success
- 100%
- Guesser cost / episode
- $0.1367
Explore full run · questions, answers & evidence
#6Claude Opus 5 (high)anthropic
- Question score
- 18.07
- 95% CI
- 16.50–19.63
- Success
- 93%
- Guesser cost / episode
- $0.0934
Explore full run · questions, answers & evidence
#6GPT-5.6 Sol (high)openai
- Question score
- 18.07
- 95% CI
- 14.38–21.75
- Success
- 97%
- Guesser cost / episode
- $0.0557
Explore full run · questions, answers & evidence
#8GPT-5.6 Luna (high)openai
- Question score
- 21.03
- 95% CI
- 18.51–23.56
- Success
- 90%
- Guesser cost / episode
- $0.0063
Explore full run · questions, answers & evidence
#9gpt-oss-120B (high)cerebras
- Question score
- 22.23
- 95% CI
- 19.58–24.89
- Success
- 67%
- Guesser cost / episode
- $0.0682
Explore full run · questions, answers & evidence
#10GLM-5.3-Flash (high)z-ai
- Question score
- 23.27
- 95% CI
- 19.58–26.96
- Success
- 73%
- Guesser cost / episode
- $0.0016
Explore full run · questions, answers & evidence
#11MiniMax M3 (high)coreweave
- Question score
- 29.17
- 95% CI
- 24.71–33.62
- Success
- 50%
- Guesser cost / episode
- $0.0100
Explore full run · questions, answers & evidence
The 95% CI uses repeated seeded trials on this edition's fixed subjects. The three CI width bands divide the displayed scale into equal ranges. They are not fixed quality thresholds. The 95% CI does not cover different subjects, model versions, or providers.