Official comparison
Results 1
Question score, stability, cost, and time for the current official runs.
- Models
- 18
- Episodes / model
- 35
- Total benchmark cost
- $175.27
- Total model time
- 33h 50m
Primary result
Question score.
Lower is better. The blue marker is the average question score. The colored line is its 95% confidence interval (CI). The companion plot shows each exact CI width. Its three bands divide the displayed width scale into equal ranges.
Question scorelower is better95% CI · color = CI widthTighterMiddleWider
CI width stabilityRows follow question-score order · lower is better
- 1. Claude Fable 5 (high): 12.06 questions. Best score. 95% confidence interval of the average: 10.13–13.98 questions. CI width: 3.85 questions; tighter band on the displayed scale. View full run for Claude Fable 5 (high)
- 2. Claude Fable 5.1 (high): 12.11 questions. 95% confidence interval of the average: 8.81–15.41 questions. CI width: 6.60 questions; middle band on the displayed scale. View full run for Claude Fable 5.1 (high)
- 3. Claude Opus 5 (high): 12.34 questions. 95% confidence interval of the average: 11.41–13.27 questions. CI width: 1.86 questions; tighter band on the displayed scale. Smallest CI width. View full run for Claude Opus 5 (high)
- 4. Kimi K3 (high): 12.74 questions. 95% confidence interval of the average: 9.16–16.33 questions. CI width: 7.18 questions; middle band on the displayed scale. View full run for Kimi K3 (high)
- 5. Gemini 3.8 Flash (high): 13.14 questions. 95% confidence interval of the average: 11.52–14.77 questions. CI width: 3.25 questions; tighter band on the displayed scale. View full run for Gemini 3.8 Flash (high)
- 6. gpt-oss-120B (high): 13.63 questions. 95% confidence interval of the average: 9.79–17.46 questions. CI width: 7.67 questions; middle band on the displayed scale. View full run for gpt-oss-120B (high)
- 7. Gemini 3.7 Flash (high): 14.03 questions. 95% confidence interval of the average: 10.78–17.28 questions. CI width: 6.49 questions; middle band on the displayed scale. View full run for Gemini 3.7 Flash (high)
- 8. Grok 4.6 (high): 14.29 questions. 95% confidence interval of the average: 12.36–16.21 questions. CI width: 3.85 questions; tighter band on the displayed scale. View full run for Grok 4.6 (high)
- 9. GPT-5.6 Sol (high): 14.40 questions. 95% confidence interval of the average: 12.84–15.96 questions. CI width: 3.11 questions; tighter band on the displayed scale. View full run for GPT-5.6 Sol (high)
- 10. Claude Sonnet 5 (high): 14.74 questions. 95% confidence interval of the average: 12.78–16.71 questions. CI width: 3.93 questions; tighter band on the displayed scale. View full run for Claude Sonnet 5 (high)
- 11. GPT-6 Astra (high): 15.09 questions. 95% confidence interval of the average: 11.88–18.30 questions. CI width: 6.42 questions; middle band on the displayed scale. View full run for GPT-6 Astra (high)
- 12. Grok 4.5 (high): 15.17 questions. 95% confidence interval of the average: 11.70–18.64 questions. CI width: 6.94 questions; middle band on the displayed scale. View full run for Grok 4.5 (high)
- 13. GPT-5 Nano (medium): 15.49 questions. 95% confidence interval of the average: 13.84–17.13 questions. CI width: 3.29 questions; tighter band on the displayed scale. View full run for GPT-5 Nano (medium)
- 14. Gemini 3.6 Flash (high): 16.89 questions. 95% confidence interval of the average: 14.08–19.69 questions. CI width: 5.61 questions; middle band on the displayed scale. View full run for Gemini 3.6 Flash (high)
- 15. Ox Alpha (high): 17.60 questions. 95% confidence interval of the average: 14.60–20.60 questions. CI width: 5.99 questions; middle band on the displayed scale. View full run for Ox Alpha (high)
- 16. GPT-5.6 Luna (high): 17.74 questions. 95% confidence interval of the average: 14.74–20.75 questions. CI width: 6.01 questions; middle band on the displayed scale. View full run for GPT-5.6 Luna (high)
- 17. Mistral Medium 3.5 (high): 20.06 questions. 95% confidence interval of the average: 16.11–24.01 questions. CI width: 7.90 questions; middle band on the displayed scale. View full run for Mistral Medium 3.5 (high)
- 18. Llama 4 Maverick (non-thinking): 32.23 questions. 95% confidence interval of the average: 26.82–37.63 questions. CI width: 10.81 questions; wider band on the displayed scale. View full run for Llama 4 Maverick (non-thinking)
| Rank | Model | Question score95% CI | Success | Contract | Guesser costper episode | Model timeper episode |
|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 (high) | Question score12.06questions · lower is better10.13–13.98 | 100% | 100% | $0.1776 | 1m 59s |
| 2 | Claude Fable 5.1 (high) | Question score12.11questions · lower is better8.81–15.41 | 97% | >99% | $0.5014 | 3m 17s |
| 3 | Claude Opus 5 (high) | Question score12.34questions · lower is better11.41–13.27 | 100% | 99% | $0.0705 | 1m 13s |
| 4 | Kimi K3 (high) | Question score12.74questions · lower is better9.16–16.33 | 94% | 98% | $0.2032 | 6m 54s |
| 5 | Gemini 3.8 Flash (high) | Question score13.14questions · lower is better11.52–14.77 | 100% | 100% | $0.1128 | 2m 6s |
| 6 | gpt-oss-120B (high) | Question score13.63questions · lower is better9.79–17.46 | 94% | 91% | $0.0363 | 1m 32s |
| 7 | Gemini 3.7 Flash (high) | Question score14.03questions · lower is better10.78–17.28 | 97% | 100% | $0.0653 | 1m 17s |
| 8 | Grok 4.6 (high) | Question score14.29questions · lower is better12.36–16.21 | 100% | 100% | $0.0653 | 3m 8s |
| 9 | GPT-5.6 Sol (high) | Question score14.40questions · lower is better12.84–15.96 | 100% | 100% | $0.1711 | 2m 16s |
| 10 | Claude Sonnet 5 (high) | Question score14.74questions · lower is better12.78–16.71 | 100% | 97% | $0.0493 | 1m 24s |
| 11 | GPT-6 Astra (high) | Question score15.09questions · lower is better11.88–18.30 | 94% | 100% | $0.3279 | 2m 42s |
| 12 | Grok 4.5 (high) | Question score15.17questions · lower is better11.70–18.64 | 94% | 98% | $0.0286 | 1m 10s |
| 13 | GPT-5 Nano (medium) | Question score15.49questions · lower is better13.84–17.13 | 94% | 100% | $0.0138 | 5m 32s |
| 14 | Gemini 3.6 Flash (high) | Question score16.89questions · lower is better14.08–19.69 | 97% | >99% | $0.1705 | 1m 59s |
| 15 | Ox Alpha (high) | Question score17.60questions · lower is better14.60–20.60 | 91% | 93% | $0.0000 | 5m 14s |
| 16 | GPT-5.6 Luna (high) | Question score17.74questions · lower is better14.74–20.75 | 91% | >99% | $0.0184 | 1m 17s |
| 17 | Mistral Medium 3.5 (high) | Question score20.06questions · lower is better16.11–24.01 | 89% | >99% | $0.4113 | 14m 32s |
| 18 | Llama 4 Maverick (non-thinking) | Question score32.23questions · lower is better26.82–37.63 | 49% | 100% | $0.0086 | 27.9 s |
#1Claude Fable 5 (high)anthropic
- Question score
- 12.06
- 95% CI
- 10.13–13.98
- Success
- 100%
- Guesser cost / episode
- $0.1776
Explore full run · questions, answers & evidence
#2Claude Fable 5.1 (high)anthropic
- Question score
- 12.11
- 95% CI
- 8.81–15.41
- Success
- 97%
- Guesser cost / episode
- $0.5014
Explore full run · questions, answers & evidence
#3Claude Opus 5 (high)anthropic
- Question score
- 12.34
- 95% CI
- 11.41–13.27
- Success
- 100%
- Guesser cost / episode
- $0.0705
Explore full run · questions, answers & evidence
#4Kimi K3 (high)moonshotai
- Question score
- 12.74
- 95% CI
- 9.16–16.33
- Success
- 94%
- Guesser cost / episode
- $0.2032
Explore full run · questions, answers & evidence
#5Gemini 3.8 Flash (high)google-ai-studio
- Question score
- 13.14
- 95% CI
- 11.52–14.77
- Success
- 100%
- Guesser cost / episode
- $0.1128
Explore full run · questions, answers & evidence
#6gpt-oss-120B (high)cerebras
- Question score
- 13.63
- 95% CI
- 9.79–17.46
- Success
- 94%
- Guesser cost / episode
- $0.0363
Explore full run · questions, answers & evidence
#7Gemini 3.7 Flash (high)google-ai-studio
- Question score
- 14.03
- 95% CI
- 10.78–17.28
- Success
- 97%
- Guesser cost / episode
- $0.0653
Explore full run · questions, answers & evidence
#8Grok 4.6 (high)xai
- Question score
- 14.29
- 95% CI
- 12.36–16.21
- Success
- 100%
- Guesser cost / episode
- $0.0653
Explore full run · questions, answers & evidence
#9GPT-5.6 Sol (high)openai
- Question score
- 14.40
- 95% CI
- 12.84–15.96
- Success
- 100%
- Guesser cost / episode
- $0.1711
Explore full run · questions, answers & evidence
#10Claude Sonnet 5 (high)anthropic
- Question score
- 14.74
- 95% CI
- 12.78–16.71
- Success
- 100%
- Guesser cost / episode
- $0.0493
Explore full run · questions, answers & evidence
#11GPT-6 Astra (high)openai
- Question score
- 15.09
- 95% CI
- 11.88–18.30
- Success
- 94%
- Guesser cost / episode
- $0.3279
Explore full run · questions, answers & evidence
#12Grok 4.5 (high)xai
- Question score
- 15.17
- 95% CI
- 11.70–18.64
- Success
- 94%
- Guesser cost / episode
- $0.0286
Explore full run · questions, answers & evidence
#13GPT-5 Nano (medium)openai
- Question score
- 15.49
- 95% CI
- 13.84–17.13
- Success
- 94%
- Guesser cost / episode
- $0.0138
Explore full run · questions, answers & evidence
#14Gemini 3.6 Flash (high)google-vertex
- Question score
- 16.89
- 95% CI
- 14.08–19.69
- Success
- 97%
- Guesser cost / episode
- $0.1705
Explore full run · questions, answers & evidence
#15Ox Alpha (high)stealth
- Question score
- 17.60
- 95% CI
- 14.60–20.60
- Success
- 91%
- Guesser cost / episode
- $0.0000
Explore full run · questions, answers & evidence
#16GPT-5.6 Luna (high)openai
- Question score
- 17.74
- 95% CI
- 14.74–20.75
- Success
- 91%
- Guesser cost / episode
- $0.0184
Explore full run · questions, answers & evidence
#17Mistral Medium 3.5 (high)mistral
- Question score
- 20.06
- 95% CI
- 16.11–24.01
- Success
- 89%
- Guesser cost / episode
- $0.4113
Explore full run · questions, answers & evidence
#18Llama 4 Maverick (non-thinking)parasail
- Question score
- 32.23
- 95% CI
- 26.82–37.63
- Success
- 49%
- Guesser cost / episode
- $0.0086
Explore full run · questions, answers & evidence
The 95% CI uses repeated seeded trials on this edition's fixed subjects. The three CI width bands divide the displayed scale into equal ranges. They are not fixed quality thresholds. The 95% CI does not cover different subjects, model versions, or providers.