Repeated-trial stability

Stability 1

Whether a model produces similar scores across repeated trials on the same fixed subjects.

Models
18
Smallest CI width
1.86
Largest CI width
10.81
More stable
Smaller CI width

Two-dimensional view

Stability.

Each dot compares two results. Lower means a better average question score. Further left means a smaller CI width, so the model produced more consistent results across repeated trials on the fixed subjects.

Lower-left is betterSelect a model point to view its full run
  1. Claude Opus 5 (high): score 12.34 questions; CI width 1.86 questions; stability rank 1. View full run for Claude Opus 5 (high)
  2. GPT-5.6 Sol (high): score 14.40 questions; CI width 3.11 questions; stability rank 2. View full run for GPT-5.6 Sol (high)
  3. Gemini 3.8 Flash (high): score 13.14 questions; CI width 3.25 questions; stability rank 3. View full run for Gemini 3.8 Flash (high)
  4. GPT-5 Nano (medium): score 15.49 questions; CI width 3.29 questions; stability rank 4. View full run for GPT-5 Nano (medium)
  5. Claude Fable 5 (high): score 12.06 questions; CI width 3.85 questions; stability rank 5. View full run for Claude Fable 5 (high)
  6. Grok 4.6 (high): score 14.29 questions; CI width 3.85 questions; stability rank 6. View full run for Grok 4.6 (high)
  7. Claude Sonnet 5 (high): score 14.74 questions; CI width 3.93 questions; stability rank 7. View full run for Claude Sonnet 5 (high)
  8. Gemini 3.6 Flash (high): score 16.89 questions; CI width 5.61 questions; stability rank 8. View full run for Gemini 3.6 Flash (high)
  9. Ox Alpha (high): score 17.60 questions; CI width 5.99 questions; stability rank 9. View full run for Ox Alpha (high)
  10. GPT-5.6 Luna (high): score 17.74 questions; CI width 6.01 questions; stability rank 10. View full run for GPT-5.6 Luna (high)
  11. GPT-6 Astra (high): score 15.09 questions; CI width 6.42 questions; stability rank 11. View full run for GPT-6 Astra (high)
  12. Gemini 3.7 Flash (high): score 14.03 questions; CI width 6.49 questions; stability rank 12. View full run for Gemini 3.7 Flash (high)
  13. Claude Fable 5.1 (high): score 12.11 questions; CI width 6.60 questions; stability rank 13. View full run for Claude Fable 5.1 (high)
  14. Grok 4.5 (high): score 15.17 questions; CI width 6.94 questions; stability rank 14. View full run for Grok 4.5 (high)
  15. Kimi K3 (high): score 12.74 questions; CI width 7.18 questions; stability rank 15. View full run for Kimi K3 (high)
  16. gpt-oss-120B (high): score 13.63 questions; CI width 7.67 questions; stability rank 16. View full run for gpt-oss-120B (high)
  17. Mistral Medium 3.5 (high): score 20.06 questions; CI width 7.90 questions; stability rank 17. View full run for Mistral Medium 3.5 (high)
  18. Llama 4 Maverick (non-thinking): score 32.23 questions; CI width 10.81 questions; stability rank 18. View full run for Llama 4 Maverick (non-thinking)
Stability ranking
Stability rankModelCIwidthQuestion rankQuestion score95% CISuccess
1Claude Opus 5 (high)anthropic1.86312.3411.41–13.27100%
2GPT-5.6 Sol (high)openai3.11914.4012.84–15.96100%
3Gemini 3.8 Flash (high)google-ai-studio3.25513.1411.52–14.77100%
4GPT-5 Nano (medium)openai3.291315.4913.84–17.1394%
5Claude Fable 5 (high)anthropic3.85112.0610.13–13.98100%
6Grok 4.6 (high)xai3.85814.2912.36–16.21100%
7Claude Sonnet 5 (high)anthropic3.931014.7412.78–16.71100%
8Gemini 3.6 Flash (high)google-vertex5.611416.8914.08–19.6997%
9Ox Alpha (high)stealth5.991517.6014.60–20.6091%
10GPT-5.6 Luna (high)openai6.011617.7414.74–20.7591%
11GPT-6 Astra (high)openai6.421115.0911.88–18.3094%
12Gemini 3.7 Flash (high)google-ai-studio6.49714.0310.78–17.2897%
13Claude Fable 5.1 (high)anthropic6.60212.118.81–15.4197%
14Grok 4.5 (high)xai6.941215.1711.70–18.6494%
15Kimi K3 (high)moonshotai7.18412.749.16–16.3394%
16gpt-oss-120B (high)cerebras7.67613.639.79–17.4694%
17Mistral Medium 3.5 (high)mistral7.901720.0616.11–24.0189%
18Llama 4 Maverick (non-thinking)parasail10.811832.2326.82–37.6349%
#1Claude Opus 5 (high)anthropic
CI width
1.86
Question score
12.34
95% CI
11.41–13.27
Success
100%
Explore full run · questions, answers & evidence
#2GPT-5.6 Sol (high)openai
CI width
3.11
Question score
14.40
95% CI
12.84–15.96
Success
100%
Explore full run · questions, answers & evidence
#3Gemini 3.8 Flash (high)google-ai-studio
CI width
3.25
Question score
13.14
95% CI
11.52–14.77
Success
100%
Explore full run · questions, answers & evidence
#4GPT-5 Nano (medium)openai
CI width
3.29
Question score
15.49
95% CI
13.84–17.13
Success
94%
Explore full run · questions, answers & evidence
#5Claude Fable 5 (high)anthropic
CI width
3.85
Question score
12.06
95% CI
10.13–13.98
Success
100%
Explore full run · questions, answers & evidence
#6Grok 4.6 (high)xai
CI width
3.85
Question score
14.29
95% CI
12.36–16.21
Success
100%
Explore full run · questions, answers & evidence
#7Claude Sonnet 5 (high)anthropic
CI width
3.93
Question score
14.74
95% CI
12.78–16.71
Success
100%
Explore full run · questions, answers & evidence
#8Gemini 3.6 Flash (high)google-vertex
CI width
5.61
Question score
16.89
95% CI
14.08–19.69
Success
97%
Explore full run · questions, answers & evidence
#9Ox Alpha (high)stealth
CI width
5.99
Question score
17.60
95% CI
14.60–20.60
Success
91%
Explore full run · questions, answers & evidence
#10GPT-5.6 Luna (high)openai
CI width
6.01
Question score
17.74
95% CI
14.74–20.75
Success
91%
Explore full run · questions, answers & evidence
#11GPT-6 Astra (high)openai
CI width
6.42
Question score
15.09
95% CI
11.88–18.30
Success
94%
Explore full run · questions, answers & evidence
#12Gemini 3.7 Flash (high)google-ai-studio
CI width
6.49
Question score
14.03
95% CI
10.78–17.28
Success
97%
Explore full run · questions, answers & evidence
#13Claude Fable 5.1 (high)anthropic
CI width
6.60
Question score
12.11
95% CI
8.81–15.41
Success
97%
Explore full run · questions, answers & evidence
#14Grok 4.5 (high)xai
CI width
6.94
Question score
15.17
95% CI
11.70–18.64
Success
94%
Explore full run · questions, answers & evidence
#15Kimi K3 (high)moonshotai
CI width
7.18
Question score
12.74
95% CI
9.16–16.33
Success
94%
Explore full run · questions, answers & evidence
#16gpt-oss-120B (high)cerebras
CI width
7.67
Question score
13.63
95% CI
9.79–17.46
Success
94%
Explore full run · questions, answers & evidence
#17Mistral Medium 3.5 (high)mistral
CI width
7.90
Question score
20.06
95% CI
16.11–24.01
Success
89%
Explore full run · questions, answers & evidence
#18Llama 4 Maverick (non-thinking)parasail
CI width
10.81
Question score
32.23
95% CI
26.82–37.63
Success
49%
Explore full run · questions, answers & evidence

How it is calculated

CI width.

A smaller CI width means the model produced more consistent aggregate results across the current repeated trials.

FormulaCI width = upper 95% CI bound − lower 95% CI bound
Calculation detailsSteps, example, interpretation, and limits
  1. Calculate the fixed-subject repeated-trial 95% CI.
  2. Subtract its lower bound from its upper bound.
  3. Sort exact widths from smallest to largest.

Example: 17.46 − 9.79 = 7.67 questions.

Every model uses the same 95% confidence level, so the level itself cannot define the order. CI width is the comparison measure.

Stable does not mean good. A model can produce a poor score consistently and rank well here. A strong average with a large CI width ranks lower because its repeated results vary more.

This is an approximate stability measure for the current fixed subjects. It is not a prediction interval for individual trials or a pairwise significance test.