Repeated-trial stability

Stability 1.1

Whether a model produces similar scores across repeated trials on the same fixed subjects.

Models
11
Smallest CI width
2.36
Largest CI width
8.90
More stable
Smaller CI width

Two-dimensional view

Stability.

Each dot compares two results. Lower means a better average question score. Further left means a smaller CI width, so the model produced more consistent results across repeated trials on the fixed subjects.

Lower-left is betterSelect a model point to view its full run
  1. GPT-6 Astra (high): score 13.67 questions; CI width 2.36 questions; stability rank 1. View full run for GPT-6 Astra (high)
  2. Claude Fable 5.1 (high): score 15.10 questions; CI width 2.69 questions; stability rank 2. View full run for Claude Fable 5.1 (high)
  3. Claude Opus 5 (high): score 18.07 questions; CI width 3.13 questions; stability rank 3. View full run for Claude Opus 5 (high)
  4. Grok 4.6 (high): score 17.13 questions; CI width 3.19 questions; stability rank 4. View full run for Grok 4.6 (high)
  5. Gemini 3.8 Flash (high): score 14.87 questions; CI width 3.35 questions; stability rank 5. View full run for Gemini 3.8 Flash (high)
  6. Gemini 3.7 Flash (high): score 16.57 questions; CI width 4.44 questions; stability rank 6. View full run for Gemini 3.7 Flash (high)
  7. GPT-5.6 Luna (high): score 21.03 questions; CI width 5.05 questions; stability rank 7. View full run for GPT-5.6 Luna (high)
  8. gpt-oss-120B (high): score 22.23 questions; CI width 5.31 questions; stability rank 8. View full run for gpt-oss-120B (high)
  9. GPT-5.6 Sol (high): score 18.07 questions; CI width 7.37 questions; stability rank 9. View full run for GPT-5.6 Sol (high)
  10. GLM-5.3-Flash (high): score 23.27 questions; CI width 7.38 questions; stability rank 10. View full run for GLM-5.3-Flash (high)
  11. MiniMax M3 (high): score 29.17 questions; CI width 8.90 questions; stability rank 11. View full run for MiniMax M3 (high)
Stability ranking
Stability rankModelCIwidthQuestion rankQuestion score95% CISuccess
1GPT-6 Astra (high)openai2.36113.6712.49–14.85100%
2Claude Fable 5.1 (high)anthropic2.69315.1013.76–16.4493%
3Claude Opus 5 (high)anthropic3.13618.0716.50–19.6393%
4Grok 4.6 (high)xai3.19517.1315.54–18.73100%
5Gemini 3.8 Flash (high)google-ai-studio3.35214.8713.19–16.54100%
6Gemini 3.7 Flash (high)google-ai-studio4.44416.5714.35–18.7993%
7GPT-5.6 Luna (high)openai5.05821.0318.51–23.5690%
8gpt-oss-120B (high)cerebras5.31922.2319.58–24.8967%
9GPT-5.6 Sol (high)openai7.37618.0714.38–21.7597%
10GLM-5.3-Flash (high)z-ai7.381023.2719.58–26.9673%
11MiniMax M3 (high)coreweave8.901129.1724.71–33.6250%
#1GPT-6 Astra (high)openai
CI width
2.36
Question score
13.67
95% CI
12.49–14.85
Success
100%
Explore full run · questions, answers & evidence
#2Claude Fable 5.1 (high)anthropic
CI width
2.69
Question score
15.10
95% CI
13.76–16.44
Success
93%
Explore full run · questions, answers & evidence
#3Claude Opus 5 (high)anthropic
CI width
3.13
Question score
18.07
95% CI
16.50–19.63
Success
93%
Explore full run · questions, answers & evidence
#4Grok 4.6 (high)xai
CI width
3.19
Question score
17.13
95% CI
15.54–18.73
Success
100%
Explore full run · questions, answers & evidence
#5Gemini 3.8 Flash (high)google-ai-studio
CI width
3.35
Question score
14.87
95% CI
13.19–16.54
Success
100%
Explore full run · questions, answers & evidence
#6Gemini 3.7 Flash (high)google-ai-studio
CI width
4.44
Question score
16.57
95% CI
14.35–18.79
Success
93%
Explore full run · questions, answers & evidence
#7GPT-5.6 Luna (high)openai
CI width
5.05
Question score
21.03
95% CI
18.51–23.56
Success
90%
Explore full run · questions, answers & evidence
#8gpt-oss-120B (high)cerebras
CI width
5.31
Question score
22.23
95% CI
19.58–24.89
Success
67%
Explore full run · questions, answers & evidence
#9GPT-5.6 Sol (high)openai
CI width
7.37
Question score
18.07
95% CI
14.38–21.75
Success
97%
Explore full run · questions, answers & evidence
#10GLM-5.3-Flash (high)z-ai
CI width
7.38
Question score
23.27
95% CI
19.58–26.96
Success
73%
Explore full run · questions, answers & evidence
#11MiniMax M3 (high)coreweave
CI width
8.90
Question score
29.17
95% CI
24.71–33.62
Success
50%
Explore full run · questions, answers & evidence

How it is calculated

CI width.

A smaller CI width means the model produced more consistent aggregate results across the current repeated trials.

FormulaCI width = upper 95% CI bound − lower 95% CI bound
Calculation detailsSteps, example, interpretation, and limits
  1. Calculate the fixed-subject repeated-trial 95% CI.
  2. Subtract its lower bound from its upper bound.
  3. Sort exact widths from smallest to largest.

Example: 17.46 − 9.79 = 7.67 questions.

Every model uses the same 95% confidence level, so the level itself cannot define the order. CI width is the comparison measure.

Stable does not mean good. A model can produce a poor score consistently and rank well here. A strong average with a large CI width ranks lower because its repeated results vary more.

This is an approximate stability measure for the current fixed subjects. It is not a prediction interval for individual trials or a pairwise significance test.