Recorded costs

Cost 1.1

Recorded provider costs for the tested model and benchmark support models.

Models
11
Total benchmark cost
$56.79
Total Guesser cost
$28.94
Support share
49%

Guesser cost

Guesser cost across the run.

Each bar adds the recorded provider cost of Guesser calls in the retained terminal attempts. Superseded infrastructure attempts and support costs are excluded.

Guesser cost · lower is betterSelect a model row to view its full run
  1. 1. GLM-5.3-Flash (high): $0.05. View full run for GLM-5.3-Flash (high)
  2. 2. GPT-5.6 Luna (high): $0.19. View full run for GPT-5.6 Luna (high)
  3. 3. MiniMax M3 (high): $0.30. View full run for MiniMax M3 (high)
  4. 4. Gemini 3.7 Flash (high): $1.58. View full run for Gemini 3.7 Flash (high)
  5. 5. GPT-5.6 Sol (high): $1.67. View full run for GPT-5.6 Sol (high)
  6. 6. gpt-oss-120B (high): $2.04. View full run for gpt-oss-120B (high)
  7. 7. Gemini 3.8 Flash (high): $2.62. View full run for Gemini 3.8 Flash (high)
  8. 8. Claude Opus 5 (high): $2.80. View full run for Claude Opus 5 (high)
  9. 9. Grok 4.6 (high): $4.10. View full run for Grok 4.6 (high)
  10. 10. GPT-6 Astra (high): $5.27. View full run for GPT-6 Astra (high)
  11. 11. Claude Fable 5.1 (high): $8.31. View full run for Claude Fable 5.1 (high)

Total benchmark cost

Where the total cost came from.

Bar length shows the retained terminal attempts' benchmark cost. Color separates the Guesser, Primary Oracle, and Adjudication. Expand the exact breakdown to compare Reviewer, Judge, and Validator cost.

Total benchmark cost by component
  • Guesser
  • Primary Oracle
  • Adjudication

Select a model row to view its full run.

Exact adjudication breakdown
Exact adjudication costs by model
ModelReviewerJudgeValidator
GLM-5.3-Flash (high)$0.69$0.58$0.01
GPT-5.6 Luna (high)$0.83$0.75$0.01
Gemini 3.7 Flash (high)$0.60$0.42$0.01
GPT-5.6 Sol (high)$0.67$0.54$0.01
MiniMax M3 (high)$0.86$1.13$0.02
Gemini 3.8 Flash (high)$0.55$0.25$0.01
gpt-oss-120B (high)$0.66$0.38$0.01
Claude Opus 5 (high)$0.60$0.72$0.01
Grok 4.6 (high)$0.65$0.47$0.01
GPT-6 Astra (high)$0.36$0.30$0.01
Claude Fable 5.1 (high)$0.51$0.36$0.01
  1. GLM-5.3-Flash (high): $2.88. Guesser $0.05. Primary Oracle $1.55. Adjudication $1.28. Reviewer $0.69. Judge $0.58. Validator $0.01. View full run for GLM-5.3-Flash (high)
  2. GPT-5.6 Luna (high): $3.38. Guesser $0.19. Primary Oracle $1.61. Adjudication $1.59. Reviewer $0.83. Judge $0.75. Validator $0.01. View full run for GPT-5.6 Luna (high)
  3. Gemini 3.7 Flash (high): $3.80. Guesser $1.58. Primary Oracle $1.19. Adjudication $1.03. Reviewer $0.60. Judge $0.42. Validator $0.01. View full run for Gemini 3.7 Flash (high)
  4. GPT-5.6 Sol (high): $4.27. Guesser $1.67. Primary Oracle $1.37. Adjudication $1.22. Reviewer $0.67. Judge $0.54. Validator $0.01. View full run for GPT-5.6 Sol (high)
  5. MiniMax M3 (high): $4.35. Guesser $0.30. Primary Oracle $2.04. Adjudication $2.01. Reviewer $0.86. Judge $1.13. Validator $0.02. View full run for MiniMax M3 (high)
  6. Gemini 3.8 Flash (high): $4.52. Guesser $2.62. Primary Oracle $1.09. Adjudication $0.80. Reviewer $0.55. Judge $0.25. Validator $0.01. View full run for Gemini 3.8 Flash (high)
  7. gpt-oss-120B (high): $4.57. Guesser $2.04. Primary Oracle $1.48. Adjudication $1.05. Reviewer $0.66. Judge $0.38. Validator $0.01. View full run for gpt-oss-120B (high)
  8. Claude Opus 5 (high): $5.55. Guesser $2.80. Primary Oracle $1.43. Adjudication $1.32. Reviewer $0.60. Judge $0.72. Validator $0.01. View full run for Claude Opus 5 (high)
  9. Grok 4.6 (high): $6.56. Guesser $4.10. Primary Oracle $1.34. Adjudication $1.12. Reviewer $0.65. Judge $0.47. Validator $0.01. View full run for Grok 4.6 (high)
  10. GPT-6 Astra (high): $6.68. Guesser $5.27. Primary Oracle $0.75. Adjudication $0.67. Reviewer $0.36. Judge $0.30. Validator $0.01. View full run for GPT-6 Astra (high)
  11. Claude Fable 5.1 (high): $10.22. Guesser $8.31. Primary Oracle $1.03. Adjudication $0.88. Reviewer $0.51. Judge $0.36. Validator $0.01. View full run for Claude Fable 5.1 (high)
Cost comparison
Benchmark cost rankModelGuesser costper episodeBenchmark costper episodeSupport costper episodeSupport shareBenchmark run cost
1GLM-5.3-Flash (high)M-0023$0.0016$0.0959$0.094498%
$2.88 Excluded repair overhead: $0.02
2GPT-5.6 Luna (high)M-0001$0.0063$0.1128$0.106594%
$3.38 Excluded repair overhead: $0.03
3Gemini 3.7 Flash (high)M-0016$0.0525$0.1267$0.074259%
$3.80
4GPT-5.6 Sol (high)M-0010$0.0557$0.1423$0.086561%
$4.27 Excluded repair overhead: $0.11
5MiniMax M3 (high)M-0024$0.0100$0.1450$0.135093%
$4.35 Excluded repair overhead: $0.18
6Gemini 3.8 Flash (high)M-0021$0.0875$0.1505$0.063042%
$4.52 Excluded repair overhead: $0.03
7gpt-oss-120B (high)M-0002$0.0682$0.1523$0.084255%
$4.57 Excluded repair overhead: $0.49
8Claude Opus 5 (high)M-0006$0.0934$0.1852$0.091750%
$5.55 Excluded repair overhead: $0.76
9Grok 4.6 (high)M-0015$0.1367$0.2187$0.082138%
$6.56 Excluded repair overhead: $0.53
10GPT-6 Astra (high)M-0022$0.1756$0.2228$0.047221%
$6.68 Excluded repair overhead: $0.17
11Claude Fable 5.1 (high)M-0020$0.2772$0.3408$0.063719%
$10.22 Excluded repair overhead: $1.54
#1GLM-5.3-Flash (high)z-ai
Guesser cost / episode
$0.0016
Benchmark cost / episode
$0.0959
Benchmark run cost
$2.88
Excluded repair overhead
$0.02
Explore full run · questions, answers & evidence
#2GPT-5.6 Luna (high)openai
Guesser cost / episode
$0.0063
Benchmark cost / episode
$0.1128
Benchmark run cost
$3.38
Excluded repair overhead
$0.03
Explore full run · questions, answers & evidence
#3Gemini 3.7 Flash (high)google-ai-studio
Guesser cost / episode
$0.0525
Benchmark cost / episode
$0.1267
Benchmark run cost
$3.80
Explore full run · questions, answers & evidence
#4GPT-5.6 Sol (high)openai
Guesser cost / episode
$0.0557
Benchmark cost / episode
$0.1423
Benchmark run cost
$4.27
Excluded repair overhead
$0.11
Explore full run · questions, answers & evidence
#5MiniMax M3 (high)coreweave
Guesser cost / episode
$0.0100
Benchmark cost / episode
$0.1450
Benchmark run cost
$4.35
Excluded repair overhead
$0.18
Explore full run · questions, answers & evidence
#6Gemini 3.8 Flash (high)google-ai-studio
Guesser cost / episode
$0.0875
Benchmark cost / episode
$0.1505
Benchmark run cost
$4.52
Excluded repair overhead
$0.03
Explore full run · questions, answers & evidence
#7gpt-oss-120B (high)cerebras
Guesser cost / episode
$0.0682
Benchmark cost / episode
$0.1523
Benchmark run cost
$4.57
Excluded repair overhead
$0.49
Explore full run · questions, answers & evidence
#8Claude Opus 5 (high)anthropic
Guesser cost / episode
$0.0934
Benchmark cost / episode
$0.1852
Benchmark run cost
$5.55
Excluded repair overhead
$0.76
Explore full run · questions, answers & evidence
#9Grok 4.6 (high)xai
Guesser cost / episode
$0.1367
Benchmark cost / episode
$0.2187
Benchmark run cost
$6.56
Excluded repair overhead
$0.53
Explore full run · questions, answers & evidence
#10GPT-6 Astra (high)openai
Guesser cost / episode
$0.1756
Benchmark cost / episode
$0.2228
Benchmark run cost
$6.68
Excluded repair overhead
$0.17
Explore full run · questions, answers & evidence
#11Claude Fable 5.1 (high)anthropic
Guesser cost / episode
$0.2772
Benchmark cost / episode
$0.3408
Benchmark run cost
$10.22
Excluded repair overhead
$1.54
Explore full run · questions, answers & evidence

Costs are provider-reported values for retained terminal attempts in the selected official runs. Superseded infrastructure attempts are excluded. A missing or unreported provider price can affect the comparison.