Recorded costs

Cost 1

Recorded provider costs for the tested model and benchmark support models.

Models
18
Total benchmark cost
$175.27
Total Guesser cost
$85.11
Support share
51%

Guesser cost

Guesser cost across the run.

Each bar adds the recorded provider cost of Guesser calls in the retained terminal attempts. Superseded infrastructure attempts and support costs are excluded.

Guesser cost · lower is betterSelect a model row to view its full run
  1. 1. Ox Alpha (high): $0.00. View full run for Ox Alpha (high)
  2. 2. Llama 4 Maverick (non-thinking): $0.30. View full run for Llama 4 Maverick (non-thinking)
  3. 3. GPT-5 Nano (medium): $0.48. View full run for GPT-5 Nano (medium)
  4. 4. GPT-5.6 Luna (high): $0.64. View full run for GPT-5.6 Luna (high)
  5. 5. Grok 4.5 (high): $1.00. View full run for Grok 4.5 (high)
  6. 6. gpt-oss-120B (high): $1.27. View full run for gpt-oss-120B (high)
  7. 7. Claude Sonnet 5 (high): $1.72. View full run for Claude Sonnet 5 (high)
  8. 8. Gemini 3.7 Flash (high): $2.28. View full run for Gemini 3.7 Flash (high)
  9. 9. Grok 4.6 (high): $2.29. View full run for Grok 4.6 (high)
  10. 10. Claude Opus 5 (high): $2.47. View full run for Claude Opus 5 (high)
  11. 11. Gemini 3.8 Flash (high): $3.95. View full run for Gemini 3.8 Flash (high)
  12. 12. Gemini 3.6 Flash (high): $5.97. View full run for Gemini 3.6 Flash (high)
  13. 13. GPT-5.6 Sol (high): $5.99. View full run for GPT-5.6 Sol (high)
  14. 14. Claude Fable 5 (high): $6.22. View full run for Claude Fable 5 (high)
  15. 15. Kimi K3 (high): $7.11. View full run for Kimi K3 (high)
  16. 16. GPT-6 Astra (high): $11.48. View full run for GPT-6 Astra (high)
  17. 17. Mistral Medium 3.5 (high): $14.39. View full run for Mistral Medium 3.5 (high)
  18. 18. Claude Fable 5.1 (high): $17.55. View full run for Claude Fable 5.1 (high)

Total benchmark cost

Where the total cost came from.

Bar length shows the retained terminal attempts' benchmark cost. Color separates the Guesser, Primary Oracle, and Adjudication. Expand the exact breakdown to compare Reviewer, Judge, and Validator cost.

Total benchmark cost by component
  • Guesser
  • Primary Oracle
  • Adjudication

Select a model row to view its full run.

Exact adjudication breakdown
Exact adjudication costs by model
ModelReviewerJudgeValidator
Llama 4 Maverick (non-thinking)$0.31$0.60$0.25
GPT-5 Nano (medium)$0.55$0.75$0.02
gpt-oss-120B (high)$0.46$0.63$0.01
Grok 4.5 (high)$0.53$0.65$0.01
Claude Sonnet 5 (high)$0.47$0.49$0.01
Claude Opus 5 (high)$0.41$0.49$0.01
GPT-5.6 Luna (high)$0.67$0.91$0.01
Claude Fable 5 (high)$0.35$0.38$0.00
Grok 4.6 (high)$0.54$1.01$0.00
Gemini 3.7 Flash (high)$0.48$0.85$0.00
Ox Alpha (high)$0.65$1.16$0.01
Gemini 3.6 Flash (high)$0.48$0.46$0.01
GPT-5.6 Sol (high)$0.41$0.35$0.01
Kimi K3 (high)$0.37$0.37$0.02
Gemini 3.8 Flash (high)$0.46$0.75$0.01
Mistral Medium 3.5 (high)$0.50$0.47$0.06
GPT-6 Astra (high)$0.49$0.62$0.01
Claude Fable 5.1 (high)$0.48$0.86$0.01
  1. Llama 4 Maverick (non-thinking): $3.73. Guesser $0.30. Primary Oracle $2.27. Adjudication $1.16. Reviewer $0.31. Judge $0.60. Validator $0.25. View full run for Llama 4 Maverick (non-thinking)
  2. GPT-5 Nano (medium): $4.59. Guesser $0.48. Primary Oracle $2.78. Adjudication $1.32. Reviewer $0.55. Judge $0.75. Validator $0.02. View full run for GPT-5 Nano (medium)
  3. gpt-oss-120B (high): $4.64. Guesser $1.27. Primary Oracle $2.26. Adjudication $1.11. Reviewer $0.46. Judge $0.63. Validator $0.01. View full run for gpt-oss-120B (high)
  4. Grok 4.5 (high): $4.99. Guesser $1.00. Primary Oracle $2.80. Adjudication $1.19. Reviewer $0.53. Judge $0.65. Validator $0.01. View full run for Grok 4.5 (high)
  5. Claude Sonnet 5 (high): $5.01. Guesser $1.72. Primary Oracle $2.31. Adjudication $0.97. Reviewer $0.47. Judge $0.49. Validator $0.01. View full run for Claude Sonnet 5 (high)
  6. Claude Opus 5 (high): $5.47. Guesser $2.47. Primary Oracle $2.08. Adjudication $0.92. Reviewer $0.41. Judge $0.49. Validator $0.01. View full run for Claude Opus 5 (high)
  7. GPT-5.6 Luna (high): $5.88. Guesser $0.64. Primary Oracle $3.65. Adjudication $1.59. Reviewer $0.67. Judge $0.91. Validator $0.01. View full run for GPT-5.6 Luna (high)
  8. Claude Fable 5 (high): $7.87. Guesser $6.22. Primary Oracle $0.91. Adjudication $0.74. Reviewer $0.35. Judge $0.38. Validator $0.00. View full run for Claude Fable 5 (high)
  9. Grok 4.6 (high): $8.25. Guesser $2.29. Primary Oracle $4.41. Adjudication $1.56. Reviewer $0.54. Judge $1.01. Validator $0.00. View full run for Grok 4.6 (high)
  10. Gemini 3.7 Flash (high): $9.58. Guesser $2.28. Primary Oracle $5.97. Adjudication $1.33. Reviewer $0.48. Judge $0.85. Validator $0.00. View full run for Gemini 3.7 Flash (high)
  11. Ox Alpha (high): $9.78. Guesser $0.00. Primary Oracle $7.96. Adjudication $1.82. Reviewer $0.65. Judge $1.16. Validator $0.01. View full run for Ox Alpha (high)
  12. Gemini 3.6 Flash (high): $9.86. Guesser $5.97. Primary Oracle $2.94. Adjudication $0.95. Reviewer $0.48. Judge $0.46. Validator $0.01. View full run for Gemini 3.6 Flash (high)
  13. GPT-5.6 Sol (high): $9.97. Guesser $5.99. Primary Oracle $3.21. Adjudication $0.77. Reviewer $0.41. Judge $0.35. Validator $0.01. View full run for GPT-5.6 Sol (high)
  14. Kimi K3 (high): $10.09. Guesser $7.11. Primary Oracle $2.22. Adjudication $0.75. Reviewer $0.37. Judge $0.37. Validator $0.02. View full run for Kimi K3 (high)
  15. Gemini 3.8 Flash (high): $10.82. Guesser $3.95. Primary Oracle $5.66. Adjudication $1.21. Reviewer $0.46. Judge $0.75. Validator $0.01. View full run for Gemini 3.8 Flash (high)
  16. Mistral Medium 3.5 (high): $18.55. Guesser $14.39. Primary Oracle $3.12. Adjudication $1.03. Reviewer $0.50. Judge $0.47. Validator $0.06. View full run for Mistral Medium 3.5 (high)
  17. GPT-6 Astra (high): $22.60. Guesser $11.48. Primary Oracle $10.01. Adjudication $1.11. Reviewer $0.49. Judge $0.62. Validator $0.01. View full run for GPT-6 Astra (high)
  18. Claude Fable 5.1 (high): $23.61. Guesser $17.55. Primary Oracle $4.71. Adjudication $1.35. Reviewer $0.48. Judge $0.86. Validator $0.01. View full run for Claude Fable 5.1 (high)
Cost comparison
Benchmark cost rankModelGuesser costper episodeBenchmark costper episodeSupport costper episodeSupport shareBenchmark run cost
1Llama 4 Maverick (non-thinking)M-0009$0.0086$0.1064$0.097992%
$3.73
2GPT-5 Nano (medium)M-0003$0.0138$0.1311$0.117389%
$4.59
3gpt-oss-120B (high)M-0002$0.0363$0.1326$0.096373%
$4.64
4Grok 4.5 (high)M-0008$0.0286$0.1426$0.114080%
$4.99
5Claude Sonnet 5 (high)M-0005$0.0493$0.1431$0.093866%
$5.01
6Claude Opus 5 (high)M-0006$0.0705$0.1562$0.085755%
$5.47
7GPT-5.6 Luna (high)M-0001$0.0184$0.1681$0.149789%
$5.88
8Claude Fable 5 (high)M-0014$0.1776$0.2248$0.047221%
$7.87
9Grok 4.6 (high)M-0015$0.0653$0.2357$0.170472%
$8.25 Excluded repair overhead: $1.25
10Gemini 3.7 Flash (high)M-0016$0.0653$0.2737$0.208476%
$9.58
11Ox Alpha (high)M-0017$0.0000$0.2793$0.2793100%
$9.78 Excluded repair overhead: $2.53
12Gemini 3.6 Flash (high)M-0004$0.1705$0.2816$0.111139%
$9.86
13GPT-5.6 Sol (high)M-0010$0.1711$0.2849$0.113940%
$9.97
14Kimi K3 (high)M-0007$0.2032$0.2882$0.085029%
$10.09
15Gemini 3.8 Flash (high)M-0021$0.1128$0.3093$0.196564%
$10.82
16Mistral Medium 3.5 (high)M-0012$0.4113$0.5300$0.118722%
$18.55 Excluded repair overhead: $1.72
17GPT-6 Astra (high)M-0022$0.3279$0.6456$0.317749%
$22.60 Excluded repair overhead: $3.15
18Claude Fable 5.1 (high)M-0020$0.5014$0.6747$0.173326%
$23.61
#1Llama 4 Maverick (non-thinking)parasail
Guesser cost / episode
$0.0086
Benchmark cost / episode
$0.1064
Benchmark run cost
$3.73
Explore full run · questions, answers & evidence
#2GPT-5 Nano (medium)openai
Guesser cost / episode
$0.0138
Benchmark cost / episode
$0.1311
Benchmark run cost
$4.59
Explore full run · questions, answers & evidence
#3gpt-oss-120B (high)cerebras
Guesser cost / episode
$0.0363
Benchmark cost / episode
$0.1326
Benchmark run cost
$4.64
Explore full run · questions, answers & evidence
#4Grok 4.5 (high)xai
Guesser cost / episode
$0.0286
Benchmark cost / episode
$0.1426
Benchmark run cost
$4.99
Explore full run · questions, answers & evidence
#5Claude Sonnet 5 (high)anthropic
Guesser cost / episode
$0.0493
Benchmark cost / episode
$0.1431
Benchmark run cost
$5.01
Explore full run · questions, answers & evidence
#6Claude Opus 5 (high)anthropic
Guesser cost / episode
$0.0705
Benchmark cost / episode
$0.1562
Benchmark run cost
$5.47
Explore full run · questions, answers & evidence
#7GPT-5.6 Luna (high)openai
Guesser cost / episode
$0.0184
Benchmark cost / episode
$0.1681
Benchmark run cost
$5.88
Explore full run · questions, answers & evidence
#8Claude Fable 5 (high)anthropic
Guesser cost / episode
$0.1776
Benchmark cost / episode
$0.2248
Benchmark run cost
$7.87
Explore full run · questions, answers & evidence
#9Grok 4.6 (high)xai
Guesser cost / episode
$0.0653
Benchmark cost / episode
$0.2357
Benchmark run cost
$8.25
Excluded repair overhead
$1.25
Explore full run · questions, answers & evidence
#10Gemini 3.7 Flash (high)google-ai-studio
Guesser cost / episode
$0.0653
Benchmark cost / episode
$0.2737
Benchmark run cost
$9.58
Explore full run · questions, answers & evidence
#11Ox Alpha (high)stealth
Guesser cost / episode
$0.0000
Benchmark cost / episode
$0.2793
Benchmark run cost
$9.78
Excluded repair overhead
$2.53
Explore full run · questions, answers & evidence
#12Gemini 3.6 Flash (high)google-vertex
Guesser cost / episode
$0.1705
Benchmark cost / episode
$0.2816
Benchmark run cost
$9.86
Explore full run · questions, answers & evidence
#13GPT-5.6 Sol (high)openai
Guesser cost / episode
$0.1711
Benchmark cost / episode
$0.2849
Benchmark run cost
$9.97
Explore full run · questions, answers & evidence
#14Kimi K3 (high)moonshotai
Guesser cost / episode
$0.2032
Benchmark cost / episode
$0.2882
Benchmark run cost
$10.09
Explore full run · questions, answers & evidence
#15Gemini 3.8 Flash (high)google-ai-studio
Guesser cost / episode
$0.1128
Benchmark cost / episode
$0.3093
Benchmark run cost
$10.82
Explore full run · questions, answers & evidence
#16Mistral Medium 3.5 (high)mistral
Guesser cost / episode
$0.4113
Benchmark cost / episode
$0.5300
Benchmark run cost
$18.55
Excluded repair overhead
$1.72
Explore full run · questions, answers & evidence
#17GPT-6 Astra (high)openai
Guesser cost / episode
$0.3279
Benchmark cost / episode
$0.6456
Benchmark run cost
$22.60
Excluded repair overhead
$3.15
Explore full run · questions, answers & evidence
#18Claude Fable 5.1 (high)anthropic
Guesser cost / episode
$0.5014
Benchmark cost / episode
$0.6747
Benchmark run cost
$23.61
Explore full run · questions, answers & evidence

Costs are provider-reported values for retained terminal attempts in the selected official runs. Superseded infrastructure attempts are excluded. A missing or unreported provider price can affect the comparison.