Recorded costs
Cost 1
Recorded provider costs for the tested model and benchmark support models.
- Models
- 18
- Total benchmark cost
- $175.27
- Total Guesser cost
- $85.11
- Support share
- 51%
Guesser cost
Guesser cost across the run.
Each bar adds the recorded provider cost of Guesser calls in the retained terminal attempts. Superseded infrastructure attempts and support costs are excluded.
- 1. Ox Alpha (high): $0.00. View full run for Ox Alpha (high)
- 2. Llama 4 Maverick (non-thinking): $0.30. View full run for Llama 4 Maverick (non-thinking)
- 3. GPT-5 Nano (medium): $0.48. View full run for GPT-5 Nano (medium)
- 4. GPT-5.6 Luna (high): $0.64. View full run for GPT-5.6 Luna (high)
- 5. Grok 4.5 (high): $1.00. View full run for Grok 4.5 (high)
- 6. gpt-oss-120B (high): $1.27. View full run for gpt-oss-120B (high)
- 7. Claude Sonnet 5 (high): $1.72. View full run for Claude Sonnet 5 (high)
- 8. Gemini 3.7 Flash (high): $2.28. View full run for Gemini 3.7 Flash (high)
- 9. Grok 4.6 (high): $2.29. View full run for Grok 4.6 (high)
- 10. Claude Opus 5 (high): $2.47. View full run for Claude Opus 5 (high)
- 11. Gemini 3.8 Flash (high): $3.95. View full run for Gemini 3.8 Flash (high)
- 12. Gemini 3.6 Flash (high): $5.97. View full run for Gemini 3.6 Flash (high)
- 13. GPT-5.6 Sol (high): $5.99. View full run for GPT-5.6 Sol (high)
- 14. Claude Fable 5 (high): $6.22. View full run for Claude Fable 5 (high)
- 15. Kimi K3 (high): $7.11. View full run for Kimi K3 (high)
- 16. GPT-6 Astra (high): $11.48. View full run for GPT-6 Astra (high)
- 17. Mistral Medium 3.5 (high): $14.39. View full run for Mistral Medium 3.5 (high)
- 18. Claude Fable 5.1 (high): $17.55. View full run for Claude Fable 5.1 (high)
Total benchmark cost
Where the total cost came from.
Bar length shows the retained terminal attempts' benchmark cost. Color separates the Guesser, Primary Oracle, and Adjudication. Expand the exact breakdown to compare Reviewer, Judge, and Validator cost.
- Guesser
- Primary Oracle
- Adjudication
Select a model row to view its full run.
Exact adjudication breakdown
| Model | Reviewer | Judge | Validator |
|---|---|---|---|
| Llama 4 Maverick (non-thinking) | $0.31 | $0.60 | $0.25 |
| GPT-5 Nano (medium) | $0.55 | $0.75 | $0.02 |
| gpt-oss-120B (high) | $0.46 | $0.63 | $0.01 |
| Grok 4.5 (high) | $0.53 | $0.65 | $0.01 |
| Claude Sonnet 5 (high) | $0.47 | $0.49 | $0.01 |
| Claude Opus 5 (high) | $0.41 | $0.49 | $0.01 |
| GPT-5.6 Luna (high) | $0.67 | $0.91 | $0.01 |
| Claude Fable 5 (high) | $0.35 | $0.38 | $0.00 |
| Grok 4.6 (high) | $0.54 | $1.01 | $0.00 |
| Gemini 3.7 Flash (high) | $0.48 | $0.85 | $0.00 |
| Ox Alpha (high) | $0.65 | $1.16 | $0.01 |
| Gemini 3.6 Flash (high) | $0.48 | $0.46 | $0.01 |
| GPT-5.6 Sol (high) | $0.41 | $0.35 | $0.01 |
| Kimi K3 (high) | $0.37 | $0.37 | $0.02 |
| Gemini 3.8 Flash (high) | $0.46 | $0.75 | $0.01 |
| Mistral Medium 3.5 (high) | $0.50 | $0.47 | $0.06 |
| GPT-6 Astra (high) | $0.49 | $0.62 | $0.01 |
| Claude Fable 5.1 (high) | $0.48 | $0.86 | $0.01 |
- Llama 4 Maverick (non-thinking): $3.73. Guesser $0.30. Primary Oracle $2.27. Adjudication $1.16. Reviewer $0.31. Judge $0.60. Validator $0.25. View full run for Llama 4 Maverick (non-thinking)
- GPT-5 Nano (medium): $4.59. Guesser $0.48. Primary Oracle $2.78. Adjudication $1.32. Reviewer $0.55. Judge $0.75. Validator $0.02. View full run for GPT-5 Nano (medium)
- gpt-oss-120B (high): $4.64. Guesser $1.27. Primary Oracle $2.26. Adjudication $1.11. Reviewer $0.46. Judge $0.63. Validator $0.01. View full run for gpt-oss-120B (high)
- Grok 4.5 (high): $4.99. Guesser $1.00. Primary Oracle $2.80. Adjudication $1.19. Reviewer $0.53. Judge $0.65. Validator $0.01. View full run for Grok 4.5 (high)
- Claude Sonnet 5 (high): $5.01. Guesser $1.72. Primary Oracle $2.31. Adjudication $0.97. Reviewer $0.47. Judge $0.49. Validator $0.01. View full run for Claude Sonnet 5 (high)
- Claude Opus 5 (high): $5.47. Guesser $2.47. Primary Oracle $2.08. Adjudication $0.92. Reviewer $0.41. Judge $0.49. Validator $0.01. View full run for Claude Opus 5 (high)
- GPT-5.6 Luna (high): $5.88. Guesser $0.64. Primary Oracle $3.65. Adjudication $1.59. Reviewer $0.67. Judge $0.91. Validator $0.01. View full run for GPT-5.6 Luna (high)
- Claude Fable 5 (high): $7.87. Guesser $6.22. Primary Oracle $0.91. Adjudication $0.74. Reviewer $0.35. Judge $0.38. Validator $0.00. View full run for Claude Fable 5 (high)
- Grok 4.6 (high): $8.25. Guesser $2.29. Primary Oracle $4.41. Adjudication $1.56. Reviewer $0.54. Judge $1.01. Validator $0.00. View full run for Grok 4.6 (high)
- Gemini 3.7 Flash (high): $9.58. Guesser $2.28. Primary Oracle $5.97. Adjudication $1.33. Reviewer $0.48. Judge $0.85. Validator $0.00. View full run for Gemini 3.7 Flash (high)
- Ox Alpha (high): $9.78. Guesser $0.00. Primary Oracle $7.96. Adjudication $1.82. Reviewer $0.65. Judge $1.16. Validator $0.01. View full run for Ox Alpha (high)
- Gemini 3.6 Flash (high): $9.86. Guesser $5.97. Primary Oracle $2.94. Adjudication $0.95. Reviewer $0.48. Judge $0.46. Validator $0.01. View full run for Gemini 3.6 Flash (high)
- GPT-5.6 Sol (high): $9.97. Guesser $5.99. Primary Oracle $3.21. Adjudication $0.77. Reviewer $0.41. Judge $0.35. Validator $0.01. View full run for GPT-5.6 Sol (high)
- Kimi K3 (high): $10.09. Guesser $7.11. Primary Oracle $2.22. Adjudication $0.75. Reviewer $0.37. Judge $0.37. Validator $0.02. View full run for Kimi K3 (high)
- Gemini 3.8 Flash (high): $10.82. Guesser $3.95. Primary Oracle $5.66. Adjudication $1.21. Reviewer $0.46. Judge $0.75. Validator $0.01. View full run for Gemini 3.8 Flash (high)
- Mistral Medium 3.5 (high): $18.55. Guesser $14.39. Primary Oracle $3.12. Adjudication $1.03. Reviewer $0.50. Judge $0.47. Validator $0.06. View full run for Mistral Medium 3.5 (high)
- GPT-6 Astra (high): $22.60. Guesser $11.48. Primary Oracle $10.01. Adjudication $1.11. Reviewer $0.49. Judge $0.62. Validator $0.01. View full run for GPT-6 Astra (high)
- Claude Fable 5.1 (high): $23.61. Guesser $17.55. Primary Oracle $4.71. Adjudication $1.35. Reviewer $0.48. Judge $0.86. Validator $0.01. View full run for Claude Fable 5.1 (high)
| Benchmark cost rank | Model | Guesser costper episode | Benchmark costper episode | Support costper episode | Support share | Benchmark run cost |
|---|---|---|---|---|---|---|
| 1 | Llama 4 Maverick (non-thinking) | $0.0086 | $0.1064 | $0.0979 | 92% | $3.73 |
| 2 | GPT-5 Nano (medium) | $0.0138 | $0.1311 | $0.1173 | 89% | $4.59 |
| 3 | gpt-oss-120B (high) | $0.0363 | $0.1326 | $0.0963 | 73% | $4.64 |
| 4 | Grok 4.5 (high) | $0.0286 | $0.1426 | $0.1140 | 80% | $4.99 |
| 5 | Claude Sonnet 5 (high) | $0.0493 | $0.1431 | $0.0938 | 66% | $5.01 |
| 6 | Claude Opus 5 (high) | $0.0705 | $0.1562 | $0.0857 | 55% | $5.47 |
| 7 | GPT-5.6 Luna (high) | $0.0184 | $0.1681 | $0.1497 | 89% | $5.88 |
| 8 | Claude Fable 5 (high) | $0.1776 | $0.2248 | $0.0472 | 21% | $7.87 |
| 9 | Grok 4.6 (high) | $0.0653 | $0.2357 | $0.1704 | 72% | $8.25 Excluded repair overhead: $1.25 |
| 10 | Gemini 3.7 Flash (high) | $0.0653 | $0.2737 | $0.2084 | 76% | $9.58 |
| 11 | Ox Alpha (high) | $0.0000 | $0.2793 | $0.2793 | 100% | $9.78 Excluded repair overhead: $2.53 |
| 12 | Gemini 3.6 Flash (high) | $0.1705 | $0.2816 | $0.1111 | 39% | $9.86 |
| 13 | GPT-5.6 Sol (high) | $0.1711 | $0.2849 | $0.1139 | 40% | $9.97 |
| 14 | Kimi K3 (high) | $0.2032 | $0.2882 | $0.0850 | 29% | $10.09 |
| 15 | Gemini 3.8 Flash (high) | $0.1128 | $0.3093 | $0.1965 | 64% | $10.82 |
| 16 | Mistral Medium 3.5 (high) | $0.4113 | $0.5300 | $0.1187 | 22% | $18.55 Excluded repair overhead: $1.72 |
| 17 | GPT-6 Astra (high) | $0.3279 | $0.6456 | $0.3177 | 49% | $22.60 Excluded repair overhead: $3.15 |
| 18 | Claude Fable 5.1 (high) | $0.5014 | $0.6747 | $0.1733 | 26% | $23.61 |
- Guesser cost / episode
- $0.0086
- Benchmark cost / episode
- $0.1064
- Benchmark run cost
- $3.73
- Guesser cost / episode
- $0.0138
- Benchmark cost / episode
- $0.1311
- Benchmark run cost
- $4.59
- Guesser cost / episode
- $0.0363
- Benchmark cost / episode
- $0.1326
- Benchmark run cost
- $4.64
- Guesser cost / episode
- $0.0286
- Benchmark cost / episode
- $0.1426
- Benchmark run cost
- $4.99
- Guesser cost / episode
- $0.0493
- Benchmark cost / episode
- $0.1431
- Benchmark run cost
- $5.01
- Guesser cost / episode
- $0.0705
- Benchmark cost / episode
- $0.1562
- Benchmark run cost
- $5.47
- Guesser cost / episode
- $0.0184
- Benchmark cost / episode
- $0.1681
- Benchmark run cost
- $5.88
- Guesser cost / episode
- $0.1776
- Benchmark cost / episode
- $0.2248
- Benchmark run cost
- $7.87
- Guesser cost / episode
- $0.0653
- Benchmark cost / episode
- $0.2357
- Benchmark run cost
- $8.25
- Excluded repair overhead
- $1.25
- Guesser cost / episode
- $0.0653
- Benchmark cost / episode
- $0.2737
- Benchmark run cost
- $9.58
- Guesser cost / episode
- $0.0000
- Benchmark cost / episode
- $0.2793
- Benchmark run cost
- $9.78
- Excluded repair overhead
- $2.53
- Guesser cost / episode
- $0.1705
- Benchmark cost / episode
- $0.2816
- Benchmark run cost
- $9.86
- Guesser cost / episode
- $0.1711
- Benchmark cost / episode
- $0.2849
- Benchmark run cost
- $9.97
- Guesser cost / episode
- $0.2032
- Benchmark cost / episode
- $0.2882
- Benchmark run cost
- $10.09
- Guesser cost / episode
- $0.1128
- Benchmark cost / episode
- $0.3093
- Benchmark run cost
- $10.82
- Guesser cost / episode
- $0.4113
- Benchmark cost / episode
- $0.5300
- Benchmark run cost
- $18.55
- Excluded repair overhead
- $1.72
- Guesser cost / episode
- $0.3279
- Benchmark cost / episode
- $0.6456
- Benchmark run cost
- $22.60
- Excluded repair overhead
- $3.15
- Guesser cost / episode
- $0.5014
- Benchmark cost / episode
- $0.6747
- Benchmark run cost
- $23.61
Costs are provider-reported values for retained terminal attempts in the selected official runs. Superseded infrastructure attempts are excluded. A missing or unreported provider price can affect the comparison.