Recorded costs
Cost 1.1
Recorded provider costs for the tested model and benchmark support models.
- Models
- 11
- Total benchmark cost
- $56.79
- Total Guesser cost
- $28.94
- Support share
- 49%
Guesser cost
Guesser cost across the run.
Each bar adds the recorded provider cost of Guesser calls in the retained terminal attempts. Superseded infrastructure attempts and support costs are excluded.
- 1. GLM-5.3-Flash (high): $0.05. View full run for GLM-5.3-Flash (high)
- 2. GPT-5.6 Luna (high): $0.19. View full run for GPT-5.6 Luna (high)
- 3. MiniMax M3 (high): $0.30. View full run for MiniMax M3 (high)
- 4. Gemini 3.7 Flash (high): $1.58. View full run for Gemini 3.7 Flash (high)
- 5. GPT-5.6 Sol (high): $1.67. View full run for GPT-5.6 Sol (high)
- 6. gpt-oss-120B (high): $2.04. View full run for gpt-oss-120B (high)
- 7. Gemini 3.8 Flash (high): $2.62. View full run for Gemini 3.8 Flash (high)
- 8. Claude Opus 5 (high): $2.80. View full run for Claude Opus 5 (high)
- 9. Grok 4.6 (high): $4.10. View full run for Grok 4.6 (high)
- 10. GPT-6 Astra (high): $5.27. View full run for GPT-6 Astra (high)
- 11. Claude Fable 5.1 (high): $8.31. View full run for Claude Fable 5.1 (high)
Total benchmark cost
Where the total cost came from.
Bar length shows the retained terminal attempts' benchmark cost. Color separates the Guesser, Primary Oracle, and Adjudication. Expand the exact breakdown to compare Reviewer, Judge, and Validator cost.
- Guesser
- Primary Oracle
- Adjudication
Select a model row to view its full run.
Exact adjudication breakdown
| Model | Reviewer | Judge | Validator |
|---|---|---|---|
| GLM-5.3-Flash (high) | $0.69 | $0.58 | $0.01 |
| GPT-5.6 Luna (high) | $0.83 | $0.75 | $0.01 |
| Gemini 3.7 Flash (high) | $0.60 | $0.42 | $0.01 |
| GPT-5.6 Sol (high) | $0.67 | $0.54 | $0.01 |
| MiniMax M3 (high) | $0.86 | $1.13 | $0.02 |
| Gemini 3.8 Flash (high) | $0.55 | $0.25 | $0.01 |
| gpt-oss-120B (high) | $0.66 | $0.38 | $0.01 |
| Claude Opus 5 (high) | $0.60 | $0.72 | $0.01 |
| Grok 4.6 (high) | $0.65 | $0.47 | $0.01 |
| GPT-6 Astra (high) | $0.36 | $0.30 | $0.01 |
| Claude Fable 5.1 (high) | $0.51 | $0.36 | $0.01 |
- GLM-5.3-Flash (high): $2.88. Guesser $0.05. Primary Oracle $1.55. Adjudication $1.28. Reviewer $0.69. Judge $0.58. Validator $0.01. View full run for GLM-5.3-Flash (high)
- GPT-5.6 Luna (high): $3.38. Guesser $0.19. Primary Oracle $1.61. Adjudication $1.59. Reviewer $0.83. Judge $0.75. Validator $0.01. View full run for GPT-5.6 Luna (high)
- Gemini 3.7 Flash (high): $3.80. Guesser $1.58. Primary Oracle $1.19. Adjudication $1.03. Reviewer $0.60. Judge $0.42. Validator $0.01. View full run for Gemini 3.7 Flash (high)
- GPT-5.6 Sol (high): $4.27. Guesser $1.67. Primary Oracle $1.37. Adjudication $1.22. Reviewer $0.67. Judge $0.54. Validator $0.01. View full run for GPT-5.6 Sol (high)
- MiniMax M3 (high): $4.35. Guesser $0.30. Primary Oracle $2.04. Adjudication $2.01. Reviewer $0.86. Judge $1.13. Validator $0.02. View full run for MiniMax M3 (high)
- Gemini 3.8 Flash (high): $4.52. Guesser $2.62. Primary Oracle $1.09. Adjudication $0.80. Reviewer $0.55. Judge $0.25. Validator $0.01. View full run for Gemini 3.8 Flash (high)
- gpt-oss-120B (high): $4.57. Guesser $2.04. Primary Oracle $1.48. Adjudication $1.05. Reviewer $0.66. Judge $0.38. Validator $0.01. View full run for gpt-oss-120B (high)
- Claude Opus 5 (high): $5.55. Guesser $2.80. Primary Oracle $1.43. Adjudication $1.32. Reviewer $0.60. Judge $0.72. Validator $0.01. View full run for Claude Opus 5 (high)
- Grok 4.6 (high): $6.56. Guesser $4.10. Primary Oracle $1.34. Adjudication $1.12. Reviewer $0.65. Judge $0.47. Validator $0.01. View full run for Grok 4.6 (high)
- GPT-6 Astra (high): $6.68. Guesser $5.27. Primary Oracle $0.75. Adjudication $0.67. Reviewer $0.36. Judge $0.30. Validator $0.01. View full run for GPT-6 Astra (high)
- Claude Fable 5.1 (high): $10.22. Guesser $8.31. Primary Oracle $1.03. Adjudication $0.88. Reviewer $0.51. Judge $0.36. Validator $0.01. View full run for Claude Fable 5.1 (high)
| Benchmark cost rank | Model | Guesser costper episode | Benchmark costper episode | Support costper episode | Support share | Benchmark run cost |
|---|---|---|---|---|---|---|
| 1 | GLM-5.3-Flash (high) | $0.0016 | $0.0959 | $0.0944 | 98% | $2.88 Excluded repair overhead: $0.02 |
| 2 | GPT-5.6 Luna (high) | $0.0063 | $0.1128 | $0.1065 | 94% | $3.38 Excluded repair overhead: $0.03 |
| 3 | Gemini 3.7 Flash (high) | $0.0525 | $0.1267 | $0.0742 | 59% | $3.80 |
| 4 | GPT-5.6 Sol (high) | $0.0557 | $0.1423 | $0.0865 | 61% | $4.27 Excluded repair overhead: $0.11 |
| 5 | MiniMax M3 (high) | $0.0100 | $0.1450 | $0.1350 | 93% | $4.35 Excluded repair overhead: $0.18 |
| 6 | Gemini 3.8 Flash (high) | $0.0875 | $0.1505 | $0.0630 | 42% | $4.52 Excluded repair overhead: $0.03 |
| 7 | gpt-oss-120B (high) | $0.0682 | $0.1523 | $0.0842 | 55% | $4.57 Excluded repair overhead: $0.49 |
| 8 | Claude Opus 5 (high) | $0.0934 | $0.1852 | $0.0917 | 50% | $5.55 Excluded repair overhead: $0.76 |
| 9 | Grok 4.6 (high) | $0.1367 | $0.2187 | $0.0821 | 38% | $6.56 Excluded repair overhead: $0.53 |
| 10 | GPT-6 Astra (high) | $0.1756 | $0.2228 | $0.0472 | 21% | $6.68 Excluded repair overhead: $0.17 |
| 11 | Claude Fable 5.1 (high) | $0.2772 | $0.3408 | $0.0637 | 19% | $10.22 Excluded repair overhead: $1.54 |
- Guesser cost / episode
- $0.0016
- Benchmark cost / episode
- $0.0959
- Benchmark run cost
- $2.88
- Excluded repair overhead
- $0.02
- Guesser cost / episode
- $0.0063
- Benchmark cost / episode
- $0.1128
- Benchmark run cost
- $3.38
- Excluded repair overhead
- $0.03
- Guesser cost / episode
- $0.0525
- Benchmark cost / episode
- $0.1267
- Benchmark run cost
- $3.80
- Guesser cost / episode
- $0.0557
- Benchmark cost / episode
- $0.1423
- Benchmark run cost
- $4.27
- Excluded repair overhead
- $0.11
- Guesser cost / episode
- $0.0100
- Benchmark cost / episode
- $0.1450
- Benchmark run cost
- $4.35
- Excluded repair overhead
- $0.18
- Guesser cost / episode
- $0.0875
- Benchmark cost / episode
- $0.1505
- Benchmark run cost
- $4.52
- Excluded repair overhead
- $0.03
- Guesser cost / episode
- $0.0682
- Benchmark cost / episode
- $0.1523
- Benchmark run cost
- $4.57
- Excluded repair overhead
- $0.49
- Guesser cost / episode
- $0.0934
- Benchmark cost / episode
- $0.1852
- Benchmark run cost
- $5.55
- Excluded repair overhead
- $0.76
- Guesser cost / episode
- $0.1367
- Benchmark cost / episode
- $0.2187
- Benchmark run cost
- $6.56
- Excluded repair overhead
- $0.53
- Guesser cost / episode
- $0.1756
- Benchmark cost / episode
- $0.2228
- Benchmark run cost
- $6.68
- Excluded repair overhead
- $0.17
- Guesser cost / episode
- $0.2772
- Benchmark cost / episode
- $0.3408
- Benchmark run cost
- $10.22
- Excluded repair overhead
- $1.54
Costs are provider-reported values for retained terminal attempts in the selected official runs. Superseded infrastructure attempts are excluded. A missing or unreported provider price can affect the comparison.