Model and benchmark time
Time 1.1
Provider-reported model-call time and end-to-end benchmark runtime.
- Models
- 11
- Median model time
- 41m 23s
- Total model time
- 11h 2m
- Total end-to-end time
- 70h 34m
Model time
Model time across the run.
Each bar adds the provider-reported time of every call to the model under test. Shorter is faster. This is not the wall-clock benchmark runtime.
- 1. Gemini 3.7 Flash (high): 25m 46s. 51.5 s per episode · 525 calls. View full run for Gemini 3.7 Flash (high)
- 2. gpt-oss-120B (high): 32m 30s. 1m 5s per episode · 662 calls. View full run for gpt-oss-120B (high)
- 3. MiniMax M3 (high): 35m 46s. 1m 12s per episode · 890 calls. View full run for MiniMax M3 (high)
- 4. GPT-6 Astra (high): 40m 43s. 1m 21s per episode · 440 calls. View full run for GPT-6 Astra (high)
- 5. Claude Opus 5 (high): 41m 2s. 1m 22s per episode · 570 calls. View full run for Claude Opus 5 (high)
- 6. GPT-5.6 Luna (high): 41m 23s. 1m 23s per episode · 658 calls. View full run for GPT-5.6 Luna (high)
- 7. Gemini 3.8 Flash (high): 51m 12s. 1m 42s per episode · 476 calls. View full run for Gemini 3.8 Flash (high)
- 8. GLM-5.3-Flash (high): 51m 31s. 1m 43s per episode · 720 calls. View full run for GLM-5.3-Flash (high)
- 9. GPT-5.6 Sol (high): 57m 25s. 1m 55s per episode · 571 calls. View full run for GPT-5.6 Sol (high)
- 10. Claude Fable 5.1 (high): 1h 19m. 2m 39s per episode · 481 calls. View full run for Claude Fable 5.1 (high)
- 11. Grok 4.6 (high): 3h 25m. 6m 51s per episode · 544 calls. View full run for Grok 4.6 (high)
End-to-end time
End-to-end benchmark time.
Each bar is the wall-clock time from run creation to final status. It includes model calls, adjudication, scheduling, concurrency, and other benchmark work.
- 1. Gemini 3.7 Flash (high): 1h 53m. 3m 47s per episode · 25m 46s model time. View full run for Gemini 3.7 Flash (high)
- 2. GPT-6 Astra (high): 1h 58m. 3m 56s per episode · 40m 43s model time. View full run for GPT-6 Astra (high)
- 3. gpt-oss-120B (high): 2h 53m. 5m 46s per episode · 32m 30s model time. View full run for gpt-oss-120B (high)
- 4. GLM-5.3-Flash (high): 2h 58m. 5m 56s per episode · 51m 31s model time. View full run for GLM-5.3-Flash (high)
- 5. Gemini 3.8 Flash (high): 2h 58m. 5m 56s per episode · 51m 12s model time. View full run for Gemini 3.8 Flash (high)
- 6. Claude Fable 5.1 (high): 3h 8m. 6m 17s per episode · 1h 19m model time. View full run for Claude Fable 5.1 (high)
- 7. Grok 4.6 (high): 7h 2m. 14m 3s per episode · 3h 25m model time. View full run for Grok 4.6 (high)
- 8. MiniMax M3 (high): 8h 54m. 17m 48s per episode · 35m 46s model time. View full run for MiniMax M3 (high)
- 9. GPT-5.6 Sol (high): 9h 2m. 18m 4s per episode · 57m 25s model time. View full run for GPT-5.6 Sol (high)
- 10. GPT-5.6 Luna (high): 9h 38m. 19m 17s per episode · 41m 23s model time. View full run for GPT-5.6 Luna (high)
- 11. Claude Opus 5 (high): 20h 9m. 40m 18s per episode · 41m 2s model time. View full run for Claude Opus 5 (high)
| Model-time rank | Model | Model time | End-to-end time | Model timeper episode | Model timeper call | Model calls |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.7 Flash (high) | 25m 46s | 1h 53m | 51.5 s | 2.9 s | 525 |
| 2 | gpt-oss-120B (high) | 32m 30s | 2h 53m | 1m 5s | 2.9 s | 662 |
| 3 | MiniMax M3 (high) | 35m 46s | 8h 54m | 1m 12s | 2.4 s | 890 |
| 4 | GPT-6 Astra (high) | 40m 43s | 1h 58m | 1m 21s | 5.6 s | 440 |
| 5 | Claude Opus 5 (high) | 41m 2s | 20h 9m | 1m 22s | 4.3 s | 570 |
| 6 | GPT-5.6 Luna (high) | 41m 23s | 9h 38m | 1m 23s | 3.8 s | 658 |
| 7 | Gemini 3.8 Flash (high) | 51m 12s | 2h 58m | 1m 42s | 6.5 s | 476 |
| 8 | GLM-5.3-Flash (high) | 51m 31s | 2h 58m | 1m 43s | 4.3 s | 720 |
| 9 | GPT-5.6 Sol (high) | 57m 25s | 9h 2m | 1m 55s | 6.0 s | 571 |
| 10 | Claude Fable 5.1 (high) | 1h 19m | 3h 8m | 2m 39s | 9.9 s | 481 |
| 11 | Grok 4.6 (high) | 3h 25m | 7h 2m | 6m 51s | 22.7 s | 544 |
#1Gemini 3.7 Flash (high)google-ai-studio
- Model time
- 25m 46s
- End-to-end time
- 1h 53m
- Model time / episode
- 51.5 s
Explore full run · questions, answers & evidence
#2gpt-oss-120B (high)cerebras
- Model time
- 32m 30s
- End-to-end time
- 2h 53m
- Model time / episode
- 1m 5s
Explore full run · questions, answers & evidence
#3MiniMax M3 (high)coreweave
- Model time
- 35m 46s
- End-to-end time
- 8h 54m
- Model time / episode
- 1m 12s
Explore full run · questions, answers & evidence
#4GPT-6 Astra (high)openai
- Model time
- 40m 43s
- End-to-end time
- 1h 58m
- Model time / episode
- 1m 21s
Explore full run · questions, answers & evidence
#5Claude Opus 5 (high)anthropic
- Model time
- 41m 2s
- End-to-end time
- 20h 9m
- Model time / episode
- 1m 22s
Explore full run · questions, answers & evidence
#6GPT-5.6 Luna (high)openai
- Model time
- 41m 23s
- End-to-end time
- 9h 38m
- Model time / episode
- 1m 23s
Explore full run · questions, answers & evidence
#7Gemini 3.8 Flash (high)google-ai-studio
- Model time
- 51m 12s
- End-to-end time
- 2h 58m
- Model time / episode
- 1m 42s
Explore full run · questions, answers & evidence
#8GLM-5.3-Flash (high)z-ai
- Model time
- 51m 31s
- End-to-end time
- 2h 58m
- Model time / episode
- 1m 43s
Explore full run · questions, answers & evidence
#9GPT-5.6 Sol (high)openai
- Model time
- 57m 25s
- End-to-end time
- 9h 2m
- Model time / episode
- 1m 55s
Explore full run · questions, answers & evidence
#10Claude Fable 5.1 (high)anthropic
- Model time
- 1h 19m
- End-to-end time
- 3h 8m
- Model time / episode
- 2m 39s
Explore full run · questions, answers & evidence
#11Grok 4.6 (high)xai
- Model time
- 3h 25m
- End-to-end time
- 7h 2m
- Model time / episode
- 6m 51s
Explore full run · questions, answers & evidence
The first chart ranks model time. The second ranks end-to-end time, so the order can change.