Model and benchmark time
Time 1
Provider-reported model-call time and end-to-end benchmark runtime.
- Models
- 18
- Median model time
- 1h 11m
- Total model time
- 33h 50m
- Total end-to-end time
- 78h 60m
Model time
Model time across the run.
Each bar adds the provider-reported time of every call to the model under test. Shorter is faster. This is not the wall-clock benchmark runtime.
- 1. Llama 4 Maverick (non-thinking): 16m 16s. 27.9 s per episode · 1,132 calls. View full run for Llama 4 Maverick (non-thinking)
- 2. Grok 4.5 (high): 40m 43s. 1m 10s per episode · 564 calls. View full run for Grok 4.5 (high)
- 3. Claude Opus 5 (high): 42m 27s. 1m 13s per episode · 467 calls. View full run for Claude Opus 5 (high)
- 4. Gemini 3.7 Flash (high): 44m 51s. 1m 17s per episode · 525 calls. View full run for Gemini 3.7 Flash (high)
- 5. GPT-5.6 Luna (high): 45m 12s. 1m 17s per episode · 653 calls. View full run for GPT-5.6 Luna (high)
- 6. Claude Sonnet 5 (high): 49m 5s. 1m 24s per episode · 551 calls. View full run for Claude Sonnet 5 (high)
- 7. gpt-oss-120B (high): 53m 48s. 1m 32s per episode · 490 calls. View full run for gpt-oss-120B (high)
- 8. Gemini 3.6 Flash (high): 1h 9m. 1m 59s per episode · 625 calls. View full run for Gemini 3.6 Flash (high)
- 9. Claude Fable 5 (high): 1h 9m. 1m 59s per episode · 457 calls. View full run for Claude Fable 5 (high)
- 10. Gemini 3.8 Flash (high): 1h 14m. 2m 6s per episode · 495 calls. View full run for Gemini 3.8 Flash (high)
- 11. GPT-5.6 Sol (high): 1h 19m. 2m 16s per episode · 539 calls. View full run for GPT-5.6 Sol (high)
- 12. GPT-6 Astra (high): 1h 35m. 2m 42s per episode · 561 calls. View full run for GPT-6 Astra (high)
- 13. Grok 4.6 (high): 1h 49m. 3m 8s per episode · 535 calls. View full run for Grok 4.6 (high)
- 14. Claude Fable 5.1 (high): 1h 55m. 3m 17s per episode · 458 calls. View full run for Claude Fable 5.1 (high)
- 15. Ox Alpha (high): 3h 3m. 5m 14s per episode · 648 calls. View full run for Ox Alpha (high)
- 16. GPT-5 Nano (medium): 3h 14m. 5m 32s per episode · 575 calls. View full run for GPT-5 Nano (medium)
- 17. Kimi K3 (high): 4h 2m. 6m 54s per episode · 479 calls. View full run for Kimi K3 (high)
- 18. Mistral Medium 3.5 (high): 8h 29m. 14m 32s per episode · 733 calls. View full run for Mistral Medium 3.5 (high)
End-to-end time
End-to-end benchmark time.
Each bar is the wall-clock time from run creation to final status. It includes model calls, adjudication, scheduling, concurrency, and other benchmark work.
- 1. Llama 4 Maverick (non-thinking): 1h 53m. 3m 14s per episode · 16m 16s model time. View full run for Llama 4 Maverick (non-thinking)
- 2. Claude Opus 5 (high): 2h 4m. 3m 33s per episode · 42m 27s model time. View full run for Claude Opus 5 (high)
- 3. gpt-oss-120B (high): 2h 13m. 3m 49s per episode · 53m 48s model time. View full run for gpt-oss-120B (high)
- 4. Grok 4.5 (high): 2h 15m. 3m 51s per episode · 40m 43s model time. View full run for Grok 4.5 (high)
- 5. Claude Sonnet 5 (high): 2h 16m. 3m 52s per episode · 49m 5s model time. View full run for Claude Sonnet 5 (high)
- 6. Claude Fable 5 (high): 2h 40m. 4m 34s per episode · 1h 9m model time. View full run for Claude Fable 5 (high)
- 7. Gemini 3.7 Flash (high): 2h 40m. 4m 35s per episode · 44m 51s model time. View full run for Gemini 3.7 Flash (high)
- 8. GPT-5.6 Luna (high): 2h 47m. 4m 47s per episode · 45m 12s model time. View full run for GPT-5.6 Luna (high)
- 9. Gemini 3.6 Flash (high): 2h 52m. 4m 55s per episode · 1h 9m model time. View full run for Gemini 3.6 Flash (high)
- 10. Gemini 3.8 Flash (high): 2h 53m. 4m 56s per episode · 1h 14m model time. View full run for Gemini 3.8 Flash (high)
- 11. GPT-5.6 Sol (high): 2h 57m. 5m 3s per episode · 1h 19m model time. View full run for GPT-5.6 Sol (high)
- 12. Claude Fable 5.1 (high): 3h 24m. 5m 50s per episode · 1h 55m model time. View full run for Claude Fable 5.1 (high)
- 13. Grok 4.6 (high): 4h 40m. 8m 0s per episode · 1h 49m model time. View full run for Grok 4.6 (high)
- 14. GPT-5 Nano (medium): 4h 55m. 8m 26s per episode · 3h 14m model time. View full run for GPT-5 Nano (medium)
- 15. Kimi K3 (high): 5h 17m. 9m 3s per episode · 4h 2m model time. View full run for Kimi K3 (high)
- 16. Ox Alpha (high): 7h 33m. 12m 57s per episode · 3h 3m model time. View full run for Ox Alpha (high)
- 17. GPT-6 Astra (high): 10h 41m. 18m 19s per episode · 1h 35m model time. View full run for GPT-6 Astra (high)
- 18. Mistral Medium 3.5 (high): 14h 58m. 25m 40s per episode · 8h 29m model time. View full run for Mistral Medium 3.5 (high)
| Model-time rank | Model | Model time | End-to-end time | Model timeper episode | Model timeper call | Model calls |
|---|---|---|---|---|---|---|
| 1 | Llama 4 Maverick (non-thinking) | 16m 16s | 1h 53m | 27.9 s | 0.9 s | 1,132 |
| 2 | Grok 4.5 (high) | 40m 43s | 2h 15m | 1m 10s | 4.3 s | 564 |
| 3 | Claude Opus 5 (high) | 42m 27s | 2h 4m | 1m 13s | 5.5 s | 467 |
| 4 | Gemini 3.7 Flash (high) | 44m 51s | 2h 40m | 1m 17s | 5.1 s | 525 |
| 5 | GPT-5.6 Luna (high) | 45m 12s | 2h 47m | 1m 17s | 4.2 s | 653 |
| 6 | Claude Sonnet 5 (high) | 49m 5s | 2h 16m | 1m 24s | 5.3 s | 551 |
| 7 | gpt-oss-120B (high) | 53m 48s | 2h 13m | 1m 32s | 6.6 s | 490 |
| 8 | Gemini 3.6 Flash (high) | 1h 9m | 2h 52m | 1m 59s | 6.6 s | 625 |
| 9 | Claude Fable 5 (high) | 1h 9m | 2h 40m | 1m 59s | 9.1 s | 457 |
| 10 | Gemini 3.8 Flash (high) | 1h 14m | 2h 53m | 2m 6s | 8.9 s | 495 |
| 11 | GPT-5.6 Sol (high) | 1h 19m | 2h 57m | 2m 16s | 8.8 s | 539 |
| 12 | GPT-6 Astra (high) | 1h 35m | 10h 41m | 2m 42s | 10.1 s | 561 |
| 13 | Grok 4.6 (high) | 1h 49m | 4h 40m | 3m 8s | 12.3 s | 535 |
| 14 | Claude Fable 5.1 (high) | 1h 55m | 3h 24m | 3m 17s | 15.1 s | 458 |
| 15 | Ox Alpha (high) | 3h 3m | 7h 33m | 5m 14s | 17.0 s | 648 |
| 16 | GPT-5 Nano (medium) | 3h 14m | 4h 55m | 5m 32s | 20.2 s | 575 |
| 17 | Kimi K3 (high) | 4h 2m | 5h 17m | 6m 54s | 30.3 s | 479 |
| 18 | Mistral Medium 3.5 (high) | 8h 29m | 14h 58m | 14m 32s | 41.6 s | 733 |
#1Llama 4 Maverick (non-thinking)parasail
- Model time
- 16m 16s
- End-to-end time
- 1h 53m
- Model time / episode
- 27.9 s
Explore full run · questions, answers & evidence
#2Grok 4.5 (high)xai
- Model time
- 40m 43s
- End-to-end time
- 2h 15m
- Model time / episode
- 1m 10s
Explore full run · questions, answers & evidence
#3Claude Opus 5 (high)anthropic
- Model time
- 42m 27s
- End-to-end time
- 2h 4m
- Model time / episode
- 1m 13s
Explore full run · questions, answers & evidence
#4Gemini 3.7 Flash (high)google-ai-studio
- Model time
- 44m 51s
- End-to-end time
- 2h 40m
- Model time / episode
- 1m 17s
Explore full run · questions, answers & evidence
#5GPT-5.6 Luna (high)openai
- Model time
- 45m 12s
- End-to-end time
- 2h 47m
- Model time / episode
- 1m 17s
Explore full run · questions, answers & evidence
#6Claude Sonnet 5 (high)anthropic
- Model time
- 49m 5s
- End-to-end time
- 2h 16m
- Model time / episode
- 1m 24s
Explore full run · questions, answers & evidence
#7gpt-oss-120B (high)cerebras
- Model time
- 53m 48s
- End-to-end time
- 2h 13m
- Model time / episode
- 1m 32s
Explore full run · questions, answers & evidence
#8Gemini 3.6 Flash (high)google-vertex
- Model time
- 1h 9m
- End-to-end time
- 2h 52m
- Model time / episode
- 1m 59s
Explore full run · questions, answers & evidence
#9Claude Fable 5 (high)anthropic
- Model time
- 1h 9m
- End-to-end time
- 2h 40m
- Model time / episode
- 1m 59s
Explore full run · questions, answers & evidence
#10Gemini 3.8 Flash (high)google-ai-studio
- Model time
- 1h 14m
- End-to-end time
- 2h 53m
- Model time / episode
- 2m 6s
Explore full run · questions, answers & evidence
#11GPT-5.6 Sol (high)openai
- Model time
- 1h 19m
- End-to-end time
- 2h 57m
- Model time / episode
- 2m 16s
Explore full run · questions, answers & evidence
#12GPT-6 Astra (high)openai
- Model time
- 1h 35m
- End-to-end time
- 10h 41m
- Model time / episode
- 2m 42s
Explore full run · questions, answers & evidence
#13Grok 4.6 (high)xai
- Model time
- 1h 49m
- End-to-end time
- 4h 40m
- Model time / episode
- 3m 8s
Explore full run · questions, answers & evidence
#14Claude Fable 5.1 (high)anthropic
- Model time
- 1h 55m
- End-to-end time
- 3h 24m
- Model time / episode
- 3m 17s
Explore full run · questions, answers & evidence
#15Ox Alpha (high)stealth
- Model time
- 3h 3m
- End-to-end time
- 7h 33m
- Model time / episode
- 5m 14s
Explore full run · questions, answers & evidence
#16GPT-5 Nano (medium)openai
- Model time
- 3h 14m
- End-to-end time
- 4h 55m
- Model time / episode
- 5m 32s
Explore full run · questions, answers & evidence
#17Kimi K3 (high)moonshotai
- Model time
- 4h 2m
- End-to-end time
- 5h 17m
- Model time / episode
- 6m 54s
Explore full run · questions, answers & evidence
#18Mistral Medium 3.5 (high)mistral
- Model time
- 8h 29m
- End-to-end time
- 14h 58m
- Model time / episode
- 14m 32s
Explore full run · questions, answers & evidence
The first chart ranks model time. The second ranks end-to-end time, so the order can change.