Model and benchmark time

Time 1

Provider-reported model-call time and end-to-end benchmark runtime.

Models
18
Median model time
1h 11m
Total model time
33h 50m
Total end-to-end time
78h 60m

Model time

Model time across the run.

Each bar adds the provider-reported time of every call to the model under test. Shorter is faster. This is not the wall-clock benchmark runtime.

Model time · lower is fasterSelect a model row to view its full run
  1. 1. Llama 4 Maverick (non-thinking): 16m 16s. 27.9 s per episode · 1,132 calls. View full run for Llama 4 Maverick (non-thinking)
  2. 2. Grok 4.5 (high): 40m 43s. 1m 10s per episode · 564 calls. View full run for Grok 4.5 (high)
  3. 3. Claude Opus 5 (high): 42m 27s. 1m 13s per episode · 467 calls. View full run for Claude Opus 5 (high)
  4. 4. Gemini 3.7 Flash (high): 44m 51s. 1m 17s per episode · 525 calls. View full run for Gemini 3.7 Flash (high)
  5. 5. GPT-5.6 Luna (high): 45m 12s. 1m 17s per episode · 653 calls. View full run for GPT-5.6 Luna (high)
  6. 6. Claude Sonnet 5 (high): 49m 5s. 1m 24s per episode · 551 calls. View full run for Claude Sonnet 5 (high)
  7. 7. gpt-oss-120B (high): 53m 48s. 1m 32s per episode · 490 calls. View full run for gpt-oss-120B (high)
  8. 8. Gemini 3.6 Flash (high): 1h 9m. 1m 59s per episode · 625 calls. View full run for Gemini 3.6 Flash (high)
  9. 9. Claude Fable 5 (high): 1h 9m. 1m 59s per episode · 457 calls. View full run for Claude Fable 5 (high)
  10. 10. Gemini 3.8 Flash (high): 1h 14m. 2m 6s per episode · 495 calls. View full run for Gemini 3.8 Flash (high)
  11. 11. GPT-5.6 Sol (high): 1h 19m. 2m 16s per episode · 539 calls. View full run for GPT-5.6 Sol (high)
  12. 12. GPT-6 Astra (high): 1h 35m. 2m 42s per episode · 561 calls. View full run for GPT-6 Astra (high)
  13. 13. Grok 4.6 (high): 1h 49m. 3m 8s per episode · 535 calls. View full run for Grok 4.6 (high)
  14. 14. Claude Fable 5.1 (high): 1h 55m. 3m 17s per episode · 458 calls. View full run for Claude Fable 5.1 (high)
  15. 15. Ox Alpha (high): 3h 3m. 5m 14s per episode · 648 calls. View full run for Ox Alpha (high)
  16. 16. GPT-5 Nano (medium): 3h 14m. 5m 32s per episode · 575 calls. View full run for GPT-5 Nano (medium)
  17. 17. Kimi K3 (high): 4h 2m. 6m 54s per episode · 479 calls. View full run for Kimi K3 (high)
  18. 18. Mistral Medium 3.5 (high): 8h 29m. 14m 32s per episode · 733 calls. View full run for Mistral Medium 3.5 (high)

End-to-end time

End-to-end benchmark time.

Each bar is the wall-clock time from run creation to final status. It includes model calls, adjudication, scheduling, concurrency, and other benchmark work.

End-to-end time · lower is fasterSelect a model row to view its full run
  1. 1. Llama 4 Maverick (non-thinking): 1h 53m. 3m 14s per episode · 16m 16s model time. View full run for Llama 4 Maverick (non-thinking)
  2. 2. Claude Opus 5 (high): 2h 4m. 3m 33s per episode · 42m 27s model time. View full run for Claude Opus 5 (high)
  3. 3. gpt-oss-120B (high): 2h 13m. 3m 49s per episode · 53m 48s model time. View full run for gpt-oss-120B (high)
  4. 4. Grok 4.5 (high): 2h 15m. 3m 51s per episode · 40m 43s model time. View full run for Grok 4.5 (high)
  5. 5. Claude Sonnet 5 (high): 2h 16m. 3m 52s per episode · 49m 5s model time. View full run for Claude Sonnet 5 (high)
  6. 6. Claude Fable 5 (high): 2h 40m. 4m 34s per episode · 1h 9m model time. View full run for Claude Fable 5 (high)
  7. 7. Gemini 3.7 Flash (high): 2h 40m. 4m 35s per episode · 44m 51s model time. View full run for Gemini 3.7 Flash (high)
  8. 8. GPT-5.6 Luna (high): 2h 47m. 4m 47s per episode · 45m 12s model time. View full run for GPT-5.6 Luna (high)
  9. 9. Gemini 3.6 Flash (high): 2h 52m. 4m 55s per episode · 1h 9m model time. View full run for Gemini 3.6 Flash (high)
  10. 10. Gemini 3.8 Flash (high): 2h 53m. 4m 56s per episode · 1h 14m model time. View full run for Gemini 3.8 Flash (high)
  11. 11. GPT-5.6 Sol (high): 2h 57m. 5m 3s per episode · 1h 19m model time. View full run for GPT-5.6 Sol (high)
  12. 12. Claude Fable 5.1 (high): 3h 24m. 5m 50s per episode · 1h 55m model time. View full run for Claude Fable 5.1 (high)
  13. 13. Grok 4.6 (high): 4h 40m. 8m 0s per episode · 1h 49m model time. View full run for Grok 4.6 (high)
  14. 14. GPT-5 Nano (medium): 4h 55m. 8m 26s per episode · 3h 14m model time. View full run for GPT-5 Nano (medium)
  15. 15. Kimi K3 (high): 5h 17m. 9m 3s per episode · 4h 2m model time. View full run for Kimi K3 (high)
  16. 16. Ox Alpha (high): 7h 33m. 12m 57s per episode · 3h 3m model time. View full run for Ox Alpha (high)
  17. 17. GPT-6 Astra (high): 10h 41m. 18m 19s per episode · 1h 35m model time. View full run for GPT-6 Astra (high)
  18. 18. Mistral Medium 3.5 (high): 14h 58m. 25m 40s per episode · 8h 29m model time. View full run for Mistral Medium 3.5 (high)
Time comparison
Model-time rankModelModel timeEnd-to-end timeModel timeper episodeModel timeper callModel calls
1Llama 4 Maverick (non-thinking)M-000916m 16s1h 53m27.9 s0.9 s1,132
2Grok 4.5 (high)M-000840m 43s2h 15m1m 10s4.3 s564
3Claude Opus 5 (high)M-000642m 27s2h 4m1m 13s5.5 s467
4Gemini 3.7 Flash (high)M-001644m 51s2h 40m1m 17s5.1 s525
5GPT-5.6 Luna (high)M-000145m 12s2h 47m1m 17s4.2 s653
6Claude Sonnet 5 (high)M-000549m 5s2h 16m1m 24s5.3 s551
7gpt-oss-120B (high)M-000253m 48s2h 13m1m 32s6.6 s490
8Gemini 3.6 Flash (high)M-00041h 9m2h 52m1m 59s6.6 s625
9Claude Fable 5 (high)M-00141h 9m2h 40m1m 59s9.1 s457
10Gemini 3.8 Flash (high)M-00211h 14m2h 53m2m 6s8.9 s495
11GPT-5.6 Sol (high)M-00101h 19m2h 57m2m 16s8.8 s539
12GPT-6 Astra (high)M-00221h 35m10h 41m2m 42s10.1 s561
13Grok 4.6 (high)M-00151h 49m4h 40m3m 8s12.3 s535
14Claude Fable 5.1 (high)M-00201h 55m3h 24m3m 17s15.1 s458
15Ox Alpha (high)M-00173h 3m7h 33m5m 14s17.0 s648
16GPT-5 Nano (medium)M-00033h 14m4h 55m5m 32s20.2 s575
17Kimi K3 (high)M-00074h 2m5h 17m6m 54s30.3 s479
18Mistral Medium 3.5 (high)M-00128h 29m14h 58m14m 32s41.6 s733
#1Llama 4 Maverick (non-thinking)parasail
Model time
16m 16s
End-to-end time
1h 53m
Model time / episode
27.9 s
Explore full run · questions, answers & evidence
#2Grok 4.5 (high)xai
Model time
40m 43s
End-to-end time
2h 15m
Model time / episode
1m 10s
Explore full run · questions, answers & evidence
#3Claude Opus 5 (high)anthropic
Model time
42m 27s
End-to-end time
2h 4m
Model time / episode
1m 13s
Explore full run · questions, answers & evidence
#4Gemini 3.7 Flash (high)google-ai-studio
Model time
44m 51s
End-to-end time
2h 40m
Model time / episode
1m 17s
Explore full run · questions, answers & evidence
#5GPT-5.6 Luna (high)openai
Model time
45m 12s
End-to-end time
2h 47m
Model time / episode
1m 17s
Explore full run · questions, answers & evidence
#6Claude Sonnet 5 (high)anthropic
Model time
49m 5s
End-to-end time
2h 16m
Model time / episode
1m 24s
Explore full run · questions, answers & evidence
#7gpt-oss-120B (high)cerebras
Model time
53m 48s
End-to-end time
2h 13m
Model time / episode
1m 32s
Explore full run · questions, answers & evidence
#8Gemini 3.6 Flash (high)google-vertex
Model time
1h 9m
End-to-end time
2h 52m
Model time / episode
1m 59s
Explore full run · questions, answers & evidence
#9Claude Fable 5 (high)anthropic
Model time
1h 9m
End-to-end time
2h 40m
Model time / episode
1m 59s
Explore full run · questions, answers & evidence
#10Gemini 3.8 Flash (high)google-ai-studio
Model time
1h 14m
End-to-end time
2h 53m
Model time / episode
2m 6s
Explore full run · questions, answers & evidence
#11GPT-5.6 Sol (high)openai
Model time
1h 19m
End-to-end time
2h 57m
Model time / episode
2m 16s
Explore full run · questions, answers & evidence
#12GPT-6 Astra (high)openai
Model time
1h 35m
End-to-end time
10h 41m
Model time / episode
2m 42s
Explore full run · questions, answers & evidence
#13Grok 4.6 (high)xai
Model time
1h 49m
End-to-end time
4h 40m
Model time / episode
3m 8s
Explore full run · questions, answers & evidence
#14Claude Fable 5.1 (high)anthropic
Model time
1h 55m
End-to-end time
3h 24m
Model time / episode
3m 17s
Explore full run · questions, answers & evidence
#15Ox Alpha (high)stealth
Model time
3h 3m
End-to-end time
7h 33m
Model time / episode
5m 14s
Explore full run · questions, answers & evidence
#16GPT-5 Nano (medium)openai
Model time
3h 14m
End-to-end time
4h 55m
Model time / episode
5m 32s
Explore full run · questions, answers & evidence
#17Kimi K3 (high)moonshotai
Model time
4h 2m
End-to-end time
5h 17m
Model time / episode
6m 54s
Explore full run · questions, answers & evidence
#18Mistral Medium 3.5 (high)mistral
Model time
8h 29m
End-to-end time
14h 58m
Model time / episode
14m 32s
Explore full run · questions, answers & evidence

The first chart ranks model time. The second ranks end-to-end time, so the order can change.