Model and benchmark time

Time 1.1

Provider-reported model-call time and end-to-end benchmark runtime.

Models
11
Median model time
41m 23s
Total model time
11h 2m
Total end-to-end time
70h 34m

Model time

Model time across the run.

Each bar adds the provider-reported time of every call to the model under test. Shorter is faster. This is not the wall-clock benchmark runtime.

Model time · lower is fasterSelect a model row to view its full run
  1. 1. Gemini 3.7 Flash (high): 25m 46s. 51.5 s per episode · 525 calls. View full run for Gemini 3.7 Flash (high)
  2. 2. gpt-oss-120B (high): 32m 30s. 1m 5s per episode · 662 calls. View full run for gpt-oss-120B (high)
  3. 3. MiniMax M3 (high): 35m 46s. 1m 12s per episode · 890 calls. View full run for MiniMax M3 (high)
  4. 4. GPT-6 Astra (high): 40m 43s. 1m 21s per episode · 440 calls. View full run for GPT-6 Astra (high)
  5. 5. Claude Opus 5 (high): 41m 2s. 1m 22s per episode · 570 calls. View full run for Claude Opus 5 (high)
  6. 6. GPT-5.6 Luna (high): 41m 23s. 1m 23s per episode · 658 calls. View full run for GPT-5.6 Luna (high)
  7. 7. Gemini 3.8 Flash (high): 51m 12s. 1m 42s per episode · 476 calls. View full run for Gemini 3.8 Flash (high)
  8. 8. GLM-5.3-Flash (high): 51m 31s. 1m 43s per episode · 720 calls. View full run for GLM-5.3-Flash (high)
  9. 9. GPT-5.6 Sol (high): 57m 25s. 1m 55s per episode · 571 calls. View full run for GPT-5.6 Sol (high)
  10. 10. Claude Fable 5.1 (high): 1h 19m. 2m 39s per episode · 481 calls. View full run for Claude Fable 5.1 (high)
  11. 11. Grok 4.6 (high): 3h 25m. 6m 51s per episode · 544 calls. View full run for Grok 4.6 (high)

End-to-end time

End-to-end benchmark time.

Each bar is the wall-clock time from run creation to final status. It includes model calls, adjudication, scheduling, concurrency, and other benchmark work.

End-to-end time · lower is fasterSelect a model row to view its full run
  1. 1. Gemini 3.7 Flash (high): 1h 53m. 3m 47s per episode · 25m 46s model time. View full run for Gemini 3.7 Flash (high)
  2. 2. GPT-6 Astra (high): 1h 58m. 3m 56s per episode · 40m 43s model time. View full run for GPT-6 Astra (high)
  3. 3. gpt-oss-120B (high): 2h 53m. 5m 46s per episode · 32m 30s model time. View full run for gpt-oss-120B (high)
  4. 4. GLM-5.3-Flash (high): 2h 58m. 5m 56s per episode · 51m 31s model time. View full run for GLM-5.3-Flash (high)
  5. 5. Gemini 3.8 Flash (high): 2h 58m. 5m 56s per episode · 51m 12s model time. View full run for Gemini 3.8 Flash (high)
  6. 6. Claude Fable 5.1 (high): 3h 8m. 6m 17s per episode · 1h 19m model time. View full run for Claude Fable 5.1 (high)
  7. 7. Grok 4.6 (high): 7h 2m. 14m 3s per episode · 3h 25m model time. View full run for Grok 4.6 (high)
  8. 8. MiniMax M3 (high): 8h 54m. 17m 48s per episode · 35m 46s model time. View full run for MiniMax M3 (high)
  9. 9. GPT-5.6 Sol (high): 9h 2m. 18m 4s per episode · 57m 25s model time. View full run for GPT-5.6 Sol (high)
  10. 10. GPT-5.6 Luna (high): 9h 38m. 19m 17s per episode · 41m 23s model time. View full run for GPT-5.6 Luna (high)
  11. 11. Claude Opus 5 (high): 20h 9m. 40m 18s per episode · 41m 2s model time. View full run for Claude Opus 5 (high)
Time comparison
Model-time rankModelModel timeEnd-to-end timeModel timeper episodeModel timeper callModel calls
1Gemini 3.7 Flash (high)M-001625m 46s1h 53m51.5 s2.9 s525
2gpt-oss-120B (high)M-000232m 30s2h 53m1m 5s2.9 s662
3MiniMax M3 (high)M-002435m 46s8h 54m1m 12s2.4 s890
4GPT-6 Astra (high)M-002240m 43s1h 58m1m 21s5.6 s440
5Claude Opus 5 (high)M-000641m 2s20h 9m1m 22s4.3 s570
6GPT-5.6 Luna (high)M-000141m 23s9h 38m1m 23s3.8 s658
7Gemini 3.8 Flash (high)M-002151m 12s2h 58m1m 42s6.5 s476
8GLM-5.3-Flash (high)M-002351m 31s2h 58m1m 43s4.3 s720
9GPT-5.6 Sol (high)M-001057m 25s9h 2m1m 55s6.0 s571
10Claude Fable 5.1 (high)M-00201h 19m3h 8m2m 39s9.9 s481
11Grok 4.6 (high)M-00153h 25m7h 2m6m 51s22.7 s544
#1Gemini 3.7 Flash (high)google-ai-studio
Model time
25m 46s
End-to-end time
1h 53m
Model time / episode
51.5 s
Explore full run · questions, answers & evidence
#2gpt-oss-120B (high)cerebras
Model time
32m 30s
End-to-end time
2h 53m
Model time / episode
1m 5s
Explore full run · questions, answers & evidence
#3MiniMax M3 (high)coreweave
Model time
35m 46s
End-to-end time
8h 54m
Model time / episode
1m 12s
Explore full run · questions, answers & evidence
#4GPT-6 Astra (high)openai
Model time
40m 43s
End-to-end time
1h 58m
Model time / episode
1m 21s
Explore full run · questions, answers & evidence
#5Claude Opus 5 (high)anthropic
Model time
41m 2s
End-to-end time
20h 9m
Model time / episode
1m 22s
Explore full run · questions, answers & evidence
#6GPT-5.6 Luna (high)openai
Model time
41m 23s
End-to-end time
9h 38m
Model time / episode
1m 23s
Explore full run · questions, answers & evidence
#7Gemini 3.8 Flash (high)google-ai-studio
Model time
51m 12s
End-to-end time
2h 58m
Model time / episode
1m 42s
Explore full run · questions, answers & evidence
#8GLM-5.3-Flash (high)z-ai
Model time
51m 31s
End-to-end time
2h 58m
Model time / episode
1m 43s
Explore full run · questions, answers & evidence
#9GPT-5.6 Sol (high)openai
Model time
57m 25s
End-to-end time
9h 2m
Model time / episode
1m 55s
Explore full run · questions, answers & evidence
#10Claude Fable 5.1 (high)anthropic
Model time
1h 19m
End-to-end time
3h 8m
Model time / episode
2m 39s
Explore full run · questions, answers & evidence
#11Grok 4.6 (high)xai
Model time
3h 25m
End-to-end time
7h 2m
Model time / episode
6m 51s
Explore full run · questions, answers & evidence

The first chart ranks model time. The second ranks end-to-end time, so the order can change.