Official run
MiniMax M3High
In this Twenty Questions LLM benchmark, MiniMax M3 (high) ranks 11 of 11 with a 29.17 question score and 50% success. Lower is better. The run contains 30 scored episodes across 10 subjects with 3 trials per subject. Choose a subject to inspect its attempts.
Rank 11 of 11coreweaveHigh
- Success
- 50%
- Contract compliance
- 99%
- Guesser cost
- $0.30
- Guesser time
- 35m 46s
Drill down
Subjects.
Complete benchmark
Run ledger.
All Guesser and benchmark-support activity.
- Total tokens
- 8,285,795
- Wall-clock runtime
- 8h 54m
- Guesser calls
- 890
- Completed
Reliability
Output contract breached.
8 invalid outputs affected 8 episodes and consumed 7 counted turns.
Provenance
- Execution
BX-20260908-B-0003-experimental-M0024-001- Model ID
M-0024- Benchmark
B-0003- Base seed
0- Git commit
eceee31eef2e
Recorded cost
By role.
- Guesser6.9%$0.30
- Primary Oracle46.8%$2.04
- Reviewer19.9%$0.86
- Judge26%$1.13
- Validator0.4%$0.02
Roles in this cost:
- Guesser
- Asks the questions and submits the scored guess.
- Primary Oracle
- Searches for evidence and proposes an answer.
- Reviewer
- Checks each Oracle YES or NO independently.
- Judge
- Decides when the Oracle and Reviewer disagree.
- Validator
- Checks a submitted guess against the trusted subject.
Run configuration
Models.
Model identity, prompt versions, routing, calls, and recorded cost for the full run.
minimax/minimax-m3
Asks the questions and submits the scored guess.
- Calls
- 890
- Cost
- $0.3002
- Reasoning
- High
- Routing
- Exact provider · coreweave
- Resolved model
- minimax/minimax-m3
- Resolved provider
- CoreWeave
- Prompt contract
stateful-category-guesser-v16-five-answer-category-guide- Configuration
M-0024
Resolved provider details
- Fallback calls
- 0
- Provider unreported
- 0
Game support
Oracle and adjudication.
These models support the game. They are fixed across the run and are not under test.
openai/gpt-5.6-luna
Searches for evidence and proposes an answer.
- Calls
- 648
- Cost
- $2.0367
- Reasoning
- Medium
- Routing
- Exact provider · openai
- Resolved model
- openai/gpt-5.6-luna
- Resolved provider
- OpenAI
- Prompt contract
live-web-oracle-v17-labelled-source-context
Resolved provider details
- Fallback calls
- 0
- Provider unreported
- 0
google/gemini-3.5-flash-lite
Checks each Oracle YES or NO independently.
- Calls
- 602
- Cost
- $0.8647
- Reasoning
- Medium
- Routing
- Exact provider · google-ai-studio
- Resolved provider
- Google AI Studio
Resolved provider details
- Fallback calls
- 0
- Provider unreported
- 0
anthropic/claude-opus-5
Decides when the Oracle and Reviewer disagree.
- Calls
- 53
- Cost
- $1.1303
- Reasoning
- Medium
- Routing
- OpenRouter automatic routing
- Resolved provider
- Claude Platform on AWS
Resolved provider details
- Fallback calls
- 0
- Provider unreported
- 0
openai/gpt-5.6-luna
Checks a submitted guess against the trusted subject.
- Calls
- 91
- Cost
- $0.0173
- Reasoning
- Medium
- Routing
- Exact provider · openai
- Resolved model
- openai/gpt-5.6-luna
- Resolved provider
- OpenAI
- Prompt contract
strict-guess-validator-v2-generic-kinds- Configuration
gpt-5.6-luna-validator
Resolved provider details
- Fallback calls
- 0
- Provider unreported
- 0