Official run
Grok 4.6High
In this Twenty Questions LLM benchmark, Grok 4.6 (high) ranks 8 of 18 with a 14.29 question score and 100% success. Lower is better. The run contains 35 scored episodes across 7 subjects with 5 trials per subject. Choose a subject to inspect its attempts.
Rank 8 of 18xaiHigh
- Success
- 100%
- Contract compliance
- 100%
- Guesser cost
- $2.29
- Guesser time
- 1h 49m
Drill down
Subjects.
Complete benchmark
Run ledger.
All Guesser and benchmark-support activity.
- Total tokens
- 4,936,122
- Wall-clock runtime
- 4h 40m
- Guesser calls
- 535
- Completed
Reliability
Output contract clean.
All 535 evaluated outputs matched the public structured-action contract.
Provenance
- Execution
BX-20260814-official-M0015-013- Model ID
M-0015- Benchmark
B-0001- Base seed
0- Git commit
4e197a516062
Recorded cost
By role.
- Guesser27.7%$2.29
- Primary Oracle53.4%$4.41
- Reviewer6.6%$0.54
- Judge12.2%$1.01
- Validator0%$0.00
Roles in this cost:
- Guesser
- Asks the questions and submits the scored guess.
- Primary Oracle
- Searches for evidence and proposes an answer.
- Reviewer
- Checks each Oracle YES or NO independently.
- Judge
- Decides when the Oracle and Reviewer disagree.
- Validator
- Checks a submitted guess against the trusted subject.
Run configuration
Models.
Model identity, prompt versions, routing, calls, and recorded cost for the full run.
x-ai/grok-4.6
Asks the questions and submits the scored guess.
- Calls
- 535
- Cost
- $2.2862
- Reasoning
- High
- Routing
- Exact provider · xai
- Resolved model
- x-ai/grok-4.6
- Resolved provider
- xAI
- Prompt contract
stateful-category-guesser-v10-unknown-evidence-guidance- Configuration
M-0015
Resolved provider details
- Fallback calls
- 0
- Provider unreported
- 0
Game support
Oracle and adjudication.
These models support the game. They are fixed across the run and are not under test.
openai/gpt-5.6-luna
Searches for evidence and proposes an answer.
- Calls
- 499
- Cost
- $4.4071
- Reasoning
- Medium
- Routing
- Exact provider · openai
- Resolved model
- openai/gpt-5.6-luna
- Resolved provider
- OpenAI
- Prompt contract
live-web-oracle-v7-direct-negative-evidence
Resolved provider details
- Fallback calls
- 0
- Provider unreported
- 0
google/gemini-3.5-flash-lite
Checks each Oracle YES or NO independently.
- Calls
- 483
- Cost
- $0.5438
- Reasoning
- Medium
- Routing
- Exact provider · google-ai-studio
- Resolved provider
- Google AI Studio
Resolved provider details
- Fallback calls
- 0
- Provider unreported
- 0
anthropic/claude-opus-5
Decides when the Oracle and Reviewer disagree.
- Calls
- 81
- Cost
- $1.0093
- Reasoning
- Medium
- Routing
- OpenRouter automatic routing
- Resolved provider
- Amazon Bedrock, Anthropic
Resolved provider details
- Fallback calls
- 0
- Provider unreported
- 0
openai/gpt-5.6-luna
Checks a submitted guess against the trusted subject.
- Calls
- 36
- Cost
- $0.0022
- Reasoning
- Medium
- Routing
- Exact provider · openai
- Resolved model
- openai/gpt-5.6-luna
- Resolved provider
- OpenAI
- Prompt contract
strict-guess-validator-v1- Configuration
gpt-5.6-luna-validator
Resolved provider details
- Fallback calls
- 0
- Provider unreported
- 0