Official run
Llama 4 MaverickNon Thinking
In this Twenty Questions LLM benchmark, Llama 4 Maverick (non-thinking) ranks 18 of 18 with a 32.23 question score and 49% success. Lower is better. The run contains 35 scored episodes across 7 subjects with 5 trials per subject. Choose a subject to inspect its attempts.
Rank 18 of 18parasailNone
- Success
- 49%
- Contract compliance
- 100%
- Guesser cost
- $0.30
- Guesser time
- 16m 16s
Drill down
Subjects.
Complete benchmark
Run ledger.
All Guesser and benchmark-support activity.
- Total tokens
- 5,365,223
- Wall-clock runtime
- 1h 53m
- Guesser calls
- 1,132
- Completed
Reliability
Output contract clean.
All 1132 evaluated outputs matched the public structured-action contract.
Provenance
- Execution
BX-20260728-official-M0009-010- Model ID
M-0009- Benchmark
B-0001- Base seed
0- Git commit
dc907f364e71
Recorded cost
By role.
- Guesser8%$0.30
- Primary Oracle60.9%$2.27
- Reviewer8.4%$0.31
- Judge16%$0.60
- Validator6.7%$0.25
Roles in this cost:
- Guesser
- Asks the questions and submits the scored guess.
- Primary Oracle
- Searches for evidence and proposes an answer.
- Reviewer
- Checks each Oracle YES or NO independently.
- Judge
- Decides when the Oracle and Reviewer disagree.
- Validator
- Checks a submitted guess against the trusted subject.
Run configuration
Models.
Model identity, prompt versions, routing, calls, and recorded cost for the full run.
meta-llama/llama-4-maverick
Asks the questions and submits the scored guess.
- Calls
- 1,132
- Cost
- $0.2995
- Reasoning
- None
- Routing
- Exact provider · parasail
- Resolved model
- meta-llama/llama-4-maverick
- Resolved provider
- Parasail
- Prompt contract
stateful-category-guesser-v10-unknown-evidence-guidance- Configuration
M-0009
Resolved provider details
Legacy run. Per-call routing totals were not retained. The configured provider was parasail.
Game support
Oracle and adjudication.
These models support the game. They are fixed across the run and are not under test.
openai/gpt-5.6-luna
Searches for evidence and proposes an answer.
- Calls
- 341
- Cost
- $2.2692
- Reasoning
- Medium
- Routing
- Exact provider · openai
- Resolved model
- openai/gpt-5.6-luna
- Resolved provider
- OpenAI
- Prompt contract
live-web-oracle-v7-direct-negative-evidence
Resolved provider details
Legacy run. Per-call routing totals were not retained. The configured provider was openai.
google/gemini-3.5-flash-lite
Checks each Oracle YES or NO independently.
- Calls
- 293
- Cost
- $0.3127
- Reasoning
- Medium
- Routing
- Exact provider · google-ai-studio
Resolved provider details
Legacy run. Per-call routing totals were not retained. The configured provider was google-ai-studio.
anthropic/claude-opus-5
Decides when the Oracle and Reviewer disagree.
- Calls
- 43
- Cost
- $0.5957
- Reasoning
- Medium
- Routing
- Exact provider · anthropic
Resolved provider details
Legacy run. Per-call routing totals were not retained. The configured provider was anthropic.
openai/gpt-5.6-luna
Checks a submitted guess against the trusted subject.
- Calls
- 791
- Cost
- $0.2481
- Reasoning
- Medium
- Routing
- Exact provider · openai
- Resolved model
- openai/gpt-5.6-luna
- Resolved provider
- OpenAI
- Prompt contract
strict-guess-validator-v1- Configuration
gpt-5.6-luna-validator
Resolved provider details
Legacy run. Per-call routing totals were not retained. The configured provider was openai.