Official run
Gemini 3.6 FlashHigh
In this Twenty Questions LLM benchmark, Gemini 3.6 Flash (high) ranks 14 of 18 with a 16.89 question score and 97% success. Lower is better. The run contains 35 scored episodes across 7 subjects with 5 trials per subject. Choose a subject to inspect its attempts.
Rank 14 of 18google-vertexHigh
- Success
- 97%
- Contract compliance
- >99%
- Guesser cost
- $5.97
- Guesser time
- 1h 9m
Drill down
Subjects.
Complete benchmark
Run ledger.
All Guesser and benchmark-support activity.
- Total tokens
- 5,594,609
- Wall-clock runtime
- 2h 52m
- Guesser calls
- 625
- Completed
Reliability
Output contract breached.
1 invalid output affected 1 episode and consumed 1 counted turn.
Provenance
- Execution
BX-20260728-official-M0004-010- Model ID
M-0004- Benchmark
B-0001- Base seed
0- Git commit
aa131c32c065
Recorded cost
By role.
- Guesser60.6%$5.97
- Primary Oracle29.8%$2.94
- Reviewer4.8%$0.48
- Judge4.6%$0.46
- Validator0.1%$0.01
Roles in this cost:
- Guesser
- Asks the questions and submits the scored guess.
- Primary Oracle
- Searches for evidence and proposes an answer.
- Reviewer
- Checks each Oracle YES or NO independently.
- Judge
- Decides when the Oracle and Reviewer disagree.
- Validator
- Checks a submitted guess against the trusted subject.
Run configuration
Models.
Model identity, prompt versions, routing, calls, and recorded cost for the full run.
google/gemini-3.6-flash
Asks the questions and submits the scored guess.
- Calls
- 625
- Cost
- $5.9676
- Reasoning
- High
- Routing
- Exact provider · google-vertex
- Resolved model
- google/gemini-3.6-flash
- Resolved provider
- Prompt contract
stateful-category-guesser-v10-unknown-evidence-guidance- Configuration
M-0004
Resolved provider details
Legacy run. Per-call routing totals were not retained. The configured provider was google-vertex.
Game support
Oracle and adjudication.
These models support the game. They are fixed across the run and are not under test.
openai/gpt-5.6-luna
Searches for evidence and proposes an answer.
- Calls
- 582
- Cost
- $2.9414
- Reasoning
- Medium
- Routing
- Exact provider · openai
- Resolved model
- openai/gpt-5.6-luna
- Resolved provider
- OpenAI
- Prompt contract
live-web-oracle-v7-direct-negative-evidence
Resolved provider details
Legacy run. Per-call routing totals were not retained. The configured provider was openai.
google/gemini-3.5-flash-lite
Checks each Oracle YES or NO independently.
- Calls
- 559
- Cost
- $0.4772
- Reasoning
- Medium
- Routing
- Exact provider · google-ai-studio
Resolved provider details
Legacy run. Per-call routing totals were not retained. The configured provider was google-ai-studio.
anthropic/claude-opus-5
Decides when the Oracle and Reviewer disagree.
- Calls
- 38
- Cost
- $0.4565
- Reasoning
- Medium
- Routing
- Exact provider · anthropic
Resolved provider details
Legacy run. Per-call routing totals were not retained. The configured provider was anthropic.
openai/gpt-5.6-luna
Checks a submitted guess against the trusted subject.
- Calls
- 41
- Cost
- $0.0129
- Reasoning
- Medium
- Routing
- Exact provider · openai
- Resolved model
- openai/gpt-5.6-luna
- Resolved provider
- OpenAI
- Prompt contract
strict-guess-validator-v1- Configuration
gpt-5.6-luna-validator
Resolved provider details
Legacy run. Per-call routing totals were not retained. The configured provider was openai.