Run workspace

Official run

Grok 4.7High

In this Twenty Questions LLM benchmark, Grok 4.7 (high) ranks 15 of 18 with a 22.23 question score and 73% success. Lower is better. The run contains 30 scored episodes across 10 subjects with 3 trials per subject. Choose a subject to inspect its attempts.

Rank 15 of 18xaiHigh

Question score22.23questions · lower is better95% CI 20.80–23.67
Success
73%
Contract compliance
100%
Guesser cost
$1.30
Guesser time
20m 31s

Drill down

Subjects.

01Albert Einstein11.00 avg02Albert Schweitzer38.33 avg03Garfield13.67 avg04Achilles8.00 avg05Genghis Khan14.00 avg06Bike pump41.00 avg07Spider web26.67 avg08Eyebrow19.33 avg09Moon9.33 avg10Door handle41.00 avg

Complete benchmark

Run ledger.

All Guesser and benchmark-support activity.

Total tokens
6,111,846
Wall-clock runtime
4h 41m
Guesser calls
689
Completed

Reliability

Output contract clean.

All 689 evaluated outputs matched the public structured-action contract.

Run details

Execution
BX-20260926-B-0003-official-M0027-001
Model ID
M-0027
Benchmark
B-0003
Base seed
0
Git commit
6e46fb003ebf

Recorded cost

By role.

Cost composition.Share of full-run cost
  1. Guesser31.7%$1.30
  2. Primary Oracle31.8%$1.30
  3. Reviewer15.1%$0.62
  4. Judge21.2%$0.87
  5. Validator0.1%$0.00
Excluded repair overhead$0.111 superseded attempt across 1 affected trial. Not included above.

Roles in this cost:

Guesser
Asks the questions and submits the scored guess.
Primary Oracle
Searches for evidence and proposes an answer.
Reviewer
Checks each Oracle YES or NO independently.
Judge
Decides when the Oracle and Reviewer disagree.
Validator
Checks a submitted guess against the trusted subject.
Read the role and answer-checking method →

Run configuration

Models.

Model identity, prompt versions, routing, calls, and recorded cost for the full run.

Game support

Answering and checking models.

These models support the game. They are fixed across the run and are not under test.

Primary Oracle

openai/gpt-5.6-luna

Searches for evidence and proposes an answer.

Calls
475
Cost
$1.3029
Reasoning
Medium
Routing
Exact provider · openai
Resolved model
openai/gpt-5.6-luna
Resolved provider
OpenAI
Prompt contract
live-web-oracle-v17-labelled-source-context
Resolved provider details
OpenAI475 calls$1.30293,718.62 s
Fallback calls
0
Provider unreported
0
Reviewer

google/gemini-3.5-flash-lite

Checks each Oracle YES or NO independently.

Calls
458
Cost
$0.6189
Reasoning
Medium
Routing
Exact provider · google-ai-studio
Resolved provider
Google AI Studio
Resolved provider details
Google AI Studio458 calls$0.6189780.07 s
Fallback calls
0
Provider unreported
0
Judge

anthropic/claude-opus-5

Decides when the Oracle and Reviewer disagree.

Calls
38
Cost
$0.8685
Reasoning
Medium
Routing
OpenRouter automatic routing
Resolved provider
Claude Platform on AWS
Resolved provider details
Claude Platform on AWS38 calls$0.8685238.61 s
Fallback calls
0
Provider unreported
0
Guess Validator

openai/gpt-5.6-luna

Checks a submitted guess against the trusted subject.

Calls
22
Cost
$0.0042
Reasoning
Medium
Routing
Exact provider · openai
Resolved model
openai/gpt-5.6-luna
Resolved provider
OpenAI
Prompt contract
strict-guess-validator-v2-generic-kinds
Configuration
gpt-5.6-luna-validator
Resolved provider details
OpenAI22 calls$0.004232.35 s
Fallback calls
0
Provider unreported
0