Run workspace

Official run

Claude Opus 5High

In this Twenty Questions LLM benchmark, Claude Opus 5 (high) ranks 6 of 11 with a 18.07 question score and 93% success. Lower is better. The run contains 30 scored episodes across 10 subjects with 3 trials per subject. Choose a subject to inspect its attempts.

Rank 6 of 11anthropicHigh

Question score18.07questions · lower is better95% CI 16.50–19.63
Success
93%
Contract compliance
>99%
Guesser cost
$2.80
Guesser time
41m 2s

Drill down

Subjects.

01Albert Einstein11.00 avg02Albert Schweitzer24.00 avg03Garfield20.33 avg04Achilles10.00 avg05Genghis Khan11.67 avg06Bike pump40.33 avg07Spider web15.67 avg08Eyebrow16.00 avg09Moon10.00 avg10Door handle21.67 avg

Complete benchmark

Run ledger.

All Guesser and benchmark-support activity.

Total tokens
5,842,602
Wall-clock runtime
20h 9m
Guesser calls
570
Completed

Reliability

Output contract breached.

2 invalid outputs affected 1 episode and consumed 2 counted turns.

Provenance

Execution
BX-20260907-B-0003-experimental-M0006-002
Model ID
M-0006
Benchmark
B-0003
Base seed
0
Git commit
04dde57c3906

Recorded cost

By role.

Cost composition.Share of full-run cost
  1. Guesser50.5%$2.80
  2. Primary Oracle25.8%$1.43
  3. Reviewer10.8%$0.60
  4. Judge12.9%$0.72
  5. Validator0.1%$0.01
Excluded repair overhead$0.767 superseded attempts across 7 affected trials. Not included above.

Roles in this cost:

Guesser
Asks the questions and submits the scored guess.
Primary Oracle
Searches for evidence and proposes an answer.
Reviewer
Checks each Oracle YES or NO independently.
Judge
Decides when the Oracle and Reviewer disagree.
Validator
Checks a submitted guess against the trusted subject.
Read the role and answer-checking method →

Run configuration

Models.

Model identity, prompt versions, routing, calls, and recorded cost for the full run.

Game support

Oracle and adjudication.

These models support the game. They are fixed across the run and are not under test.

Primary Oracle

openai/gpt-5.6-luna

Searches for evidence and proposes an answer.

Calls
535
Cost
$1.4306
Reasoning
Medium
Routing
Exact provider · openai
Resolved model
openai/gpt-5.6-luna
Resolved provider
OpenAI
Prompt contracts
live-web-oracle-v16-search-headroom
live-web-oracle-v17-labelled-source-context
Resolved provider details
OpenAI535 calls$1.43064,646.67 s
Fallback calls
0
Provider unreported
0
Reviewer

google/gemini-3.5-flash-lite

Checks each Oracle YES or NO independently.

Calls
485
Cost
$0.5992
Reasoning
Medium
Routing
Exact provider · google-ai-studio
Resolved provider
Google AI Studio
Resolved provider details
Google AI Studio485 calls$0.5992768.83 s
Fallback calls
0
Provider unreported
0
Judge

anthropic/claude-opus-5

Decides when the Oracle and Reviewer disagree.

Calls
36
Cost
$0.7164
Reasoning
Medium
Routing
OpenRouter automatic routing
Resolved provider
Claude Platform on AWS
Resolved provider details
Claude Platform on AWS36 calls$0.7164218.35 s
Fallback calls
0
Provider unreported
0
Guess Validator

openai/gpt-5.6-luna

Checks a submitted guess against the trusted subject.

Calls
32
Cost
$0.0062
Reasoning
Medium
Routing
Exact provider · openai
Resolved model
openai/gpt-5.6-luna
Resolved provider
OpenAI
Prompt contract
strict-guess-validator-v2-generic-kinds
Configuration
gpt-5.6-luna-validator
Resolved provider details
OpenAI32 calls$0.006251.04 s
Fallback calls
0
Provider unreported
0