Run workspace

Official run

Claude Fable 5High

In this Twenty Questions LLM benchmark, Claude Fable 5 (high) ranks 1 of 18 with a 12.06 question score and 100% success. Lower is better. The run contains 35 scored episodes across 7 subjects with 5 trials per subject. Choose a subject to inspect its attempts.

Rank 1 of 18anthropicHigh

Question score12.06questions · lower is better95% CI 10.13–13.98
Success
100%
Contract compliance
100%
Guesser cost
$6.22
Guesser time
1h 9m

Drill down

Subjects.

01Albert Einstein10.00 avg02Albert Schweitzer30.20 avg03Stephen King10.80 avg04Garfield9.00 avg05Achilles7.80 avg06Genghis Khan9.40 avg07Mario7.20 avg

Complete benchmark

Run ledger.

All Guesser and benchmark-support activity.

Total tokens
3,830,215
Wall-clock runtime
2h 40m
Guesser calls
457
Completed

Reliability

Output contract clean.

All 457 evaluated outputs matched the public structured-action contract.

Provenance

Execution
BX-20260805-official-M0014-011
Model ID
M-0014
Benchmark
B-0001
Base seed
0
Git commit
966cbf712380

Recorded cost

By role.

Cost composition.Share of full-run cost
  1. Guesser79%$6.22
  2. Primary Oracle11.6%$0.91
  3. Reviewer4.5%$0.35
  4. Judge4.8%$0.38
  5. Validator0%$0.00

Roles in this cost:

Guesser
Asks the questions and submits the scored guess.
Primary Oracle
Searches for evidence and proposes an answer.
Reviewer
Checks each Oracle YES or NO independently.
Judge
Decides when the Oracle and Reviewer disagree.
Validator
Checks a submitted guess against the trusted subject.
Read the role and answer-checking method →

Run configuration

Models.

Model identity, prompt versions, routing, calls, and recorded cost for the full run.

Game support

Oracle and adjudication.

These models support the game. They are fixed across the run and are not under test.

Primary Oracle

openai/gpt-5.6-luna

Searches for evidence and proposes an answer.

Calls
397
Cost
$0.9140
Reasoning
Medium
Routing
Exact provider · openai
Resolved model
openai/gpt-5.6-luna
Resolved provider
OpenAI
Prompt contract
live-web-oracle-v7-direct-negative-evidence
Resolved provider details
OpenAI397 calls$0.91404,534.58 s
Fallback calls
0
Provider unreported
0
Reviewer

google/gemini-3.5-flash-lite

Checks each Oracle YES or NO independently.

Calls
379
Cost
$0.3531
Reasoning
Medium
Routing
Exact provider · google-ai-studio
Resolved provider
Google AI Studio
Resolved provider details
Google AI Studio379 calls$0.3531568.48 s
Fallback calls
0
Provider unreported
0
Judge

anthropic/claude-opus-5

Decides when the Oracle and Reviewer disagree.

Calls
33
Cost
$0.3797
Reasoning
Medium
Routing
OpenRouter automatic routing
Resolved provider
Amazon Bedrock, Anthropic, Claude Platform on AWS
Resolved provider details
Amazon Bedrock30 calls$0.3462143.69 s
Anthropic2 calls$0.022730.55 s
Claude Platform on AWS1 calls$0.010726.41 s
Fallback calls
3
Provider unreported
0
Guess Validator

openai/gpt-5.6-luna

Checks a submitted guess against the trusted subject.

Calls
60
Cost
$0.0038
Reasoning
Medium
Routing
Exact provider · openai
Resolved model
openai/gpt-5.6-luna
Resolved provider
OpenAI
Prompt contract
strict-guess-validator-v1
Configuration
gpt-5.6-luna-validator
Resolved provider details
OpenAI60 calls$0.0038117.96 s
Fallback calls
0
Provider unreported
0