Run workspace

Official run

Ox AlphaHigh

In this Twenty Questions LLM benchmark, Ox Alpha (high) ranks 15 of 18 with a 17.60 question score and 91% success. Lower is better. The run contains 35 scored episodes across 7 subjects with 5 trials per subject. Choose a subject to inspect its attempts.

Rank 15 of 18stealthHigh

Question score17.60questions · lower is better95% CI 14.60–20.60
Success
91%
Contract compliance
93%
Guesser cost
$0.00
Guesser time
3h 3m

Drill down

Subjects.

01Albert Einstein9.80 avg02Albert Schweitzer41.20 avg03Stephen King17.40 avg04Garfield18.80 avg05Achilles10.40 avg06Genghis Khan19.20 avg07Mario6.40 avg

Complete benchmark

Run ledger.

All Guesser and benchmark-support activity.

Total tokens
8,394,071
Wall-clock runtime
7h 33m
Guesser calls
648
Completed

Reliability

Output contract breached.

47 invalid outputs affected 35 episodes and consumed 47 counted turns.

Provenance

Execution
BX-20260823-official-M0017-018
Model ID
M-0017
Benchmark
B-0001
Base seed
0
Git commit
d9c61bce0822

Recorded cost

By role.

Cost composition.Share of full-run cost
  1. Guesser0%$0.00
  2. Primary Oracle81.4%$7.96
  3. Reviewer6.6%$0.65
  4. Judge11.9%$1.16
  5. Validator0.1%$0.01
Excluded repair overhead$2.534 superseded attempts across 4 affected trials. Not included above.

Roles in this cost:

Guesser
Asks the questions and submits the scored guess.
Primary Oracle
Searches for evidence and proposes an answer.
Reviewer
Checks each Oracle YES or NO independently.
Judge
Decides when the Oracle and Reviewer disagree.
Validator
Checks a submitted guess against the trusted subject.
Read the role and answer-checking method →

Run configuration

Models.

Model identity, prompt versions, routing, calls, and recorded cost for the full run.

Game support

Oracle and adjudication.

These models support the game. They are fixed across the run and are not under test.

Primary Oracle

openai/gpt-5.6-luna

Searches for evidence and proposes an answer.

Calls
553
Cost
$7.9601
Reasoning
Medium
Routing
Exact provider · openai
Resolved model
openai/gpt-5.6-luna
Resolved provider
OpenAI
Prompt contract
live-web-oracle-v8-classified-research
Resolved provider details
OpenAI578 calls$7.96016,541.73 s
Fallback calls
0
Provider unreported
0
Reviewer

google/gemini-3.5-flash-lite

Checks each Oracle YES or NO independently.

Calls
530
Cost
$0.6486
Reasoning
Medium
Routing
Exact provider · google-ai-studio
Resolved provider
Google AI Studio
Resolved provider details
Google AI Studio530 calls$0.6486872.43 s
Fallback calls
0
Provider unreported
0
Judge

anthropic/claude-opus-5

Decides when the Oracle and Reviewer disagree.

Calls
97
Cost
$1.1614
Reasoning
Medium
Routing
OpenRouter automatic routing
Resolved provider
Claude Platform on AWS
Resolved provider details
Claude Platform on AWS97 calls$1.1614313.39 s
Fallback calls
0
Provider unreported
0
Guess Validator

openai/gpt-5.6-luna

Checks a submitted guess against the trusted subject.

Calls
45
Cost
$0.0057
Reasoning
Medium
Routing
Exact provider · openai
Resolved model
openai/gpt-5.6-luna
Resolved provider
OpenAI
Prompt contract
strict-guess-validator-v1
Configuration
gpt-5.6-luna-validator
Resolved provider details
OpenAI45 calls$0.005762.70 s
Fallback calls
0
Provider unreported
0