Run workspace

Official run

Mistral Medium 3.5High

In this Twenty Questions LLM benchmark, Mistral Medium 3.5 (high) ranks 17 of 18 with a 20.06 question score and 89% success. Lower is better. The run contains 35 scored episodes across 7 subjects with 5 trials per subject. Choose a subject to inspect its attempts.

Rank 17 of 18mistralHigh

Question score20.06questions · lower is better95% CI 16.11–24.01
Success
89%
Contract compliance
>99%
Guesser cost
$14.39
Guesser time
8h 29m

Drill down

Subjects.

01Albert Einstein12.20 avg02Albert Schweitzer42.60 avg03Stephen King16.80 avg04Garfield11.20 avg05Achilles8.60 avg06Genghis Khan31.60 avg07Mario17.40 avg

Complete benchmark

Run ledger.

All Guesser and benchmark-support activity.

Total tokens
7,228,806
Wall-clock runtime
14h 58m
Guesser calls
733
Completed

Reliability

Output contract breached.

2 invalid outputs affected 1 episode and consumed 2 counted turns.

Provenance

Execution
BX-20260729-official-M0012-011
Model ID
M-0012
Benchmark
B-0001
Base seed
0
Git commit
e90f98eeaed0

Recorded cost

By role.

Cost composition.Share of full-run cost
  1. Guesser77.6%$14.39
  2. Primary Oracle16.8%$3.12
  3. Reviewer2.7%$0.50
  4. Judge2.5%$0.47
  5. Validator0.3%$0.06
Excluded repair overhead$1.727 superseded attempts across 7 affected trials. Not included above.

Roles in this cost:

Guesser
Asks the questions and submits the scored guess.
Primary Oracle
Searches for evidence and proposes an answer.
Reviewer
Checks each Oracle YES or NO independently.
Judge
Decides when the Oracle and Reviewer disagree.
Validator
Checks a submitted guess against the trusted subject.
Read the role and answer-checking method →

Run configuration

Models.

Model identity, prompt versions, routing, calls, and recorded cost for the full run.

Game support

Oracle and adjudication.

These models support the game. They are fixed across the run and are not under test.

Primary Oracle

openai/gpt-5.6-luna

Searches for evidence and proposes an answer.

Calls
545
Cost
$3.1247
Reasoning
Medium
Routing
Exact provider · openai
Resolved model
openai/gpt-5.6-luna
Resolved provider
OpenAI
Prompt contract
live-web-oracle-v7-direct-negative-evidence
Resolved provider details

Legacy run. Per-call routing totals were not retained. The configured provider was openai.

Reviewer

google/gemini-3.5-flash-lite

Checks each Oracle YES or NO independently.

Calls
508
Cost
$0.5018
Reasoning
Medium
Routing
Exact provider · google-ai-studio
Resolved provider details

Legacy run. Per-call routing totals were not retained. The configured provider was google-ai-studio.

Judge

anthropic/claude-opus-5

Decides when the Oracle and Reviewer disagree.

Calls
35
Cost
$0.4668
Reasoning
Medium
Routing
Exact provider · anthropic
Resolved provider details

Legacy run. Per-call routing totals were not retained. The configured provider was anthropic.

Guess Validator

openai/gpt-5.6-luna

Checks a submitted guess against the trusted subject.

Calls
186
Cost
$0.0606
Reasoning
Medium
Routing
Exact provider · openai
Resolved model
openai/gpt-5.6-luna
Resolved provider
OpenAI
Prompt contract
strict-guess-validator-v1
Configuration
gpt-5.6-luna-validator
Resolved provider details

Legacy run. Per-call routing totals were not retained. The configured provider was openai.