Methodology

Edition 1.1 method.

Start with one model playing Twenty Questions. Then repeat the same controlled game across fixed subjects and convert the completed trials into one score.

Benchmark editions

What changed in 1.1

Version 1.0 uses Yes, No, and Unknown for factual questions. Version 1.1 adds Rather yes and Rather no, so an answer can preserve a supported direction while acknowledging incomplete evidence.

Comparison of published benchmark definitions
SettingVersion 1.1Version 1
Question answersYes, Rather yes, Rather no, No, UnknownYes, No, Unknown
Subjects107
Rounds per subject35
Question limit4050
Failure score4151
Review policyEvidence or knowledge under the recorded policy; every exact-token disagreement goes to the JudgeEvidence first, with a limited stable-knowledge fallback
Guess validationYes, No, UnknownYes, No, Unknown

Why add Rather yes and Rather no?

Web evidence can favor an answer without fully establishing the exact claim. With three answers, Unknown loses that useful direction, while a firm Yes or No can overstate what the evidence supports. The two additional answers make that gap explicit.

Rather yes
Relevant evidence favors the claim, but a material gap prevents a firm Yes.
Rather no
Relevant evidence favors rejecting the claim, but a material gap prevents a firm No.
Unknown
There is no reliable direction, or unresolved ambiguity or conflicting evidence changes the answer.

These labels express uncertainty about the exact claim. They are not numerical probabilities or a measure of how often a property applies. A failed search alone does not justify No or Rather no.

The Guesser is instructed to use qualified answers as clues, keep alternatives possible, and check key assumptions with a different property before narrowing heavily. Every directional answer still receives blind review, and any exact-token disagreement goes to the Judge. Identity guesses still use only Yes, No, or Unknown.

The aim is to retain useful partial evidence without treating it as certainty. Whether this improves gameplay needs evaluation; adding answer classes alone does not establish better accuracy.

Edition 1.1 also revises the question strategy instructions and adds a fixed guide to category meanings. Scores, uncertainty, costs, and efficiency are calculated within each edition. Both use the same subject-balanced question-score formula. Because subjects, trial counts, limits, and prompts also differ, a score difference across editions cannot isolate the effect of the new answers or establish model improvement.

01 · One round

Guesser = model under test

One hidden subject. One adaptive conversation.

The Guesser receives a broad category, but not the subject. It asks one yes-or-no question at a time, uses every prior answer, and eventually makes an exact guess.

In this illustrative round, three questions count. The correct guess does not. An incorrect guess before the limit consumes one counted question and play continues.

Starting clue
Broad category
Answer tokens
YES · RATHER_YES · RATHER_NO · NO · UNKNOWN
Question limit
40
Final opportunity
Exact guess only

02 · Answer checks

Separate, blind roles

Questions and guesses follow separate paths.

Factual questions go through independent adjudication. Exact guesses go to a separate Validator. The Guesser receives only the final answer token.

01 · Oracle

Search and cite evidence.

For a fresh answer, the Oracle searches the live web and retains relevant source context. The current policy also permits its own knowledge, with a private supporting statement. It proposes one of this edition’s question-answer tokens.

02 · Reviewer

Make a blind second decision.

For every directional Oracle answer, the no-web Reviewer uses the subject, question, and evidence without seeing the Oracle answer.

03 · Judge

Resolve disagreement.

If the first two decisions differ, including Reviewer UNKNOWN, the no-web Judge decides from the same limited material without seeing either answer.

04 · Guess Validator

Check the proposed identity.

The separate no-web Validator receives only the trusted subject and the structured guess, then returns YES, NO, or UNKNOWN.

ASK pathOracle → Reviewer for a directional answer → Judge on disagreement → finalOracle UNKNOWN is final and bypasses review.
GUESS pathGuess Validator → finalThe factual-answer roles never evaluate the identity.

Rather yes and Rather no indicate a direction supported by evidence or knowledge with a material gap. They are not numerical probabilities. Current runs allow the Oracle, Reviewer, and Judge to use evidence or their own knowledge, with a private supporting statement. The earlier accepted revision limits the Reviewer to supplied evidence and permits a bounded Judge knowledge fallback. Each run retains its actual policy and prompt versions. Any exact-token disagreement, including Yes versus Rather yes, invokes the blind Judge. Guess validation still uses only Yes, No, or Unknown.

For a particular entity, the Validator requires the same identity. For a general kind, it also accepts a recognized subtype or design variant that preserves the defining kind and satisfies the subject’s explicit restrictions. Extra detail alone does not invalidate a match.

The Guesser is fully isolated from adjudication.

It never sees the hidden subject, searches, evidence, citations, adjudicator prompts or decisions, provider traces, or private artifacts. Its visible history contains only the broad category, its own prior actions, final YES, RATHER_YES, RATHER_NO, NO, UNKNOWN tokens, and the fixed format reminder after its own invalid output.

03 · Repetition

Core subjects - qualified answers

One round becomes 30 isolated trials.

Every model plays the same 10 subjects in 3 fresh trials per subject.

Each trial starts a new Guesser conversation, without access to another trial's transcript, evidence, or private adjudicator state. A versioned, subject-independent variation token changes the repeated-call condition without revealing anything about the hidden subject.

Subjects
10
Trials / subject
3
Trials / model
30
Base seed
0

Repetition matters because model outputs vary across fresh calls. Multiple trials show whether a question strategy is consistently effective or succeeds only in some rounds.

Subject design and contamination

The current subject set is small.

Each subject has a canonical identity, accepted aliases, a clear description, and a public reference. Every model in an edition uses the same explicit subject list. The subjects and their categories are listed in each published run. This selection is not random, balanced, or representative.

This small subject set does not support broad conclusions. The size is mainly a cost constraint: every additional subject adds repeated Guesser turns, live Oracle searches, Reviewer calls, and sometimes Judge calls. Repeated trials per subject help measure variation, but repetition does not make the small subject set more representative.

Versioning and contamination

The cohort and protocol are versioned together. Official comparisons use the same fixed cohort and protocol. Subject identities and transcripts become public after publication, so a later model may have seen the subject list or earlier runs. Deep20Bench does not claim that this public cohort is resistant to benchmark contamination.

Future cohorts

Future cohorts will aim to include more subjects and broader entity types, including places and objects. Their selection rules and identities will be fixed before evaluation. A changed cohort or protocol receives a new version, and its results will be reported separately instead of merged with the current leaderboard.

04 · Scoring

average-then-average-v1

Question score is the average counted questions.

Lower is better. A model failure counts as 41, one above the 40-question limit.

Successful trialCounted questions used
Model failure41 questions
Infrastructure failureNot scored · incomplete run
Trial scorequestions used · failed trial = 41
Subject average1Tt=1Ttrial scoretT = number of trials for the subject
Final model score1Ss=1Ssubject averagesS = number of subjects

Uncertainty across repeated trials

Each model score includes a 95% confidence interval for repeated seeded trials on these fixed subjects. The calculation estimates the trial variance separately for each subject, divides it by that subject’s trial count, and combines the equally weighted variance estimates. It uses a Welch–Satterthwaite t interval so subjects may have different trial variance.

Standard errorSE=1Ss=1Sσˆs2nsS = subjects · ns = trials for subject s · σ̂2s = sample trial variance for subject s
95% intervalmodel score ± t critical value × standard error

A wider interval means the repeated trials were less consistent. The interval does not cover new subjects, model or provider changes, or future benchmark versions. It assumes separate seeded calls act as independent repetitions within each subject. A unique seed supports that assumption but does not prove it. The interval describes the mean score, not the range of individual trials. Individual model intervals are not a pairwise significance test.

Reused adjudicated answers can create shared conditions across trials. For runs using answer reuse, interpret the interval conditional on that policy and its recorded source history. The calculation does not correct for dependence from shared answers or estimate fresh-research variation for cached questions.

The Stability result view ranks the exact interval width from narrowest to widest. Every model uses the same 95% confidence level. Question score remains visible but does not affect this rank, so a consistently poor model can still be highly repeatable.

05 · Reliability

Structured action contract

Success does not erase a broken contract.

The Guesser must return exactly one valid ASK or GUESS action. Every invalid response remains visible, even when the model later finds the subject.

Turn consequenceOne counted turn before the limit
Semantic feedbackNone · format only
Published measureValid ÷ evaluated outputs

Before the question limit, an invalid response consumes one counted turn and receives the same fixed format reminder. The reminder contains no parser detail, correctness feedback, evidence, or subject information.

Episode, subject, run, and leaderboard pages report compliance, violations, affected trials, and counted penalties. The turn already affects the question total, so reliability adds no second score penalty.

06 · Official runs

Comparable evidence

Only complete runs with accepted settings enter the leaderboard.

Edition 1.1 fixes the ten subjects, three rounds per subject, question limit, game rules, and scoring policy. It includes explicitly accepted revisions of adjudication and subject descriptions. These differences can affect scores as well as the Guesser model; each run records its settings.

  • Signed run files pass integrity checks.
  • The run is terminal and contains every subject declared for this edition.
  • Every subject has every configured completed trial.
  • Completed model failures remain valid scored trials.
  • Missing or infrastructure-failed trials prevent qualification until an explicitly requested resume or repair completes them.

Edition 1.1 also checks the declared prompt revisions, subject identities, seed, game rules, support configurations, and retained role audits against one complete accepted release contract. Unlisted revisions and incomplete diagnostic runs do not qualify. Experimental execution provenance remains visible in each published run.

If several current runs qualify for one model, the newest completed run is used. The publisher never selects the best score. Invalid discovered input stops the build.

Published cost comparisons use only each trial's retained terminal attempt. Superseded infrastructure attempts remain in the signed repair ledger and gross execution total, but do not increase public model or benchmark costs.

07 · Publication

One-way reporting

Publication happens after play is finished.

The static site reads completed, signed run artifacts. It never participates in a trial and never sends published information back to the Guesser.

Model callsSigned artifactsPublic projectionStatic site

The publisher is a separate package. It does not import provider, prompt, session, retry, or credential code. Published data never returns to the Guesser.

Public pages connect each score to model runs, subjects, episodes, transcripts, answer evidence, contract violations, usage, cost, and timing. Private prompts, hidden reasoning, provider traces, and credentials remain excluded.