Origin · news · prior work

A shared idea, built into a benchmark.

Patrick Heusser and Markus Tuor came up with Deep20Bench while playing Twenty Questions with the kids. Patrick then designed and built the benchmark.

Latest news ↓

Updates

Project news.

Project updates, releases, and notes.

GPT-6 Astra (high) added.

It scored 15.1 questions with 33 of 35 successful trials.

View run →

Gemini 3.8 Flash (high) added.

It scored 13.1 questions with 35 of 35 successful trials.

View run →

Claude Fable 5.1 (high) added.

It scored 12.1 questions with 34 of 35 successful trials.

View run →

Behind Deep20Bench: the hard part wasn’t the game.

A first-person account of the Oracle errors, blind review, format failures, and design choices behind the benchmark.

Read article ↗

OpenRouter’s Ox Alpha (high) tested.

The Stealth-routed model won 32 of 35 trials and scored 17.6 questions, placing 12th of 15.

View run →

Gemini 3.7 Flash (high) added.

It scored 14.0 questions with 34 of 35 successful trials.

View run →

Grok 4.6 (high) added.

It scored 14.3 questions with 35 of 35 successful trials.

View run →

Claude Fable 5 (high) added.

We added Claude Fable 5 (high) to the official Deep20Bench results.

View run →

Prior work

Related research.

Deep20Bench was developed independently. These projects address related problems.

Research related to Deep20Bench
YearFocusPublicationSource
2018Question strategy

Learning-to-Ask: Knowledge Acquisition via 20 Questions

Yihong Chen et al. · KDD

Learns question strategies for finding an entity and gathering knowledge.

Microsoft Research ↗
2022World knowledge

20Q: Overlap-Free World Knowledge Benchmark for Language Models

Maxime De Bruyn et al. · GEM

Uses Twenty Questions to test world knowledge without train–test overlap.

ACL Anthology ↗
2024Planning benchmark

The Entity-Deduction Arena

Yizhe Zhang, Jiarui Lu & Navdeep Jaitly · Apple · ACL

Tests LLM planning and state tracking through a hidden-entity game.

Apple Machine Learning Research ↗
2025Adaptive elicitation

Adaptive Elicitation of Latent Information Using Natural Language

Jimmy Wang et al. · ICML

Chooses natural-language questions by reducing uncertainty.

OpenReview ↗
2025Information gain

BED-LLM: Intelligent Information Gathering with LLMs

Deepro Choudhury et al. · Apple · ICLR

Uses expected information gain to choose questions.

Apple Machine Learning Research ↗

Next

Method and results.

Read the rules or inspect the current model runs.