GPT-6 Astra (high) added.
It scored 15.1 questions with 33 of 35 successful trials.
Origin · news · prior work
Patrick Heusser and Markus Tuor came up with Deep20Bench while playing Twenty Questions with the kids. Patrick then designed and built the benchmark.
Updates
Project updates, releases, and notes.
It scored 15.1 questions with 33 of 35 successful trials.
It scored 13.1 questions with 35 of 35 successful trials.
It scored 12.1 questions with 34 of 35 successful trials.
A first-person account of the Oracle errors, blind review, format failures, and design choices behind the benchmark.
The Stealth-routed model won 32 of 35 trials and scored 17.6 questions, placing 12th of 15.
It scored 14.0 questions with 34 of 35 successful trials.
It scored 14.3 questions with 35 of 35 successful trials.
We added Claude Fable 5 (high) to the official Deep20Bench results.
Prior work
Deep20Bench was developed independently. These projects address related problems.
| Year | Focus | Publication | Source |
|---|---|---|---|
| 2018 | Question strategy | Learning-to-Ask: Knowledge Acquisition via 20 QuestionsLearns question strategies for finding an entity and gathering knowledge. | Microsoft Research ↗ |
| 2022 | World knowledge | 20Q: Overlap-Free World Knowledge Benchmark for Language ModelsUses Twenty Questions to test world knowledge without train–test overlap. | ACL Anthology ↗ |
| 2024 | Planning benchmark | The Entity-Deduction ArenaTests LLM planning and state tracking through a hidden-entity game. | Apple Machine Learning Research ↗ |
| 2025 | Adaptive elicitation | Adaptive Elicitation of Latent Information Using Natural LanguageChooses natural-language questions by reducing uncertainty. | OpenReview ↗ |
| 2025 | Information gain | BED-LLM: Intelligent Information Gathering with LLMsUses expected information gain to choose questions. | Apple Machine Learning Research ↗ |
Next
Read the rules or inspect the current model runs.