Benchmark editions
What changed in 1.1
Version 1.0 uses Yes, No, and Unknown for factual questions. Version 1.1 adds Rather yes and Rather no, so an answer can preserve a supported direction while acknowledging incomplete evidence.
Comparison of published benchmark definitions| Setting | Version 1.1 | Version 1 |
|---|
| Question answers | Yes, Rather yes, Rather no, No, Unknown | Yes, No, Unknown |
|---|
| Subjects | 10 | 7 |
|---|
| Rounds per subject | 3 | 5 |
|---|
| Question limit | 40 | 50 |
|---|
| Failure score | 41 | 51 |
|---|
| Review policy | Evidence or knowledge under the recorded policy; every exact-token disagreement goes to the Judge | Evidence first, with a limited stable-knowledge fallback |
|---|
| Guess validation | Yes, No, Unknown | Yes, No, Unknown |
|---|
Why add Rather yes and Rather no?
Web evidence can favor an answer without fully establishing the exact claim. With three answers, Unknown loses that useful direction, while a firm Yes or No can overstate what the evidence supports. The two additional answers make that gap explicit.
- Rather yes
- Relevant evidence favors the claim, but a material gap prevents a firm Yes.
- Rather no
- Relevant evidence favors rejecting the claim, but a material gap prevents a firm No.
- Unknown
- There is no reliable direction, or unresolved ambiguity or conflicting evidence changes the answer.
These labels express uncertainty about the exact claim. They are not numerical probabilities or a measure of how often a property applies. A failed search alone does not justify No or Rather no.
The Guesser is instructed to use qualified answers as clues, keep alternatives possible, and check key assumptions with a different property before narrowing heavily. Every directional answer still receives blind review, and any exact-token disagreement goes to the Judge. Identity guesses still use only Yes, No, or Unknown.
The aim is to retain useful partial evidence without treating it as certainty. Whether this improves gameplay needs evaluation; adding answer classes alone does not establish better accuracy.
Edition 1.1 also revises the question strategy instructions and adds a fixed guide to category meanings. Scores, uncertainty, costs, and efficiency are calculated within each edition. Both use the same subject-balanced question-score formula. Because subjects, trial counts, limits, and prompts also differ, a score difference across editions cannot isolate the effect of the new answers or establish model improvement.