Human Baseline

Wason selection task

View original ↗
Reasoning →Wason selection task19686 items →Select cards

At a glance

cards selected · framing · response time

Four cards, a conditional rule, and the question of which cards must be turned to test it. In the abstract version, typically fewer than one person in ten picks the right cards; most choose the card that would confirm the rule rather than the one that could break it.

Framed as a social rule (e.g. a drinking-age check), most people solve it. The textbook version is all over the internet, so model scores on it say little.

Best AI
n/a
Humans
< 10%
Our items
6
Reasoning gap
-26.3

Human vs AI

What the benchmark reports, and where the faculty stands overall

Best published AI

n/a

no clean held-out AI measurement (people: < 10%)

Wason (1968) ↗

Humans

< 10%

abstract version (Wason, 1968)

Live figure on our items below, once n ≥ 30.

Reasoning, faculty level · 0–100

Human ref.
69.7
Best AI
96.0

Human: GPQA Diamond, PhD experts. AI: GPQA Diamond, GPT-6 Astra (vendor-reported).

How we mirror it

Fresh cards and rules in both abstract and social-contract framings, so the classic framing effect can be measured on a large modern sample.

Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.

What we record

cards selectedframingresponse timepastestab switches

Items

6 rows
ItemFormatDifficultyTime limitResponsesHumans passed
Check the rule
wason-001 · v1
Select cards—0collecting…
Check the rule
wason-002 · v1
Select cards—0collecting…
Check the rule
wason-003 · v1
Select cards—0collecting…
Check the rule
wason-004 · v1
Select cards—0collecting…
Check the rule
wason-005 · v1
Select cards—0collecting…
Check the rule
wason-006 · v1
Select cards—0collecting…
Difficulty is the author’s prior; it is replaced by estimates from the data.0 responses · pass rates shown from n ≥ 30