Human Baseline
Executive functions →Kaggle × Google DeepMind hackathon, Executive Functions track winner20264 items →Card sorting

At a glance

rules found · perseverative errors · cards used

Turn Bench secretly flips a game's rules mid-match. Claude Haiku 4.5 re-adapted in 80% of games, Claude Opus 4.6 in only 42%.

Both noticed the flip; Opus tried to reconcile it with its old picture of the game instead of starting fresh. It extends the Wisconsin Card Sorting Test (Grant & Berg, 1948).

Best AI
80%
Humans
—
Our items
4
Executive functions gap
+17.5

Human vs AI

What the benchmark reports, and where the faculty stands overall

Best published AI

80%

best re-adaptation rate; Opus 4.6 at 42%

Turn Bench, Executive Functions track winner, 2026 ↗

Humans

—

No human figure published on this benchmark. Ours will be the first.

Live figure on our items below, once n ≥ 30.

Executive functions, faculty level · 0–100

Human ref.
92.0
Best AI
74.5

Human: GAIA, university-educated testers. AI: GAIA, Claude Sonnet 4.5 (HAL).

How we mirror it

A card-sorting task: you sort cards onto four piles with feedback only (right/wrong). The hidden rule (colour, shape, number) switches without warning after a streak of correct sorts.

Perseverative errors (sorting by the rule that just stopped working) are logged separately, as in the clinical test.

Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.

What we record

rules foundperseverative errorscards usedpastestab switches

Items

4 rows
ItemFormatDifficultyTime limitResponsesHumans passed
Catch the rule change
cardsort-001 · v1
Card sorting—0collecting…
Catch the rule change
cardsort-002 · v1
Card sorting—0collecting…
Catch the rule change
cardsort-003 · v1
Card sorting—0collecting…
Catch the rule change
cardsort-004 · v1
Card sorting—0collecting…
Difficulty is the author’s prior; it is replaced by estimates from the data.0 responses · pass rates shown from n ≥ 30