Turn Bench
View original ↗At a glance
rules found · perseverative errors · cards usedTurn Bench secretly flips a game's rules mid-match. Claude Haiku 4.5 re-adapted in 80% of games, Claude Opus 4.6 in only 42%.
Both noticed the flip; Opus tried to reconcile it with its old picture of the game instead of starting fresh. It extends the Wisconsin Card Sorting Test (Grant & Berg, 1948).
- Best AI
- 80%
- Humans
- —
- Our items
- 4
- Executive functions gap
- +17.5
Human vs AI
What the benchmark reports, and where the faculty stands overall
Best published AI
80%
best re-adaptation rate; Opus 4.6 at 42%
Humans
—
No human figure published on this benchmark. Ours will be the first.
Live figure on our items below, once n ≥ 30.
Executive functions, faculty level · 0–100
Human: GAIA, university-educated testers. AI: GAIA, Claude Sonnet 4.5 (HAL).
How we mirror it
A card-sorting task: you sort cards onto four piles with feedback only (right/wrong). The hidden rule (colour, shape, number) switches without warning after a streak of correct sorts.
Perseverative errors (sorting by the rule that just stopped working) are logged separately, as in the clinical test.
Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.
What we record
Items
| Item | Format | Difficulty | Time limit | Responses | Humans passed |
|---|---|---|---|---|---|
Catch the rule change cardsort-001 · v1 | Card sorting | — | 0 | collecting… | |
Catch the rule change cardsort-002 · v1 | Card sorting | — | 0 | collecting… | |
Catch the rule change cardsort-003 · v1 | Card sorting | — | 0 | collecting… | |
Catch the rule change cardsort-004 · v1 | Card sorting | — | 0 | collecting… |