Human Baseline
Problem solving →SimpleBench20247 items →Multiple choice

At a glance

per-question choice · response time

SimpleBench asks everyday spatial, temporal and social trick questions that are easy for people who read carefully. Nine people without special training scored 83.7%.

The best model reached 79.6% in early 2026: close to the human baseline, still below it.

Best AI
79.6%
Humans
83.7%
Our items
7
Problem solving gap
+1.8

Human vs AI

What the benchmark reports, and where the faculty stands overall

Best published AI

79.6%

best model (people: 83.7%)

SimpleBench ↗

Humans

83.7%

nine non-specialists

Live figure on our items below, once n ≥ 30.

Problem solving, faculty level · 0–100

Human ref.
83.7
Best AI
81.9

Human: SimpleBench, nine non-specialists. AI: SimpleBench, Claude Fable 5.

How we mirror it

Pairs of everyday questions with a twist (melting ice, turning on the spot, who can see what) written fresh so they can't be looked up.

Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.

What we record

per-question choiceresponse timepastestab switches

Items

7 rows
ItemFormatDifficultyTime limitResponsesHumans passed
Mind the trap
mcq-002 · v1
Multiple choice—0collecting…
Mind the trap
mcq-009 · v1
Multiple choice—0collecting…
Mind the trap
mcq-010 · v1
Multiple choice—0collecting…
Mind the trap
mcq-011 · v1
Multiple choice—0collecting…
Mind the trap
mcq-012 · v1
Multiple choice—0collecting…
Mind the trap
mcq-013 · v1
Multiple choice—0collecting…
Mind the trap
mcq-014 · v1
Multiple choice—0collecting…
Difficulty is the author’s prior; it is replaced by estimates from the data.0 responses · pass rates shown from n ≥ 30