SimpleBench
View original ↗At a glance
per-question choice · response timeSimpleBench asks everyday spatial, temporal and social trick questions that are easy for people who read carefully. Nine people without special training scored 83.7%.
The best model reached 79.6% in early 2026: close to the human baseline, still below it.
- Best AI
- 79.6%
- Humans
- 83.7%
- Our items
- 7
- Problem solving gap
- +1.8
Human vs AI
What the benchmark reports, and where the faculty stands overall
Humans
83.7%
nine non-specialists
Live figure on our items below, once n ≥ 30.
Problem solving, faculty level · 0–100
Human: SimpleBench, nine non-specialists. AI: SimpleBench, Claude Fable 5.
How we mirror it
Pairs of everyday questions with a twist (melting ice, turning on the spot, who can see what) written fresh so they can't be looked up.
Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.
What we record
Items
| Item | Format | Difficulty | Time limit | Responses | Humans passed |
|---|---|---|---|---|---|
Mind the trap mcq-002 · v1 | Multiple choice | — | 0 | collecting… | |
Mind the trap mcq-009 · v1 | Multiple choice | — | 0 | collecting… | |
Mind the trap mcq-010 · v1 | Multiple choice | — | 0 | collecting… | |
Mind the trap mcq-011 · v1 | Multiple choice | — | 0 | collecting… | |
Mind the trap mcq-012 · v1 | Multiple choice | — | 0 | collecting… | |
Mind the trap mcq-013 · v1 | Multiple choice | — | 0 | collecting… | |
Mind the trap mcq-014 · v1 | Multiple choice | — | 0 | collecting… |