ClockBench
View original ↗At a glance
error in minutes · response timeClockBench pitted 11 models against five people on 180 analog clocks. People read the time correctly 89.1% of the time; the best model, Gemini 2.5 Pro, managed 13.3%.
When models missed, they were typically off by one to three hours; people's median error was three minutes.
- Best AI
- 13.3%
- Humans
- 89.1%
- Our items
- 7
- Perception gap
- +13.0
Human vs AI
What the benchmark reports, and where the faculty stands overall
Humans
89.1%
five people on 180 clocks
Live figure on our items below, once n ≥ 30.
Perception, faculty level · 0–100
Human: Perception Test, untrained adults. AI: Perception Test, Gemini 2.5 Pro.
How we mirror it
Two or three clocks per item, mixing Roman and Arabic numerals, dials without ticks and rotated dials. Within two minutes counts.
We store the error in minutes for every clock, so the human error distribution can be compared with ClockBench's directly.
Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.
What we record
Items
| Item | Format | Difficulty | Time limit | Responses | Humans passed |
|---|---|---|---|---|---|
Read the clock clock-001 · v1 | Read clocks | — | 0 | collecting… | |
Read the clock clock-002 · v1 | Read clocks | — | 0 | collecting… | |
Read the clock clock-003 · v1 | Read clocks | — | 0 | collecting… | |
Read the clock clock-004 · v1 | Read clocks | — | 0 | collecting… | |
Read the clock clock-005 · v1 | Read clocks | — | 0 | collecting… | |
Read the clock clock-006 · v1 | Read clocks | — | 0 | collecting… | |
Read the clock clock-007 · v1 | Read clocks | — | 0 | collecting… |