Human Baseline
Perception →ClockBench (Safar)20257 items →Read clocks

At a glance

error in minutes · response time

ClockBench pitted 11 models against five people on 180 analog clocks. People read the time correctly 89.1% of the time; the best model, Gemini 2.5 Pro, managed 13.3%.

When models missed, they were typically off by one to three hours; people's median error was three minutes.

Best AI
13.3%
Humans
89.1%
Our items
7
Perception gap
+13.0

Human vs AI

What the benchmark reports, and where the faculty stands overall

Best published AI

13.3%

best model (people: 89.1%)

ClockBench, 2025 ↗

Humans

89.1%

five people on 180 clocks

Live figure on our items below, once n ≥ 30.

Perception, faculty level · 0–100

Human ref.
91.4
Best AI
78.4

Human: Perception Test, untrained adults. AI: Perception Test, Gemini 2.5 Pro.

How we mirror it

Two or three clocks per item, mixing Roman and Arabic numerals, dials without ticks and rotated dials. Within two minutes counts.

We store the error in minutes for every clock, so the human error distribution can be compared with ClockBench's directly.

Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.

What we record

error in minutesresponse timepastestab switches

Items

7 rows
ItemFormatDifficultyTime limitResponsesHumans passed
Read the clock
clock-001 · v1
Read clocks—0collecting…
Read the clock
clock-002 · v1
Read clocks—0collecting…
Read the clock
clock-003 · v1
Read clocks—0collecting…
Read the clock
clock-004 · v1
Read clocks—0collecting…
Read the clock
clock-005 · v1
Read clocks—0collecting…
Read the clock
clock-006 · v1
Read clocks—0collecting…
Read the clock
clock-007 · v1
Read clocks—0collecting…
Difficulty is the author’s prior; it is replaced by estimates from the data.0 responses · pass rates shown from n ≥ 30