Human Baseline
Attention →Kaggle × Google DeepMind hackathon, Attention track winner20265 items →Numeric answer

At a glance

response time · timeouts

ABC (Attention Benchmark for Cognition) tested 15 models on selective attention. Accuracy was highest when the right items could be spotted from local cues, and lowest when the answer depended on belonging to the right group, line or region.

Many models flattened nested structure, treating every matching line the same regardless of where it sat.

Best AI
weakest
Humans
—
Our items
5
Attention gap
+44.0

Human vs AI

What the benchmark reports, and where the faculty stands overall

Best published AI

weakest

on answers that depend on grouping

ABC, Attention track winner, 2026 ↗

Humans

—

No human figure published on this benchmark. Ours will be the first.

Live figure on our items below, once n ≥ 30.

Attention, faculty level · 0–100

Human ref.
~95.0
Best AI
51.0

Human: Our estimate, untimed reading task. AI: RIAC, one planted lure, pooled.

How we mirror it

Our items are nested checklists: projects with tasks and sub-tasks. The question asks for a count inside one group, with an exclusion rule, so the answer depends on structure rather than keywords.

Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.

What we record

response timetimeoutspastestab switches

Items

5 rows
ItemFormatDifficultyTime limitResponsesHumans passed
Stay in the right group
numeric-002 · v1
Numeric answer25s0collecting…
Stay in the right group
numeric-009 · v1
Numeric answer20s0collecting…
Stay in the right group
numeric-010 · v1
Numeric answer25s0collecting…
Stay in the right group
numeric-011 · v1
Numeric answer25s0collecting…
Stay in the right group
numeric-012 · v1
Numeric answer30s0collecting…
Difficulty is the author’s prior; it is replaced by estimates from the data.0 responses · pass rates shown from n ≥ 30