At a glance
response time · timeoutsABC (Attention Benchmark for Cognition) tested 15 models on selective attention. Accuracy was highest when the right items could be spotted from local cues, and lowest when the answer depended on belonging to the right group, line or region.
Many models flattened nested structure, treating every matching line the same regardless of where it sat.
- Best AI
- weakest
- Humans
- —
- Our items
- 5
- Attention gap
- +44.0
Human vs AI
What the benchmark reports, and where the faculty stands overall
Humans
—
No human figure published on this benchmark. Ours will be the first.
Live figure on our items below, once n ≥ 30.
Attention, faculty level · 0–100
Human: Our estimate, untimed reading task. AI: RIAC, one planted lure, pooled.
How we mirror it
Our items are nested checklists: projects with tasks and sub-tasks. The question asks for a count inside one group, with an exclusion rule, so the answer depends on structure rather than keywords.
Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.
What we record
Items
| Item | Format | Difficulty | Time limit | Responses | Humans passed |
|---|---|---|---|---|---|
Stay in the right group numeric-002 · v1 | Numeric answer | 25s | 0 | collecting… | |
Stay in the right group numeric-009 · v1 | Numeric answer | 20s | 0 | collecting… | |
Stay in the right group numeric-010 · v1 | Numeric answer | 25s | 0 | collecting… | |
Stay in the right group numeric-011 · v1 | Numeric answer | 25s | 0 | collecting… | |
Stay in the right group numeric-012 · v1 | Numeric answer | 30s | 0 | collecting… |