Human Baseline
Attention →Kaggle × Google DeepMind hackathon20267 items →Numeric answer

At a glance

lure capture · response time · timeouts

RIAC asks a model to report a number from a target sentence while a look-alike decoy number sits nearby.

With one planted lure, frontier models fell to 51% accuracy and every miss picked the lure. Repeating the lure a hundred times restored a perfect score: a single decoy is harder to ignore than a crowd of them.

Best AI
51%
Humans
—
Our items
7
Attention gap
+44.0

Human vs AI

What the benchmark reports, and where the faculty stands overall

Best published AI

51%

frontier models with one planted lure

RIAC, Kaggle × Google DeepMind hackathon, 2026 ↗

Humans

—

No human figure published on this benchmark. Ours will be the first.

Live figure on our items below, once n ≥ 30.

Attention, faculty level · 0–100

Human ref.
~95.0
Best AI
51.0

Human: Our estimate, untimed reading task. AI: RIAC, one planted lure, pooled.

How we mirror it

Each item is a short, realistic note with one target number and one planted decoy of similar size and shape. You have a few seconds to read it and type the number.

Picking the decoy is logged separately from other errors, so we can measure lure capture in people directly, not just accuracy.

Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.

What we record

lure captureresponse timetimeoutspastestab switches

Items

7 rows
ItemFormatDifficultyTime limitResponsesHumans passed
Ignore the lure
numeric-001 · v1
Numeric answer12s0collecting…
Ignore the lure
numeric-003 · v1
Numeric answer12s0collecting…
Ignore the lure
numeric-004 · v1
Numeric answer15s0collecting…
Ignore the lure
numeric-005 · v1
Numeric answer18s0collecting…
Ignore the lure
numeric-006 · v1
Numeric answer20s0collecting…
Ignore the lure
numeric-007 · v1
Numeric answer25s0collecting…
Ignore the lure
numeric-008 · v1
Numeric answer25s0collecting…
Difficulty is the author’s prior; it is replaced by estimates from the data.0 responses · pass rates shown from n ≥ 30