Human Baseline

Hendrycks et al. (2025)

View original ↗
Memory →Hendrycks et al., A definition of AGI20255 items →Study, then recall

At a glance

words recalled · intrusions

Hendrycks et al. scored GPT-4 and GPT-5 at 0% on long-term memory storage: what a model learns in one conversation is gone in the next, unless it is given an external notes file.

Within one conversation, holding eight words is trivial for a model. The gap is across sessions, not within one.

Best AI
0%
Humans
—
Our items
5
Memory gap
+21.4

Human vs AI

What the benchmark reports, and where the faculty stands overall

Best published AI

0%

long-term memory across sessions

Hendrycks et al., A definition of AGI, 2025 ↗

Humans

—

No human figure published on this benchmark. Ours will be the first.

Live figure on our items below, once n ≥ 30.

Memory, faculty level · 0–100

Human ref.
87.9
Best AI
66.5

Human: LoCoMo F1 (Maharana et al., 2024). AI: LoCoMo F1, Claude Sonnet + memory system.

How we mirror it

You study eight unrelated words for 20 seconds before the first test and recall them after all the others, several minutes later.

Intrusions (words you recall that weren't on the list) are logged, as in classic free-recall studies.

Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.

What we record

words recalledintrusionspastestab switches

Items

5 rows
ItemFormatDifficultyTime limitResponsesHumans passed
Remember the words
memory-001 · v1
Study, then recall—0collecting…
Remember the words
memory-002 · v1
Study, then recall—0collecting…
Remember the words
memory-003 · v1
Study, then recall—0collecting…
Remember the words
memory-004 · v1
Study, then recall—0collecting…
Remember the words
memory-005 · v1
Study, then recall—0collecting…
Difficulty is the author’s prior; it is replaced by estimates from the data.0 responses · pass rates shown from n ≥ 30