Hendrycks et al. (2025)
View original ↗At a glance
words recalled · intrusionsHendrycks et al. scored GPT-4 and GPT-5 at 0% on long-term memory storage: what a model learns in one conversation is gone in the next, unless it is given an external notes file.
Within one conversation, holding eight words is trivial for a model. The gap is across sessions, not within one.
- Best AI
- 0%
- Humans
- —
- Our items
- 5
- Memory gap
- +21.4
Human vs AI
What the benchmark reports, and where the faculty stands overall
Humans
—
No human figure published on this benchmark. Ours will be the first.
Live figure on our items below, once n ≥ 30.
Memory, faculty level · 0–100
Human: LoCoMo F1 (Maharana et al., 2024). AI: LoCoMo F1, Claude Sonnet + memory system.
How we mirror it
You study eight unrelated words for 20 seconds before the first test and recall them after all the others, several minutes later.
Intrusions (words you recall that weren't on the list) are logged, as in classic free-recall studies.
Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.
What we record
Items
| Item | Format | Difficulty | Time limit | Responses | Humans passed |
|---|---|---|---|---|---|
Remember the words memory-001 · v1 | Study, then recall | — | 0 | collecting… | |
Remember the words memory-002 · v1 | Study, then recall | — | 0 | collecting… | |
Remember the words memory-003 · v1 | Study, then recall | — | 0 | collecting… | |
Remember the words memory-004 · v1 | Study, then recall | — | 0 | collecting… | |
Remember the words memory-005 · v1 | Study, then recall | — | 0 | collecting… |