Human Baseline

Can you beat AI?

View methodology
10 faculties →15 benchmarks mirrored →~5 min · no sign-upOpen dataset ↗Updated Oct 4, 2026

At a glance

10 faculties · 2 estimated human refs · live sample

AI now aces exams that stump most people. But intelligence isn’t one number. Google DeepMind’s 2026 framework splits it into ten faculties, from perception to social cognition. On published benchmarks, people still lead on 7 of them, and frontier models lead on 3.

This study measures where you sit. Seven short items, each modelled on a published AI benchmark, show your answer next to how frontier models did. Every answer joins an open dataset comparing human and machine cognition on the same tasks.

Humans tested
—
Humans lead on
7/10
Largest human lead
Attention +44.0
Largest AI lead
Metacognition -55.0

Cognitive profile

Best published AI vs human reference on each faculty, 0–100. Take the test to add your own line.

255075100Perception+13Generation-35Attention+44Learning0Memory+21Reasoning-26Metacognition-55Executivefunctions+17Problemsolving+2Socialcognition+4
Human referenceOur estimateBest published AILabel = human − AI, points

Both shapes are lopsided, and the bumps are in different places. AI pulls ahead on expert reasoning, producing work and flagging false premises.

People stay ahead on attention, memory, perception and executive control, the faculties that everyday life leans on.

Hollow points are our own estimates where no published human figure exists.

Benchmark map

Each square is one test item in the bank; its colour shows who leads on that faculty today.AI leadsHumans leadParity

Perception+13.0 · 7

Generation-34.9 · 5

Attention+44.0 · 12

Learning+0.1 · 12

Memory+21.4 · 5

Reasoning-26.3 · 6

Metacognition-55.0 · 18

Executive functions+17.5 · 4

Problem solving+1.8 · 14

Social cognition+4.4 · 7

Human vs AI

Head-to-head by faculty, best published result on each side

FacultyHumanBest AIGapHuman referenceAI reference
Perception
Taking in information from the world and making sense of it.
91.478.4+13.0Perception Test, untrained adultsPerception Test, Gemini 2.5 Pro
Generation
Producing words or actions, and executing them well, including under constraints.
50.084.9-34.9GDPval, expert vs expert (parity by design)GDPval wins or ties, GPT-5.5 (vendor-reported)
Attention
Focusing on what matters and filtering out distractors that look relevant.
~95.051.0+44.0Our estimate, untimed reading taskRIAC, one planted lure, pooled
Learning
Picking up new rules or skills from examples or feedback, then applying them.
100.099.9+0.1ARC-AGI-3, members of the publicARC-AGI-3, GPT-6 Astra with custom harness
Memory
Holding on to information over time and retrieving it when needed.
87.966.5+21.4LoCoMo F1 (Maharana et al., 2024)LoCoMo F1, Claude Sonnet + memory system
Reasoning
Drawing valid conclusions: deduction, induction, analogy and maths.
69.796.0-26.3GPQA Diamond, PhD expertsGPQA Diamond, GPT-6 Astra (vendor-reported)
Metacognition
Knowing what you know: spotting broken questions, judging your confidence.
~45.0100.0-55.0Our estimate (Moses illusion + invented names)MetaBound, best model; two models at 0%
Executive functions
Planning, switching strategy when rules change, inhibiting old habits.
92.074.5+17.5GAIA, university-educated testersGAIA, Claude Sonnet 4.5 (HAL)
Problem solving
Combining faculties to solve problems, from novel puzzles to common sense.
83.781.9+1.8SimpleBench, nine non-specialistsSimpleBench, Claude Fable 5
Social cognition
Understanding what people think, feel and mean without saying it.
84.480.0+4.4Social IQa (Sap et al., 2019)CogToM, GPT-5.1 (a different test)
Red = AI leads~ = our estimateGreen = humans lead
Largest human leads7 faculties
Attention95 vs 51 +44.0 pts
Memory88 vs 67 +21.4 pts
Executive functions92 vs 75 +17.5 pts
Largest AI leads3 faculties
Metacognition45 vs 100 -55.0 pts
Generation50 vs 85 -34.9 pts
Reasoning70 vs 96 -26.3 pts

Benchmarks mirrored

Every item is original, written in the spirit of a published test

All benchmark pages →
BenchmarkFacultyItemsAI resultWhat it means
RIACAttention751%frontier models with one planted lure
ARC-AGIProblem solving784.6%best model on ARC-AGI-2 (humans: 100%)
LearningBenchLearning60.85best model; 11 of 14 models under 0.50
EphLangBenchLearning689%best model; GPT-5.4 at 36%
MetaBoundMetacognition7100% → 0%false premises flagged, best vs worst model
GAUGEMetacognition50 foldsmost accurate model never folded in 270 problems
Turn BenchExecutive functions480%best re-adaptation rate; Opus 4.6 at 42%
HedgeDecodeSocial cognition7≤ 0.60every model on the hardest delivery task
MEDLEY-BENCHMetacognition6splitsome models weigh arguments, others count heads
ABCAttention5weakeston answers that depend on grouping
ClockBenchPerception713.3%best model (people: 89.1%)
Constrained writingGeneration5n/ano direct human-vs-AI benchmark yet
Wason selection taskReasoning6n/ano clean held-out AI measurement (people: < 10%)
SimpleBenchProblem solving779.6%best model (people: 83.7%)
Hendrycks et al. (2025)Memory50%long-term memory across sessions

Can you beat AI?

Seven items, about five minutes, no sign-up. Your answers stay anonymous.

Start the test