Can you beat AI?
View methodologyAt a glance
10 faculties · 2 estimated human refs · live sampleAI now aces exams that stump most people. But intelligence isn’t one number. Google DeepMind’s 2026 framework splits it into ten faculties, from perception to social cognition. On published benchmarks, people still lead on 7 of them, and frontier models lead on 3.
This study measures where you sit. Seven short items, each modelled on a published AI benchmark, show your answer next to how frontier models did. Every answer joins an open dataset comparing human and machine cognition on the same tasks.
- Humans tested
- —
- Humans lead on
- 7/10
- Largest human lead
- Attention +44.0
- Largest AI lead
- Metacognition -55.0
Cognitive profile
Best published AI vs human reference on each faculty, 0–100. Take the test to add your own line.
Both shapes are lopsided, and the bumps are in different places. AI pulls ahead on expert reasoning, producing work and flagging false premises.
People stay ahead on attention, memory, perception and executive control, the faculties that everyday life leans on.
Hollow points are our own estimates where no published human figure exists.
Benchmark map
Each square is one test item in the bank; its colour shows who leads on that faculty today.AI leadsHumans leadParity
Perception+13.0 · 7
Generation-34.9 · 5
Attention+44.0 · 12
Learning+0.1 · 12
Memory+21.4 · 5
Reasoning-26.3 · 6
Metacognition-55.0 · 18
Executive functions+17.5 · 4
Problem solving+1.8 · 14
Social cognition+4.4 · 7
Human vs AI
Head-to-head by faculty, best published result on each side
| Faculty | Human | Best AI | Gap | Human reference | AI reference |
|---|---|---|---|---|---|
Perception Taking in information from the world and making sense of it. | 91.4 | 78.4 | +13.0 | Perception Test, untrained adults | Perception Test, Gemini 2.5 Pro |
Generation Producing words or actions, and executing them well, including under constraints. | 50.0 | 84.9 | -34.9 | GDPval, expert vs expert (parity by design) | GDPval wins or ties, GPT-5.5 (vendor-reported) |
Attention Focusing on what matters and filtering out distractors that look relevant. | ~95.0 | 51.0 | +44.0 | Our estimate, untimed reading task | RIAC, one planted lure, pooled |
Learning Picking up new rules or skills from examples or feedback, then applying them. | 100.0 | 99.9 | +0.1 | ARC-AGI-3, members of the public | ARC-AGI-3, GPT-6 Astra with custom harness |
Memory Holding on to information over time and retrieving it when needed. | 87.9 | 66.5 | +21.4 | LoCoMo F1 (Maharana et al., 2024) | LoCoMo F1, Claude Sonnet + memory system |
Reasoning Drawing valid conclusions: deduction, induction, analogy and maths. | 69.7 | 96.0 | -26.3 | GPQA Diamond, PhD experts | GPQA Diamond, GPT-6 Astra (vendor-reported) |
Metacognition Knowing what you know: spotting broken questions, judging your confidence. | ~45.0 | 100.0 | -55.0 | Our estimate (Moses illusion + invented names) | MetaBound, best model; two models at 0% |
Executive functions Planning, switching strategy when rules change, inhibiting old habits. | 92.0 | 74.5 | +17.5 | GAIA, university-educated testers | GAIA, Claude Sonnet 4.5 (HAL) |
Problem solving Combining faculties to solve problems, from novel puzzles to common sense. | 83.7 | 81.9 | +1.8 | SimpleBench, nine non-specialists | SimpleBench, Claude Fable 5 |
Social cognition Understanding what people think, feel and mean without saying it. | 84.4 | 80.0 | +4.4 | Social IQa (Sap et al., 2019) | CogToM, GPT-5.1 (a different test) |
Benchmarks mirrored
Every item is original, written in the spirit of a published test
| Benchmark | Faculty | Items | AI result | What it means |
|---|---|---|---|---|
| RIAC | Attention | 7 | 51% | frontier models with one planted lure |
| ARC-AGI | Problem solving | 7 | 84.6% | best model on ARC-AGI-2 (humans: 100%) |
| LearningBench | Learning | 6 | 0.85 | best model; 11 of 14 models under 0.50 |
| EphLangBench | Learning | 6 | 89% | best model; GPT-5.4 at 36% |
| MetaBound | Metacognition | 7 | 100% → 0% | false premises flagged, best vs worst model |
| GAUGE | Metacognition | 5 | 0 folds | most accurate model never folded in 270 problems |
| Turn Bench | Executive functions | 4 | 80% | best re-adaptation rate; Opus 4.6 at 42% |
| HedgeDecode | Social cognition | 7 | ≤ 0.60 | every model on the hardest delivery task |
| MEDLEY-BENCH | Metacognition | 6 | split | some models weigh arguments, others count heads |
| ABC | Attention | 5 | weakest | on answers that depend on grouping |
| ClockBench | Perception | 7 | 13.3% | best model (people: 89.1%) |
| Constrained writing | Generation | 5 | n/a | no direct human-vs-AI benchmark yet |
| Wason selection task | Reasoning | 6 | n/a | no clean held-out AI measurement (people: < 10%) |
| SimpleBench | Problem solving | 7 | 79.6% | best model (people: 83.7%) |
| Hendrycks et al. (2025) | Memory | 5 | 0% | long-term memory across sessions |
Can you beat AI?
Seven items, about five minutes, no sign-up. Your answers stay anonymous.