Human Baseline
Metacognition →Kaggle × Google DeepMind hackathon, Metacognition track winner20267 items →Multiple choice

At a glance

per-question choice · false-premise detection · response time

MetaBound asked models 140 questions built on fake drugs, invented laws and impossible facts. The right move is to flag the false premise instead of answering.

Results span the full range: GLM-5, Qwen 3 Next and Gemma 4 flagged every one; Claude Opus 4.6 and GPT-5.4 nano answered all 140 confidently, inventing details.

Best AI
100% → 0%
Humans
~20%
Our items
7
Metacognition gap
-55.0

Human vs AI

What the benchmark reports, and where the faculty stands overall

Best published AI

100% → 0%

false premises flagged, best vs worst model

MetaBound, Metacognition track winner, 2026 ↗

Humans

~20%

detect subtle false premises (Moses illusion)

Live figure on our items below, once n ≥ 30.

Metacognition, faculty level · 0–100

Human ref.
~45.0
Best AI
100.0

Human: Our estimate (Moses illusion + invented names). AI: MetaBound, best model; two models at 0%.

How we mirror it

Each set mixes genuine questions with subtle false premises (Moses-illusion style swaps) and invented entities. Every question offers “The question is false” as an option.

Separating subtle swaps from invented names lets us test our own human estimate (~45%), which is currently a guess.

Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.

What we record

per-question choicefalse-premise detectionresponse timepastestab switches

Items

7 rows
ItemFormatDifficultyTime limitResponsesHumans passed
Spot the fake
mcq-001 · v1
Multiple choice—0collecting…
Spot the fake
mcq-003 · v1
Multiple choice—0collecting…
Spot the fake
mcq-004 · v1
Multiple choice—0collecting…
Spot the fake
mcq-005 · v1
Multiple choice—0collecting…
Spot the fake
mcq-006 · v1
Multiple choice—0collecting…
Spot the fake
mcq-007 · v1
Multiple choice—0collecting…
Spot the fake
mcq-008 · v1
Multiple choice—0collecting…
Difficulty is the author’s prior; it is replaced by estimates from the data.0 responses · pass rates shown from n ≥ 30