Human Baseline

MEDLEY-BENCH

View original ↗
Metacognition →Kaggle × Google DeepMind hackathon, grand prize20266 items →Answer, see crowd, revise

At a glance

first vs final answer · switch to majority · response time

MEDLEY-BENCH shows a model a wrong majority while keeping each analyst's argument visible, then asks whether it keeps or changes its answer.

Some models (Claude Haiku 4.5, Gemini 3 Flash, GPT-4.1 Nano) weighed the arguments and held steady; others (Mistral Small, Qwen 3 8B) tracked the headcount.

Best AI
split
Humans
—
Our items
6
Metacognition gap
-55.0

Human vs AI

What the benchmark reports, and where the faculty stands overall

Best published AI

split

some models weigh arguments, others count heads

MEDLEY-BENCH, grand prize, 2026 ↗

Humans

—

No human figure published on this benchmark. Ours will be the first.

Live figure on our items below, once n ≥ 30.

Metacognition, faculty level · 0–100

Human ref.
~45.0
Best AI
100.0

Human: Our estimate (Moses illusion + invented names). AI: MetaBound, best model; two models at 0%.

How we mirror it

You answer a problem with a counter-intuitive solution, then see eight analysts: a confident wrong majority and one analyst with the sound argument. You lock a final answer.

We log whether you switched, and whether you switched to the majority or to the argument.

Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.

What we record

first vs final answerswitch to majorityresponse timepastestab switches

Items

6 rows
ItemFormatDifficultyTime limitResponsesHumans passed
Hold your ground
crowd-001 · v1
Answer, see crowd, revise—0collecting…
Hold your ground
crowd-002 · v1
Answer, see crowd, revise—0collecting…
Hold your ground
crowd-003 · v1
Answer, see crowd, revise—0collecting…
Hold your ground
crowd-004 · v1
Answer, see crowd, revise—0collecting…
Hold your ground
crowd-005 · v1
Answer, see crowd, revise—0collecting…
Hold your ground
crowd-006 · v1
Answer, see crowd, revise—0collecting…
Difficulty is the author’s prior; it is replaced by estimates from the data.0 responses · pass rates shown from n ≥ 30