MEDLEY-BENCH
View original ↗At a glance
first vs final answer · switch to majority · response timeMEDLEY-BENCH shows a model a wrong majority while keeping each analyst's argument visible, then asks whether it keeps or changes its answer.
Some models (Claude Haiku 4.5, Gemini 3 Flash, GPT-4.1 Nano) weighed the arguments and held steady; others (Mistral Small, Qwen 3 8B) tracked the headcount.
- Best AI
- split
- Humans
- —
- Our items
- 6
- Metacognition gap
- -55.0
Human vs AI
What the benchmark reports, and where the faculty stands overall
Best published AI
split
some models weigh arguments, others count heads
Humans
—
No human figure published on this benchmark. Ours will be the first.
Live figure on our items below, once n ≥ 30.
Metacognition, faculty level · 0–100
Human: Our estimate (Moses illusion + invented names). AI: MetaBound, best model; two models at 0%.
How we mirror it
You answer a problem with a counter-intuitive solution, then see eight analysts: a confident wrong majority and one analyst with the sound argument. You lock a final answer.
We log whether you switched, and whether you switched to the majority or to the argument.
Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.
What we record
Items
| Item | Format | Difficulty | Time limit | Responses | Humans passed |
|---|---|---|---|---|---|
Hold your ground crowd-001 · v1 | Answer, see crowd, revise | — | 0 | collecting… | |
Hold your ground crowd-002 · v1 | Answer, see crowd, revise | — | 0 | collecting… | |
Hold your ground crowd-003 · v1 | Answer, see crowd, revise | — | 0 | collecting… | |
Hold your ground crowd-004 · v1 | Answer, see crowd, revise | — | 0 | collecting… | |
Hold your ground crowd-005 · v1 | Answer, see crowd, revise | — | 0 | collecting… | |
Hold your ground crowd-006 · v1 | Answer, see crowd, revise | — | 0 | collecting… |