GAUGE
View original ↗At a glance
bets vs folds · good folds · pointsGAUGE pays models to fold when unsure. Gemini 2.5 Pro was the most accurate model (94%) yet never folded once across 270 problems.
Claude Haiku 4.5 was less accurate but folded 25 times, and 84% of those answers would have been wrong: good calibration can beat raw accuracy.
- Best AI
- 0 folds
- Humans
- —
- Our items
- 5
- Metacognition gap
- -55.0
Human vs AI
What the benchmark reports, and where the faculty stands overall
Humans
—
No human figure published on this benchmark. Ours will be the first.
Live figure on our items below, once n ≥ 30.
Metacognition, faculty level · 0–100
Human: Our estimate (Moses illusion + invented names). AI: MetaBound, best model; two models at 0%.
How we mirror it
Three problems; after each answer you bet or fold. A correct bet earns 3, a wrong bet costs 1, folding earns 1 if you were wrong. One problem is deliberately hard and timed, so folding is the rational move.
Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.
What we record
Items
| Item | Format | Difficulty | Time limit | Responses | Humans passed |
|---|---|---|---|---|---|
Bet or fold betfold-001 · v1 | Answer, then bet or fold | — | 0 | collecting… | |
Bet or fold betfold-002 · v1 | Answer, then bet or fold | — | 0 | collecting… | |
Bet or fold betfold-003 · v1 | Answer, then bet or fold | — | 0 | collecting… | |
Bet or fold betfold-004 · v1 | Answer, then bet or fold | — | 0 | collecting… | |
Bet or fold betfold-005 · v1 | Answer, then bet or fold | — | 0 | collecting… |