Human Baseline
Social cognition →Kaggle × Google DeepMind hackathon, Social Cognition track winner20267 items →Pick a reply

At a glance

option chosen · response time

HedgeDecode found that models usually pick the right social strategy but fumble the delivery. On the hardest delivery task, no model scored above 0.60.

The most common slip is naming the hedge out loud. Overall, Claude Sonnet 4.6 led at 0.84.

Best AI
≤ 0.60
Humans
—
Our items
7
Social cognition gap
+4.4

Human vs AI

What the benchmark reports, and where the faculty stands overall

Best published AI

≤ 0.60

every model on the hardest delivery task

HedgeDecode, Social Cognition track winner, 2026 ↗

Humans

—

No human figure published on this benchmark. Ours will be the first.

Live figure on our items below, once n ≥ 30.

Social cognition, faculty level · 0–100

Human ref.
84.4
Best AI
80.0

Human: Social IQa (Sap et al., 2019). AI: CogToM, GPT-5.1 (a different test).

How we mirror it

A short scenario and an indirect message (face-saving, polite refusal, a hint). You pick the best of four replies; each wrong option embodies a typical failure, such as taking it literally or calling out the hedge.

Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.

What we record

option chosenresponse timepastestab switches

Items

7 rows
ItemFormatDifficultyTime limitResponsesHumans passed
Read between the lines
social-001 · v1
Pick a reply—0collecting…
Read between the lines
social-002 · v1
Pick a reply—0collecting…
Read between the lines
social-003 · v1
Pick a reply—0collecting…
Read between the lines
social-004 · v1
Pick a reply—0collecting…
Read between the lines
social-005 · v1
Pick a reply—0collecting…
Read between the lines
social-006 · v1
Pick a reply—0collecting…
Read between the lines
social-007 · v1
Pick a reply—0collecting…
Difficulty is the author’s prior; it is replaced by estimates from the data.0 responses · pass rates shown from n ≥ 30