HedgeDecode
View original ↗At a glance
option chosen · response timeHedgeDecode found that models usually pick the right social strategy but fumble the delivery. On the hardest delivery task, no model scored above 0.60.
The most common slip is naming the hedge out loud. Overall, Claude Sonnet 4.6 led at 0.84.
- Best AI
- ≤ 0.60
- Humans
- —
- Our items
- 7
- Social cognition gap
- +4.4
Human vs AI
What the benchmark reports, and where the faculty stands overall
Best published AI
≤ 0.60
every model on the hardest delivery task
Humans
—
No human figure published on this benchmark. Ours will be the first.
Live figure on our items below, once n ≥ 30.
Social cognition, faculty level · 0–100
Human: Social IQa (Sap et al., 2019). AI: CogToM, GPT-5.1 (a different test).
How we mirror it
A short scenario and an indirect message (face-saving, polite refusal, a hint). You pick the best of four replies; each wrong option embodies a typical failure, such as taking it literally or calling out the hedge.
Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.
What we record
Items
| Item | Format | Difficulty | Time limit | Responses | Humans passed |
|---|---|---|---|---|---|
Read between the lines social-001 · v1 | Pick a reply | — | 0 | collecting… | |
Read between the lines social-002 · v1 | Pick a reply | — | 0 | collecting… | |
Read between the lines social-003 · v1 | Pick a reply | — | 0 | collecting… | |
Read between the lines social-004 · v1 | Pick a reply | — | 0 | collecting… | |
Read between the lines social-005 · v1 | Pick a reply | — | 0 | collecting… | |
Read between the lines social-006 · v1 | Pick a reply | — | 0 | collecting… | |
Read between the lines social-007 · v1 | Pick a reply | — | 0 | collecting… |