Human Baseline

Constrained writing

View original ↗
Generation →Constrained writing (lipogram)20265 items →Constrained writing

At a glance

sentence text · word count · response time

No benchmark compares people and models on letter-level constraints directly. They are a known weak spot for models, which read text as tokens rather than letters; DeepMind's framework paper flags spelling trouble as a quirk of how AI perceives text.

Best AI
n/a
Humans
—
Our items
5
Generation gap
-34.9

Human vs AI

What the benchmark reports, and where the faculty stands overall

Best published AI

n/a

no direct human-vs-AI benchmark yet

Burnell et al., Measuring progress toward AGI, 2026 ↗

Humans

—

No human figure published on this benchmark. Ours will be the first.

Live figure on our items below, once n ≥ 30.

Generation, faculty level · 0–100

Human ref.
50.0
Best AI
84.9

Human: GDPval, expert vs expert (parity by design). AI: GDPval wins or ties, GPT-5.5 (vendor-reported).

How we mirror it

Write one sentence of a set length that never uses a given letter, with a minimum number of distinct words, against the clock.

This is one of the items where same-item model runs will create the first comparison, not just mirror one.

Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.

What we record

sentence textword countresponse timepastestab switches

Items

5 rows
ItemFormatDifficultyTime limitResponsesHumans passed
Write without E
lipogram-001 · v1
Constrained writing90s0collecting…
Write without A
lipogram-002 · v1
Constrained writing90s0collecting…
Write without T
lipogram-003 · v1
Constrained writing120s0collecting…
Write without O
lipogram-004 · v1
Constrained writing75s0collecting…
Write without E
lipogram-005 · v1
Constrained writing120s0collecting…
Difficulty is the author’s prior; it is replaced by estimates from the data.0 responses · pass rates shown from n ≥ 30