Human Baseline

LearningBench

View original ↗
Learning →Kaggle × Google DeepMind hackathon, grand prize20266 items →Classify strings

At a glance

per-string accuracy · response time

LearningBench tests whether models can pick up rules they have never seen from a handful of examples. Only one of 14 models cleared 0.70 (Gemini 3.1 Pro, 0.85); 11 scored under 0.50.

It descends from Reber's artificial-grammar experiments (1967), where people classify new letter strings correctly about 60% of the time without being able to state the rule.

Best AI
0.85
Humans
~60%
Our items
6
Learning gap
+0.1

Human vs AI

What the benchmark reports, and where the faculty stands overall

Best published AI

0.85

best model; 11 of 14 models under 0.50

LearningBench, grand prize, 2026 ↗

Humans

~60%

classic artificial-grammar studies

Live figure on our items below, once n ≥ 30.

Learning, faculty level · 0–100

Human ref.
100.0
Best AI
99.9

Human: ARC-AGI-3, members of the public. AI: ARC-AGI-3, GPT-6 Astra with custom harness.

How we mirror it

You study 7–9 strings generated by a hidden finite-state rule, then classify 6–8 new strings as following or breaking it.

Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.

What we record

per-string accuracyresponse timepastestab switches

Items

6 rows
ItemFormatDifficultyTime limitResponsesHumans passed
Learn an alien grammar
grammar-001 · v1
Classify strings—0collecting…
Learn an alien grammar
grammar-002 · v1
Classify strings—0collecting…
Learn an alien grammar
grammar-003 · v1
Classify strings—0collecting…
Learn an alien grammar
grammar-004 · v1
Classify strings—0collecting…
Learn an alien grammar
grammar-005 · v1
Classify strings—0collecting…
Learn an alien grammar
grammar-006 · v1
Classify strings—0collecting…
Difficulty is the author’s prior; it is replaced by estimates from the data.0 responses · pass rates shown from n ≥ 30