LearningBench
View original ↗At a glance
per-string accuracy · response timeLearningBench tests whether models can pick up rules they have never seen from a handful of examples. Only one of 14 models cleared 0.70 (Gemini 3.1 Pro, 0.85); 11 scored under 0.50.
It descends from Reber's artificial-grammar experiments (1967), where people classify new letter strings correctly about 60% of the time without being able to state the rule.
- Best AI
- 0.85
- Humans
- ~60%
- Our items
- 6
- Learning gap
- +0.1
Human vs AI
What the benchmark reports, and where the faculty stands overall
Humans
~60%
classic artificial-grammar studies
Live figure on our items below, once n ≥ 30.
Learning, faculty level · 0–100
Human: ARC-AGI-3, members of the public. AI: ARC-AGI-3, GPT-6 Astra with custom harness.
How we mirror it
You study 7–9 strings generated by a hidden finite-state rule, then classify 6–8 new strings as following or breaking it.
Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.
What we record
Items
| Item | Format | Difficulty | Time limit | Responses | Humans passed |
|---|---|---|---|---|---|
Learn an alien grammar grammar-001 · v1 | Classify strings | — | 0 | collecting… | |
Learn an alien grammar grammar-002 · v1 | Classify strings | — | 0 | collecting… | |
Learn an alien grammar grammar-003 · v1 | Classify strings | — | 0 | collecting… | |
Learn an alien grammar grammar-004 · v1 | Classify strings | — | 0 | collecting… | |
Learn an alien grammar grammar-005 · v1 | Classify strings | — | 0 | collecting… | |
Learn an alien grammar grammar-006 · v1 | Classify strings | — | 0 | collecting… |