Human Baseline
Problem solving →ARC Prize Foundation20197 items →Grid painting

At a glance

tries used · final grid · response time

ARC-AGI presents a few input→output grid pairs and asks for the output of a new input. Each puzzle has its own rule, so memorised knowledge doesn't help: the rule must be induced from two or three examples.

On ARC-AGI-2 a human panel solves every puzzle; the best model as of March 2026 reached 84.6%. The interactive ARC-AGI-3 widens the gap further.

Best AI
84.6%
Humans
100%
Our items
7
Problem solving gap
+1.8

Human vs AI

What the benchmark reports, and where the faculty stands overall

Best published AI

84.6%

best model on ARC-AGI-2 (humans: 100%)

ARC Prize leaderboard ↗

Humans

100%

ARC-AGI-2 human panel

Live figure on our items below, once n ≥ 30.

Problem solving, faculty level · 0–100

Human ref.
83.7
Best AI
81.9

Human: SimpleBench, nine non-specialists. AI: SimpleBench, Claude Fable 5.

How we mirror it

Three worked examples and one test grid. You paint the answer cell by cell and get two tries.

Every painted grid is stored, not just pass/fail, so near-misses and common wrong rules can be analysed.

Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.

What we record

tries usedfinal gridresponse timepastestab switches

Items

7 rows
ItemFormatDifficultyTime limitResponsesHumans passed
Find the rule
arc-001 · v1
Grid painting—0collecting…
Find the rule
arc-002 · v1
Grid painting—0collecting…
Find the rule
arc-003 · v1
Grid painting—0collecting…
Find the rule
arc-004 · v1
Grid painting—0collecting…
Find the rule
arc-005 · v1
Grid painting—0collecting…
Find the rule
arc-006 · v1
Grid painting—0collecting…
Find the rule
arc-007 · v1
Grid painting—0collecting…
Difficulty is the author’s prior; it is replaced by estimates from the data.0 responses · pass rates shown from n ≥ 30