ARC-AGI
View original ↗At a glance
tries used · final grid · response timeARC-AGI presents a few input→output grid pairs and asks for the output of a new input. Each puzzle has its own rule, so memorised knowledge doesn't help: the rule must be induced from two or three examples.
On ARC-AGI-2 a human panel solves every puzzle; the best model as of March 2026 reached 84.6%. The interactive ARC-AGI-3 widens the gap further.
- Best AI
- 84.6%
- Humans
- 100%
- Our items
- 7
- Problem solving gap
- +1.8
Human vs AI
What the benchmark reports, and where the faculty stands overall
Humans
100%
ARC-AGI-2 human panel
Live figure on our items below, once n ≥ 30.
Problem solving, faculty level · 0–100
Human: SimpleBench, nine non-specialists. AI: SimpleBench, Claude Fable 5.
How we mirror it
Three worked examples and one test grid. You paint the answer cell by cell and get two tries.
Every painted grid is stored, not just pass/fail, so near-misses and common wrong rules can be analysed.
Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.
What we record
Items
| Item | Format | Difficulty | Time limit | Responses | Humans passed |
|---|---|---|---|---|---|
Find the rule arc-001 · v1 | Grid painting | — | 0 | collecting… | |
Find the rule arc-002 · v1 | Grid painting | — | 0 | collecting… | |
Find the rule arc-003 · v1 | Grid painting | — | 0 | collecting… | |
Find the rule arc-004 · v1 | Grid painting | — | 0 | collecting… | |
Find the rule arc-005 · v1 | Grid painting | — | 0 | collecting… | |
Find the rule arc-006 · v1 | Grid painting | — | 0 | collecting… | |
Find the rule arc-007 · v1 | Grid painting | — | 0 | collecting… |