Human Baseline

EphLangBench

View original ↗
Learning →Kaggle × Google DeepMind hackathon, Learning track winner20266 items →Evaluate expressions

At a glance

per-problem accuracy · response time

EphLangBench makes models solve problems in programming languages generated on the spot. Every model dropped sharply compared with plain Python, losing between 24 and 78 points.

Pass rates: Gemini 3.1 Pro 89%, Claude Opus 4.6 63%, GPT-5.4 36%, GPT-5.4 Mini 7%.

Best AI
89%
Humans
—
Our items
6
Learning gap
+0.1

Human vs AI

What the benchmark reports, and where the faculty stands overall

Best published AI

89%

best model; GPT-5.4 at 36%

EphLangBench, Learning track winner, 2026 ↗

Humans

—

No human figure published on this benchmark. Ours will be the first.

Live figure on our items below, once n ≥ 30.

Learning, faculty level · 0–100

Human ref.
100.0
Best AI
99.9

Human: ARC-AGI-3, members of the public. AI: ARC-AGI-3, GPT-6 Astra with custom harness.

How we mirror it

Each item defines a tiny invented language (prefix operators, reversed digits, renamed symbols) with three worked examples, then asks you to evaluate three new expressions.

Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.

What we record

per-problem accuracyresponse timepastestab switches

Items

6 rows
ItemFormatDifficultyTime limitResponsesHumans passed
Speak Zeltan
minilang-001 · v1
Evaluate expressions—0collecting…
Speak Brovan
minilang-002 · v1
Evaluate expressions—0collecting…
Speak Kalo
minilang-003 · v1
Evaluate expressions—0collecting…
Speak Pento
minilang-004 · v1
Evaluate expressions—0collecting…
Speak Morv
minilang-005 · v1
Evaluate expressions—0collecting…
Speak Vetic
minilang-006 · v1
Evaluate expressions—0collecting…
Difficulty is the author’s prior; it is replaced by estimates from the data.0 responses · pass rates shown from n ≥ 30