EphLangBench
View original ↗At a glance
per-problem accuracy · response timeEphLangBench makes models solve problems in programming languages generated on the spot. Every model dropped sharply compared with plain Python, losing between 24 and 78 points.
Pass rates: Gemini 3.1 Pro 89%, Claude Opus 4.6 63%, GPT-5.4 36%, GPT-5.4 Mini 7%.
- Best AI
- 89%
- Humans
- —
- Our items
- 6
- Learning gap
- +0.1
Human vs AI
What the benchmark reports, and where the faculty stands overall
Humans
—
No human figure published on this benchmark. Ours will be the first.
Live figure on our items below, once n ≥ 30.
Learning, faculty level · 0–100
Human: ARC-AGI-3, members of the public. AI: ARC-AGI-3, GPT-6 Astra with custom harness.
How we mirror it
Each item defines a tiny invented language (prefix operators, reversed digits, renamed symbols) with three worked examples, then asks you to evaluate three new expressions.
Items are original, versioned, and assigned at random, so a posted answer spoils only one variant.
What we record
Items
| Item | Format | Difficulty | Time limit | Responses | Humans passed |
|---|---|---|---|---|---|
Speak Zeltan minilang-001 · v1 | Evaluate expressions | — | 0 | collecting… | |
Speak Brovan minilang-002 · v1 | Evaluate expressions | — | 0 | collecting… | |
Speak Kalo minilang-003 · v1 | Evaluate expressions | — | 0 | collecting… | |
Speak Pento minilang-004 · v1 | Evaluate expressions | — | 0 | collecting… | |
Speak Morv minilang-005 · v1 | Evaluate expressions | — | 0 | collecting… | |
Speak Vetic minilang-006 · v1 | Evaluate expressions | — | 0 | collecting… |