Claude Fable 5.1’s four largest public leads
On the boards that separate it most from other named models, Fable 5.1 is not a rounding error. It is the gap.
Anthropic’s Fable 5.1 launch table, Vals AI’s 1 September 2026 writeup, BenchLM’s CursorBench snapshot, and Artificial Analysis’s Terminal-Bench v2.1 page were read as a set. Each row keeps Claude Fable 5.1 and the other named models published on that same row. The page does not average unrelated boards, and it does not hide the benches Fable 5.1 loses.
A board favors Fable 5.1 when the score is a percent and Fable 5.1 is strictly ahead of every listed comparator. Lead is Fable 5.1 minus the strongest of those comparators. The four figures below are the four largest leads in that sourced set — currently 4 boards. Elo rows, ties, empty comparator lists, and losses (including Harvey’s Legal Agent Benchmark) stay in the table and out of the figures.
Terminal-Bench-Science 0.1
What it tests.
How it is scored.
AutomationBench
What it tests.
How it is scored.
Terminal-Bench 4.0
What it tests.
How it is scored.
Humanity's Last Exam no-tools
What it tests.
How it is scored.
Orchestration
Three of those four boards already measure orchestration: a scientific research workflow, a SaaS business workflow, and a long-horizon terminal agent. The figures below add the other public rows in the sourced set where Fable 5.1 is ahead and the job is to run a loop — knowledge-work deliverables, a desktop, an IDE agent — rather than answer an exam question.
GDPval-AA v2
What it tests.
How it is scored.
OSWorld 2.0 (partial)
What it tests.
How it is scored.
CursorBench 3.2.0
What it tests.
How it is scored.
Sources
- Introducing Claude Fable 5.1 and Claude Mythos 5.1 — Anthropic launch comparison table.
- Claude Fable 5.1 on Vals AI — 1 September 2026 independent suite, including ProofBench, EMB, and LiveCodeBench.
- CursorBench on BenchLM — public snapshot of CursorBench 3.2.0, 1 September 2026.
- GDPval-AA v2 — Artificial Analysis — occupational deliverables, pairwise Elo.
- Terminal-Bench-Science 0.1 — Harbor / Stanford scientific-workflow board.
- AutomationBench — Zapier — simulated SaaS workflows, state-based grading.
- OSWorld 2.0 — long-horizon computer-use on a desktop VM.
- CursorBench 3.2 — Cursor’s first-party IDE-agent eval.
- Humanity’s Last Exam — 2,500 closed-ended expert questions.