Woodfield

/ Research

Claude Fable 5.1’s four largest public leads

On the boards that separate it most from other named models, Fable 5.1 is not a rounding error. It is the gap.

Folded cream paper, a black ink bottle, and a steel ruling pen on a white table
Paper, ink, and a ruling pen. The scores below are copied from public launch and leaderboard pages, then ranked by lead.

Anthropic’s Fable 5.1 launch table, Vals AI’s 1 September 2026 writeup, BenchLM’s CursorBench snapshot, and Artificial Analysis’s Terminal-Bench v2.1 page were read as a set. Each row keeps Claude Fable 5.1 and the other named models published on that same row. The page does not average unrelated boards, and it does not hide the benches Fable 5.1 loses.

A board favors Fable 5.1 when the score is a percent and Fable 5.1 is strictly ahead of every listed comparator. Lead is Fable 5.1 minus the strongest of those comparators. The four figures below are the four largest leads in that sourced set — currently 4 boards. Elo rows, ties, empty comparator lists, and losses (including Harvey’s Legal Agent Benchmark) stay in the table and out of the figures.

Largest lead

Terminal-Bench-Science 0.1

+23.6pts vs next

What it tests.

How it is scored.

Claude Fable 5.1 compared with named models on this board.
Second-largest lead

AutomationBench

+4.5pts vs next

What it tests.

How it is scored.

Claude Fable 5.1 compared with named models on this board.
Third-largest lead

Terminal-Bench 4.0

+3.5pts vs next

What it tests.

How it is scored.

Claude Fable 5.1 compared with named models on this board.
Fourth-largest lead

Humanity's Last Exam no-tools

+3.1pts vs next

What it tests.

How it is scored.

Claude Fable 5.1 compared with named models on this board.

Orchestrating other agents

That is a different job from running a terminal, a desktop, or a SaaS workflow alone. Here the lead model plans, dispatches workers, and synthesizes what comes back. Fable 5.1’s launch table does not include a dedicated planner/executor board, so the figures below use the public Fable 5 numbers that actually measure that seat: Morph’s mixed-model app builds, Anthropic’s Claude Managed Agents run on BrowseComp, and the Fable 5 system-card five-agent ProgramBench team.

Planner vs planner

Morph planner accuracy

+5.0pts vs next planner

What it tests.

How it is scored.

Claude Fable 5.1 compared with named models on this board.
Fable 5 + Sonnet 5 workers

BrowseComp, Fable 5 + Sonnet workers

96%of Fable 5 solo quality

What it tests.

How it is scored.

Claude Fable 5.1 compared with named models on this board.
Same-model worker team

ProgramBench 5-agent team

+7.9pts vs single Fable 5

What it tests.

How it is scored.

Claude Fable 5.1 compared with named models on this board.

Sources

  1. Introducing Claude Fable 5.1 and Claude Mythos 5.1 — Anthropic launch comparison table.
  2. Claude Fable 5.1 on Vals AI — 1 September 2026 independent suite, including ProofBench, EMB, and LiveCodeBench.
  3. CursorBench on BenchLM — public snapshot of CursorBench 3.2.0, 1 September 2026.
  4. Morph Multi-Agent — planner/executor pairs on 40 app builds.
  5. ClaudeDevs on BrowseComp CMA — Fable 5 orchestrator, Sonnet 5 workers, 7 Jul 2026.
  6. Claude Fable 5 and Claude Mythos 5 — system card §8.15 multi-agent harnesses.
  7. GDPval-AA v2 — Artificial Analysis — occupational deliverables, pairwise Elo.
  8. Terminal-Bench-Science 0.1 — Harbor / Stanford scientific-workflow board.
  9. AutomationBench — Zapier — simulated SaaS workflows, state-based grading.
  10. OSWorld 2.0 — long-horizon computer-use on a desktop VM.
  11. CursorBench 3.2 — Cursor’s first-party IDE-agent eval.
  12. Humanity’s Last Exam — 2,500 closed-ended expert questions.