Claude Fable 5 tops APEX-SWE with a 20-point lead

Anthropic's Claude Fable 5 was released this week. We tested it on both APEX-Agents and APEX-SWE ahead of the launch. Here's how it performed.

APEX-SWE

Pass@1, at release

Fable 5 — 65.5% ± 6.2%

Opus 4.8 (High) — 45.3% ± 6.3%

GPT 5.3 Codex (High) — 41.5% ± 6.3%

Opus 4.7 (Max) — 41.3% ± 6.3%

GPT 5.5 (xHigh) — 40.8% ± 6.5%

Fable 5 by APEX-SWE domain

Observability — 69.7%

Integration — 61.3%

APEX-Agents

Score, at release

Gemini 3.5 Flash (High) — 49.6% ± 3.9%

Fable 5 (Max) — 45.0% ± 4.1%

Opus 4.8 (Max) — 42.5% ± 4.0%

GPT 5.5 (xHigh) — 38.4% ± 3.9%

GPT 5.4 (xHigh) — 36.0% ± 3.8%

Average token usage

Claude Fable 5 (Max) — 924k tokens per APEX-Agents run

By domain

Corporate Lawyer Pass@1 — Claude Fable 5 (Max) 40.9%