Fable 5 is back. Here's what changed.

Fable 5's re-release scores 54.8% Pass@1 on APEX-SWE. That's roughly 10 points below its June debut. The model still beats Opus 4.8 by 9.5 points (45.3%) and reclaims the top spot on the APEX-SWE leaderboard.

The biggest change is that Anthropic added a new cybersecurity safety classifier, which flags more benign coding and debugging requests and routes blocked prompts to Opus 4.8. As the numbers below show, that tradeoff lands on coding-heavy software work while leaving agentic professional-services tasks essentially unchanged. Keep reading to see how Fable 5 performed across our APEX benchmarks.

Fable 5 re-release scores across APEX benchmarks

How Fable 5 scored on APEX-SWE

01 · Still ranked #1 on APEX-SWE

Even after the drop, the re-release's 54.8% Pass@1 keeps Fable 5 atop the leaderboard: 9.5 points ahead of Opus 4.8 (45.3%) and 11.2 ahead of Claude Sonnet 5 (43.6%). The added safeguards cost points, but not the lead.

02 · Regression in Observability

Integration held nearly steady (61.33% → 59.33%, a 2-point dip). Observability (debugging from production-style telemetry) carried almost the entire loss, sliding from 69.67% to 50.33% (−19.3 points). The re-release seems to struggle with diagnosis, as opposed to construction.

How Fable 5 scored on APEX-Agents

01 · Agentic work: Mostly unchanged

On APEX-Agents, the re-release is statistically flat against the original: 43.3% Pass@1 versus 43.6%, well inside the margin of error, with 243 of 480 tasks solved to the original's 246. Mean criteria passed is identical at 59.2%. Whatever the new classifier costs, it doesn't land on professional-services agent work.

02 · Same work, fewer tokens

The re-release reached that flat result while running leaner: about 1.29 million tokens per run against the original's 1.37 million, roughly 6% fewer, across 14.1 tool steps to the original's 14.3.