The model improved its rankings on APEX-SWE and APEX-Agents. Its 43.7% Pass@1 score on APEX-SWE ranks third overall, just behind Opus 4.8 (45.3%) and Anthropic's own Fable 5 (65.5%). On APEX-Agents, Sonnet 5 climbed from #15 to #10 in Pass@1.
Here's how Sonnet 5 did across individual domains: For Integration tasks on APEX-SWE, Sonnet 5 now ranks #2 at 54.3% Pass@1. Observability remains the harder domain for nearly every model we've tested — Sonnet 5 is no exception, though its 33.0% score marks a substantial gain from Sonnet 4.5's 18.7%. For APEX-Agents long-horizon agentic work in investment banking and management consulting, the model's Pass@1 rose 9–16 points over Sonnet 4.6.

Four takeaways from the Sonnet 5 model release
01 — A generational jump on APEX-Agents
Sonnet 5 gained 8.8 points in Pass@1 and 7.8 points in mean score (40.7% → 48.5%) over Sonnet 4.6. That moved the model from #15 to #10 in Pass@1 and #13 to #11 in mean score on the overall APEX-Agents leaderboard.
02 — Investment banking and consulting lead the gains
Two APEX-Agents domains showed the biggest improvement. Investment Banking Analyst Pass@1 rose 9.3 points (26.2% → 35.5%), climbing from #18 to #10. Management Consultant Pass@1 rose 15.9 points (24.0% → 39.9%). This was the largest domain gain in this release — moving the model from #20 to #13.
03 — Third on APEX-SWE
Sonnet 5 posted 43.7% Pass@1 on APEX-SWE, landing in third place on our updated leaderboard — behind Fable 5 (65.5%) and Opus 4.8 (45.3%), and ahead of GPT 5.3 Codex and the rest of the field. It's the highest APEX-SWE score we've recorded from a Sonnet-class model to date.
04 — Integration up, Observability narrowing
Both APEX-SWE domains improved over Sonnet 4.5. Integration Pass@1 rose to 54.3% (vs. 43.3%), good for #2 on that domain's leaderboard. Observability climbed to 33.0% (vs. 18.7%) — still behind Fable 5 (69.7%) and Opus 4.8 (43.3%), but a meaningful step. Fable 5 remains the only model on the leaderboard where Observability outscores Integration.

