APEX-SWE: Observability

Evaluates a model's ability to diagnose and remediate real-world software engineering production failures.
Opus 5
Opus 5Max
63.5%
Fable 5
Fable 5Max
54.2%
Grok 4.6
Grok 4.6High
50.0%
Grok 4.5
Grok 4.5High
47.0%
Kimi K3
Kimi K3Max
36.0%
Opus 4.8
Opus 4.8Max
35.7%
Opus 4.7
Opus 4.7Max
33.0%
Sonnet 5
Sonnet 5Max
32.5%
DeepSeek-V4-Flash
DeepSeek-V4-FlashMax
32.0%
GPT-5.6 Sol
GPT-5.6 SolxHigh
31.5%
Kimi K2.7 Code
Kimi K2.7 CodeHigh
24.5%
MiniMax-M3
MiniMax-M3High
22.0%
Sonnet 4.6
Sonnet 4.6High
20.0%
GLM-5.2
GLM-5.2Max
18.0%
GPT-5.5
GPT-5.5xHigh
17.0%
DeepSeek-V4-Pro
DeepSeek-V4-ProMax
17.0%
GPT-5.4
GPT-5.4xHigh
16.3%
Qwen 3.5
Qwen 3.5
15.0%
Gemini 3.5 Flash
Gemini 3.5 FlashHigh
14.5%
Gemini 3.6 Flash
Gemini 3.6 FlashHigh
14.5%
Gemini 3.1 Pro
Gemini 3.1 ProHigh
12.0%
Inkling
InklingHigh
10.2%
MiniMax-M2.7
MiniMax-M2.7High
9.5%
DeepSeek-V3.2
DeepSeek-V3.2
9.2%
Nemotron 3 Ultra
Nemotron 3 UltraHigh
8.3%
Kimi K2
Kimi K2HighThinking
8.0%
GPT-OSS-120B
GPT-OSS-120BHigh
3.0%
0%
10%
20%
30%
40%
50%
60%
70%
80%
90%
100%

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.