APEX-SWE: Integration

Evaluates a model's ability to orchestrate end-to-end workflows and synchronize data across heterogeneous services.
Gemini 3.6 Flash
Gemini 3.6 FlashHigh
64.2%
Opus 5
Opus 5Max
64.0%
Fable 5
Fable 5Max
63.5%
Grok 4.6
Grok 4.6High
62.7%
Opus 4.7
Opus 4.7Max
62.3%
DeepSeek-V4-Flash
DeepSeek-V4-FlashMax
61.8%
Grok 4.5
Grok 4.5High
60.3%
Sonnet 5
Sonnet 5Max
60.3%
GPT-5.6 Sol
GPT-5.6 SolxHigh
60.0%
Kimi K3
Kimi K3Max
60.0%
Gemini 3.5 Flash
Gemini 3.5 FlashHigh
57.8%
GPT-5.5
GPT-5.5xHigh
57.0%
GLM-5.2
GLM-5.2Max
56.8%
Kimi K2.7 Code
Kimi K2.7 CodeHigh
56.0%
Gemini 3.1 Pro
Gemini 3.1 ProHigh
55.7%
GPT-5.4
GPT-5.4xHigh
52.2%
Opus 4.8
Opus 4.8Max
52.0%
DeepSeek-V4-Pro
DeepSeek-V4-ProMax
50.5%
MiniMax-M3
MiniMax-M3High
49.5%
Sonnet 4.6
Sonnet 4.6High
49.0%
DeepSeek-V3.2
DeepSeek-V3.2
42.0%
Inkling
InklingHigh
42.0%
Qwen 3.5
Qwen 3.5
35.2%
MiniMax-M2.7
MiniMax-M2.7High
33.5%
Kimi K2
Kimi K2HighThinking
26.8%
Nemotron 3 Ultra
Nemotron 3 UltraHigh
24.5%
GPT-OSS-120B
GPT-OSS-120BHigh
3.5%
0%
10%
20%
30%
40%
50%
60%
70%
80%
90%
100%

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.