Real-world code issue resolution.
Tests whether LLMs can resolve real-world GitHub issues by editing codebases, spanning 2,294 problems from 12 popular Python repos and requiring multi-file, long-context changes.
Model
Score
Fable 5Max
93.5%
Opus 5Max
Grok 4.5High
81.7%
Gemini 3.6 FlashHigh
81.1%
GPT-5.5xHigh
78.8%
APEX NEWSLETTER
New benchmarks, leaderboard shifts, and research from the APEX team.
By subscribing you agree to receive updates from Mercor.Unsubscribe anytime.