Real-world code issue resolution.
Tests whether LLMs can resolve real-world GitHub issues by editing codebases, spanning 2,294 problems from 12 popular Python repos and requiring multi-file, long-context changes.
Model
Score
Opus 5Max
Score on the original open-source benchmark: 93.5% / Score on this Mercor-extended benchmark: 82.0%
Fable 5Max
Score on the original open-source benchmark: 95.9% / Score on this Mercor-extended benchmark: 78.7%
Grok 4.5High
Score on the original open-source benchmark: 81.7% / Score on this Mercor-extended benchmark: 70.7%
Sonnet 5Max
Score on the original open-source benchmark: 82.8% / Score on this Mercor-extended benchmark: 67.3%
Opus 4.8Max
Score on the original open-source benchmark: 88.5% / Score on this Mercor-extended benchmark: 63.3%
APEX NEWSLETTER
New benchmarks, leaderboard shifts, and research from the APEX team.
By subscribing you agree to receive updates from Mercor.Unsubscribe anytime.