SWE-bench Verified

Real-world code issue resolution.

500Public tasks
93.5%Highest score

The SWE-bench Verified leaderboard

Tests whether LLMs can resolve real-world GitHub issues by editing codebases, spanning 2,294 problems from 12 popular Python repos and requiring multi-file, long-context changes.

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.