SWE Atlas - Codebase QnA Extended

Codebase comprehension, testing, and refactoring.

50Mercor tasks
124Public tasks
40.0%Highest score

The SWE Atlas - Codebase QnA Extended leaderboard

Benchmarks coding agents beyond issue resolution across Codebase Q&A (124 tasks), test writing (90), and refactoring (70), combining programmatic validation with rubric-based software-quality scoring (maintainability, abstractions, hygiene). Uses under-specified, agentic task formulations.

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.