Get the latest updates about APEX benchmarks – model releases, leaderboard changes, and what they mean for economically valuable work.
APEX is our family of benchmarks that assess a broader shift in how the workforce is evolving by measuring whether AI can complete professional, economically valuable work. Each benchmark tests a different dimension of professional capability. All tasks are built with Mercor experts and leading industry partners.
Long-horizon, cross-application tasks in professional services
Fable 5Max
43.3% ±4.1%
Muse Spark 1.1xHigh
41.9% ±3.9%
GPT 5.6 Sol (Max + Pro)
40.0% ±4.1%
GPT 5.6 SolMax
39.9% ±4.0%
Opus 4.8Max
39.4% ±4.2%
Work with the team behind APEX — get custom evaluations for your agents and source expert-built training data across 30+ domains.