Complex multi-turn instruction following.
Evaluates complex, multi-turn, and system-level instruction following via expert-curated rubrics over 1,600+ prompts. Paired with a reinforcement-learning method for improving instruction following.
Model
Score
Gemini 3.1 ProHigh
86.7%
Gemini 3.6 FlashHigh
85.3%
GPT-5.6 SolMax • Pro
82.1%
GPT-5.5xHigh
81.3%
GPT-5.4xHigh
81.0%
APEX NEWSLETTER
New benchmarks, leaderboard shifts, and research from the APEX team.
By subscribing you agree to receive updates from Mercor.Unsubscribe anytime.