Long-horizon command-line agent execution.
Evaluates AI agents on hard, realistic long-horizon command-line tasks — 89 curated tasks with unique environments, human-written solutions, and verification tests.
Model
Score
GPT-5.5xHigh
83.1%
Kimi K3Max
82.0%
Grok 4.5High
79.0%
Sonnet 5Max
77.9%
Gemini 3.6 FlashHigh
74.5%
APEX NEWSLETTER
New benchmarks, leaderboard shifts, and research from the APEX team.
By subscribing you agree to receive updates from Mercor.Unsubscribe anytime.