Long-horizon command-line agent execution.
Evaluates AI agents on hard, realistic long-horizon command-line tasks — 89 curated tasks with unique environments, human-written solutions, and verification tests.
Model
Score
GPT-5.5xHigh
Score on the original open-source benchmark: 83.1% / Score on this Mercor-extended benchmark: 38.0%
Gemini 3.1 ProHigh
Score on the original open-source benchmark: 72.7% / Score on this Mercor-extended benchmark: 35.0%
Opus 5High
34.0%
Grok 4.5High
Score on the original open-source benchmark: 79.0% / Score on this Mercor-extended benchmark: 33.3%
GPT-5.4xHigh
Score on the original open-source benchmark: 77.2% / Score on this Mercor-extended benchmark: 31.3%
APEX NEWSLETTER
New benchmarks, leaderboard shifts, and research from the APEX team.
By subscribing you agree to receive updates from Mercor.Unsubscribe anytime.