Terminal-Bench 2.1 Extended

Long-horizon command-line agent execution.

99Mercor tasks
89Public tasks
38.0%Highest score

The Terminal-Bench 2.1 Extended leaderboard

Evaluates AI agents on hard, realistic long-horizon command-line tasks — 89 curated tasks with unique environments, human-written solutions, and verification tests.

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.