Terminal-Bench 2.1

Long-horizon command-line agent execution.

89Public tasks
83.1%Highest score

The Terminal-Bench 2.1 leaderboard

Evaluates AI agents on hard, realistic long-horizon command-line tasks — 89 curated tasks with unique environments, human-written solutions, and verification tests.

APEX NEWSLETTER

The latest on frontier AI performance, straight to your inbox.

New benchmarks, leaderboard shifts, and research from the APEX team.

By subscribing you agree to receive updates from Mercor.
Unsubscribe anytime.