Best AI models right now

Best AI models right now

New models launch almost weekly, and most rankings still judge them on chat quality or trivia-style questions.

This guide takes a different angle: it ranks the leading models on how well they handle real professional work.

It draws on Mercor's APEX AI Productivity Index: a set of benchmarks built and graded by practicing domain experts, including lawyers, investment bankers, consultants, physicians, and software engineers, who write real professional tasks and the rubric used to grade every model and agent. Rather than a single score, it measures different kinds of work

  • APEX-Agents measures whether an AI agent can complete real professional tasks that span multiple steps, tools, and applications.
  • APEX-SWE measures whether a model can resolve real software-engineering issues, from integration to production debugging.

Best AI models: The quick verdict

Based on APEX benchmarks results, the models that lead where the work is hardest, autonomous agents and software engineering, are Fable 5 and Opus 5. The table below maps the best model to each need.

Best forModelWhy
Best AI agentFable 5 (Anthropic)Leads APEX-Agents at 43.6
Best for codingFable 5 (Anthropic)Leads APEX-SWE at 54.8
Best open-weightKimi K3 (Moonshot)44.7 on APEX-SWE, 39.3 on APEX-Agents
Best valueSonnet 5 (Anthropic)43.6 on APEX-SWE at roughly $2 / $20 per 1M tokens

Every figure below comes from the default leaderboard shown on each benchmark page, which reports Pass@1, with APEX-Agents on the Loop harness.

Best AI agent models across corporate law, investment banking, and management consulting

APEX-Agents measures whether an AI agent can complete real professional tasks that span multiple steps and applications, across corporate lawyer, investment banking analyst, and management consultant work. These are the top 3 on the default leaderboard, Loop harness, Pass@1.

RankAgentPass@1 (Loop)Who it is for
1Fable 5 (Anthropic)43.6Autonomous agents on the highest-stakes professional work
2Opus 5 (Anthropic)43.5The same agentic strength at standard Opus pricing
3Muse Spark 1.1 (Meta)41.9Teams that need an open-weight agent they can self-host

View full APEX-Agents leaderboard →

Opus 5 and Fable 5 are separated by just 0.2 percentage points at the top, so for agentic professional work the 2 are effectively co-leaders, with Muse Spark 1.1 close behind.

Best AI model and agent across software engineering

APEX-SWE measures whether a model can resolve real software engineering issues, the closest of the 3 families to autonomous coding work. These are the top 3 on the default leaderboard, Pass@1.

RankModelPass@1Who it is for
1Fable 5 (Anthropic)54.8Flagship coding tools and the hardest engineering issues
2Opus 5 (Anthropic)54.7Production coding and debugging at standard Opus pricing
3Grok 4.5 (xAI)51.2The strongest non-Anthropic coder, for xAI-aligned stacks

View full APEX-SWE leaderboard →

Fable 5 and Opus 5 are near-tied at the top (0.1 percentage points apart), with Grok 4.5 the strongest non-Anthropic model at 51.2%.

How should you read these rankings based on your workload?

The right model depends on which one matches your workload.

  • If you are deploying autonomous agents on professional work, choose based on APEX-Agents. Opus 5 and Fable 5 lead there, followed by Muse Spark 1.1.
  • If you are building coding tools or resolving engineering issues, choose based on APEX-SWE. Fable 5 and Opus 5 lead there, Grok 4.5 is the strongest alternative.
  • Cost or self-hosting matters most: Kimi K3 for open weights, Sonnet 5 for a low-cost hosted model with strong coding.

Notable new and rising models

The frontier is no longer just 2 US labs. 4 models stand out this cycle:

  • Gemini 3.1 Pro (Google) is the strongest non-Anthropic, non-OpenAI model, scoring 33.4 on agents, and it is the only model outside Anthropic and OpenAI in the all-around top 4.
  • Grok 4.5 (xAI) is a genuine coding contender at 51.2 on APEX-SWE, 3rd overall, though it trails on agentic work at 29.4.
  • Kimi K3 (Moonshot) is the best open-weight model, at 44.7 on software engineering and 39.3 on agents, competitive with proprietary flagships on coding.
  • Muse Spark 1.1 (Meta) is an agentic standout at 41.9 on APEX-Agents, 3rd on that board, though it has not been scored on software engineering yet.

Get the full APEX benchmark data for your workflows

The public leaderboards show the aggregate scores. The full datasets are what enterprise technology teams and AI labs use to make deployment decisions, including per-domain breakdowns, harness comparisons, rubric details, and agent trajectories.

For teams that need more:

Evaluate AI model and agent performance on your own workflows with custom evaluations graded by Mercor's network of domain experts.

Get in touch →

Frequently Asked Questions

What is the best AI model right now?+

On individual leaderboards the newest models lead: Opus 5 and Fable 5 top the agentic and software-engineering boards.

What is the best AI model for coding and software engineering?+

On APEX-SWE scores, Fable 5 leads at 54.8, with Opus 5 close behind at 54.7 and Grok 4.5 the strongest non-Anthropic model at 51.2.

How often are the APEX leaderboards updated?+

The leaderboards are updated as new frontier models are evaluated, and rankings shift with each major release. Verify scores at the time of your decision on the live APEX AI productivity leaderboards.

Which scores does this article use?+

Every figure uses the (default) leaderboard shown on each benchmark page, which reports Pass@1 and the Loop harness. Per-domain scores, such as the corporate lawyer or investment banking subcategories, are not used here; those are covered in the domain-specific articles.

What is the best AI model for agentic tasks?+

On APEX-Agents, which measures agents completing multi-step professional tasks with tools, Fable 5 leads, with Opus 5 a fraction behind followed by Muse Spark 1.1.