AI impact on jobs: What the benchmark data shows

AI impact on jobs: What the benchmark data shows

Professionals are increasingly using AI to complete work tasks, such as drafting the 1st version of a document, summarizing research, reconciling spreadsheets, or writing code.

Most jobs remain fundamentally the same in spite of AI adoption, but the work people do within them is beginning to shift alongside broader changes in the economy, technology, and employer expectations. Employment statistics and predictions from economists can tell you where hiring is changing, but they do not address what AI can actually do.

Mercor’s APEX benchmarks evaluate how frontier models and AI agents perform real professional work and provide a useful way to understand where those changes are most likely to occur.

This article examines what those results indicate about the future of work and compares them with broader research on hiring trends, job displacement, and the changing value of human expertise.

AI's impact on jobs by profession: What Mercor's AI productivity benchmarks reveal

Labor statistics show where employment is changing, but they can't explain which parts of a job AI can actually perform. Professional benchmarks provide a more detailed view by testing how well AI models complete tasks that are part of real professional work.

Mercor's APEX benchmarks score frontier models on tasks that are created and assessed by experienced practitioners in each profession. Rather than relying on speculation, they offer real insight into where AI performs well or is less helpful. That’s because these benchmarks were written and graded by real experts in the Mercor network, across domains like software engineering, management consulting, corporate law, and investment banking.

A strong benchmark score shows that a model completed a particular task successfully, not that a professional role is being replaced. It’s an indication of where frontier AI models are most capable of assisting working professionals. Additionally, performance can change as new models are released, so results should always be verified at the source against current benchmarks.

Here’s what that looks like across 5 professions:

ProfessionHow AI is changing the workWhere AI performs bestWhere humans still lead
Software engineeringSpeeds up development and increases the value of diagnosisIntegration, code generation, and implementationProduction debugging and root-cause analysis
Primary care medicineSupports clinical reasoning without assuming responsibility for careCase summaries, differential generation, and documentationExamination, patient communication, and accountability
Big LawSupports research and drafting workResearch, document review, and first draftsLong-running matters and incomplete context
Investment bankingAccelerates analysis but still requires close reviewComparable-company analysis, first-pass models, and reportingDeal judgment, reconciliation, and regulatory sign-off
Management consultingSpeeds up analysis and shifts value toward problem selectionResearch, synthesis, and deliverable productionProblem framing and executive persuasion

Software engineering

APEX-SWE, developed in partnership with Cognition, separates software work into 2 categories: building systems across services and diagnosing failures from production data. Results show that models perform better when tasks begin with a clear objective and a well-defined starting point.

Performance is less consistent when the system is already running, the evidence is incomplete, and the cause of failure must be inferred from logs or telemetry.

Findings indicated that AI is compressing more greenfield development work than production diagnosis. As a result, engineers who can trace failures, understand system behavior, and make decisions under uncertainty are increasingly valuable.

Medical practitioner

APEX's primary care benchmark evaluates the written clinical reasoning of AI models using real case files. The results show that today's most advanced models can organize clinical symptoms, identify possible diagnoses, and produce useful clinical documentation when the relevant evidence is already contained in the medical record.

However, delivering clinical care involves more than just medical records. Models don't examine patients, notice subtle changes in their conditions, ask follow-up questions to resolve uncertainty, or take responsibility for clinical decisions and patient outcomes.

AI may reduce the time clinicians spend documenting and structuring their thinking, giving them more time to focus on patients, but it doesn't replace the broader work of delivering care.

Big Law

The implications for the legal workforce may go beyond faster document drafting. Junior associates have traditionally built their expertise and developed judgment through the research and document work AI now compresses. As firms adopt these tools, they must find new ways to preserve that training while benefiting from the AI productivity gains.

The APEX Big Law benchmark highlights a clear gap in legal work: producing a high-quality legal document is not the same as managing an entire legal matter.

AI models perform well on basic assignments, such as researching questions, reviewing documents, and drafting polished written content. However, their performance is less reliable on long-horizon tasks when a matter unfolds over several days, involves several systems or applications, or depends on incomplete or changing context.

Finance: banking and accounting

The APEX leaderboard for investment banking tasks show that AI performs well on standardized financial work, such as gathering comparable company data, preparing first-pass analyses, and supporting recurring reporting. These tasks follow repeatable processes and produce outputs that experienced professionals can review against known standards.

The limitations are more apparent when precision is critical. Even small errors in formulas, reconciliations, or assumptions can carry significant consequences. Financial work often requires a named professional to approve the result.

As a result, AI is likely to reduce the manual effort involved in financial analysis, but human reviewers remain essential to provide the judgment, verification, and regulatory accountability that make the work usable.

Management consulting

According to Mercor’s consulting associate benchmarks, the field appears highly exposed to the impacts of AI because much of the visible output consists of research, synthesis, and presentation development. Models can accelerate each of those activities once the problem has been framed and the expected deliverable is clear.

AI handles less of the work that happens before and after a document is produced.

Consultants still have to determine which questions are important, interpret organizational dynamics, and persuade leaders to act. As AI lowers the time and cost of analysis, the value of consulting professionals shifts away from output volume and toward problem selection, stakeholder judgment, and implementation.

What these results mean for professional roles

The same pattern emerges across every profession. AI performs best on work that is structured, digital, and easy to evaluate. Its performance is less reliable when tasks depend on context, professional judgment, and accountability.

The immediate impact of AI is therefore less about eliminating entire professions and more about changing how work is distributed within them.

What does broader research show about AI’s impact on jobs?

Mercor’s benchmarks show which professional tasks AI can perform, but capability does not translate automatically into job losses. Employers still have to redesign workflows, adopt new systems, and decide whether productivity gains justify changing staffing levels.

Research on the broader labor market suggests that the process of transition is underway, but its effects remain modest, uneven, and difficult to separate from other economic forces.

The International Labour Organization (ILO) and World Bank released a joint study for the World Development Report 2026, which covered labor markets across 135 countries and roughly two-thirds of global employment. It found that exposure to generative AI is highly uneven across economies, with the greatest potential for rapid job losses in occupations that are already digital and automation-prone.

A separate ILO brief from April 2026 highlighted that exposure indicators are early warning signals, not a direct measure of job loss, and should be read alongside actual data on employment, wages, and job transitions.

Anthropic's labor market research, also published in 2026, found limited evidence that AI has affected employment to date. Unemployment rates in the most exposed occupations haven't moved, though there's tentative evidence that hiring has slowed slightly for workers aged 22 to 25 in those fields.

You can see how AI agents handle real professional work for a closer look at the task-level capability behind these numbers.

Where AI is adding the most value

AI adds the most immediate value when work is routine, fully digital, high volume, and governed by clear rules. Across these tasks, the technology can reduce manual effort, accelerate first-pass output, and give employees more time for higher-judgment work:

  • Customer service: AI can resolve structured inquiries, allowing employees to focus on escalations and unusual cases.
  • Data entry and administration: Models can process observable inputs when the rules and expected outputs are clearly documented.
  • Document production and review: AI can generate or assess first drafts when teams can check quality against a defined rubric.
  • Entry-level analysis: Models can accelerate narrow research and reporting tasks that are easy to specify and review.

People who are early in their careers face a different challenge than outright replacement. AI can absorb many of the training tasks traditionally undertaken by entry-level roles. This includes the research, drafting, reconciliation, and first-pass analysis that once helped new employees build judgment through repetition.

Companies that rely on AI to complete these tasks may gain efficiency now but risk weakening the pipeline that produces experienced professionals later.

Where AI is creating new work

As AI reduces the volume of routine tasks, time and effort can be redirected to expert judgment that directs, corrects, and evaluates these systems. This expertise can provide judgment that AI systems can't generate themselves.

LinkedIn data reported by the World Economic Forum in January 2026 attributes roughly 1.3 million new positions to AI, driven by AI engineers, forward-deployed engineers, and evaluation specialists. These are new roles emerging because of AI, not existing roles that are being renamed.

What can enterprises do to prepare teams for AI in the workplace?

The greatest risk may not come from adopting AI too slowly but from moving quickly without understanding what the technology can actually improve or impact. Companies that begin with headcount cuts can lose the expertise, training pipelines, and accountability structures they will need later.

A more effective approach is to begin by examining the work itself, then redesigning teams around the tasks AI can support:

  1. Audit tasks, not roles: Exposure exists at the task level, so role-level headcount cuts rely on assumptions the data can't support.
  2. Identify where AI increases productivity: For tasks where someone must answer for the outcome, use AI to augment their work. Where accountability is lower, redesign the workflow before making headcount decisions.
  3. Source domain expertise: Evaluating AI in a specialized field requires practitioners from that field. Some teams build that bench internally, while others partner with expert platforms like Mercor, where AI trainer roles represent one of several judgment-based career paths now emerging.
  4. Redesign workflows and prepare employees for the change: AI adoption requires more than giving workers access to new tools. Enterprises must retrain employees, redefine responsibilities, establish review processes, and help teams adjust as routine tasks shift and new judgment-based work emerges.

No role is insulated from AI, but companies still control whether the transition erodes their workforce or strengthens it.

See how Mercor can help your team turn AI's impact into an opportunity

Mercor partners with enterprises that are building and adopting AI, connecting them with domain experts whose knowledge and judgment help make these systems more accurate and reliable.

Find out what it looks like for your team:

Get in touch

Frequently Asked Questions

Will AI create more jobs than it destroys?+

Most institutional projections say yes, though the estimates are contested. The World Economic Forum's Future of Jobs Report 2025 projects a net gain of 78 million jobs globally by 2030, with 170 million created against 92 million displaced. The caveat is that displaced and created positions rarely involve the same people in the same places, which makes any net figure a poor guide to individual risk.

Is AI actually causing job losses yet?+

Economy-wide, the evidence is thin. The Budget Lab at Yale finds no discernible disruption to the US occupational mix. Rather, roles are moving toward work that requires judgment and accountability, exactly where the benchmark data shows models still falling short.

Is AI different from past waves of automation?+

Two things are genuinely new. Generative AI reaches cognitive and creative work earlier than past waves of automation, and it impacts tasks rather than whole occupations, reshaping more roles than it removes.

Economists disagree about how much that matters. Some read current data as an ordinary diffusion curve that will look unremarkable in hindsight. Others argue the breadth of task exposure makes the historical comparison misleading. Both camps work from the same 3 years of evidence.