Introducing the AI Productivity Index for Accounting










Introducing APEX-Accounting, a new benchmark built in collaboration with Ramp.
Accounting benchmarks often test whether a model can produce the right answer once. Closing the books demands more. An agent must reconcile conflicting files, apply company-specific context, carry conclusions across multiple steps, and produce the correct result consistently. A model that drops a correct intermediate answer can still create a bad journal entry. APEX-Accounting builds on the evaluation system Ramp developed while building Stack, their AI operating platform for accountants. Ramp and Mercor built APEX-Accounting to test whether AI agents can complete real accounting work to professional standards. The benchmark includes 160 tasks across 10 simulated companies. Experts with prior experience at firms like Deloitte, PwC, EY, and KPMG hired through Mercor's platform created the tasks and grading rubrics.
The headline result is not which model leads, but how rarely any model succeeds consistently. We ran every model on every task eight times because accounting work must be repeatedly correct, not correct once. Even the most consistent model solved just 2.6% of tasks correctly in all eight runs.
This is a very different result from performance on accounting exams. Models can produce the right answer in a controlled test and still fail to carry accounting judgment through a real workflow.
Leaderboard results
Claude Fable 5 tops the leaderboard at 56.4%, followed by Meta's Muse Spark 1.1 at 52.6%, and GPT-5.6 Sol at 51.5%. The best models manage to complete just over half of the work that a professional would.
For Pass@8, which measures whether a model is successful in at least one of eight attempts, Muse Spark 1.1 is highest at 21.5%, just ahead of Fable 5 (20.1%).
Models frequently earned partial credit. At least one model earned some credit on more than 95% of tasks. Full solutions were more difficult. 58% of tasks were never fully solved by any model on any run.

Performance vs. cost
We tested models at spending budgets of $1, $5, $10, and $50 per task. A larger budget allows models to use more tokens, which generally improves results. But the effect varies enormously. Fable 5 was extremely budget sensitive, scoring 11.8% with a $1 budget and improving to 55.2% with a $50 budget. Muse Spark 1.1 is the opposite. It's already strong on a tight budget and barely improves with more money. This is because above the $1 cap, models typically do not fully utilize the token limits. There is only a moderate increase in token usage when going from $10 to $50 and, on average, models use just 64.7% of the maximum budget available at $50.
Models differ in how they utilize the capped budget. At the $50 maximum spending budget, Fable 5 actually spends ~$32 per run while Muse Spark 1.1 spends ~$5, yet their scores are within 4 percentage points.

What the leaderboard measures
The canonical leaderboard runs every model in the Loop Harness, Mercor's standard model-and-tools architecture. The paper separately compares it with a purpose-built Ramp Harness, created with Ramp to mirror parts of a specialized accounting-agent architecture. The Ramp Harness improves Mean Criteria@3 by 1.2 percentage points on average.
This is a harness ablation, not an evaluation of the complete Ramp Stack product. It does not measure Stack's production integrations, accounting skills, memory, controls, auditability, or recent spreadsheet tooling.
Why do models fail?
Working with accounting and bookkeeping experts, we developed a taxonomy of the mistakes that prevented models from completing each task. These included failures in reasoning, information gathering, instruction following, and planning and reflection.
The top three models all fail in much the same way. Roughly seven in ten failures came from flawed reasoning, not an inability to find the right information. A model might correctly identify a discrepancy early in the workflow, then omit or contradict that finding in its final journal entry. Better retrieval alone will not solve this. Models need stronger accounting judgment and more discipline in carrying conclusions through an entire workflow.

Methodology
More than 40 accounting professionals authored and solved the benchmark tasks. They had a median of 11 years of experience, and more than half had worked at a Big Four accounting firm. Each task takes place in a self-contained world: a fictional company frozen at month-end close, with its own accounts, records, business history, accounting software, spreadsheets, PDFs, and other files.
The companies themselves are fictional but every transaction, discrepancy, and edge case inside was written by accountants who do this work for a living. The experts authored a rubric for each task with, on average, 13.7 criteria. An open-source AI judge performs the grading and achieves 97% agreement with expert graders.
APEX-Accounting focuses on month-end close and bookkeeping workflows. It does not evaluate tax, audit, consolidation, multi-entity or multi-currency accounting, external reporting, or how agents respond when they need clarification. Its findings should be understood within that scope.
The APEX-Accounting leaderboard comprises a heldout test set of n=160 tasks (associated with 10 worlds), kept hidden to resist contamination. We have released an open-source sample set with 1 world and 10 tasks on Hugging Face, and Archipelago, our internal framework for running agent evaluations, is available on GitHub.
Read the APEX-Accounting technical report to learn more about the methodology and results.
APEX benchmarks
APEX-Accounting joins Mercor's family of APEX benchmarks, which evaluate AI models' ability to do economically valuable work. Other APEX benchmarks include APEX-SWE for software engineering tasks across integration and observability, and APEX-Agents for professional services domains including corporate law, management consulting, and investment banking.
We thank Ramp and all the accountants on the Mercor marketplace who contributed their time to creating APEX-Accounting.
As APEX-Accounting is a closed benchmark, leaderboard evals can be run for any frontier model on request. Reach out here to find out more.
