Can AI Replace Accountants? Full Success Reaches Just 2.6% on 160 Tasks

APEX-Accounting tests systems, spreadsheets, and PDFs across realistic accounting work, revealing a large gap between average rubric scores and end-to-end success.

A benchmark built from accounting work

APEX-Accounting, submitted July 29, tests whether frontier models can complete realistic accounting tasks. Mercor and Ramp worked with accounting and bookkeeping experts to create 160 private tasks across ten worlds, including reconciliation, accruals, transaction posting, and reporting. Each world includes an accounting system, spreadsheets, PDFs, and other files.

Partial credit did not mean complete success

The authors report that Claude Fable 5 Max led nine models at 56.4% Mean Criteria@3, followed by Muse Spark 1.1 xHigh at 52.6%. Yet the best Pass^8—meeting every criterion across repeated runs—was only 2.6% for GPT-5.6 Sol Max+Pro. The highest Pass@8, succeeding completely at least once in eight runs, was 21.5% for Muse Spark 1.1 xHigh.

Increasing the budget from $1 to $50 raised aggregate scores, while within a fixed harness the tasks consuming more tokens tended to score lower, an instance the authors describe as Simpson's paradox.

Practical meaning and limitations

Accounting automation should track transaction-level checks, balance reconciliation, approvals, and recovery rather than one average score. AI can prepare evidence and drafts while qualified people retain final posting, close, and tax decisions.

This is a preprint and the private test set limits independent reproduction. Results depend on the authors' tool environment, and model names and performance can change quickly.

Primary source