All posts
2 min readby Romiel Inolino

Frontier models now beat licensed CPAs on bounded month-end close tasks. The full close is still unsolved

AIaccountingfinance automationbenchmarkshuman in the loop

The interesting number in this accounting study is not the AI score. It is the gap between a bounded task and the whole job.

What happened

Mercor hired 12 licensed CPAs, averaging about five and a half years of experience, to complete simplified tasks from its APEX-Accounting benchmark. Each of four month-end close scenarios meant digging through a company's working files, finding the right numbers, doing the math and delivering a table of results.

Mercor reports that frontier models were faster and more accurate than every accountant in the study, including the best one. Eighteen months ago, the best models fell short of the accountants' average of about 37 percent. Today, Mercor says, models ace the same tasks. It also says frontier models were more than an order of magnitude cheaper per task criterion.

The caveats matter. Mercor says the tasks tested what AI is best at: detail work, searching files and following instructions closely. There were no coworkers to ask and no client communication. On the full benchmark of 160 tasks across 10 simulated companies, The Decoder reports Claude Opus 5.5 leads with 61.8 percent of grading criteria met, and Mercor says no model fully solved almost 60 percent of the tasks. The Decoder's read: AI cannot close the books without oversight yet.

My take

This is the clearest map I have seen of where finance automation should start and stop.

Bounded steps with clear inputs and a defined output, like pulling figures from working files, reconciling them and producing a table, are now strong candidates for a model. The open ended part, where records are incomplete or a judgment call needs context, still belongs to a person.

So the design is not "replace the accountant." It is:

  • The model prepares each close step and shows its working
  • A checklist validates totals against the source before anything posts
  • A CPA reviews exceptions and signs off

That shifts the human from typing numbers to reviewing them, which is where their experience is worth the most.

I design and build AI automations that take on real jobs like this: reconciliation prep, reporting, lead routing, data entry, content pipelines and agent workflows, always with a review step where it counts. See examples at romielwillautomate.dev.

More posts