AI News

AI Is Acing Accounting Tests: What the Latest Evidence Shows

Published October 2, 2026. Research and leaderboard checks current to this date.

AI’s accounting abilities have improved fast enough that firms should reassess which tasks people prepare from scratch and which they supervise.

On October 1, Mercor researcher Aden Barton reported that frontier models had overtaken licensed accountants on four defined work scenarios. Mercor’s study summary reports that Claude Opus 5 earned full marks on all 20 attempts, each in under ten minutes; unassisted accountants averaged roughly 37% of grading criteria.

That is a consequential result. It makes a credible case for automating certain document-heavy accounting tasks. It leaves much more uncertain the questions of dependable month-end closing, audit judgment, and employment.

Those distinctions matter because accounting errors can survive a convincing explanation. A model can produce a polished reconciliation while carrying the wrong balance into the ledger. Businesses need to know both how rapidly the technology is improving and where its performance still breaks down.

What the October 1 study measured

The full research paper describes 12 practicing US CPAs, averaging 5.4 years of accounting experience. Seven held senior roles and three were managers or above, a useful qualification to the paper’s “junior accountants” label. Participants completed adapted month-end tasks involving nonprofit budget allocations, hotel revenue and commissions, event-revenue analysis, and lease schedules.

They had three hours per task. The comparison pooled 23 unassisted human attempts against 20 AI attempts. Models used tools; this was an agent evaluation, rather than a chatbot answering isolated questions. Rubrics were written by accountants. The AI grader agreed with expert grading on 98% of criterion decisions in eight checked submissions. Participants lacked coworkers and accumulated company context, and five of twelve rated the setting unrealistic.

Read the result at that scale: four tasks, a small professional sample, and a specific testing environment. Repeated success on the same scenarios provides evidence about consistency within those scenarios. It provides less information about unfamiliar clients, new industries, missing documents, or a difficult conversation with management.

The graders’ agreement is reassuring, but eight checked submissions cannot establish that an automated judge will handle every accounting disagreement correctly. Independent replication with different tasks, reviewers, and firms would make the conclusion substantially stronger.

There is also a commercial context. Mercor markets expert data for AI training, while its accounting-benchmark partner Ramp sells accounting software. That does not invalidate their measurements. It does make transparent methods and outside replication especially valuable.

The rate of improvement deserves attention

The historical evidence shows how quickly yesterday’s assessment can become outdated.

In a large 2023 accounting-education study led by David A. Wood, researchers examined 28,085 questions from 186 institutions in 14 countries. On assessments weighted by question points, students averaged 76.7%; ChatGPT averaged 47.5% without partial credit and 56.5% with it. Early general-purpose AI was plainly struggling.

A subsequent peer-reviewed certification study published in June 2024 reported a 53.1% average for GPT-3.5 across its tested material. GPT-4, supplied with ten examples and reasoning-and-tool support, reached 85.1% across multiple-choice material for CPA, CMA, CIA, and Enrolled Agent certifications. The improvement reflected model capability, prompting, and access to tools together.

That was already a major advance. Still, these were research evaluations of exam material. The AICPA explains that the actual CPA Exam’s passing score of 75 is a scaled score, not 75% correct; its sections combine multiple-choice questions and task-based simulations. Raw research accuracy cannot be converted directly into an official exam result or a professional license.

The latest work moves the evidence closer to doing the job. Mercor’s retrospective model comparison places the crossover above the human average during 2025, followed by near-perfect frontier results in 2026. Its stated span from below-average performance to acing these tasks is about eighteen months.

These results should not be stitched into one percentage-growth curve. The older studies used different questions, scoring, and tools. Within the recent study, historical models were tested retrospectively against one contemporary human baseline. The graph measures performance by model release date, not changing accountant productivity over time.

Even with those qualifications, a procurement decision based on a 2023 chatbot trial may now be badly out of date. Firms need periodic evaluations of the actual tasks they care about. They also need to specify which model, tools, instructions, and review process produced a result. An AI product name alone tells them too little.

The harder benchmark exposes the remaining gap

Mercor and Ramp introduced APEX-Accounting in July 2026: 160 tasks across ten simulated companies, authored by 42 accounting experts with a median eleven years of experience. The tasks cover reconciliation, data entry, variance analysis, and schedules and accruals.

In the original report, the leading mean-criteria score was 56.4%. The best result for solving a task at least once in eight attempts was 21.5%. For solving it correctly in all eight attempts, the best result was 2.6%. These are different metrics, achieved by different model configurations.

The live mean-score table checked October 2 now lists Opus 5.5 Max at 61.8%, with a reported 95% interval of ±4.0 percentage points. Its Pass@1 view lists that configuration at 14.8%, with ±4.9 points. The current table has advanced beyond the original report; older figures elsewhere on the page should be read as the launch results.

Different accounting metrics answer different business questions.
MeasureResultMeaning
Four adapted tasks, October human-baseline studyOpus 5: 100% across 20 attemptsPerfect rubric performance on a narrow task set.
Full APEX, current mean criteriaOpus 5.5 Max: 61.8%Average share of requirements satisfied; partial credit counts.
Full APEX, current Pass@1Opus 5.5 Max: 14.8%Share of tasks completed perfectly in a single attempt.
Full APEX, original best all-eight success2.6%Share of tasks solved perfectly on every one of eight runs.

Sources: the October study summary, July technical report, and current official leaderboard linked above. The task sets, models, and evaluation settings differ. These rows are not a controlled comparison of one model across time.

A 61.8% mean score does not mean a firm can safely automate 61.8% of its close. Requirements have different financial consequences. Missing a formatting detail and misstating revenue may each lose a rubric point, while creating very different risks for a business.

The benchmark’s launch analysis attributes roughly seven in ten examined failures among its top three models to reasoning problems. A model could find a discrepancy and then fail to preserve the finding in its final entry. It also explicitly excludes tax, audit, consolidation, and external reporting from the evaluation’s scope.

That is why both sets of findings can be true. A bounded task with a clear deliverable can be ready for substantial automation even while a longer workflow remains unreliable. Progress is uneven across the work inside one profession.

What real workplace evidence adds

A particularly useful complement is Jung Ho Choi and Chloe L. Xie’s Human + AI in Accounting, first published online in April and included in the June 2026 Journal of Accounting Research. It combines a survey of 277 accountants with platform data covering more than 200,000 transactions at 79 small and medium-sized enterprises.

AI adoption was associated with greater productivity, more detailed ledgers, faster closing, and a shift from data entry toward communication and quality assurance. Accountants intervened selectively when AI confidence was low. A separate framed experiment found that AI improved classification accuracy on average, while reliance on conflicting AI recommendations could increase errors.

The operational associations do not establish a universal causal effect. Adopters may differ from non-adopters, and one platform’s users cannot represent every accounting firm. The experiment supports a narrower conclusion about transaction classification. Together, the findings provide evidence for assisted production and targeted professional review.

This is an important distinction for adoption. A benchmark asks whether the model can produce an answer under specified conditions. A firm asks whether the complete process produces better records, with less total effort, across successive reporting periods.

The October experiment’s assisted arm was slower and slightly less accurate than AI alone. With AI already scoring perfectly, the rubric could not measure further quality gains from review. Human involvement therefore needs a defined purpose: checking evidence, resolving uncertainty, or taking responsibility for a consequential decision. Merely adding a person to the process does not guarantee an improvement.

Intuit’s 2026 AI Impact Report provides broader context. It draws on over 34,000 survey responses and anonymized payment records from more than 5.3 million businesses. In its US survey, 77% of small and midsize businesses reported regular AI use. Across observed businesses, about one in ten paid for dedicated AI tools during 2021–2025.

Those figures measure different populations and behaviors. Regular use can include free tools or embedded features; dedicated subscription payments capture a narrower activity. Neither figure tells us how many companies have handed accounting decisions to an autonomous agent. Adoption statistics need the same care as benchmark scores.

Recent products are connecting AI to the records

The practical change in 2026 is increasingly about connecting models to business systems, policies, and recurring work.

Ramp launched Stack on June 3 as a platform for accounting firms. Its current support documentation describes bookkeeping review, bank reconciliation, and monthly-close workflows, with present support focused on QuickBooks Online businesses with straightforward bookkeeping processes. That scope matters when assessing a complex group, unusual revenue arrangements, or a business operating across jurisdictions.

On July 28, Intuit announced expanded QuickBooks integrations in Claude and ChatGPT, including invoice actions and payroll queries. Its August 27 Intuit Intelligence documentation lists report creation, transaction searches, recurring automations, and accounting-policy templates in specified plans. Availability depends on the product, plan, region, and feature.

These are vendor descriptions of capabilities. They demonstrate where products are heading and which workflows are being packaged. They do not independently establish error rates or prove that a deployment will save a particular firm money. The APEX harness experiment also did not evaluate the complete production Stack product.

Connected software can remove tedious file transfers and give the AI more relevant context. It can also turn an incorrect suggestion into a ledger change. The business value depends on both the quality of the answer and the controls around acting on it.

Where firms can put the gains to work

The strongest starting point is a recurring task with an identifiable source of truth and a result someone can check. The following is a practical assessment of candidate workflows, not a claim that Kingy.ai has tested these products.

WorkflowUseful AI contributionCritical review
Bank and ledger reconciliationMatch records and prepare an exception list with source references.Investigate timing differences, duplicates, and unexplained balances.
Accruals and supporting schedulesExtract terms and draft calculations linked to documents.Check periods, assumptions, formulas, and agreement with the ledger.
Budget and revenue variancesCalculate differences and identify transactions behind changes.Verify the drivers before turning an association into an explanation.
Expense and invoice codingPropose categories using approved accounting policies.Review unusual items, tax treatment, and departures from policy.
Audit document reviewLocate contract clauses or summarize minutes for further examination.Check original documents, completeness, and relevant audit assertions.
Client reportingDraft explanations from approved financial statements.Check every number and distinguish recorded facts from forecasts.

Consider a hypothetical reconciliation: the bank statement shows $99,500 and the cash ledger shows $100,000. Calculating the $500 difference is easy. Establishing whether it reflects a deposit in transit, a duplicated receipt, or another issue requires evidence.

An effective AI workflow would identify the relevant transactions, attach supporting records, and leave unresolved explanations visible. An ineffective one would create an adjustment merely to make the totals agree. Both could produce an attractive spreadsheet. Only the first gives the reviewer a defensible basis for a decision.

That example also shows why traceability matters. A reviewer should be able to move from the output to the input records and reproduce the calculation. Otherwise, the person may have to redo the whole task, consuming much of the supposed saving.

The economics extend beyond cheap model output

Mercor reports $0.21 per satisfied rubric criterion for Claude, against $10.35 for unassisted accountants: roughly a 49-fold difference. Its calculation uses token prices and a median US accountant wage. It is a benchmark cost comparison, not a measured reduction in a firm’s operating costs.

A business has to pay for data preparation, software, integration, review, corrections, and maintaining the process. The right unit is the cost of an accepted result. Measuring only model charges omits precisely the work that becomes expensive when the output is wrong.

The distinction has commercial consequences. Faster preparation can create capacity for more clients, earlier management reporting, or deeper review. Turning that capacity into revenue requires demand. Turning it into lower prices requires a pricing decision. Turning it into fewer staff requires that the work and responsibilities can actually be redistributed.

Firms billing by the hour face a particular tension: shorter preparation time may reduce billable hours even as productivity improves. Fixed-fee services may capture more of the saving, provided review and exception handling do not expand. These are plausible business effects, not outcomes established by the four-task study.

There is a quality opportunity as well. If searching records becomes cheaper, firms may investigate more anomalies or update analysis more frequently. That benefit should be measured directly through fewer unresolved differences, better supporting documentation, or earlier detection of mistakes.

Regulators are addressing the transition

On March 30, 2026, the UK Financial Reporting Council published guidance for generative and agentic AI in audit. Its examples include summarizing board minutes and reviewing contracts for revenue-recognition testing. The FRC explicitly states that firms and responsible individuals remain accountable for audit quality. The guidance supports adoption while retaining professional judgment about confidence in outputs.

In the US, the PCAOB’s April 2026 report on audit-committee conversations describes questions about AI’s effects on financial reporting and associated controls. Committee chairs also discussed how audit firms use technology and ensure that controls operate consistently.

These developments concern audit, which has its own obligations and differs from routine bookkeeping. They show why professional accountability deserves as much attention as the ability to generate an answer.

For a firm designing an accounting workflow, a useful implementation would record the source files, model version, instructions, calculations, proposed changes, and reviewer decisions. Access should match the task. Posting entries and approving them should remain distinct actions when the firm’s controls require separation.

Confidence scores alone cannot substitute for evidence. A confident explanation may still depend on a missing document or an incorrect assumption. Review needs to test the work against the records and policy, with a clear route for questions the system cannot resolve.

The implications for jobs and training

Broad claims about the disappearance of accountants outrun the available evidence. So do assurances that automation will leave every role intact.

The latest US Bureau of Labor Statistics outlook projects 5% employment growth for accountants and auditors from 2025 to 2035. For bookkeeping, accounting, and auditing clerks, it projects a 6% decline. The latter outlook explicitly discusses software automation reducing demand. These are occupational projections, not a causal forecast of the October study’s impact.

The difference is useful. Accounting includes work that can be compressed through automation and work whose value depends on investigation, communication, judgment, and responsibility. One firm might use AI to expand its client base; another might reduce preparation roles. The same technical advance can produce different staffing decisions.

The AICPA’s June 2026 CPA Firm Top Issues Survey identified technology and AI change management as the leading anticipated issue over the next five years across firm sizes. Implementation is already a management concern.

Training deserves particular attention. Junior accountants often develop judgment through preparing work, receiving corrections, and learning which discrepancies matter. If AI produces most of the first draft, firms need another way to provide that practice.

A practical response would give trainees structured responsibility for reviewing source evidence, explaining exceptions, rebuilding selected calculations, and defending a conclusion before seeing the automated answer. Merely asking them to approve generated work risks producing familiarity with software without equivalent understanding of accounting.

This creates an opportunity for education as well. Assessments can require students to explain assumptions, document corrections, and identify unsupported conclusions in AI outputs. The objective should be to develop an accountant who can investigate a difficult case and assess an automated result.

What would justify more autonomy

The next persuasive advance will be measured across complete processes. A firm evaluating accounting AI should look for repeated success on unfamiliar records, timely escalation when information is missing, and a lower total burden of preparation and review.

A proportionate pilot could run a bounded task alongside the existing process for several closes. Compare accepted-output accuracy, material errors, reviewer minutes, corrections, and the time taken to finish. Include awkward records and exceptions. Keep proposed changes reviewable while the evidence accumulates.

A model upgrade should earn its place through the same task evaluation. Improved general benchmarks do not establish that it handles a firm’s particular lease schedule, revenue policy, or file structure better. Cost and speed matter only after the required quality is met.

AI has already made enough progress to warrant serious use in accounting. Firms can act on that opportunity now by choosing a defined workflow, measuring the complete result, and expanding the system’s responsibility when repeated evidence supports it.

Research note: Kingy.ai reviewed the original Mercor study and technical report, checked both live APEX scoring views, and consulted peer-reviewed research, regulator publications, official labor statistics, and vendor documentation. We did not independently rerun the benchmarks or conduct hands-on product trials for this article. Historical results retain their original dates and scope. The practical workflow assessments and business implications are editorial analysis.