THEAccounting EducatorEvidence, ideas and practice for accounting teachers
Two matching smooth ivory tiles, one lying flat and the other tilted to expose its interlocking coloured layers.
AI & Technology

AI can complete the accounting task. What should students learn from doing it?

A small benchmark study found strong AI performance on four selected accounting tasks, but also reported correct figures disappearing before final submission in several human-AI sessions. For lecturers, it raises a useful question: are students learning to produce an answer, justify reliance on it, or investigate the problem behind it?

By The Accounting Educator · Published · Updated · 7 minute read

Save this article to return to when it is useful in your teaching.

A student submits a polished month-end analysis. The figures reconcile, the explanation sounds convincing and AI helped produce it. What does the submission tell you about the student’s accounting competence?

Perhaps they understood the problem and checked every important figure. Perhaps they accepted the output without examining its basis. The finished answer alone may not distinguish those possibilities.

A Mercor Economics report by Aden Barton, Human Baselines for Benchmarks: AI Now Outperforms Junior Accountants, makes that teaching question more concrete. Its most interesting educational detail is not simply that AI scored highly. It is that, in several assisted sessions, a correct figure appeared during the interaction but disappeared before submission.

A strong result on a narrow test

The study compared 12 practising US accountants, all reporting active CPA licences, with AI on four simulated month-end-close tasks. These involved nonprofit budget analysis, hotel revenue reconciliation, event revenue analysis and lease schedules. Each participant was assigned two tasks with Claude assistance and two without, with task order and assistance allocation randomised. The report includes 24 assisted and 23 unaided attempts and does not explain the missing unaided attempt.

Despite the report’s title, this was not a sample of fresh graduates. Participants averaged 5.4 years of accounting experience; seven held senior accountant or senior auditor roles and three were managers or above.

Claude Opus 5 received a 100% score on the study rubric in the 20 standalone AI attempts. Unaided accountants averaged approximately 37% of rubric requirements met. Those percentages describe performance against task-specific marking criteria, not the proportion of accounting knowledge possessed. Missing one input could affect several criteria and sharply reduce a score. The report itself notes that mistakes could cascade across rubric items.

Accountants using Claude scored substantially better than those working unaided, but Claude alone was slightly more accurate, while accountants working with Claude took about 15 times as long per task as Claude alone. The assisted condition used Claude Cowork, while standalone AI runs used a tool-using API agent. The report does not establish that these operating conditions were identical, so the time comparison should not be read as a clean estimate of the cost of human review. That comparison also has an important ceiling: once the standalone model scored 100%, the marking scheme could not register any additional quality from human review.

The tasks also had an unusual history. They originated in a benchmark deliberately developed to make AI models fail. The authors say that this process created a high density of realistic but hard-to-spot accounting conditions and edge cases. They later acknowledge that this density pushed the tasks beyond what they would expect even strong accountants to catch in ordinary work.

The report is company-authored, uses only four selected tasks and does not establish general accounting reliability or professional replaceability. The tasks concentrated on dense file searches, detail and close instruction-following. They excluded colleagues, client clarification and knowledge accumulated within an organisation.

This is accounting-practice evidence, not accounting-education research. No student learning or teaching method was evaluated, so the classroom ideas that follow are proposals to investigate, not demonstrated learning benefits.

This commentary is based on the published report; the underlying task files, complete rubrics and interaction logs were not available for examination. Peer-review status is not established. Grading relied on an AI judge, with only eight submissions manually checked by task authors. The report states that these checks showed 98% agreement with the automated grading. No accounting-education literature search was conducted for this article.

Review is a decision, not just an extra step

Four AI-assisted sessions fell short of perfect scores. In three, Claude had produced a correct figure at some point, but it was absent from the final submission. Twice, Claude revised its answer incorrectly and the participant accepted the revision. In the other session, the participant overrode correct AI advice.

These observations do not identify why the decisions went wrong. They do, however, complicate a simple instruction to “check the AI”. Both accepting and rejecting advice contributed to incorrect final selections. The educational target, then, is not blanket trust or blanket scepticism, but deciding when the evidence supports relying on an answer.

For teaching, the useful question becomes:

What evidence warrants accepting, rejecting or investigating this answer?

That question also helps separate the purposes of an accounting task. Preparing an analysis can provide procedural practice. Completing it without assistance can provide evidence of independent competence. Reviewing an AI-generated analysis can reveal how a student evaluates evidence. These purposes overlap, but they are not interchangeable.

AI’s success at completing a task does not, by itself, settle whether performing it remains educationally useful. Equally, retaining a manual exercise does not explain what it is intended to develop. A lecturer could make that purpose explicit: is this activity about calculating an average, understanding what belongs in its numerator and denominator, or defending a reported variance?

Make the basis for reliance visible

Consider a lecturer-designed illustration inspired by the benchmark’s event-analysis task:

A venue’s December event register lists 12 completed events. Its revenue ledger records $24,000 in venue fees for those same events. The budget specifies 10 events and $22,000 in venue fees. For this exercise, average revenue per event means venue fees divided by event count; catering and other charges are excluded. Assume the supplied records are complete and contain no cancellations or duplicate events.

Actual average revenue is $2,000 per event, compared with a budgeted $2,200. Total venue-fee revenue is $2,000 above budget, while average revenue per event is $200 below budget.

A proposed review exercise could present two prepared AI-style analyses. One uses the correct actual event count of 12. The other divides actual revenue by the budgeted count of 10, reporting an actual average of $2,400. These would be designed teaching materials, not reconstructions of the study’s interactions.

Rather than asking students only to select the correct answer, ask them to identify the supporting records, show which figures belong in the calculation and explain why they accept one analysis and reject the other. Include correct advice as well as incorrect advice, so that “find the AI error” is not always the task.

The distinction matters because a student might select $2,000 for the wrong reason, or obtain it without checking that the revenue and event count cover the same population. A brief review note can expose reasoning that a final figure leaves hidden.

Numerical correctness and evidential justification could therefore be marked separately. For students who have only recently learned the calculation, keep the source set small and demonstrate a source-based check before asking them to review independently. A dense benchmark file set is not a ready-made introductory activity.

Let students investigate what the files cannot settle

The benchmark also raises a question about what makes an accounting task resemble work. Seven participants rated the tasks realistic and five did not. Their comments often distinguished realistic spreadsheet activities from an unrealistic setting. Participants said that, in real work, they would often have months or years of accumulated organisational context, access to supervisors or clients for clarification and experience within a particular specialisation.

A classroom case can contain plausible company records while still excluding the activity of deciding what information to request.

Once students can handle the bounded event example, a later version could introduce a discrepancy between the event register and a management summary. Instead of immediately demanding a definitive average, ask what needs clarification, whom they would ask and how the response would affect the calculation. The lecturer could then release additional information and require a revised conclusion.

That would assess something different from the original benchmark: investigation before calculation, followed by a justified response to new evidence. It should not be treated as inherently superior for every learner or lesson.

A modest first step is to adapt one existing task rather than redesign an entire module. Add a short requirement asking students to document why they relied on an answer, then use a later independent task to examine whether they can justify the calculation with changed facts.

The benchmark does not tell lecturers how much manual practice students need or which review exercise works best. It gives them a sharper design question:

What should this task reveal about the student that an AI-generated answer cannot reveal on its own?

Sources and further reading

Reader account

Continue with email

A free account keeps your saved articles in one place and available on any device. It can make it easier to revisit ideas, evidence and practical examples as your courses develop.

The link signs you in, or creates a free account if you're new. It does not subscribe you to the newsletter.