
Separate AI-assisted performance from what students can do alone
A meta-analysis in programming education highlights a distinction that accounting assessment cannot afford to blur: the quality of work produced with AI access is not the same as competence demonstrated without it. A paired assessment can make both visible.
Save this article to return to when it is useful in your teaching.
A student submits a polished analysis of an accounting problem. The calculations reconcile, the explanation is fluent and the recommendations appear sensible. Generative AI was permitted, so the work may also demonstrate effective use of a contemporary tool.
But what, exactly, has been assessed?
The submission could show that the student can direct AI, evaluate its response and produce a useful final document. It does not automatically show that the student can independently identify the accounting issue, detect an incorrect assumption or apply the reasoning to a changed problem.
A meta-analysis in Educational Psychology Review helps clarify this assessment problem. It distinguishes performance achieved with generative AI from performance demonstrated after access to AI has been removed. That distinction is more useful for accounting educators than a general argument about whether AI is good or bad for learning.
The evidence comes entirely from programming education. The researchers synthesised 35 peer-reviewed studies published between 2022 and 2025, contributing 131 effect sizes. Most involved higher education, lasted no more than three months and used ChatGPT-based systems. No accounting education studies were included.
In the main statistical models, groups learning with generative AI had better average results across four categories: AI-assisted programming performance, independent programming performance, higher-order skills, and motivational or emotional outcomes. Comparison groups varied, however. They included students receiving no additional support, human support or support from non-generative AI tools.
The pooled estimate for assisted performance was Hedges' g of 0.36, with a 95% credible interval from 0.07 to 0.66. For independent performance measured without AI, it was 0.25, with an interval from 0.01 to 0.47. Hedges' g expresses the difference between groups in standard deviation units. These point estimates represent modest average advantages, but the analysis did not establish a credible difference between the outcome categories.
The independent-performance result was also fragile. Its interval came close to zero and included zero under some alternative assumptions about relationships among results from the same study. Few studies measured assisted and independent performance in the same students, and none used a delayed test. The synthesis therefore cannot tell us how much assisted performance remained when AI was removed or whether any advantage lasted.
Results also varied substantially across studies and measures. Adjustments for possible publication bias left only the higher-order-skills estimate clearly above zero, although other diagnostics did not provide clear evidence that publication bias drove the findings.
For accounting educators, the defensible lesson is not that AI improves accounting learning. It is that assisted and independent performance answer different questions and should be assessed accordingly.
Decide which performance the task is meant to reveal
An AI-permitted assignment can assess worthwhile capabilities. Students might need to formulate a useful request, supply relevant facts, interrogate an answer, check calculations, identify unsupported claims and revise the output for a professional audience. These are not trivial activities.
Yet a high-quality final product combines contributions from the student and the tool. If the assessment is also intended to support a claim about independent accounting competence, another source of evidence is needed.
This matters particularly in accounting, where conclusions may depend on ambiguous facts, alternative assumptions, technical requirements and professional judgement. A response can look convincing while resting on a hidden error. Students with limited prior knowledge may also be least able to recognise that error.
The programming meta-analysis did not identify a reliably better tool configuration, level of instructor involvement or teaching approach. That does not mean these choices are unimportant. The evidence within many of the relevant categories was sparse and overlapping, so the conditions that strengthen or weaken the effects remain uncertain.
A practical starting point is to ask four separate questions about what a task is intended to reveal:
- AI-assisted product: Can the student use permitted tools to produce and verify a useful accounting output?
- Independent explanation: Can the student explain the reasoning without the tool?
- Independent application: Can the student apply the underlying idea to a related problem?
- Later performance: Can the student still do this after time has passed?
One task need not answer all four questions. Problems arise when the mark from an assisted product is treated as though it answers them all.
Pair production with a changed problem
A lecturer could pilot a paired assessment in which students first complete a task with AI access and then undertake a short unaided follow-up. The second part might ask students to carry the reasoning into a different problem, an ability education researchers often call transfer.
This is an accounting application of the distinction identified in the programming research, not an intervention tested by the meta-analysis.
Consider a simple adjusting-entry task. Students are told that a business paid $12,000 on 1 December for 12 months of insurance coverage and initially debited Prepaid Insurance for the full amount. Its reporting date is 31 December. Assume that coverage begins on the payment date, the insurance benefit is consumed evenly over the 12 months and adjustments are calculated using whole months.
With AI permitted, students prepare the adjusting entry and a short explanation. The entry under these assumptions is:
- Debit Insurance Expense $1,000
- Credit Prepaid Insurance $1,000
The assisted phase could assess whether students verify the period covered, check the monthly calculation and ensure that the entry is consistent with the original debit to Prepaid Insurance. They might also annotate the AI output to show what they accepted, corrected or rejected.
A brief unaided follow-up could then change the facts. Suppose the business paid $18,000 on 1 October for 12 months of coverage, again recording the full payment in Prepaid Insurance. The reporting date remains 31 December and the other assumptions are unchanged. Students would need to recognise that three months have expired and record:
- Debit Insurance Expense $4,500
- Credit Prepaid Insurance $4,500
The follow-up should not merely test whether students remember the first answer. Ask them to explain why an expense is recognised, how they identified the relevant period and what additional information they would need if the original payment had been recorded differently. This samples reasoning as well as journal-entry mechanics.
The same pattern can be used without creating a second full assessment. After an AI-assisted case, students might complete one focused activity without AI:
- explain the most consequential judgement in a short oral discussion;
- diagnose a deliberately flawed AI response;
- solve a changed version of one part of the problem;
- identify missing information that prevents a conclusion;
- reconstruct a calculation from a blank page;
- justify why an apparently plausible alternative does not fit the facts given.
These follow-ups reveal different aspects of competence. An unaided calculation, an error diagnosis and an oral explanation should not be assumed to provide interchangeable evidence. This is especially relevant because the meta-analysis found considerable variation among independent-performance measures within the same programming studies.
Do not turn the follow-up into surveillance
The purpose of paired assessment should be clear to students. It need not be an attempt to catch prohibited AI use. Where AI is permitted, the first phase can recognise assisted professional performance openly. The second asks a separate question about what the student can explain or apply independently.
That distinction can also improve feedback. A strong assisted product followed by weak independent reasoning suggests a different teaching need from a weak product supported by sound unaided understanding. In the first case, the student may need more practice retrieving and applying the accounting reasoning. In the second, the difficulty may lie in directing the tool, checking its output or communicating the result.
Keep the unaided component proportionate. A five-minute explanation or one carefully chosen variant may reveal more than another lengthy submission. It should also match the course outcome. If the intended outcome concerns professional judgement, a rapid numerical exercise alone will provide limited evidence. If accurate procedural application matters, an oral explanation without any application may also be insufficient.
For higher-stakes decisions, more than one observation is preferable. A student may misunderstand a question, perform unusually poorly under time pressure or succeed through short-term rehearsal. Since the programming studies assessed independent performance at the end of their interventions rather than after a delay, a later low-stakes question, examination item or classroom problem can add evidence about whether learning has endured.
Treat the distinction as an assessment lens
The meta-analysis offers no formula for the ideal accounting assessment. It does not tell lecturers how much weighting to assign to assisted and unaided work, which AI system to permit or whether a particular follow-up format will improve accounting learning.
It does, however, expose a question that every AI-permitted assessment must answer: is the course evaluating the product students can produce with contemporary tools, the accounting competence they can demonstrate independently, or both?
Once that question is explicit, assessment choices become easier to defend. A lecturer can retain authentic AI-assisted work while adding a focused check on explanation, correction or application to a changed problem. The result is not an AI-proof assessment. It is a clearer account of what the available evidence can, and cannot, say about the student.