
AI can draft MCQs, but experts still decide what reaches the exam
A medical education study found that selected AI drafts performed similarly to licensing-exam questions on several limited measures, but only after expert rejection and revision. For accounting lecturers, the useful lesson lies in the review process, not autonomous item generation.
Save this article to return to when it is useful in your teaching.
A lecturer asks a language model for ten multiple-choice questions on a course topic. Seconds later, the screen contains polished question wording, plausible-looking incorrect options and a neatly identified correct answer for each item.
The speed is impressive. The harder question is what happens next.
Is the identified correct answer defensible? Does the item test the intended learning objective? Has the model silently assumed the wrong jurisdiction, reporting date or accounting treatment? Are the incorrect options plausible because they represent genuine student errors, or merely because they sound technical?
A study in medical education offers a useful way to think about these questions. Its most relevant contribution is not evidence that AI can independently write exam-ready items. It is a worked example of AI drafting followed by human refinement, rejection and approval.
That distinction matters. The questions that reached students were the survivors of a structured process, not raw model output.
What students actually encountered
The study took place in a graded family medicine examination at Saarland University in Germany. All 119 fifth-year students in the cohort participated. They answered 30 questions originating from ChatGPT-4o or Gemini 1.5 Pro and 30 questions selected from the German National Licensing Examination.
The researchers gave the language models the course curriculum, learning objectives and learning materials. Initial drafts were reviewed, followed by further prompting intended to elicit decision-making and clinical reasoning. An expert panel then considered 43 of the 44 generated questions. It accepted 36 as potentially suitable, and 30 were eventually used.
Retained questions could be adapted to the German primary-care and guideline context. Reviewers also clarified clinical details and adjusted wording or incorrect options, known as distractors. Items requiring substantial content or structural repair were excluded rather than extensively rewritten.
In the resulting examination, the researchers detected no statistically significant overall difference between the selected AI and licensing-exam questions in how often students answered them correctly, the measure used for item difficulty. They also detected no significant difference in whether students thought the questions matched the curriculum. A later exploratory analysis found no significant overall difference in how responses were distributed among the distractors.
Because linked questions were grouped together, the source-recognition analysis used 39 question groups rather than treating all 60 questions separately. Students identified the source above chance, with mean correct recognition of 63.9% across those groups. Recognition accuracy did not, however, differ significantly between the AI and licensing-exam questions. This is subtler than saying students could not recognise AI. They showed some ability to identify question sources, but were not significantly more successful with one source than the other. Previous exposure to licensing-exam questions through commercial learning platforms may have influenced recognition and was not fully controlled.
A separate exploratory comparison found a difficulty difference among ChatGPT, Gemini and licensing-exam items, including between Gemini and the licensing-exam questions. The authors advise caution because each source contributed relatively few question groups and one unusually difficult Gemini item influenced the result. The study does not establish that one model produces better questions than another.
Nor do the overall findings demonstrate equivalent assessment quality. Failing to detect a statistically significant difference is not the same as showing equivalence or using a non-inferiority design to test whether one approach is no more than an acceptable amount worse than another. The measured indicators also leave important questions unanswered, including score consistency, fairness, content coverage and the validity of decisions based on students’ scores.
The workflow matters more than the label
For an accounting lecturer, the most transferable feature is the allocation of responsibility within this workflow.
That process could be explored in accounting, although the study itself contains no accounting students, content or assessments. Accounting MCQs bring their own risks. A question may depend on a particular reporting framework, jurisdiction or effective date. A numerical item may omit an assumption needed to determine the answer. A distractor may be factually indefensible rather than a plausible reflection of student reasoning. Professional judgement may be compressed into a question that falsely suggests only one conclusion is possible.
A small accounting pilot could therefore begin with formative practice questions rather than a high-stakes examination. The language model might receive a defined learning objective and a limited set of current, verified course materials. Every draft would then pass through a documented review before students saw it.
The review could address five connected questions:
- Is the accounting technically sound? Check the economic event, assumptions, calculations, terminology and identified correct answer. Where the topic depends on authoritative requirements, verify the item against the applicable current source rather than relying on the model’s explanation.
- Is the context properly bounded? State the relevant jurisdiction, reporting framework and date where these affect the answer. Remove ambiguity unless dealing with ambiguity is itself the learning objective.
- Does the item assess the intended learning? A polished calculation may test recall or arithmetic when the objective concerns analysis or judgement. Conversely, adding a long scenario does not automatically create a higher-level question.
- Do the distractors do useful work? For a single-correct-answer item, the identified correct response should be uniquely defensible under the stated facts, assumptions and applicable requirements. For a single-best-answer judgement item, reviewers should be able to explain why the identified answer is better supported than each alternative. Incorrect options should still represent credible misconceptions or errors, not rely on trick wording or missing facts.
- What happened when students attempted it? Examine the proportion selecting the correct answer, the choices made by students who answered incorrectly and any signs of misunderstanding. This shows how that question performed in that cohort, not permanent proof of quality.
Keeping a simple record of the model version, source materials, prompts, rejected drafts, revisions and final approval would make the pilot easier to evaluate. Rejection is useful information. If many drafts require substantial repair, the model may be generating more checking work than it removes.
That last point remains unresolved by the medical study. It did not measure prompting time, review time, costs or workload savings. AI-assisted drafting might redistribute effort from initial composition to checking and revision, but it cannot yet be assumed to save accounting lecturers time.
Perceived alignment was associated with difficulty
The study also provides a caution about asking students whether an assessment feels aligned with the curriculum. Easier questions were more often perceived as aligned, while more difficult questions were more often judged not to be aligned.
Student views remain valuable, particularly when an item seems disconnected from teaching or uses unfamiliar language. But perceived alignment is not the same as an independent mapping of the assessment to course objectives. A demanding question may be closely aligned yet expose incomplete learning. An easy question may feel familiar while sampling only a narrow part of the curriculum.
For an accounting pilot, student perceptions could therefore sit alongside a lecturer’s curriculum map and evidence of how each question performed. If students regard an item as misaligned, the useful next question is why. The problem might be omitted content, unclear wording, unexpected difficulty or an assumption that was obvious to the writer but not to the class.
The medical study’s boundaries matter when deciding whether to test this workflow in accounting. It covered one course, one institution, one cohort and a small, selected set of questions. Its licensing-exam comparators were chosen for relevance to the course, and the study was exploratory rather than an experiment testing different item-production methods. Several analyses, including the distractor comparisons, were conducted after the main analyses had been specified.
Even within those limits, the study helps reframe the practical question. The immediate choice is not between writing every MCQ unaided and handing assessment design to a machine. A more defensible possibility is to test whether AI can contribute drafts within a process that preserves subject expertise, documented review and local evaluation.
For accounting educators, that is a modest proposition worth investigating. The model may supply the first wording of a question. Responsibility for what the question means, what it measures and whether it belongs in an assessment remains with the lecturer.