THEAccounting EducatorEvidence, ideas and practice for accounting teachers
Three overlapping aperture frames reveal different bounded sections of the same complex geometric lattice.
Curriculum

Can audit students assure an algorithm?

Early bias audits use concepts that accounting students already encounter, but they vary sharply in scope, evidence, independence and reporting. That makes them promising material for exploring assurance judgement, provided lecturers treat the case as an emerging practice rather than a proven teaching intervention.

By The Accounting Educator · Published · Updated · 8 minute read
Listen to this articleNarrated edition

Save this article to return to when it is useful in your teaching.

An audit student may be able to define suitable criteria, sufficient appropriate evidence and practitioner independence. Give that student a report labelled “bias audit”, however, and those familiar ideas become harder to apply.

What exactly has been examined: a model, its use by one employer or pooled results from several employers? Which criteria support the conclusion? What does an impact ratio reveal, and what does it leave unanswered? Does an independence statement provide enough information to assess the provider’s relationships?

These questions sit at the centre of an Accounting Horizons Perspective article on algorithmic assurance. The author analysed 13 publicly available bias-audit reports produced from 2023 to 2025 under New York City’s Local Law 144. The reports concern automated tools used in employment decisions and were coded across dimensions including criteria, independence disclosures, governance review, monitoring and conclusion wording.

The analysis is descriptive and exploratory. There is no authoritative registry of completed audits, the total population is unknown and only reports with sufficient public documentation were included. The sample may therefore favour larger providers and organisations more willing to publish their work. The paper does not report whether another reviewer independently checked and consistently reproduced the classifications, or whether anyone outside the analysis checked them. The source is also not an accounting education study: it examines no students, courses, assessments or learning outcomes.

Within those boundaries, it presents an intriguing curriculum opportunity. The 13 reports repeatedly used ideas recognisable from assurance, yet shared terminology coexisted with substantial variation in engagement design and reporting.

Similar labels concealed different work

All 13 reports identified protected groups, calculated impact ratios and provided a public data summary. The reports divided each demographic group’s selection or scoring rate by the corresponding rate for the highest-rate group. Under the framework described in the source, a ratio below 0.80 indicated potential adverse impact. A result at or above 0.80 did not by itself establish an absence of discrimination, broader fairness or legal compliance.

The reported ratio can be affected by choices about aggregation, missing data, small demographic groups, benchmarks and the unit of analysis. Some reports examined an employer’s particular use of a tool, some pooled data across employers and others focused on the underlying system or model.

The reports also differed in the conclusions and contextual information they provided. Five of the 13 contained explicit opinion-style wording, four included an explicit independence statement, three discussed governance or risk, and two specified a future audit cycle or model-update plan. The presence of opinion-style wording did not establish the governing standard, engagement type or level of assurance. These figures describe this small set of disclosed reports, not the wider population of algorithmic audits.

Length varied from one to 22 pages, with a median of six. Word counts ranged from 435 to 26,560, with a median of 2,239. Longer reports sometimes offered more detail, but the author cautions that length alone did not guarantee rigour.

The differences become clearer in the article’s comparison of three reports. One BLDS report emphasised statistical testing and sensitivity analysis but did not issue an assurance opinion.

The Perspective describes a DCI Consulting report as using an employment-assessment validation approach drawn from industrial psychology. It disclosed limitations associated with combining data across different implementations, but did not identify the work as a recognised assurance engagement or state whether it provided reasonable or limited assurance.

A BABL AI report more closely resembled a conventional assurance engagement. It defined scope, considered governance and issued an explicit opinion that the system conformed to stated criteria.

These descriptions do not establish that one provider performed higher-quality work overall. They show that reports created under the same regulatory setting can represent substantially different kinds of activity, from statistical testing to assurance-style reporting.

For an audit class, that distinction is more useful than a simple debate about whether AI is good or bad. Students can examine what a report actually permits its users to conclude.

Build a report map before asking for a verdict

A lecturer could begin with one carefully selected public report and ask students to construct a report map:

  1. What system, deployment or process is being examined?
  2. Who are the intended users of the report?
  3. How is the scope defined, including the unit of analysis?
  4. Which criteria and thresholds are applied?
  5. What evidence and procedures are described?
  6. Which methodological choices could affect the reported result?
  7. What does the report say about independence, governance and limitations?
  8. What conclusion is expressed, and how strong is its wording?

The purpose is not to make students instant specialists in employment law or data science. It is to ask whether they can use assurance reasoning when the conventional cues of a financial audit case are removed.

Missing information can become part of the task. Students might prepare a short memo separating three categories:

  • facts and methods clearly stated in the report;
  • conclusions reasonably supported by that information;
  • additional evidence needed before placing reliance on the conclusion.

This structure discourages students from filling gaps with assumptions simply because a document uses the word “audit”. It also creates space to discuss the difference between an absent disclosure and evidence of poor practice. If a public report does not contain an independence statement, for example, students can identify the resulting uncertainty. They cannot infer from that omission alone that the provider lacked independence.

A comparison task could then ask students to evaluate two reports with different units of analysis or styles of conclusion. Which report gives an intended user a clearer account of the work performed? Where do the reports make different methodological choices? Do those choices prevent direct comparison? What further information would an applicant, employer or regulator need?

Any standards-based version of the task requires additional preparation. The article refers to ISAE 3000, the AICPA Trust Services Criteria and COSO’s Internal Control Integrated Framework. The relevant primary materials would need to be retrieved before students assessed whether an engagement followed ISAE 3000, whether controls were evaluated against selected AICPA Trust Services Criteria, or how a report applied the COSO internal-control framework. The underlying public reports should also be checked for accessibility, context and comparability rather than adopted solely from the article’s summaries.

Do not let a metric stand in for the assurance question

Impact ratios offer a particularly useful opportunity to examine the relationship between a calculation and a conclusion. Students may be tempted to see a ratio at or above 0.80 and declare the system fair. The source shows why that leap is unsafe.

A ratio is produced within a defined scope and through a series of methodological choices. Results may differ depending on whether data are analysed by job category, pooled across roles or drawn from several employers. Decisions about missing observations and groups with small sample sizes may also matter. Even a correctly calculated ratio addresses only the question represented by that metric.

A classroom prompt could therefore ask students to complete the sentence: “This evidence supports a conclusion about _, but does not by itself establish _.” Their answers should connect the metric to the stated criteria while identifying claims about broader fairness, governance or legal compliance that require more evidence.

This is also where algorithmic assurance differs from a straightforward transfer of financial-audit procedures. Models can change as data evolve or systems are retrained, making monitoring and drift, meaning changes in model behaviour or performance over time, relevant. Questions about fairness, accuracy and acceptable error involve context-dependent tradeoffs that auditors are not uniquely qualified to resolve. Familiar assurance concepts provide a foundation, but not a complete template.

Keep the profession’s role open to examination

None of the 13 reports in the analysis was produced by a public accounting firm. Providers included economic consultants, human-resources specialists and algorithmic assurance firms. This is evidence about the reports studied, not proof that accounting firms are absent from every AI assurance market.

The author argues that accounting professionals are well positioned to contribute because the profession has established approaches to independence, evidence evaluation, professional judgement and public reporting. That is a professional case rather than a finding from a comparison of accountants with other providers. Institutional independence also does not guarantee high-quality work.

This unresolved issue could support a structured classroom debate. One group could make the case for accountants’ involvement by identifying relevant assurance capabilities. Another could specify the technical, legal and domain expertise that an algorithmic engagement might require. Students could then design a multidisciplinary team and explain how responsibilities should be divided.

Such a discussion keeps the focus on competence and engagement design rather than assuming that accountants should own an emerging field. It also asks students to distinguish the skills needed to build or test a model from those needed to scope an engagement, evaluate evidence and communicate a bounded conclusion.

The same restraint can shape the topic’s place in the curriculum. Algorithmic assurance need not become a standalone AI topic. A verified report could be introduced within existing teaching on criteria, evidence, independence, scope or reporting. Its value would lie in testing transfer: can students recognise the limits of an assurance conclusion when the subject matter is unfamiliar?

The source does not tell us whether they can, or whether this activity improves learning. A lecturer trying it could nevertheless look for revealing difficulties. Do students define the subject matter before judging the evidence? Do they distinguish a statistical result from a broader claim about fairness? Can they explain uncertainty without inventing missing facts?

The most useful lesson may be the simplest. A document does not provide meaningful assurance merely because it is called an audit. Students should be able to ask what was examined, against which criteria, using what evidence and with what limits, whether the object being examined is a set of financial statements or an algorithmic system.

Sources and further reading

Reader account

Continue with email

A free account keeps your saved articles in one place and available on any device. It can make it easier to revisit ideas, evidence and practical examples as your courses develop.

The link signs you in, or creates a free account if you're new. It does not subscribe you to the newsletter.