12 Accounting and Auditing
In October 1929, the New York Stock Exchange lost roughly a quarter of its value in a week. Congress spent the next several years arguing about why, and produced two statutes (the Securities Act of 1933 and the Securities Exchange Act of 1934) that required publicly traded companies to file financial statements examined by an independent auditor. The external audit went from a discretionary service a firm might commission to reassure its banker to a legally mandated act of public verification. The profession that performed it acquired, as part of the bargain, a public interest mandate. A certified public accountant’s client is the company paying the invoice. The duty of care runs to shareholders, lenders, regulators, and everyone downstream of the company’s claims about itself.
That bargain has been renegotiated since. Arthur Andersen signed off on Enron’s financial statements for years before Enron collapsed in 2001. Andersen’s consulting practice had been earning more from Enron than its audit practice, and its shredded workpapers became the emblem of audit failure. Andersen did not survive. Congress responded with the Sarbanes-Oxley Act of 2002, which tightened independence rules and created the Public Company Accounting Oversight Board to audit the auditors. The mandate was not abolished. It was recertified under closer supervision.
This chapter opens Part IV, the course’s third module, which works at the level of the state and treats accounting and engineering together as assurance. Keegan (2026) codes both as lineages of oversight: accounting formalizes “legible records, standardized controls, and independent verification,” and engineering adds testing and incident learning (Chapter 13). The pressure is exemption (Chapter 8) in a particular costume. Lee (1995) describes accountancy’s professional history as “protecting the public interest in a self-interested way,” and Baker (2005) finds the profession’s own pronouncements invoking “the public” as a legitimating abstraction while serving the paying client. Algorithmic auditing is about to repeat that history unless it studies it.
12.1 Audit-washing: exemption in oversight’s clothes
Raji and colleagues (2020) propose an internal algorithmic auditing framework borrowed from accounting, aerospace, and medicine. It names five stages (scoping, mapping, artifact collection, testing, reflection) and insists on durable, dated workpapers. A good audit produces documents a reasonable outsider could read a year later and use to reconstruct what was examined, what was found, and what was decided. A bad one produces a one-paragraph attestation.
The module’s US anchor shows why the distinction matters. New York City’s Local Law 144, enforced since July 2023, requires employers using automated employment decision tools to commission an independent “bias audit” within the year before use and to publish a summary of selection rates and impact ratios by sex and race/ethnicity (NYC Department of Consumer and Worker Protection 2023). Goodman and Trehu (2023) use LL144 to illustrate audit-washing: audits whose institutional form looks like oversight but whose content is shallow enough to credential almost any system. Some published LL144 audits are serious. Others are short PDFs that say which tool was audited and little about which data, which metric choices, or which population. A regime that accepts both as compliant has granted an exemption while appearing to deny one.
Colorado, your state case, has moved in the same direction at a larger scale. The Colorado AI Act (SB 24-205), signed in 2024, places duties on developers and deployers of “high-risk” AI systems used in consequential decisions, including risk-management programs and impact assessments, and gives the Attorney General enforcement authority (Colorado General Assembly 2024). The legislature has revisited the Act since passage, which is itself a lesson: assurance duties are negotiated, delayed, and narrowed long before anyone audits anything. Whether Colorado’s impact assessments become workpapers or attestations depends on what the people writing them can demonstrate. The European Union’s AI Act is the module’s main counter-case: its conformity assessment for high-risk systems (Article 43) borrows the EU product-safety tradition of technical files and notified bodies, and requires either internal control or third-party assessment depending on the system (European Union 2024). The same consultancies that wrote thin LL144 attestations are selling conformity packages. Chapter 14 adds Canada’s Algorithmic Impact Assessment and asks how a state buys assurance through procurement.
Oversight, as Chapter 10 developed it, needs authority and remedy that an audit cannot supply by itself. Accounting adds what an audit must supply on its own: dated workpapers, separation of audit from consulting, standards fixed before the numbers are seen, and a credential that can be revoked. An algorithmic audit that lacks all four performs oversight rather than providing it, and that performance is how exemption survives inside a regime that looks like regulation. The continuity element of the installed base (Chapter 2) matters for the same reason: an audit that cannot be rerun next quarter because the analyst’s laptop was wiped is, for public purposes, an audit that did not happen. The specific claim of this chapter is that a manifest, a thresholds file, and a changelog are the minimum installed base for independent verification.
12.2 Tutorial, part 1: a disaggregated audit of a Colorado model
Gender Shades (Buolamwini and Gebru 2018) showed that a single accuracy number can hide large disparities across intersectional groups. The disaggregated audit is the technique that refuses the single number. You will run one on folktables (Ding et al. 2021), which packages American Community Survey (ACS) microdata into prediction tasks, using fairlearn for disaggregated metrics. The task, ACSIncome, predicts whether a working adult’s income exceeds $50,000. It is deliberately banal. Restrict it to Colorado.
# pip install folktables fairlearn scikit-learn pandas pyyaml joblib
from folktables import ACSDataSource, ACSIncome
source = ACSDataSource(survey_year="2018", horizon="1-Year", survey="person")
co = source.get_data(states=["CO"], download=True)
features, label, group = ACSIncome.df_to_pandas(co)
y = label.squeeze() # True if PINCP > 50,000
print(features.shape, round(y.mean(), 2))
# => (n, 10), with n in the tens of thousands; record your n in the manifestfolktables returns race as the ACS RAC1P code. Recode it, and pull sex from the feature matrix:
RAC1P = {1: "White", 2: "Black", 3: "AIAN", 4: "AIAN", 5: "AIAN",
6: "Asian", 7: "NHPI", 8: "Some other race", 9: "Two or more"}
race = group.squeeze().map(RAC1P).fillna("Unknown")
sex = features["SEX"].map({1: "Male", 2: "Female"})Check every key of that dictionary against the ACS PUMS data dictionary for your vintage, not against a blog post. Codes 3 through 5 are three American Indian and Alaska Native categories, which this map collapses. Hispanic origin is a separate ACS variable (HISP) and is not in the task at all. These are choices. Write them down.
Fit the simplest defensible model:
from sklearn.model_selection import train_test_split
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
X_tr, X_te, y_tr, y_te, r_tr, r_te, s_tr, s_te = train_test_split(
features, y, race, sex, test_size=0.3, random_state=42, stratify=y)
model = make_pipeline(StandardScaler(), LogisticRegression(max_iter=1000))
model.fit(X_tr, y_tr)
y_pred = model.predict(X_te)The model is crude: occupation and birthplace codes are treated as numbers. That is fine. If an audit only works on a model you understand, it is a commentary on your own pipeline. The audit should be method-agnostic.
import pandas as pd
from sklearn.metrics import accuracy_score
from fairlearn.metrics import (MetricFrame, count, selection_rate,
false_positive_rate, false_negative_rate,
demographic_parity_difference, equalized_odds_difference)
sensitive = pd.DataFrame({"race": r_te, "sex": s_te})
mf = MetricFrame(
metrics={"n": count, "accuracy": accuracy_score,
"selection_rate": selection_rate,
"fpr": false_positive_rate, "fnr": false_negative_rate},
y_true=y_te, y_pred=y_pred, sensitive_features=sensitive)
print(mf.by_group.sort_values("n"))
# => one row per (race, sex) cell; the smallest cells print first
dpd = demographic_parity_difference(y_te, y_pred, sensitive_features=sensitive)
eod = equalized_odds_difference(y_te, y_pred, sensitive_features=sensitive)
print(f"DPD {dpd:.3f} EOD {eod:.3f}")
# => two different numbers; neither is "the" fairness scoreMetricFrame.by_group is the centerpiece. The n column is not optional: it tells your reader whether any of the other numbers mean anything.
fairlearn will not tell you which parity metric is defensible. Demographic parity asks whether positive predictions are allocated at the same rate across groups; equalized odds asks whether error rates are balanced given the true label. With unequal base rates they cannot both be satisfied, which is a mathematical result, not a political one. Two quieter surprises. First, demographic_parity_difference is the gap between the highest and lowest group, computed over every cell you pass in, including a cell of thirty people whose rate is mostly noise. The headline number is often set by your smallest group. Second, nothing in the library warns you about small cells. Report n beside every metric, suppress cells below a documented minimum, and say in prose which groups the audit cannot speak to.
Save the evaluation set and the model so the next step can verify them:
import joblib
X_te.assign(label=y_te.values, race=r_te.values, sex=s_te.values) \
.to_csv("data/acs_co_2018_test.csv", index=False)
joblib.dump(model, "models/logit-v1.joblib")12.3 Tutorial, part 2: from notebook to workpaper
A notebook produced those numbers once. An audit has to produce them on demand, with a record of every decision and a log of what changed between runs. That is the difference between a paragraph and a workpaper. Build this layout:
audit-pipeline/
├── CHANGELOG.md
├── Makefile
├── manifest.yaml
├── thresholds.yaml
├── audit.py
├── data/ models/ reports/
manifest.yaml is the input contract. A reviewer a year from now should be able to reconstruct the run without guessing.
audit:
name: "ACSIncome Colorado disaggregated audit"
run_id: "2027-03-04-001"
analyst: "your.name@colorado.edu"
dataset:
source: "ACS PUMS 2018 1-Year, Colorado, via folktables ACSIncome"
local_path: "data/acs_co_2018_test.csv"
sha256: "<output of sha256sum>"
retrieved_on: "2027-03-04"
recoding: "RAC1P 3-5 collapsed to AIAN; HISP not used"
model:
path: "models/logit-v1.joblib"
sha256: "<output of sha256sum>"
description: "StandardScaler + LogisticRegression, random_state=42"
protected_attributes: [race, sex]
min_cell_size: 100
thresholds_file: "thresholds.yaml"thresholds.yaml carries the decision rules separately, because they change on a different cadence and need their own authority:
demographic_parity_difference: {warn: 0.05, fail: 0.10}
equalized_odds_difference: {warn: 0.05, fail: 0.10}
adopted: "2027-03-04"
authority: >
Teaching values. A real audit cites the statute, rule, policy,
or contract clause that set them, and who may change them.Financial auditors call the authority field authority documentation. Most algorithmic audits skip it, which is how thresholds drift without anyone deciding to move them.
audit.py reads the manifest, verifies the hashes, runs the disaggregation, applies the thresholds, and writes a dated report:
import hashlib, json, pathlib, sys
from datetime import datetime, timezone
import joblib, pandas as pd, yaml
from fairlearn.metrics import (MetricFrame, count, selection_rate,
false_positive_rate, false_negative_rate,
demographic_parity_difference, equalized_odds_difference)
SUMMARIES = {"demographic_parity_difference": demographic_parity_difference,
"equalized_odds_difference": equalized_odds_difference}
def sha256_of(path):
return hashlib.sha256(pathlib.Path(path).read_bytes()).hexdigest()
def verify(m):
pairs = [("dataset", m["dataset"]["local_path"]), ("model", m["model"]["path"])]
bad = [k for k, p in pairs if sha256_of(p) != m[k]["sha256"]]
if bad:
sys.exit(f"hash mismatch: {bad}") # no report for unverified inputs
def run_audit(df, model, m):
attrs = m["protected_attributes"]
y_true = df["label"]
y_pred = model.predict(df.drop(columns=["label"] + attrs))
mf = MetricFrame(metrics={"n": count, "selection_rate": selection_rate,
"fpr": false_positive_rate, "fnr": false_negative_rate},
y_true=y_true, y_pred=y_pred, sensitive_features=df[attrs])
by_group = mf.by_group[mf.by_group["n"] >= m["min_cell_size"]]
summaries = {k: float(f(y_true, y_pred, sensitive_features=df[attrs]))
for k, f in SUMMARIES.items()}
return by_group, summaries
def evaluate(summaries, thresholds):
out = []
for name, value in summaries.items():
rule = thresholds[name]
level = ("FAIL" if value >= rule["fail"]
else "WARN" if value >= rule["warn"] else "PASS")
out.append({"metric": name, "value": round(value, 4), "level": level})
return out
def main():
m = yaml.safe_load(open("manifest.yaml"))
thresholds = yaml.safe_load(open(m["thresholds_file"]))
verify(m)
df = pd.read_csv(m["dataset"]["local_path"])
by_group, summaries = run_audit(df, joblib.load(m["model"]["path"]), m)
report = {"run_id": m["audit"]["run_id"],
"generated_at": datetime.now(timezone.utc).isoformat(),
"manifest": m, "findings": evaluate(summaries, thresholds),
"by_group": by_group.reset_index().to_dict(orient="records")}
pathlib.Path("reports").mkdir(exist_ok=True)
out = pathlib.Path(f"reports/{m['audit']['run_id']}.json")
out.write_text(json.dumps(report, indent=2, default=str))
print(json.dumps(report["findings"], indent=2))
if __name__ == "__main__":
main()
# => [{"metric": "demographic_parity_difference", "value": ..., "level": "WARN"}, ...]Notice what the script does that a notebook would not. It refuses to run on inputs whose hashes do not match the manifest, because a run against unverified inputs is not an audit. It suppresses small cells by a rule written down in advance. It stamps the run in UTC. Each behavior is borrowed from financial audit practice, where a workpaper exists to be legible to someone who was not in the room.
The Makefile is the one-command interface. If a reviewer cannot type make audit and get a report, the pipeline is not reproducible:
.PHONY: audit clean
audit:
python audit.py
clean:
rm -f reports/*.jsonclean removes generated reports, never inputs. You do not delete workpapers. You archive them. Finally, CHANGELOG.md is the narrative companion to the numbers:
## 2027-03-04-001
- First run. Dataset and model hashes as in manifest.
- Thresholds adopted 2027-03-04 (teaching values).
- Findings: see reports/2027-03-04-001.json.
## 2027-03-11-001
- Re-run, no manifest changes except run_id.
- Library versions: record fairlearn and scikit-learn versions here.
- Changes in findings, if any, and the root cause.Commit the directory, run make audit, and read the report for what it does not say. Does it record which columns were features? Whether the model was evaluated on a holdout? Both choices are defensible; neither is defensible if implicit. A week later, update run_id and run again. Usually nothing moves. Sometimes a minor library release changes a default and a metric shifts in the third decimal place. The discipline is not preventing drift. It is noticing it, logging it, and deciding whether it changes the conclusion. The occasional run where it does is why you keep a changelog.
12.4 Writing the audit for a non-technical reader
The JSON report is raw material. The audit is what you write around it, and it is the seed of the report you will write in Chapter 15. Use four sections.
- What the model is, and is not. A logistic regression trained on one year of Colorado ACS data. A teaching artifact, not a procurement recommendation.
- Who is in the data. The per-cell counts, with the smallest cells named in prose, including those below
min_cell_sizethat the audit cannot speak to. Audits that silently drop small cells pretend to a completeness they lack. - What the metrics show. The disaggregated table with a plain-language reading of each gap, its threshold level, and its direction.
- What the audit rests on. The recoding, the omitted Hispanic-origin variable, self-reported and sampled survey data, a single vintage and a single holdout. Name each choice and how a second auditor might make it differently.
Section 4 is where most real audits fail, and where audit-washing hides. If the next auditor cannot reproduce your table from your manifest, the audit is not an instrument of oversight. It is a press release.
12.5 Exercises
Exercise 12.1 (Guided, laptop, real public data). Run the Colorado ACSIncome audit above end to end. Then build the audit-pipeline/ directory, fill manifest.yaml with real hashes from sha256sum, commit it, and run make audit. Submit the repository URL and the generated report. Graded on completeness, not on whether the model passes.
Exercise 12.2 (Guided). One week later, rerun with a new run_id. Write a changelog entry stating the exact fairlearn and scikit-learn versions and what changed, even if nothing did. Commit it.
Exercise 12.3 (Analytic). From the NYC Department of Consumer and Worker Protection’s LL144 materials, find one published bias audit summary from a real employer. In 400 words, assess it against the Raji et al. (2020) stages. Which workpaper is missing? Is it audit-washing in Goodman and Trehu’s (2023) sense? Say what evidence would change your verdict.
Exercise 12.4 (Comparative). In 500 words, compare LL144, the EU AI Act’s conformity assessment, and Colorado’s SB 24-205 on three dimensions: which systems are covered, who performs the assessment, and what is published. Quote the text of each. Flag any provision of the Colorado Act whose current status you could not confirm.
Exercise 12.5 (Builds Piece 3). Extend your pipeline into the audit for Piece 3. Swap in a second folktables task (for example, ACSEmployment or ACSPublicCoverage) or a second sensitive attribute, still restricted to Colorado. Rewrite thresholds.yaml so that its authority field names a real Colorado state body and the policy instrument it could use to adopt such thresholds. Then write the four-section audit memo (about 800 words). The memo becomes the methods and findings of your Piece 3 report in Chapter 15.
12.6 Looking ahead
A workpaper records what you did. It does not prove that the code does what the workpaper says. Chapter 13 puts this pipeline under test: shape and range checks on a fixture, a regression snapshot, a deliberately introduced bug, a post-incident review, and continuous integration. Engineering also asks who keeps the record when the audited party will not, which is where worker observatories enter the module.
12.7 Further Reading and Resources
- C. Richard Baker (2005), “What is the meaning of ‘the public interest’? Examining the ideology of the American public accounting profession,” Accounting, Auditing & Accountability Journal 18(5). The close reading of AICPA language this chapter leans on.
- Tom Lee (1995), “The professionalization of accountancy: A history of protecting the public interest in a self-interested way,” Accounting, Auditing & Accountability Journal 8(4). The title is the argument.
- Inioluwa Deborah Raji et al. (2020), “Closing the AI accountability gap,” Proceedings of FAccT 2020. Audit as a practice rather than a metric.
- Bryce Goodman and Julia Trehu (2023), “AI audit washing and accountability.” The term in its useful form.
- NYC Department of Consumer and Worker Protection, Local Law 144 page: https://www.nyc.gov/site/dca/about/automated-employment-decision-tools.page. Read the final rule, not only the FAQ.
- Colorado General Assembly, SB 24-205: https://leg.colorado.gov/bills/sb24-205. Read the bill history as well as the text; the amendments are part of the lesson.
- EU AI Act, Regulation (EU) 2024/1689: https://eur-lex.europa.eu/eli/reg/2024/1689/oj. Article 43 and Annexes VI and VII cover conformity assessment.
folktablesrepository and paper: https://github.com/socialfoundations/folktables. Read the documentation on the limits of ACS tasks as benchmarks.fairlearnuser guide: https://fairlearn.org/main/user_guide/. The sections onMetricFrameand on disagreement between parity metrics.- PCAOB inspection reports: https://pcaobus.org/oversight/inspections. What detailed, public criticism of audit workpapers looks like.