13  Engineering

Most bridges do not fall down. This is unremarkable until you notice how many there are, how heavy the trucks on them are, and how few engineers are on site at any given bridge on any given day. The low background rate of collapse is a professional accomplishment. It rests on standards, inspection schedules, conservative design margins, licensure, and the slow accumulation of what Henry Petroski (1992) calls the knowledge of failure. Engineering became a profession partly by agreeing to pay close attention to its own disasters: the Tay Bridge, the Tacoma Narrows, the DC-10 cargo door. A failure taxonomy on an engineer’s bookshelf is one of the most important cultural facts about the field.

Data science does not yet have that bookshelf. It has papers on bias, a growing body of audit work, and a scattered practice of post-incident reviews at firms that run production systems. What it lacks is a professional identity organized around the question “what failed, why, and what will we change so it does not fail again.” Chapter 12 gave you accounting’s half of assurance: workpapers, independence, thresholds fixed in advance. This chapter supplies the other half. Keegan (2026) describes safety engineering as the lineage that operationalizes oversight through “standards, testing, incident learning, and high-reliability operations,” creating “feedback loops that keep infrastructures governable rather than merely optimized.” The audit you built in the last chapter is the thing under test here. Most of the tests you write will fail at first. That is the point.

13.1 The engineer’s duty to the public

Engineering’s claim to be a profession rests on codes of ethics that put the public ahead of the client. The National Society of Professional Engineers code places “hold paramount the safety, health, and welfare of the public” first among its fundamental canons. An engineer who signs off on a bridge they know to be unsafe has not disappointed a customer. They have violated a duty to the people who will drive over it, and that duty survives the contract. In the United States the duty is enforced at the state level: states license professional engineers, and in Colorado a licensing board within the Department of Regulatory Agencies can discipline or revoke. That makes engineering a natural lineage for the module that works at the state level.

When you write a pipeline that supports a public decision, you inherit a weaker version of that obligation. You cannot stamp a dataset with a license seal. You can agree that the client’s deadline does not overrule the public’s exposure to a quietly broken number.

13.2 Normalization of deviance

The second thing you borrow is a vocabulary for how safe systems go bad. Diane Vaughan’s (1996) study of the Challenger launch decision introduced normalization of deviance: over a series of flights, NASA engineers came to treat O-ring erosion as acceptable even though it was never part of the design. Each decision to fly was defensible on the data at hand. The drift was collective and gradual. By the night before the launch, the engineers who objected had to prove the shuttle was unsafe, rather than the other way around. Sidney Dekker (2016) generalizes the pattern as drift into failure: no broken component, just a long sequence of locally reasonable adjustments.

The audit pipeline from Chapter 12 has an obvious drift path. The model lands at an equalized-odds difference of 0.07 every quarter, just over the 0.05 warning line. Someone proposes raising the warning threshold to 0.08 “to reduce alert fatigue.” Nobody objects, because the number has always been 0.07 and nothing bad has happened. A year later the thresholds file is an archaeological record of a standard that moved to meet the system instead of the reverse. That is audit-washing achieved without anyone intending it. Incident reviews and tests, done well, put the question “when did we stop noticing this” on the table and keep it there.

Karl Weick and Kathleen Sutcliffe’s work on high-reliability organizations (Weick et al. 1999) describes what organizations with low failure rates share: preoccupation with failure, reluctance to simplify, sensitivity to operations, commitment to resilience, deference to expertise. Charles Perrow (1984) answered that some systems are so tightly coupled that accidents in them are “normal.” The debate is unresolved and productive. The takeaway is that reliability is organizational before it is technical. Tests do not run themselves.

13.3 Tutorial: a test suite for the audit

Software engineering has sometimes claimed this tradition and sometimes squatted on it. The Therac-25, the 737 MAX, and the UK Post Office’s Horizon system are software failures the older professions would recognize at once. Data pipelines are usually less safety-critical one at a time and often more consequential in aggregate. What converts the workpapers of Chapter 12 from a genre into a discipline is a test suite. Add it to the project:

audit-pipeline/
├── audit.py   manifest.yaml   thresholds.yaml   CHANGELOG.md
├── requirements.txt
├── tests/
│   ├── conftest.py
│   ├── fixtures/
│   │   ├── tiny_acs_co.csv
│   │   └── expected.json
│   ├── test_shape.py
│   ├── test_rules.py
│   └── test_regression.py
└── .github/workflows/ci.yml

tiny_acs_co.csv is a fifty-row synthetic file with the same columns as the evaluation set you exported from folktables. It looks like the real input but is not, so it can be committed, shared, and run in under a second with no network. Build it so every (race, sex) cell has at least one positive and one negative label; otherwise false positive and false negative rates are undefined for that cell and the snapshot fills with NaN. A rule-based stand-in model keeps the tests independent of a pickled file:

# tests/conftest.py
import pandas as pd
import pytest

class RuleModel:
    """Deterministic stand-in: predicts high income for full-time hours."""
    def predict(self, X):
        return (X["WKHP"] >= 40).to_numpy()

@pytest.fixture
def df():
    return pd.read_csv("tests/fixtures/tiny_acs_co.csv")

@pytest.fixture
def rule_model():
    return RuleModel()

13.3.1 Shape and range tests

The first tests catch the failure modes a data pipeline actually has. Not “does this function return 42,” but “does the table have the columns I expect, and do the values fall where the ACS data dictionary says they can.”

# tests/test_shape.py
import pytest

REQUIRED = ["AGEP", "SCHL", "WKHP", "SEX", "RAC1P", "label", "race", "sex"]

def test_required_columns(df):
    missing = set(REQUIRED) - set(df.columns)
    assert not missing, f"missing columns: {missing}"

def test_row_count(df):
    assert 40 <= len(df) <= 60      # a range, never an equality

@pytest.mark.parametrize("col,low,high", [
    ("AGEP", 17, 99),               # ACSIncome keeps ages over 16
    ("WKHP", 1, 99),                # usual hours worked per week
    ("SEX", 1, 2),
    ("RAC1P", 1, 9),
])
def test_bounds(df, col, low, high):
    assert df[col].between(low, high).all(), f"{col} outside [{low}, {high}]"

def test_sex_labels_match_codes(df):
    pairs = set(zip(df["SEX"], df["sex"]))
    assert pairs <= {(1, "Male"), (2, "Female")}

The last test looks trivial. It is the one that catches a mis-keyed recoding dictionary, which is among the most common silent errors in demographic analysis. An early draft of this book’s own audit code mapped several RAC1P codes to the wrong labels; a test like this one, written for race, would have caught it on the first run.

TipThe Missing Manual

Most testing tutorials teach the web-application kind: assert that the login endpoint returns 200 and the body says “Welcome.” A data test written that way (“there are exactly 18,472 rows”) is almost always wrong, because data moves. Write assertions in the shape of ranges and distributions, and reach for pytest.mark.parametrize early: without it you write fifteen near-identical bounds tests, and the fifteenth is where the bug lives. The second surprise is subtler. A regression snapshot tests the numbers. It does not test the judgment applied to them. If the code that converts a metric into PASS, WARN, or FAIL is wrong, every number can match the snapshot while the audit reports a failing system as merely worth watching. Test the decision rule directly.

13.3.2 Rule tests and a regression snapshot

# tests/test_rules.py
import pytest, yaml
from audit import evaluate

RULES = {"demographic_parity_difference": {"warn": 0.05, "fail": 0.10}}

@pytest.mark.parametrize("value,expected", [
    (0.01, "PASS"), (0.07, "WARN"), (0.10, "FAIL"), (0.25, "FAIL"),
])
def test_threshold_levels(value, expected):
    [finding] = evaluate({"demographic_parity_difference": value}, RULES)
    assert finding["level"] == expected

def test_threshold_change_is_logged():
    adopted = yaml.safe_load(open("thresholds.yaml"))["adopted"]
    assert adopted in open("CHANGELOG.md").read()

The second test is a governance rule written as code. Nobody can change thresholds.yaml and its adopted date without a matching changelog entry, or the build goes red. It does not stop drift. It makes drift visible and attributable, which is the antidote Vaughan’s account calls for.

# tests/test_regression.py
import json, pathlib
from audit import run_audit

def test_summaries_match_snapshot(df, rule_model):
    manifest = {"protected_attributes": ["race", "sex"], "min_cell_size": 1}
    _, summaries = run_audit(df, rule_model, manifest)
    expected = json.loads(pathlib.Path("tests/fixtures/expected.json").read_text())
    for name, value in expected.items():
        assert abs(summaries[name] - value) < 1e-9, name

Generate expected.json once by running run_audit on the fixture and saving the summaries. When the snapshot changes, the diff in the pull request becomes part of the documentation, and someone has to explain it in the changelog.

Run the suite:

pytest -q
# => .............                                   [100%]
# => 13 passed in 0.6s

13.4 Deliberately introducing a bug

You learn more from breaking the pipeline on purpose than from any amount of testing theory. On a branch, swap the order of the two comparisons in evaluate:

# audit.py, on branch demo/threshold-bug
level = ("WARN" if value >= rule["warn"]          # checked first: bug
         else "FAIL" if value >= rule["fail"] else "PASS")
git checkout -b demo/threshold-bug
git commit -am "demo: check warn before fail in evaluate()"
pytest -q
# => ...FF........
# => FAILED tests/test_rules.py::test_threshold_levels[0.1-FAIL]
# => FAILED tests/test_rules.py::test_threshold_levels[0.25-FAIL]

Notice what passed. The regression snapshot is green, because every metric is computed correctly. Only the rule test catches the bug, and the bug is the worst kind an audit can have: it downgrades every failure to a warning. No number is wrong. The escalation is. A reviewer skimming the report would see sensible values and a reassuring column of WARNs. Fix the order, rerun, and commit. The two-commit sequence is the evidence that the test does real work.

13.5 The post-incident review

The fix is not the end of the response. The review is. Engineering’s contribution to public documentation genres is the post-incident review (in aviation, the accident report). A useful template has five sections:

  • Summary. What broke, when it was detected, who was affected, what was done.
  • Timeline. Detection, response, each diagnostic step, mitigation, all-clear, in UTC. Timestamps are non-negotiable.
  • Impact. How many published findings were wrong, for how long, and which decisions might have relied on them.
  • Contributing factors. Plural. Design choices, process gaps, missing tests, organizational pressure, plain error. Blameless does not mean vague.
  • What changes. Specific, assigned, dated. A change assigned to nobody is not a change.

For the threshold bug, the timeline shows it introduced and caught by pytest minutes later in a pre-merge run. Impact is zero because nothing shipped. Contributing factors include a decision rule with no direct test until this chapter and a snapshot that gave false comfort. Publishing the review, even for a demo, is the habit. The NTSB, NASA, and many software firms publish theirs, and writing for an outside reader forces a different standard of clarity.

13.6 Incident learning from below: worker observatories

Every review above assumes that the organization running the system also runs the review. Often it does not, and it will not. Keegan’s (2026) second case, supporting worker observatories, is the module’s counter-case for exactly that situation. On ride-hailing and delivery platforms, pay formulas, ranking rules, and deactivation triggers change without notice, and the platform keeps the only logs (Jarrahi et al. 2021). When a pay change cuts earnings, there is no published post-incident review, because from the platform’s side nothing went wrong.

Worker observatories build the counter-record. Calacci and Pentland (2022) built a tool with gig shoppers that let them pool pay data and test whether an algorithm change had lowered earnings. Dalal (2024) frames such work as a sociotechnical audit grounded in the tradition of workers’ inquiry. In the UK, Worker Info Exchange helps drivers and couriers use data-protection access rights to obtain the data platforms hold about them, turning individual requests into collective evidence. Keegan calls the result “an installed base for contestation”: shared conventions for naming events like deactivations, documented methods so evidence cannot be dismissed as “folk statistics,” and tiered access, governed by workers, so contributors are not exposed to retaliation (Calacci and Stein 2023).

Read with this chapter’s tools, an observatory is an incident database kept by the people the incident happened to. Its logs are the timeline. Its pooled pay data are the impact section. Its organizers write the contributing factors the platform would leave out. Miceli, Posada, and Yang (2022) argue that talk of “bias” in data often hides questions of power over who defines categories and who records what; observatories answer by moving the pen. They also catch drift the platform has normalized, because the worker whose pay falls a little each month notices the gradient that a quarterly dashboard smooths away.

The state-level hook matters here. The Colorado AI Act, as enacted, includes duties to report discovered algorithmic discrimination to the Attorney General. Duties like that create a channel for incident reports, but only for the incidents a developer or deployer chooses to discover. A report from below needs a route in too, which is a question about remedy that Chapter 10 raised and Chapter 15 will ask you to address to a named state body.

NoteIn the Public Interest

Oversight is not a heroic intervention. It is routine scrutiny, organized so that the routine itself surfaces surprises. Tests and continuous integration are that routine at the pipeline layer; they are how you earn the right to make a public claim about a number. The specific claim of this chapter is that oversight fails in two predictable places: in the decision rule that converts a metric into an escalation, where a single swapped comparison launders a failure into a warning, and in the choice of who keeps the incident record. A tested audit addresses the first. Worker observatories address the second, by making the people affected into record-keepers with standing. Both serve the continuity element of the installed base (Chapter 2): the next person who inherits the system can tell, from the tests and the reviews alone, what it is for and when it broke.

13.7 Continuous integration

Tests you do not run are decorations. One YAML file makes pytest run on every push:

# .github/workflows/ci.yml
name: ci
on: [push, pull_request]
jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: "3.11"
      - run: pip install -r requirements.txt pytest
      - run: pytest -v

Pin versions in requirements.txt (fairlearn==..., scikit-learn==...) so that CI and the changelog agree about what ran. Push a branch and open a pull request. GitHub Actions marks the check green or red next to the commit, and with branch protection on, a red check blocks the merge. The cost is minutes. The return is that the whole suite runs on a clean machine, with declared dependencies, every time anyone proposes a change. A test that runs only on the author’s laptop does not, for the purposes of the installed base, exist.

13.8 Exercises

Exercise 13.1 (Guided, laptop, real public data). Add tests/ to the audit-pipeline/ you built in Chapter 12, with a shape test, a parametrized range test, a rule test, and a regression snapshot against a fifty-row synthetic fixture. The suite must run in under two seconds. Then point the df fixture at your real Colorado evaluation set and run the shape and range tests again. Record which assertions fail on real ACS data and whether each failure is a bug in the data, the code, or the test. Document in tests/README.md what each test asserts and why.

Exercise 13.2 (Guided). On a branch named demo/bug, introduce a silent error of your choice: the swapped comparison above, a mis-keyed RAC1P map, or a filter that drops a group. Commit it with an honest message, confirm a test fails, then fix it in a second commit. Do not merge. The two-commit history is the artifact. If no test failed, write the test that would have caught it first.

Exercise 13.3 (Analytic). Write a post-incident review for the bug in 13.2 using the five-section template, with commit SHAs and UTC timestamps. Write it as though a state legislative staffer might read it, because in the public version of this work one might.

Exercise 13.4 (Builds Piece 3). Add .github/workflows/ci.yml and test_threshold_change_is_logged. Push, confirm a green check, then push the bug branch and confirm a red one. Add the CI status badge to your README.md. The tested, CI-backed audit is the technical artifact for Piece 3.

Exercise 13.5 (Open-ended). Find a public worker-led data project (Worker Info Exchange or another observatory you can document from its own publications) and one corporate incident review (Cloudflare, GitLab, or a NASA mishap report). In 500 words, compare them on the five template sections. Which sections does each omit, and why? Who could act on each record, and through what Colorado or federal channel?

13.9 Looking ahead

Tests and reviews are what a careful team does for itself. Most public systems are not built by careful teams inside government. They are bought. Chapter 14 turns to the state as a buyer, where assurance is either written into a procurement contract or quietly given up, and where a requirement that a vendor run tests like these in CI and share the results is worth more than any promise in a sales deck.

13.10 Further Reading and Resources

Baker, C. Richard. 2005. “What Is the Meaning of ‘the Public Interest’? Examining the Ideology of the American Public Accounting Profession.” Accounting, Auditing & Accountability Journal 18 (5): 690–703. https://doi.org/10.1108/09513570510620510.
Calacci, Dana, and Alex Pentland. 2022. “Bargaining with the Black-Box: Designing and Deploying Worker-Centric Tools to Audit Algorithmic Management.” Proc. ACM Hum.-Comput. Interact. 6 (CSCW2): 428:1–24.
Calacci, Dana, and Jake Stein. 2023. “From Access to Understanding: Collective Data Governance for Workers.” European Labour Law Journal 14 (2): 253–82.
Dalal, Samantha. 2024. “Lessons from Workers’ Inquiry: A Sociotechnical Approach to Audits of Algorithmic Management Systems.” XRDS: Crossroads, The ACM Magazine for Students 30 (4): 36–40.
Dekker, Sidney. 2016. Drift into Failure: From Hunting Broken Components to Understanding Complex Systems. CRC Press.
Jarrahi, Mohammad Hossein, Gemma Newlands, Min Kyung Lee, Christine T. Wolf, Eliscia Kinder, and Will Sutherland. 2021. “Algorithmic Management in a Work Context.” Big Data & Society 8 (2): 20539517211020332.
Keegan, Brian C. 2026. “Public Interest Data Infrastructuring.” Under Review.
Miceli, Milagros, Julian Posada, and Tianling Yang. 2022. “Studying up Machine Learning Data: Why Talk about Bias When We Mean Power?” Proceedings of the ACM on Human-Computer Interaction 6 (GROUP). https://doi.org/10.1145/3492853.
Perrow, Charles. 1984. Normal Accidents: Living with High-Risk Technologies. Basic Books. https://press.princeton.edu/books/paperback/9780691004129/normal-accidents.
Petroski, Henry. 1992. To Engineer Is Human: The Role of Failure in Successful Design. Vintage Books. https://archive.org/details/toengineerishuma0000petr.
Vaughan, Diane. 1996. The Challenger Launch Decision: Risky Technology, Culture, and Deviance at NASA. University of Chicago Press. https://press.uchicago.edu/ucp/books/book/chicago/C/bo22781921.html.
Weick, Karl E., Kathleen M. Sutcliffe, and David Obstfeld. 1999. “Organizing for High Reliability: Processes of Collective Mindfulness.” Research in Organizational Behavior 21: 81–123.