19  Critical Data Studies, Data Justice, and Refusal

By now you have built installed-base pieces in three modules and begun a fourth. You documented the provenance of a city dataset, built an exemptions and remedy ledger from county records, ran a tested audit for a state body, and, in this module, audited the erosion of a federal dataset and drafted a plan to steward it. Each piece added one or more of the six elements: records that link, categories that can be interpreted, evidence that persists, scrutiny that is safe, authorities who can act, and people who can seek remedy (Keegan 2026). This chapter turns the framework on itself. For some of the people the installed base is supposed to serve, any one of those six can also be the mechanism by which harm arrives. A record that links is a record that surveils. A category that can be interpreted is one that was once imposed. Evidence that persists can be subpoenaed a decade later by an agency that did not exist when the data was collected. Scrutiny can be predatory. Authority can be carceral. Remedy can be a pretext for more collection.

Critical data studies holds these possibilities in view. Data justice asks who decides, under what conditions, on whose terms. Refusal is what you do when the answer is that the data should not exist in the form proposed. This chapter asks you to do two things with the same federally governed dataset: audit its classification schema, and write a refusal specification for it. The pair is the point. Neither alone does the work. For Piece 4, the refusal specification is the alternative to the stewardship plan from Chapter 18: some federal datasets deserve a mirror, and some deserve a narrower existence.

Keegan (2026) positions these fields carefully. Critical data studies supplies indispensable diagnostic language for how power shapes classification and surveillance, but it is organized more around critique than around specifying institutions that can be sustained. Data justice specifies what infrastructuring must be for: redistribution, representation, and the capacity to refuse. The installed base is only as just as the commitments it is built to carry.

19.1 The starting canon: boyd and Crawford

Begin with boyd and Crawford (2012), “Critical questions for big data,” which named six provocations every chapter of this book has implicitly answered. Big data changes the definition of knowledge. Claims to objectivity are misleading. Bigger is not always better. Out of context, big data lose meaning. Accessibility is not ethicality. Limited access creates new digital divides.

The fifth matters most here. Accessibility is not consent. That a record exists and is legally available for analysis does not settle whether you should use it. The Homeless Management Information System (HMIS), the running example, is a tight case, and a federal one. HUD requires Continuums of Care and Emergency Solutions Grants recipients to record client data in an HMIS as a condition of funding, and the HMIS Data Standards that define every field are issued by HUD with its federal partners at HHS and the VA (US Department of Housing and Urban Development 2024). The records are real, detailed, and, under some conditions, available to researchers. The people documented in them did not consent in any sense a research ethics board would recognize. They sought shelter, and documentation was the price.

19.2 The political work of classification

Bowker and Star (2000), Sorting Things Out, is the second book you need. Classification is never neutral. Every category system is a historical artifact produced by some authority, for some purpose, at some moment. The American Psychiatric Association listed homosexuality as a mental disorder until 1973. The US Census race categories have been rewritten roughly every decade since 1790, each time to serve a different administrative project. Classification does not describe the world. It imposes a shape on it, and the shape tends to serve whoever did the imposing.

When you call value_counts() on a column labeled race or veteran_status and see six buckets, you are not seeing human variation. You are seeing the surviving output of a classification pipeline that, upstream, had more, fewer, or different categories. The flattening usually happened at intake because a reporting template required a fixed list. Federal standards set that list. In March 2024 the Office of Management and Budget revised its race and ethnicity standard (Statistical Policy Directive 15) to use a single combined question and add a Middle Eastern or North African category, and federal data systems, HMIS among them, have been revising their fields to follow. Each revision changes what a Denver intake worker can write down, and each one breaks the comparability of a time series. Erosion and classification are the same problem seen from two sides.

Denton et al. (2021) make this concrete for machine learning. ImageNet, assembled in 2006 to 2009 on WordNet’s 1980s hierarchy, carried into the 2010s a taxonomy of human kinds that included slurs and categories like “loser.” When ImageNet became the benchmark for a generation of vision models, those choices propagated into systems that never revisited them. The dataset was a historical artifact; the benchmark made it load-bearing.

19.3 Data justice and the question of for whom

Taylor (2017) puts a different question on the table. The fairness audit you ran in Chapter 12 asked whether a model treats groups equivalently. Data justice asks something prior: what counts as justice here, for whom, through what mechanisms? A system can satisfy demographic parity and still be unjust, because the deeper question is whether it should exist, whether the people subject to it had any say, and whether they recognize the categories used to sort them. D’Ignazio and Klein (2020), Data Feminism, extends this into seven principles, of which “examine power” and “elevate emotion and embodiment” are the most likely to make an engineering audience uncomfortable, and therefore the most likely to be doing real work.

You met openness in Chapter 6 and ownership in Chapter 18. Data justice says neither is sufficient. Openness can make a community’s vulnerability indexable by anyone with a laptop. Ownership can be a contract that extracts data under conditions nobody could negotiate. Both are incomplete without the prior question of for whom.

19.4 Refusal as design

Zong and Matias (2024) and Garcia et al. (2022) develop the concept this chapter is built around. Refusal is not the absence of data practice. It is a design move. To refuse is to specify what would not be collected, what would be collected differently, what would be collected only under conditions the affected community controls, and who makes each call. A well-written refusal specification is as detailed as a data management plan and does similar work: it commits its authors to operations they can be held to.

The move students miss the first time is that refusal is generative. A refusal specification is not a protest document. It says: if we did not collect this, what would we do to answer the underlying question? That substitution is where the intellectual work lives.

19.5 The installed base, revisited

The installed base assumes records should link, persist, be interpretable, and be open to scrutiny. Carroll et al. (2020) propose something different. The CARE Principles (collective benefit, authority to control, responsibility, ethics) were written as a companion and corrective to FAIR (findable, accessible, interoperable, reusable), the default standard for scientific data stewardship (Wilkinson et al. 2016). CARE does not reject FAIR. It says FAIR is necessary and insufficient: a dataset can satisfy every letter of FAIR and still be in the wrong relationship with the community it describes.

Māori data sovereignty, as Te Mana Raraunga articulates it (Te Mana Raraunga 2018), begins with the premise that data about Māori are a Māori responsibility. Who holds the data, who decides what questions are asked of it, and who benefits are not concerns downstream of the research question. They are the question. The First Nations principles of OCAP (First Nations Information Governance Centre, n.d.) and the US Indigenous Data Sovereignty Network (United States Indigenous Data Sovereignty Network 2024) make parallel arguments in Canada and the United States.

This is the honest moment. The installed-base framing treats linkable, persistent, interpretable records under scrutiny as the default. CARE and Indigenous data sovereignty say that default is a political choice reflecting the settler state’s interest in legibility more than any universal principle of good governance. Keegan (2026) concedes the point in advance: critical perspectives are “internal guardrails” that keep infrastructuring from reproducing extractive visibility under the banner of accountability. You can take the critique seriously without abandoning the framework. You cannot take it seriously while pretending the framework is neutral.

19.6 Running example: HMIS in Denver

The Metro Denver Continuum of Care, like every CoC, operates an HMIS because HUD funding requires it. HMIS records client-level data for anyone entering homelessness services: intake demographics, shelter stays, exits, returns. The CoC reports aggregates upward for HUD’s annual reports to Congress and for Point-in-Time counts that drive resource allocation. A researcher working with a CoC can, under a data use agreement, obtain de-identified extracts.

The categories are set in Washington; the intake happens in Denver. Without HMIS, a continuum cannot document whom it served or qualify for funding. With it, people experiencing homelessness are among the most heavily documented populations in the country, and that documentation is potentially reachable by immigration enforcement, family courts, and any agency with a sufficiently motivated subpoena (Eubanks 2014). The audit is not academic. The refusal is not rhetorical.

TipThe Missing Manual

The HMIS Data Standards manual names the categories. It does not name the decisions that produced them, or what each revision broke. Recent standards replaced separate race and ethnicity fields with a combined “Race and Ethnicity” element, added Middle Eastern or North African and Hispanic/Latina/e/o options, and added a free-text “Additional Race and Ethnicity Detail” field. Two consequences follow. First, a multi-year extract mixes two schemas, and value_counts() will happily count both as if they were one. Second, the free-text field is where self-identification survives, and it is also the field most likely to be dropped from a research extract because it is “messy.” A client who identified as Oaxaqueño, Somali Bantu, or Diné is only legible in that field. The documentation tells you what the codes mean. It does not tell you what was lost when an intake worker had three seconds to pick one.

19.7 Technical component 1: the category audit

The first artifact inventories the categorical fields, traces each to an administrative source, and documents what the collapsing does.

import pandas as pd

df = pd.read_csv("data/hmis_extract.csv", low_memory=False)
CATEGORICAL_FIELDS = ["race_ethnicity", "gender", "veteran_status",
                      "project_type", "destination", "prior_living_situation"]

def audit_field(series):
    counts = series.value_counts(dropna=False).to_frame("n")
    counts["pct"] = (counts["n"] / len(series) * 100).round(2)
    counts["field"] = series.name
    return counts.reset_index(names="value")

audit = pd.concat([audit_field(df[f]) for f in CATEGORICAL_FIELDS])
audit.to_csv("out/category_audit.csv", index=False)
# => one row per (field, value) with counts and percentages

That gives you frequencies, not provenance. For each field, write a provenance note answering four questions: who defined the categories, when, for what administrative need, and what changed from the prior version. The Data Standards manual answers the first three; HUD’s notices for each revision cycle answer the fourth.

PROVENANCE = {
    "race_ethnicity": {
        "authority": "OMB SPD 15 (1997; revised 2024), via HMIS Data Standards",
        "known_flattening": ("Indigenous nations collapse into one option; "
                             "South, Southeast, and East Asian into 'Asian'."),
        "schema_break": "separate race and ethnicity fields before the revision",
        "purpose": "federal civil-rights and funding reporting, not service-matching",
    },
    # Repeat for each field.
}

Then cross-tabulate missingness. A provenance note is stronger when you can show how often “Data not collected” is chosen and how that rate varies by intake site.

MISSING = ["Client doesn't know", "Client prefers not to answer", "Data not collected"]
missing_by_site = (df.assign(m=df["race_ethnicity"].isin(MISSING))
                     .groupby("intake_site")["m"].mean()
                     .sort_values(ascending=False))
# => sites with high non-response often serve clients who distrust documentation;
#    the distrust is itself data

That last step treated absence as a finding. D’Ignazio and Klein (2020) call this counting what is missing, and counting why. A site with 40 percent non-response on race is not a site with poor data quality. It is a site where 40 percent of clients, faced with a form, chose not to answer. That refusal is a signal. A competent audit surfaces it.

19.8 Technical component 2: the refusal specification

The second artifact is a markdown document with a YAML preamble. The YAML commits the authors to specific operations; the markdown explains each.

---
dataset: denver_metro_hmis
governance_body: lived_experience_advisory_board
review_cycle: annual

would_not_collect:
  - field: full_legal_name
    rationale: |
      Enables linkage with criminal-legal and immigration databases.
      Replace with a salted pseudonymous ID generated at intake.
  - field: social_security_number
    rationale: |
      Collect only where a federal program requires it; defaulting
      to collection exceeds the federal minimum and creates a subpoena target.

collect_differently:
  - field: race_ethnicity
    current: closed federal list plus optional free text
    proposed: |
      Free-text self-identification first; the federal mapping is
      applied only for reporting. Raw self-identification stays under
      community control; only the mapping leaves the Continuum.

collect_under_community_governance:
  - field: case_management_notes
    governance: |
      Retained by the direct-service provider, not uploaded to the
      shared HMIS. Aggregates reviewed by the advisory board before release.

decision_process:
  who_decides: |
    The Lived Experience Advisory Board, with at least five currently
    or recently unhoused members voting equally with agency staff.
  by_what_process: Majority vote at a noticed public meeting.
  appeals: |
    A client may ask that their record be excluded from any downstream
    analytic extract; honored within 30 days and logged in aggregate.
---

The structure is load-bearing. Each section answers one question: what you would not collect, what you would collect differently, what you would collect only under community control, and who changes the specification, by what process, with what appeal. A refusal specification without the decision process is a wish list, not a governance document. Note too what the specification can and cannot reach: the would-not-collect list must respect what federal funding requires, and where it cannot, the honest move is to name the federal requirement as the thing to change. That is where a public comment, in Chapter 20, picks up.

19.9 Refusal in the wild: the Algorithmic Ecology

The Stop LAPD Spying Coalition, with the Free Radicals collective, published “The Algorithmic Ecology” in 2020 (Stop LAPD Spying Coalition and Free Radicals 2020). It maps an algorithmic system, LAPD’s predictive policing programs, as an ecology with community, institutional, and political-economic levels, including the funders and vendors behind the police department. Its refusal is not “fix the algorithm” but “dismantle the conditions that make the algorithm legible as a solution to anything.” The coalition’s earlier report “Before the Bullet Hits the Body” (Stop LAPD Spying Coalition 2018) documented how LAPD’s data-driven programs routed back to the same neighborhoods. LAPD ended its LASER program in 2019 and its PredPol contract in 2020, and the coalition’s sustained organizing was part of the pressure behind both decisions.

That is refusal as generative, in the sense Zong and Matias (2024) intend. The coalition did not settle for a seat at the table. It specified the conditions under which the table should not exist in its current form and built the political capacity to enforce the specification. Your specification is modest by comparison. The move is the same.

NoteIn the Public Interest

Chapter 6 argued that openness is interpretive legibility. Chapter 18 argued that ownership is collective, accountable governance over preservation and use. This chapter shows where they collide. A federal dataset can be maximally open and perfectly preserved and still violate CARE, because the community it describes never had authority over whether it should exist. The specific claim: ownership that can be asserted only by entities the installed base already recognizes (an agency, a university, a volunteer mirror) is not ownership by the people described. A refusal specification rebuilds ownership on the prior question of authority, and it is the one Piece 4 artifact that can argue for less data rather than more.

19.10 Exercises

Exercise 19.1 (Guided; laptop, real data). Download HUD’s Point-in-Time and Housing Inventory Count data by Continuum of Care from HUD Exchange, or another public HMIS-derived table (the Annual Homeless Assessment Report tables work). Produce a category audit: field frequencies for the Colorado CoCs, a provenance note for each demographic field, and a comparison across at least two years that shows where a federal category revision breaks the series. Commit as audit.ipynb.

Exercise 19.2 (Applied; Piece 4 component). Write a refusal specification following the YAML skeleton, for HMIS or for the federal dataset you audited in Exercise 16.3. Include at least two would_not_collect entries, two collect_differently entries, one collect_under_community_governance entry, and a full decision_process. Name any federal requirement that blocks an entry. If your Piece 4 takes the refusal path rather than the stewardship path from Exercise 18.3, this is its technical half. Commit as refusal.md.

Exercise 19.3 (Reading). Read the CARE Principles (Carroll et al. 2020) and FAIR (Wilkinson et al. 2016). In 400 words, compare them on who holds data-governance authority, what each says about reuse, and where they conflict in practice. Do not pretend they are compatible where they are not.

Exercise 19.4 (Analytic). Pick a refusal in the wild: the Algorithmic Ecology, the Detroit Community Technology Project’s work on Project Green Light (Detroit Community Technology Project 2019), an Indigenous data sovereignty initiative tracked by USIDSN, or another case your instructor approves. In 500 words, say what is refused, by whom, what alternative the refusal proposes, and what made it effective or not. Cite at least three primary sources.

Exercise 19.5 (Open-ended). Return to your project sketch from Exercise 1.5. Identify one question where refusal is the right answer. In two paragraphs, say what you would not collect or analyze and what substitutes for it: a different question, unit of analysis, governance structure, or partnership. Save as ch19_project_refusal.md.

19.11 Looking ahead

The audit and the refusal specification are internal artifacts, the kind of document a team uses to set its own terms. Chapter 20 turns them outward. A public comment on a federal rulemaking is where an argument for stewardship, or for refusal, enters the administrative record that an agency is legally obliged to consider. It is the module’s genre and the last portfolio piece.

19.12 Further Reading and Resources

Bowker, Geoffrey C., and Susan Leigh Star. 2000. Sorting Things Out: Classification and Its Consequences. MIT Press. https://mitpress.mit.edu/9780262522953/sorting-things-out/.
boyd, danah, and Kate Crawford. 2012. “Critical Questions for Big Data: Provocations for a Cultural, Technological, and Scholarly Phenomenon.” Information, Communication & Society 15 (5): 662–79.
Carroll, Stephanie Russo, Ibrahim Garba, Oscar L. Figueroa-Rodríguez, et al. 2020. “The CARE Principles for Indigenous Data Governance.” Data Science Journal 19: 43. https://doi.org/10.5334/dsj-2020-043.
D’Ignazio, Catherine, and Lauren F. Klein. 2020. Data Feminism. MIT Press. https://data-feminism.mitpress.mit.edu/.
Denton, Emily, Alex Hanna, Razvan Amironesei, Andrew Smart, and Hilary Nicole. 2021. “On the Genealogy of Machine Learning Datasets: A Critical History of ImageNet.” Big Data & Society 8 (2). https://doi.org/10.1177/20539517211035955.
Detroit Community Technology Project. 2019. A Critical Summary of Detroit’s Project Green Light and Its Greater Context. Detroit Community Technology Project. https://detroitcommunitytech.org/?q=content/critical-summary-detroits-project-green-light-and-its-greater-context.
Eubanks, Virginia. 2014. “Want to Predict the Future of Surveillance? Ask Poor Communities.” The American Prospect, January.
First Nations Information Governance Centre. n.d. The First Nations Principles of OCAP. First Nations Information Governance Centre. https://fnigc.ca/ocap-training/.
Garcia, Patricia, Tonia Sutherland, Marika Cifor, et al. 2022. No: Critical Refusal as Feminist Data Practice. Proceedings of CSCW 2022.
Keegan, Brian C. 2026. “Public Interest Data Infrastructuring.” Under Review.
Stop LAPD Spying Coalition. 2018. Before the Bullet Hits the Body: Dismantling Predictive Policing in Los Angeles. Stop LAPD Spying Coalition. https://stoplapdspying.org/before-the-bullet-hits-the-body-a-report-on-predictive-policing/.
Stop LAPD Spying Coalition, and Free Radicals. 2020. The Algorithmic Ecology: An Abolitionist Tool for Organizing Against Algorithms. Medium / Stop LAPD Spying Coalition. https://stoplapdspying.org/the-algorithmic-ecology-an-abolitionist-tool-for-organizing-against-algorithms/.
Taylor, Linnet. 2017. “What Is Data Justice? The Case for Connecting Digital Rights and Freedoms Globally.” Big Data & Society 4 (2).
Te Mana Raraunga. 2018. Principles of Māori Data Sovereignty. Māori Data Sovereignty Network. https://www.temanararaunga.maori.nz/tutohinga.
United States Indigenous Data Sovereignty Network. 2024. USIDSN: Promoting Indigenous Data Sovereignty in the United States. Native Nations Institute, University of Arizona. https://usindigenousdata.org/.
US Department of Housing and Urban Development. 2024. HMIS Data Standards Manual. HUD Exchange. https://www.hudexchange.info/resource/3824/hmis-data-dictionary/.
Wilkinson, Mark D., Michel Dumontier, IJsbrand Jan Aalbersberg, et al. 2016. “The FAIR Guiding Principles for Scientific Data Management and Stewardship.” Scientific Data 3: 160018.
Zong, Jonathan, and J. Nathan Matias. 2024. “Data Refusal from Below: A Framework for Resisting Institutional Data Harms.” ACM Journal on Responsible Computing, ahead of print. https://doi.org/10.1145/3630107.