Public Interest Data Science

Lineages, Levels, and Genres for Accountable Data Work

A textbook on doing data science for public-interest ends: diagnosing the structural pressures on public knowledge, articulating the values that push back, learning from the public-interest professions, and contributing to public conversations through the genres that make analysis accountable.

Author

Brian C. Keegan

Published

January 1, 2026

Preface

This is a book about doing data science for ends the public can name, contest, and use.

You already know, or will shortly, that most data science is not that. It produces things for an employer to ship, a platform to monetize, or a manager to track. When it produces things for the public, it tends to produce dashboards: objects that look like accountability without being accountable to anyone in particular. This book is written for the student who has noticed the gap between what data science can do and what it typically does, and who wants a way in that is neither naive about the tools nor resigned about the politics.

Who this book is for

If you are a student taking a data-science course, a journalism or policy student learning computational methods, a working practitioner considering a move into civic or public-interest work, an instructor looking for a course companion, or a public-sector analyst being asked to do “data-driven” work with no time to read the entire literature, this book is for you. It assumes introductory Python, comfort with the command line, and a willingness to read ten pages about a court case before writing the function that acts on it. It does not assume a background in science and technology studies, law, or accounting. It will give you enough of each to keep going.

You can read it straight through, pick chapters relevant to a specific project, or use it as a course text. It is the companion to INFO 4871/5871 Public Interest Data Science at the University of Colorado Boulder, and its parts follow that course’s modules: a two-week opening, four three-week modules that each end in a portfolio piece, and a closing week before the final project.

What the book does

The book rests on a framework from Keegan (2026). Three structural pressures narrow the public’s capacity to observe and govern data-intensive systems: enclosure (who can observe?), exemption (who must answer?), and erosion (what can be remembered?). Three values push back: openness, oversight, and ownership. Values last only when they are built into an installed base with six elements: linkability, interpretability, continuity, safe scrutiny, authority, and remedy.

The book teaches that framework through the professions that have already solved versions of the problem, and it climbs the ladder of government as it goes. It has six parts:

  1. Foundations (1  Introduction to Public Interest Data Science to 3  Public Interest Technology and Data for Good). The contested idea of “the public interest,” the framework and its installed base, and the neighboring fields (public interest technology, data for good, civic tech) that clarify what this work is and is not.
  2. Journalism: The City (4  Enclosure: Who Can Observe? to 7  Op-Eds). Enclosure meets the watchdog tradition. You reproduce a published investigation, collect city data responsibly, publish it openly, and write an op-ed for a local outlet.
  3. Law: The County (8  Exemption: Who Must Answer? to 11  Testimony). Exemption meets public interest law. You work with records Boulder County has already released under the Colorado Open Records Act, turn them into data, build a ledger of exemptions and remedies, and write testimony for a county hearing.
  4. Assurance: The State (12  Accounting and Auditing to 15  Reports). Audit-washing meets accounting and engineering. You run a disaggregated audit, test it like safety-critical code, examine procurement as governance, and write a report for a state body.
  5. Planning: The Nation (16  Erosion: What Can Be Remembered? to 20  Public Comments). Erosion meets advocacy planning, stewardship, sovereignty, and refusal. You audit the decay of a federal dataset, design a stewardship plan or a refusal specification, and write a public comment on a federal rulemaking.
  6. Beyond the Portfolio (21  Research Proposals to 23  The Limits of the Framework). Research proposals, archival deposits with DOIs, and the limits of the framework itself.

Each module part follows the same order: a pressure, a lineage, a value, and a genre. Each uses a national anchor case, a non-US counter-case, and a local case at its level of government. Each ends with a portfolio piece.

How the book is built

Every chapter interweaves conceptual narrative with runnable Python tutorials at roughly a 50/50 balance. You will read two or three pages of argument, then read and run two or three pages of code that makes the argument visible in a specific case, then return to the argument. This is deliberate. Concepts without technique produce earnest seminars. Technique without concepts produces dashboards that nobody trusts.

Code blocks do not run at build time. They are illustrative: you are expected to type them, adapt them, and run them yourself. The book’s repository contains test fixtures, sample data, and the occasional larger notebook that a chapter only excerpts. When a chapter uses a real API (Wikimedia pageviews, regulations.gov, the Internet Archive), the endpoints cited were live as of the book’s most recent release; if you find one that has moved, open an issue.

Every chapter has at least one “Missing Manual” callout flagging something that textbooks usually omit: a surprising library default, a real-world API friction, the politics of a “simple” choice. Every chapter has at least one “In the Public Interest” callout tying the technical material back to openness, oversight, or ownership. Every chapter ends with 4–5 graduated exercises, from guided (“follow these steps”) to open-ended (“design a pipeline that…”), and a curated list of further resources.

The genre chapter that closes each module part has one additional rule: it must end with a real, submittable portfolio piece. Not a thought experiment. A technical artifact someone else can run, a public text (an op-ed, a testimony statement, a report section, or a public comment on a live rulemaking) addressed to a named audience, and a short note on what installed base the work built. A reader who finishes the four module parts has a portfolio of four pieces at four levels of government. Part VI shows how to take one of them further and archive it with a DOI.

Running examples

To give the book coherence across 23 chapters, a small set of running examples is threaded throughout, arranged by the level of government each part works at. When a chapter needs a case, it reaches for one of these before inventing a new one:

  • City: City of Boulder and City and County of Denver records (campaign finance, meeting minutes, agendas, budgets, permits, 311 requests).
  • County: Boulder County records already released under the Colorado Open Records Act (from the Public Records Archive of its Open Records Center), and the county’s public meeting packets.
  • State: Colorado data from the American Community Survey (via folktables) and the Colorado AI Act as a state accountability regime.
  • Federal: federal environmental and scientific data (EPA, NOAA, USGS, Census) as at-risk datasets, and regulations.gov for federal rulemaking.
  • Census and FRED APIs for demographic and economic context.
  • Wikipedia and Internet Archive APIs for archival work and link-rot audits.
  • Hugging Face model cards and datasheets for audit and documentation exercises.
  • A small synthetic worker-log fixture that lets us teach audit methods without violating anyone’s privacy.
  • Public FOIA and CORA request data as a class-wide shared resource.

By the end of the book you will have built, in parts, a small public-interest data infrastructure of your own: a reproduced investigation and a documented city dataset with an op-ed; county records turned into data with a ledger of exemptions and remedies, and testimony; a tested, disaggregated audit with a report for a state body; a federal dataset’s erosion audit, a stewardship plan or refusal specification, and a public comment on a live rulemaking; and, finally, one of those pieces deepened, recast, and deposited on Zenodo with a DOI.

What the book does not do

Honesty up front. This book is partial in several ways that matter.

It is US-centric. Most of the legal and administrative infrastructure discussed is American: FOIA and CORA, the Administrative Procedure Act, NIST, NSF, NYC Local Law 144. Where European, Latin American, and Global South counterparts exist (the Dutch SyRI case, IndiaStack, Ushahidi, the CARE Principles, Te Mana Raraunga), the book names them and points you to further reading, but it does not pretend to do them justice in three pages.

It is comfortable with records and less comfortable with refusal. The installed-base framework takes linkability, durable identifiers, and retention as load-bearing. Chapter 19 takes that comfort seriously as a limit, drawing on Zong and Matias, D’Ignazio and Klein, the CARE Principles, and Indigenous data sovereignty scholarship. Readers who want to start from refusal and arrive at records (rather than the other way around) should read Chapter 19 early and often.

It is thin on the labor inside data work. Gray and Suri’s Ghost Work is cited, and Chapter 3 tries to name invisible maintenance labor explicitly, but a book that seriously treated data-labor economics would look quite different. Consider this an invitation to the reader, or to the next edition.

It is thin on measurement theory. You will not find a full psychometric critique of the audit metrics in Chapter 12, and you will not find a deep philosophy-of-science treatment of what it means for a model to “work.” Those books exist and are cited.

It over-relies on the professional-lineage framing. Treating journalism, law, accounting, planning, and engineering as lineages the field can inherit from is a useful move, but every profession has its own pathologies, and the book does not do equal-time analysis of all of them. Read the criticisms. They are in the further-reading lists.

These limits are not bugs to be patched. They are framing choices that enable the book to cohere. Knowing where they bind is part of what you are learning, and 23  The Limits of the Framework turns the framework on itself to find out.

How to contribute

This book lives in a git repository. Fork it, improve it, and open a pull request. Before you do, read claude.md, which codifies the voice, structure, and citation conventions. Typo fixes and small corrections can come in as pull requests directly. Rewrites of sections, new examples, or proposed new chapters should start as an issue so we can talk about them first.

Students in the companion course make one chapter-scale contribution each term: a new or rewritten section with narrative, a runnable tutorial, verified references, and curated resources. Students whose contributions shape the book are named in the acknowledgments. Some of the best chapters in future editions will be written by readers who encountered this one and saw what was missing.

A note on the AI-assisted drafts

The first complete draft of this book was produced with substantial assistance from Claude, Anthropic’s assistant, following a detailed editorial prompt written by the author. The 2027 restructure into module parts, including the new chapter on the limits of the framework, was also drafted with Claude’s assistance from the author’s course design and is under author review. The disclosure statement is in Appendix A. A book about public-interest data science should not hide how it was made, any more than a model card should hide its training data.

Acknowledgments

The framework owes its shape to years of conversation with colleagues in the INFO department at CU Boulder, students in previous versions of the course who asked the questions this book tries to answer, and the infrastructure studies and STS communities whose citations fill the bibliography. Specific individuals will be named once the revision cycle is complete. Any errors are the author’s, and there will be errors. File an issue.

Now to Chapter 1.

Keegan, Brian C. 2026. “Public Interest Data Infrastructuring.” Under Review.