1 Introduction to Public Interest Data Science
Open a newspaper any week and you will find a story about a data-driven system that has shaped a public decision in a way the public did not choose. A county’s child-welfare office screens families with an algorithm. A city’s benefits portal denies claims for reasons the claimants cannot discover. A federal agency’s climate database is quietly reorganized and three years of comparison data vanish. These are not exotic cases. They are ordinary governance in the 2020s.
This book is about the practice that takes those systems seriously as objects of public concern. Not as opportunities to build a better dashboard, though dashboards have their place. Not as abstract ethics problems, though ethics are unavoidable. As infrastructure: settled arrangements of data, code, authority, and labor that shape what the public can know and what it can demand. Public interest data science is the work of making that infrastructure observable, contestable, and, where necessary, reshaped.
1.1 What does “public interest” even mean
The phrase “public interest” has an emancipatory history and an exclusionary one, and they are often the same history told from opposite sides of a highway.
Start with the highway. Robert Moses spent four decades building parks, bridges, and expressways in and around New York City, and he justified nearly all of it in the name of the public. Caro (1974) documents what that justification covered, including the Cross Bronx Expressway’s path through the East Tremont neighborhood, where the route displaced a working-class community that had asked for a modest alternative alignment. Jane Jacobs (1961) answered with a different public interest: the sidewalk, the mixed-use block, the residents already living where the planners drew their lines. Moses’s public was the motorist and the regional economy; Jacobs’s was the people on the block (Dory 2018). Both said “the public.”
Now cross an ocean and half a century. India’s Aadhaar program has issued a twelve-digit identity number, tied to fingerprints and iris scans, to more than a billion residents. Its advocates framed it as inclusion: an identity for people who had no documents, and an end to “leakage” in welfare programs where benefits went to duplicate or fictitious beneficiaries. Its critics, using the same vocabulary, documented people turned away from ration shops when fingerprint authentication failed, and warned that a single number linking every transaction would make a population legible to the state on terms the population never negotiated (Rao and Nair 2019). In 2017 the Supreme Court of India recognized privacy as a fundamental right in Justice K.S. Puttaswamy (Retd.) v. Union of India, and in 2018 the Court upheld Aadhaar for welfare delivery while striking down the provision that let private companies require it. Both sides were arguing about who counts as the public, and what a data system owes the people it records.
That is why this book does not open with a tidy definition. Washington and Cheung (2024) go to history (public interest law, curb cuts, shared space) for what a public interest in technology could mean. Stapleton and colleagues (2022) ask the sharper question in their title: who has an interest in “public interest technology,” and who is left out when technologists working with local governments speak for impacted communities? “Public” is a word that asks: which public? Whose interests? Who was in the room?
The book does take a position, and it names the position explicitly. Following Keegan (2026), it treats the public interest as operational rather than abstract. The public interest is served when the six elements of what Chapter 2 calls the installed base are maintained. Records can be linked. Categories can be interpreted. Evidence persists across political cycles. Researchers and journalists can scrutinize systems without being sued or arrested. Someone has authority to act on what is learned. And the people affected have a path to remedy. When those six elements are in place, “public interest” is a working arrangement rather than a slogan. When they are missing, even the most benevolent data practice can produce what Eubanks (2018) calls automating inequality. Aadhaar is a useful test: it is superb at linkability and far weaker at remedy for the person whose fingerprint will not scan.
That is a particular theory, not the only one. Chapter 19 takes seriously the traditions (Indigenous data sovereignty, abolitionist critique, feminist refusal) that question whether “records” and “linkability” are the right starting places at all, and Chapter 23 returns to the framework’s limits. But you have to start somewhere.
1.2 Relative to neighbors
Public interest data science overlaps with several fields you may already know. Data journalism investigates public questions computationally and publishes in the press; Chapter 5 treats journalism as a lineage. Public interest technology and data for good build tools and talent pipelines for government and civil society, usually on philanthropic money; Chapter 3 sorts them out. Digital government and civic tech build inside and around the state; Chapter 14 returns to them when procurement becomes the subject. Critical data studies and data justice theorize the politics of data and, at their sharpest, refuse it; Chapter 19 argues that your practice is stronger when it can absorb that critique. Public interest data science borrows from all of them.
1.3 The arc of the book
The book has six parts. The first and last frame the work. The middle four are modules, each built on one professional lineage, set at one level of government, and ending in one piece of public writing.
Part I: Foundations (this chapter, Chapter 2, and Chapter 3) gives you the whole vocabulary at once: three pressures, three values, the installed base, five lineages, and the neighboring fields.
The four modules climb a ladder of government, from the city to the nation. Each one follows the same order: a pressure that narrows public knowledge, a lineage that has already solved a version of the problem, a value that pushes back, and a genre for reaching the public.
| Part | Lineage | Level | Pressure → value | Genre |
|---|---|---|---|---|
| II. Journalism: The City | Journalism | City | Enclosure → openness | Op-ed |
| III. Law: The County | Law | County | Exemption → oversight as remedy | Testimony |
| IV. Assurance: The State | Accounting and engineering | State | Exemption (audit-washing) → oversight as verification | Report |
| V. Planning: The Nation | Planning | Federal | Erosion → ownership | Public comment |
Part VI: Beyond the Portfolio closes the book with research proposals (Chapter 21), archival deposits (Chapter 22), and the limits of the whole framework (Chapter 23).
Each module uses a new case rather than one topic carried all term. Each pairs a well-documented national anchor case with a non-US counter-case that tests whether the framework travels, and a local case at the module’s level of government: city records in Boulder and Denver, county records Boulder County has already released under the Colorado Open Records Act, Colorado’s state AI law and procurement, and a federal dataset at risk alongside a live federal rulemaking. Chapter 2 explains why the lineages map onto those levels in that order.
1.4 A first request
Enough about the book. Do something with it. The rest of this chapter is a twenty-minute tour of the toolkit, built around a first public data retrieval; it is also the course’s first lab. The Environment setup sidebar below lists what you need.
You need Python 3.11 or newer, git, and a working command line. On macOS and Linux, python3 --version and git --version should return something reasonable. On Windows, WSL2 works well, as does the official Python installer with the “Add Python to PATH” box ticked. You will also want pip install pandas requests matplotlib jupyter before continuing, and a GitHub account if you do not have one already.
You will retrieve pageview data for the Wikipedia article “Public interest” in three language editions and see how the concept’s visibility varies across linguistic publics. The example is small, but it installs habits that recur: fetching from a real public API, handling the response carefully, moving it into a pandas DataFrame, plotting it, and asking what the data can and cannot tell you.
1.5 The Wikimedia pageviews API
Wikimedia Foundation operates a pageviews API that returns view counts for any article on any of their projects, back to 2015. The endpoints are documented at wikimedia.org/api/rest_v1. The relevant endpoint takes the form:
GET /metrics/pageviews/per-article/{project}/{access}/{agent}/{article}/{granularity}/{start}/{end}
You plug in the project (e.g., en.wikipedia), access method (all-access is fine), agent type (user filters out bots and spiders, which you want), article title, and a start and end date in YYYYMMDD or YYYYMMDDHH format.
Here is a first retrieval:
import requests
HEADERS = {
"User-Agent": "PublicInterestDataScience/0.1 (your-email@example.edu)",
"Accept": "application/json",
}
url = (
"https://wikimedia.org/api/rest_v1/metrics/pageviews/per-article/"
"en.wikipedia/all-access/user/Public_interest/monthly/"
"2020010100/2024123100"
)
response = requests.get(url, headers=HEADERS, timeout=30)
response.raise_for_status()
data = response.json()
print(data["items"][:2])
# => [{'project': 'en.wikipedia', 'article': 'Public_interest',
# 'granularity': 'monthly', 'timestamp': '2020010100',
# 'access': 'all-access', 'agent': 'user', 'views': 8442}, ...]You did three things there that you will keep doing: you sent a User-Agent header naming yourself, called raise_for_status() to turn HTTP errors into exceptions, and set a timeout so a hung server does not hang your script. Chapter 4 turns these habits into a full “responsible scraper” function. For now, notice that requests.get() does not hand you the response body directly. It hands you a Response object, and .json() is one of several methods you can call on it.
Most tutorials teach you requests.get(url).json() as a single chained call. That works until it does not. The Response object carries the status code, headers, cookies, and error context separately from the body. When a request fails (and it will), you need the status code and the headers to understand why; calling .json() on a 403 will give you an exception that hides the 403. Write your code so the Response survives long enough to inspect. You will thank yourself the first time an API silently starts returning an HTML error page in place of JSON.
1.6 Into pandas
A single JSON response is interesting. Three parallel time series for three language editions are more so. Collect “Public interest” pageviews for English, Spanish, and Hindi Wikipedia and plot them side by side.
import pandas as pd
def fetch_pageviews(project, article, start="2020010100", end="2024123100"):
url = (
f"https://wikimedia.org/api/rest_v1/metrics/pageviews/per-article/"
f"{project}/all-access/user/{article}/monthly/{start}/{end}"
)
response = requests.get(url, headers=HEADERS, timeout=30)
response.raise_for_status()
items = response.json()["items"]
df = pd.DataFrame(items)
df["timestamp"] = pd.to_datetime(df["timestamp"], format="%Y%m%d%H")
return df[["timestamp", "project", "views"]]
en = fetch_pageviews("en.wikipedia", "Public_interest")
es = fetch_pageviews("es.wikipedia", "Interés_público")
hi = fetch_pageviews("hi.wikipedia", "लोक_हित")
combined = pd.concat([en, es, hi], ignore_index=True)
combined.groupby("project")["views"].describe()
# => project-level descriptive stats across 60 monthsAlready the choices are political. The English article is titled “Public interest.” The Spanish equivalent is “Interés público.” The Hindi article is “लोक हित,” which translates roughly as “people’s good” rather than “interest.” These are not translations in the sense a machine would call translations; they are conceptually distinct articles that Wikipedia’s interlanguage links happen to join. The fact that all three exist tells you something. The fact that their pageview trajectories differ tells you something else.
Plot them:
import matplotlib.pyplot as plt
fig, ax = plt.subplots(figsize=(10, 4))
for project, group in combined.groupby("project"):
ax.plot(group["timestamp"], group["views"], label=project)
ax.set_xlabel("Month")
ax.set_ylabel("Views")
ax.set_title("Wikipedia pageviews: 'Public interest' across three editions")
ax.legend()
fig.tight_layout()
fig.savefig("ch01_pageviews.png", dpi=120)If your plot looks reasonable, you have just run an end-to-end public interest data pipeline. Not a large one, but a real one: it fetches from a public API, handles errors, reshapes JSON into a DataFrame, and produces a figure legible in a campus paper or a policy brief.
1.7 Reading the plot
Three things to notice, because they recur.
First, the series do not share a calendar. Look for months where the English series spikes and the Spanish or Hindi series does not, and ask what was in the news in each place that month. A single global “public interest” is a fiction; there are many publics, and their interests peak at different times.
Second, all three series are small relative to any major news article. “Public interest” is a reference concept, not a subject of mass curiosity. The audience for public interest work is small but reachable, if you know where it is.
Third, the data will drift. Wikipedia articles get renamed, redirected, merged. The endpoint you called today may return a different trajectory next year because the article itself has changed. Chapter 16 takes up the problem of what can be remembered; this exercise is your first tiny encounter with it.
Three line plots of Wikipedia pageviews look like a trifle. Notice what they required. A public API, documented and open. Endpoints that return consistent, structured data. A retention policy that preserves history for at least a decade. A User-Agent expectation that asks you to identify yourself rather than hiding. Wikipedia built these as deliberate commitments to openness (Chapter 6). If you were trying to do the same exercise against a commercial platform in 2026, you would find most of those commitments have been withdrawn. The pressure of enclosure (Chapter 4) makes the simple thing you just did unusually precious. Part of this book’s work is to help you notice when the simple thing is rare.
1.8 What you will build
The course this book accompanies is graded on a portfolio, not on exams. Each of the four modules ends in one portfolio piece, and every piece has the same four parts: a technical artifact (a notebook or small repository a classmate can clone and run), a public text in the module’s genre for a named audience at the module’s level of government, a short installed-base note saying which pressure you met and which of the six elements your work built or left missing, and, for graduate students, a methods memo that situates the piece in the research literature.
- Piece 1, Journalism and the city. Reproduce a published finding, such as ProPublica’s “Machine Bias” analysis, extend it with city data you collected responsibly, and write a 750-word op-ed pitched to a named local outlet (Chapter 4 through Chapter 7).
- Piece 2, Law and the county. Choose open-records requests the county has already answered, turn the released records into a table, keep a ledger of exemptions and remedies, and deliver written and oral testimony for a county commission hearing (Chapter 8 through Chapter 11).
- Piece 3, Assurance and the state. Run a disaggregated audit with a manifest, thresholds, and tests that run in continuous integration, and write a report section for a named state body (Chapter 12 through Chapter 15).
- Piece 4, Planning and the nation. Audit a federal dataset for erosion, write a stewardship plan or a refusal specification, and submit a substantive comment on a live federal rulemaking (Chapter 16 through Chapter 20).
The final project takes one piece further in three moves. You deepen it (new data, a second level of government, or a comparison with the module’s non-US counter-case). You recast it in a different genre, so the audit becomes an op-ed or the county records become a report; graduate students may write a research proposal (Chapter 21). And you archive the artifact and text with a DOI minted through Zenodo or a similar repository (Chapter 22), alongside a short memo on whose interests the work serves and where the framework fails for your case.
One more thing runs through the term: a chapter-scale pull request to this book, which is a public repository and is wrong in places you will find before the author does.
A portfolio, in other words: things you can point to when you apply for a job, a fellowship, or a seat on a commission, and say: I did this, I published it, here is the permanent link.
1.9 Exercises
Exercise 1.1 (Guided). Install Python 3.11+, git, and the packages listed in the environment setup callout. Clone the course repository. Create a branch named yourname/intro. Add a file introductions/yourname.md with three sentences: who you are, a public you belong to, and a level of government whose decisions you have felt directly. Commit with the message intro: yourname and open a pull request. The content matters less than the mechanism. Branch, commit, and pull request are your default mode from here forward, including for your contribution to this book.
Exercise 1.2 (Guided, real public data). Adapt the pageviews script to retrieve monthly views for three Wikipedia articles you care about, across three language editions each. Produce one plot comparing them. Write 150 words interpreting it. Save the plot and writeup to exercises/ch01_pageviews/. Record one choice (article, language, or date range) you had to make, and why.
Exercise 1.3 (Reflective). Write a 300-word “data obituary” for an API, dataset, or data portal that has been withdrawn, paywalled, or substantially restricted. Reddit’s Pushshift is a classic candidate. Name the source, say what it was useful for, describe what happened to it, and identify who has been worse off since. Cite two sources.
Exercise 1.4 (Analytic). Find a recent public document from a city, county, state, or federal body in which someone invokes “the public interest” (a council resolution, a court filing, an agency press release, a records denial). Write 300 words answering: which public does the document mean, who is left out, and what would have to be on the record for an outsider to check the claim? Compare your answer with the Moses or Aadhaar case.
Exercise 1.5 (Open-ended). What is a public interest data question you want to answer by the end of the term? Write two to three paragraphs: the question, who would benefit, what data would be needed, which level of government could act on the answer, and what obstacles you already see. Save it as exercises/ch01_project_sketch.md. You will revisit it as each module gives you new tools, and it may become the seed of your final project.
1.10 Looking ahead
Chapter 2 gives you the full vocabulary in one sitting: the three pressures of enclosure, exemption, and erosion; the three values of openness, oversight, and ownership; the six elements of the installed base; and the five lineages the modules borrow from. It also explains why the book climbs from the city to the nation. Then you will build a minimum viable installed base for a small civic dataset, because the vocabulary is only useful if you can see it in a file.
1.11 Further Reading and Resources
- Brian C. Keegan (2026), “Public interest data infrastructuring,” under review (Keegan 2026). The paper this book extends. Read the introduction now and the rest across the term.
- Anne L. Washington and Joanne Cheung (2024), “Towards defining the public interest in technology: Lessons from history,” Journal of Integrated Global STEM 1(2): 67–74 (Washington and Cheung 2024). The course’s framing reading for week 1.
- Logan Stapleton and colleagues (2022), “Who has an interest in ‘public interest technology’?: Critical questions for working with local governments and impacted communities,” CSCW ’22 Companion, 282–286 (Stapleton et al. 2022). The critical counterpart.
- Robert A. Caro (1974), The Power Broker: Robert Moses and the Fall of New York (Caro 1974). The chapter titled “One Mile” is the Cross Bronx Expressway story.
- Ursula Rao and Vijayanka Nair (2019), “Aadhaar: Governing with biometrics,” South Asia: Journal of South Asian Studies 42(3): 469–481 (Rao and Nair 2019). An entry point to the scholarship on Aadhaar.
- Virginia Eubanks (2018), Automating Inequality (Eubanks 2018). How data-driven systems reshape welfare, policing, and child services.
- Mike Ananny and Kate Crawford (2018), “Seeing without knowing,” New Media & Society 20(3): 973–989. A counterweight to the assumption that transparency is self-executing.
- Wikimedia Foundation’s REST API documentation: https://wikimedia.org/api/rest_v1/. The endpoint you used in this chapter, plus siblings you will find handy later.