2 The Installed Base
Chapter 1 argued that “the public interest” is contested and promised a working definition instead of a tidy one. This chapter delivers it, all at once. It gives you the six ideas the rest of the book develops one module at a time: three pressures that narrow public knowledge, three values that push back, the installed base where those values become durable, and five professional lineages that have built such bases before. It then explains why the book climbs from the city to the nation, and ends by building a small installed base in code. Treat it as a map you will keep unfolding. Each idea gets its own chapter later.
2.1 Three pressures
Keegan (2026) names three mutually reinforcing forces. Each answers a question.
Enclosure asks who can observe. It is the privatization or withdrawal of data that once held public or research value. Facebook restricted its APIs after Cambridge Analytica; Twitter revoked research access after its 2022 acquisition; Reddit ended Pushshift’s access and priced its API beyond most researchers. Freelon (2018) calls the result a “post-API age,” in which access can be revoked at any time, for any reason, without recourse. Governments enclose too, when they retire datasets on policing, health, or the environment. You meet enclosure at the city level in Chapter 4.
Exemption asks who must answer. It is the capacity of powerful firms and agencies to operate beyond meaningful scrutiny. Keegan distinguishes legal exemption (formal carve-outs, thin duties, jurisdictional gaps) from practical exemption (trade secrecy, restrictive terms of service, technical controls that make scrutiny infeasible). The result is a structural mismatch: the systems with the greatest public consequences face the weakest oversight (Pasquale 2016). You meet exemption at the county level in Chapter 8, and again at the state level, disguised as audit-washing, in Chapter 12.
Erosion asks what can be remembered. It is the decay of the institutions and infrastructures that make long-term observation possible: portals that vanish, contracts that lapse, corpora with silent holes (Gaffney and Matias 2018; Plantin et al. 2016). Erosion is rarely dramatic, which is what makes it effective. You meet it at the federal level in Chapter 16.
2.2 Three values
Each pressure has a counter-value.
Openness counters enclosure. It is the durable legibility of evidentiary traces: not only access to files, but provenance, documentation, and reproducible workflows that let more than experts evaluate how data were made. Chapter 6 develops it on city records.
Oversight counters exemption. It is routinized, testable scrutiny: audits, documentation standards, impact assessments, and institutions with authority to act on what they find. Chapter 10 treats oversight as remedy; Part IV treats it as independent verification.
Ownership counters erosion. It does not mean private property. It means collective, accountable governance over the preservation, access, and use of public-relevant data, with stewardship obligations that outlast any one grant or administration. Chapter 18 develops it, including the sovereignty claims that complicate it.
Values stated in a mission statement are cheap. They become durable only when they are built into something. That something is the installed base.
2.3 The installed base and its six elements
Star and Ruhleder (1994) used “installed base” to describe how new infrastructure never lands on a blank site. It inherits prior wiring, formats, routines, and labor, the way a neighborhood inherits its streets. Karasti and Blomberg (2018) call the ongoing work of keeping infrastructure aligned with what the installed base will tolerate infrastructuring, and Aanestad and colleagues (2017) show national health-IT programs failing, repeatedly, by ignoring it. Keegan (2026) extends the idea: for public interest work, the installed base is the socio-technical settlement that makes accountability possible. It has six elements.
Linkability. Records carry stable identifiers that let them be referenced and joined across time and systems. Linkability is also jurisdictional: someone must have the duty to assign references and carry them through reorganizations.
Interpretability. Categories, units, and transformations are documented well enough for a qualified outsider, including an affected party, to reason about them. Classification is never neutral (Bowker and Star 2000); interpretability lets you see the classification and argue with it.
Continuity. Evidence persists across political cycles, vendor changes, and budget gaps. It is a resourcing commitment more than a technical feature: retention schedules, mirrors, migration plans, and paid maintainers.
Safe scrutiny. Researchers, journalists, auditors, and communities can examine the system without being sued, arrested, or retaliated against. Tiered access belongs here: governed visibility that protects people without becoming an excuse for exemption.
Authority. A named office has the responsibility and power to act on what scrutiny reveals. An audit that lands on no one’s desk is a training exercise.
Remedy. The people affected have accessible, protected, consequential paths to contest decisions and correct the record. Remedy often has no software layer at all; it lives in appeals, ombuds offices, and organizing.
The order matters: you cannot offer remedy over records nobody can link. The elements also map onto the values. Openness operationalizes linkability, interpretability, and much of continuity. Oversight operationalizes safe scrutiny and authority. Ownership operationalizes remedy and the stewardship that keeps continuity from collapsing into somebody’s weekend hobby.
2.4 Five lineages
Data science did not invent the problem of making expertise answer to the public. Keegan (2026) identifies five professions that built public interest traditions, each by infrastructuring: turning a value into records, routines, and institutions.
Journalism maintains the evidentiary surface of public life through archives, verification, provenance, and publication (Chapter 5). Law translates values into enforceable procedures: rights to access, duties to disclose, standards of review, and remedies (Chapter 9). Accounting built a public interest mandate around independence and verification after audits failed the public (Baker 2005) (Chapter 12). Engineering made safety, testing, and incident learning into codes (Petroski 1992) (Chapter 13). The book teaches accounting and engineering together as assurance. Planning learned, through advocacy planning (Davidoff 1965) and the backlash against figures like Moses, that technical choices allocate burdens across space and time, and that affected people belong in the room (Chapter 17).
Every lineage also has an exclusionary past. Borrowing a method means borrowing its blind spots, and each module asks you to name them.
2.5 Why the book climbs the ladder
The four modules move from the city to the county, the state, and the federal government. The ladder is a teaching device, and it rests on the authority element: evidence only matters if it reaches a body that can act on it. Each lineage has a natural home on the ladder.
The city is where records are most abundant and most scattered: council packets posted as scanned PDFs, permit logs, campaign finance filings (Keegan 2026). Local news has been hollowed out, so the watchdog role is open, and an op-ed in a local outlet can be read by every council member. Journalism fits.
The county, in many US states, runs human services, child welfare, elections, and property records. Open-records statutes such as the Colorado Open Records Act give you an enforceable right to ask (Colorado General Assembly 1968), and county commissioners hold public hearings where you can testify. Law fits.
The state licenses professions, regulates employers and insurers, buys automated decision systems, and has written much of the AI legislation in the United States, including Colorado’s AI Act, SB 24-205 (Colorado General Assembly 2024). Legislative committees commission and read reports. Assurance fits.
The federal government holds the long-horizon datasets (the statistical system and the scientific agencies and laboratories, several of them in Boulder) and runs notice-and-comment rulemaking, a formal channel any member of the public can use (US Congress 1946). Erosion is most consequential there, and planning’s long time horizon fits.
The ladder also moves outward from where you live. You start where you can attend the meeting and build skills before the decision-makers get distant. Real problems cross levels, which is why the final project offers “a second level of government” as one way to deepen a piece. And the ladder is a US federal structure; each module’s non-US counter-case shows how differently other countries stack their authority.
2.6 Diagnosing a missing element
The framework earns its keep diagnostically. Take the Allegheny Family Screening Tool, which Chapter 8 and Chapter 10 treat in detail. Linkability is good: each hotline call is joined to administrative records. Interpretability is partial: features are documented, weights are harder to reason about (Vaithianathan et al. 2019). Continuity is maintained by named teams. The tail is weak. Safe scrutiny depends on the county’s willingness to cooperate with outside researchers. Authority sits with call screeners who can override the score but feel pressure not to. Remedy for a family flagged and investigated is almost nonexistent; you cannot unring a caseworker visit. The framework does not tell you what to do. It tells you where to look, so that an intervention lands somewhere that will hold it.
The installed base is where openness, oversight, and ownership stop being rhetorical. A city portal that earns a five-star “open data” rating can still fail openness in the sense this book uses, if its URLs break at every vendor rebid and nobody preserves the raw extracts, because continuity is part of legibility. Oversight fails when findings have no authority to land on, however good the audit. Ownership fails when the steward is a volunteer with no successor. When you see the three values named abstractly elsewhere, translate back to the six elements. That translation is the discipline this chapter teaches, and every portfolio piece’s installed-base note asks you to perform it.
2.7 A minimum viable installed base
Enough vocabulary. You cannot talk seriously about the six elements without building a version of them at least once. The dataset is a procurement log for a fictional small city: 500 contract awards with two deliberate failures. Some rows are missing identifiers, and some amounts went through an undocumented currency conversion. Generate it yourself so the tutorial runs anywhere:
import numpy as np
import pandas as pd
rng = np.random.default_rng(4871)
n = 500
df = pd.DataFrame({
"record_id": [f"SC-{i:05d}" for i in range(n)],
"vendor": rng.choice(["Acme Paving", "Front Range IT", "Rhein GmbH",
"Liffey Consulting", "Flatirons Print"], n),
"amount": rng.lognormal(10, 1, n).round(2),
"award_date": pd.date_range("2021-01-01", periods=n, freq="3D"),
"award_type": rng.choice(["bid", "sole_source", "emergency"], n),
"contact_email": [f"vendor{i}@example.com" for i in range(n)],
"contract_reference": [f"C-{2021 + i // 125}-{i:04d}" for i in range(n)],
})
df.loc[rng.choice(n, 28, replace=False), "record_id"] = None
df["record_id"].isna().sum()
# => 282.7.1 Linkability: persistent identifiers
Two reasonable ways to fill the 28 gaps: a random UUID, or a hash of the record’s own content. A UUID is opaque and survives corrections; a content hash detects silent changes but shifts the moment anyone fixes a typo. For civic records, where correction is expected, use both. The UUID is the canonical identifier; the content hash is a side channel you can log over time as a change history.
import hashlib
import uuid
def content_hash(row):
payload = f"{row['vendor']}|{row['amount']}|{row['award_date']}|{row['award_type']}"
return hashlib.sha256(payload.encode("utf-8")).hexdigest()[:16]
missing = df["record_id"].isna()
df.loc[missing, "record_id"] = [str(uuid.uuid4()) for _ in range(missing.sum())]
df["content_hash"] = df.apply(content_hash, axis=1)
df["record_id"].is_unique
# => TrueUUIDs are the easy part. Every language ships a generator. The hard part is durable reference, and durable reference is social work. Somebody has to maintain the resolver (the service that turns an identifier into a record), hold the naming authority (who may mint identifiers), and make the migration commitment (when systems change, identifiers move with them). DOIs work because Crossref and DataCite exist as institutions, not because the string 10.1234/... is clever. Your procurement log’s identifiers will work if, and only if, some office commits to carrying them across the next three vendor contracts. If nobody has signed up, you have a UUID, not an identifier.
2.7.2 Interpretability: a provenance document
In the real version of this log, the rows from Rhein GmbH and Liffey Consulting arrived in euros and were converted to dollars by someone who never wrote it down. An auditor who assumes the source was always in dollars will misreport. The fix is not code. It is a provenance document committed next to the data, readable by people and machines.
# data/provenance.yaml
dataset: smallcity_procurement
version: "2.0"
upstream_source:
name: Small City Department of Finance, ERP export
accessed: 2026-03-12
transformations:
- step: 01_extract
description: CSV export from the accounts-payable module.
- step: 02_currency_normalize
description: >
Rows for vendors in DE and IE carried EUR amounts. Converted to USD
using the ECB monthly reference rate for the award_date month.
rate_source: https://www.ecb.europa.eu/stats/policy_and_exchange_rates/
- step: 03_identifier_assignment
description: UUID minted for 28 rows with missing record_id; content_hash added.
known_issues:
- Monthly average rates; daily rates not preserved.2.7.3 Continuity: a retention policy
Continuity without a retention policy is an act of faith. Write it down in a form that survives its author leaving the job.
# data/retention.yaml
schedule:
raw_extracts: {retain_for: 10_years, storage: city_records_archive}
published_versions: {retain_for: indefinite, storage: zenodo_plus_ia_mirror}
derivation_scripts: {retain_for: indefinite, storage: city_git}
access_logs: {retain_for: 2_years, purpose: audit_restricted_queries}
responsible_office: Department of Finance, Data Stewardship Lead
review_cadence: annual2.7.5 Remedy: an escalation log
Access without escalation is a suggestion box nobody empties. The last file says who is notified when something happens, and how fast they must respond. Keep it boring.
# data/escalation.yaml
routes:
- trigger: correction request received
notify: [records_officer, affected_department]
sla: 10_business_days
remedy: issue corrected row, log in change_history
- trigger: removal request (privacy or safety)
notify: [city_attorney, records_officer]
sla: 5_business_days
remedy: redaction, tombstone record retained
- trigger: restricted view accessed
notify: [records_officer, audit_log]2.8 What this demonstrates
In about a hundred lines you built a minimum viable installed base: identifiers for linkability, a provenance file for interpretability, a retention policy for continuity, tiered views for safe scrutiny, and an escalation log naming the offices that hold authority and owe remedy. It is not a dashboard, a model, or an “AI governance” framework. It is the plumbing under all of them, and Keegan (2026) calls attending to it the infrastructuring move: treating what looks like a platform problem as an installed-base problem. Every portfolio piece in this book ends with an installed-base note that asks which of these files your work produced and which are still missing.
2.9 Exercises
Exercise 2.1 (Guided). Run the tutorial end to end and save the result as exercises/ch02_procurement_v2.csv. In 150 words, explain why the tutorial makes the UUID, not the content hash, the canonical record_id, and describe one situation where you would choose the opposite.
Exercise 2.2 (Guided). Extend provenance.yaml with one transformation step of your own (vendor-name normalization, deduplication) and add a contact field naming a responsible office. Then write retention.yaml for a city or organization you know, citing one real records-retention schedule (your city’s, your state archives’, or a federal agency’s) to justify at least one retention period.
Exercise 2.3 (Applied). Using the database from the tutorial, write three queries: total by vendor, lookup by contract_reference, and a full row export. For each, say which tier can run it and why. Save as exercises/ch02_access.md.
Exercise 2.4 (Diagnostic, real public data). Pick one dataset from a real open data portal: the City of Boulder’s, Denver’s, or the State of Colorado’s at data.colorado.gov. The Colorado portal runs on Socrata, which exposes dataset metadata at /api/views/<dataset-id>.json; city portals built on other platforms have their own metadata pages. Fetch the metadata with requests and record the license, the last-updated timestamp, the publishing department, and how many columns have descriptions. Then score the dataset on all six installed-base elements (0 to 2 each), with one sentence of evidence per score. Which element would you fix first, and which office has the authority to fix it? 500 words plus your script.
Exercise 2.5 (Open). Design an escalation pathway and a threat model for a small civic dataset of your choosing (school board minutes, a housing-permit feed, a campus crime log). Name at least three triggers, the offices notified, and a service-level agreement for each, plus three adversaries (not all malicious; an underfunded vendor counts) and one failure mode for each. 800 to 1,200 words. You will return to this pattern when you build a remedy ledger in Chapter 10.
2.10 Looking ahead
Chapter 3 turns from the framework to the neighbors: public interest technology, data for good, digital government, civic tech, critical data studies, and data justice. Each claims part of this territory. The chapter asks what each one builds, what it leaves unbuilt, and how a project you admire might score on the six elements five years after launch.
2.11 Further Reading and Resources
- Brian C. Keegan (2026), “Public interest data infrastructuring,” under review (Keegan 2026). Read the introduction through the section titled “Public interest data infrastructuring” for this chapter.
- Susan Leigh Star and Karen Ruhleder (1994), “Steps toward an ecology of infrastructure,” CSCW 1994, 253–264 (Star and Ruhleder 1994). The original “installed base” argument.
- Helena Karasti and Jeanette Blomberg (2018), “Studying infrastructuring ethnographically,” Computer Supported Cooperative Work 27: 233–265 (Karasti and Blomberg 2018). Infrastructuring as alignment work.
- Jean-Christophe Plantin, Carl Lagoze, Paul N. Edwards, and Christian Sandvig, “Infrastructure studies meet platform studies in the age of Google and Facebook,” New Media & Society (Plantin et al. 2016). The course’s supplementary reading for week 2.
- Deen Freelon (2018), “Computational research in the post-API age,” Political Communication 35(4): 665–668 (Freelon 2018). Two pages that name the condition of enclosure.
- Geoffrey C. Bowker and Susan Leigh Star (2000), Sorting Things Out (Bowker and Star 2000). Indispensable for interpretability.
- Margunn Aanestad and colleagues (2017), “Information infrastructures and the challenge of the installed base” (Aanestad et al. 2017). Why ignoring the installed base is not a niche failure.
- The DataCite metadata schema: https://schema.datacite.org. A production-grade persistent-identifier standard backed by an institutional community.
- Socrata Open Data API documentation: https://dev.socrata.com/. The API behind many US city and state open data portals, including the one in Exercise 2.4.