6  Openness

Chapter 4 named the pressure: who can observe? This chapter builds the value that answers it. Chapter 2 argued that the six elements of the installed base are the conditions under which “public interest” can be said to be served by a data system rather than invoked on its behalf. Openness takes three of those elements (linkability, interpretability, and the openness-facing part of continuity) and turns them into a practice. The practice has a name most readers will already have heard. It is called open data, and the name is doing too much rhetorical work.

Open data as a programmatic idea is roughly twenty years old. It has produced real goods: the federal data.gov portal, municipal 311 feeds, campaign finance extracts, the census API. It has also produced what practitioners call, with a mixture of resignation and contempt, the data dump: a civic file dropped onto a portal with a permissive license, thin documentation, inconsistent columns, and no plan for next year. It is technically open. It satisfies the ordinance. For any public that does not already work inside the agency, it is close to unusable.

The distinction this chapter insists on is between formal openness and durable legibility. Formal openness is what a compliance officer can certify: the file exists at a URL, the license permits redistribution, the API returns a 200. Durable legibility is what a non-expert can actually do with the thing six months from now: load it, understand its columns, cite it, compare it to last year’s version. Formal openness without legibility quietly lets an agency off the hook. It has disclosed, and nobody outside can use the disclosure to ask it a question. (Chapter 8 gives that failure its proper name.)

6.1 Posting is not publishing

An analogy with scholarly publication helps. A PDF uploaded to a webserver with no metadata is not a publication, though it meets some naive definition of “posted.” A publication is a document, plus a record about the document (title, author, date, version, license), plus a citable handle, plus a commitment from some institution that the handle will keep resolving. Publication is a social act, not a file-system operation.

The same holds for data. df.to_csv("out.csv") produces a file, not a publication. A published dataset requires provenance (where it came from and what was done to it), documentation (what the columns mean), a license, a citation, and a hosting commitment. The FAIR principles (findable, accessible, interoperable, reusable) name most of these. This chapter adds legibility to publics, which FAIR understates: a dataset that is machine-interoperable and human-inscrutable is FAIR and useless. Hutchinson and colleagues (2021) make the adjacent argument for machine-learning datasets, treating a dataset as an artifact with a lifecycle, not a file.

Zuiderwijk and Janssen (2014) reviewed national open-data policies and found them clearest on what agencies must post and weakest on what they must sustain. You can write an open-data ordinance in an afternoon. You cannot write, in any afternoon, the ongoing labor that keeps a civic dataset findable, legible, and citable three budget cycles from now. That labor is what Star and Ruhleder (1994) named the installed base, and what Karasti and Blomberg (2018) insisted be studied as work rather than assumed as context.

6.2 Two kinds of city openness

At the city level, “open data” usually means a portal: tables of permits, crashes, budgets, and contributions, published under an open license. The richer record of city government lives elsewhere, in agendas, packets, minutes, and meeting video. Keegan (2026) treats that second record as the first of three cases, mining municipal archives, and argues that it fails less through secrecy than through soft enclosure: records are posted but not publishable, because they cannot be followed from an agenda item to a vote to a contract to an expenditure to an outcome. The goal Keegan sets is a municipal “public memory layer” that makes the everyday operations of government followable.

Denver shows both halves at once. Its council records system exposes structured metadata through a public API:

import requests

events = requests.get(
    "https://webapi.legistar.com/v1/denver/events",
    params={"$filter": "EventBodyId eq 138 and "
                       "EventDate ge datetime'2025-01-01' and "
                       "EventDate lt datetime'2026-01-01'"},
    timeout=60,
).json()
len(events)
# => 48   (City Council meetings in 2025, as of October 2026)
sum(bool(e["EventAgendaFile"]) for e in events)
# => 48   (every agenda is a link to a PDF)

The metadata is published: dates, bodies, statuses, stable identifiers. The substance is posted: PDFs that you must download and extract one at a time before you can ask what the council discussed. Einstein, Palmer, and Glick (2019) could show who speaks at local land-use meetings only because they turned minutes into data, and projects like LocalView (Barari and Simko 2023) exist because cities rarely do that work themselves. You will practice this extraction on county records in Chapter 10. For now, notice that the line between posting and publishing runs through the middle of a single city system.

Keegan (2026) anticipates two objections that you should carry into your own work. The first is capacity: many cities lack the staff to do more than post, and proposals for metadata standards or AI-assisted search can read as techno-solutionism from people who will not maintain them. The answer is shared services (common tooling, templates, and custodial partnerships with libraries and archives) rather than heroic one-off projects. The second is vendor capture: a commercial indexing service can make records searchable while creating a new enclosure around public memory, producing what Keegan calls “transparency theater.” The test is whether indices are exportable, ranking is auditable, and the public, not the vendor, controls the memory layer. Both objections apply to you as well. A student scraper that indexes a city’s packets for one semester and then disappears is a project, not infrastructure.

6.3 Licensing as a governance choice

An open license is a small act of governance. It tells downstream users what they may do, what they owe in return, and what you expect when your dataset is recombined with others. CC-BY 4.0 permits any reuse with attribution. CC0 waives copyright, placing the dataset as close to the public domain as a legal instrument can reach; it suits data that is already the product of public work. PDDL is CC0’s database-specific cousin. For U.S. city data, CC-BY 4.0 or CC0 covers almost every case.

What to avoid: “non-commercial” restrictions. Local newsrooms sell their work. A non-commercial clause tells the local paper it cannot use the dataset without a lawyer’s letter. That is not openness. It is a defensive crouch in a compliance costume.

6.4 The running example: Boulder election contributions

The tutorial uses the City of Boulder’s Election Contributions table, published on the city’s open data portal under CC0 and served through an ArcGIS REST endpoint. Campaign finance is the most politically interesting of the common civic datasets because it sits where accountability meets mess. This one is a good teacher. When this chapter was revised in October 2026, the layer held 9,921 rows, its filings ran from September 2010 to January 2018, and its last edit was in August 2020. The portal still lists it. Nothing on the page says the series stopped.

6.4.1 Ingest

The service returns at most 2,000 records per request, so you page through it. Reuse polite_get from Chapter 4, with the documented robots override from that chapter, and save the raw pull before you touch it.

import re
import pandas as pd

LAYER = ("https://services.arcgis.com/ePKBjXrBZ2vEEgWd/arcgis/rest/"
         "services/Election_Contributions/FeatureServer/0/query")

def fetch_all(layer, page=2000):
    rows, offset = [], 0
    while True:
        payload = polite_get(layer, respect_robots=False, params={
            "where": "1=1", "outFields": "*", "f": "json",
            "orderByFields": "ObjectId",
            "resultOffset": offset, "resultRecordCount": page,
        }).json()
        rows += [f["attributes"] for f in payload["features"]]
        if not payload.get("exceededTransferLimit"):
            return pd.DataFrame(rows)
        offset += page

raw = fetch_all(LAYER)
raw.to_csv("data/boulder_contributions_raw.csv", index=False)
raw.shape
# => (9921, 20)

Treat the raw file as immutable. Do not open it in Excel, which will strip the leading zeros you are about to need.

6.4.2 Clean

df = raw.copy()
df.columns = [re.sub(r"(?<!^)(?=[A-Z])", "_", c).lower()
              for c in df.columns]           # TransactionDate -> transaction_date

# Dates arrive as epoch milliseconds
for col in ["filing_date", "amended_date", "transaction_date"]:
    df[col] = pd.to_datetime(df[col], unit="ms")
(df["transaction_date"].dt.year > 2030).sum()
# => 6    six contributions dated between 2100 and 2115; keep them, flag them
df["date_flag"] = df["transaction_date"] > df["filing_date"]

# Money arrives as text: "$50.00 " and refunds as "(100.00)"
df["amount_raw"] = df["contribution"]
df["amount"] = pd.to_numeric(
    df["contribution"].str.strip()
      .str.replace(r"[\$,]", "", regex=True)
      .str.replace(r"^\((.*)\)$", r"-\1", regex=True),
    errors="coerce",
).astype("Float64")
(df["amount"] < 0).sum()
# => 279

# One city, many spellings
key = df["city"].fillna("").str.lower().str.replace(r"[^a-z]", "", regex=True)
is_boulder = key.str.startswith("bou") & key.ne("bouldercounty")
df.loc[is_boulder, "city"].nunique()
# => 26   ("Boulder", "boulder", "BOULDER", "Bouldr", "Boul;der", ...)
df["city_is_boulder"] = is_boulder

# Names are not identifiers
df["committee_num"].nunique(), df["committee"].nunique()
# => (131, 114)   16 committee names recur across cycles with new numbers

6.4.3 Normalize and save

keep = ["object_id", "committee_num", "committee", "type", "candidate",
        "filing_date", "official_filing", "transaction_date", "date_flag",
        "last_name", "first_name", "city", "city_is_boulder", "state",
        "zip", "amount", "amount_raw", "contribution_type", "match"]
clean = df[keep].sort_values("transaction_date").reset_index(drop=True)
clean.to_csv("data/boulder_contributions_clean.csv", index=False)
clean.to_parquet("data/boulder_contributions_clean.parquet")

Two formats, intentionally. CSV is the compatibility format. Parquet preserves the types you just fought to recover, including ZIP codes as text, so whoever reads it next does not repeat your work. Notice what is missing from keep: the street column. That is a decision, not an oversight, and the end of this chapter explains it.

Three habits are worth naming. Every destructive transformation keeps the original column: amount_raw survives next to amount, and city survives next to city_is_boulder. Failures stay visible rather than dropped, so the datasheet can mention them. And every normalization is documented inline, because “we treated 26 spellings as Boulder” is a choice future you will have forgotten. It is also a choice with a caveat: a mailing city of “Boulder” is not the same as an address inside city limits, since some unincorporated addresses use a Boulder mailing city. The committee check makes the linkability point concrete. If you joined filings across elections by committee name, you would silently merge different committees from different cycles; the clerk’s year-coded number is the identifier, and it is the column you must never drop.

TipThe Missing Manual

Three things about this portal are not in its quickstart. First, the ArcGIS query endpoint silently caps each response at the layer’s maxRecordCount (2,000 here) and signals truncation only through an exceededTransferLimit flag; forget to check it and you will analyze a fifth of the data with total confidence. Second, the portal’s “last updated” date describes the item page, not the data, and neither tells you the filing series ended years earlier. Third, df.to_csv("out.csv") is still not publication. Until you have metadata, a citation, a license, and a place the file will live, you have a file.

6.5 Writing the README and the datasheet

A dataset’s README is the first thing a non-specialist reads and often the only thing. Write it as prose, with six sections: what this is (one plain paragraph); provenance (source URL, retrieval date, and the provenance log you started in Chapter 4); what was done to it (a short narrative pointing to the script); column documentation (every column, units, missing-value conventions); license; and canonical citation (a versioned URL and date until Chapter 22 gives you a DOI). One page covering those six beats an 80-page data dictionary nobody reads.

A datasheet is for auditors and future maintainers. Gebru and colleagues (2021) proposed the form; a compact version has seven sections: motivation, composition (including sensitive fields such as donor street addresses), collection process (under what legal authority), preprocessing, uses (appropriate and cautioned), distribution, and maintenance. For this dataset, “maintenance” is where you write down that the series stopped in 2018 and nobody said so. Two pages is plenty. The point is that a reader who will never meet you can decide, from the datasheet alone, whether to trust the file for their purpose. Chapter 16 returns to datasheets as a defense against forgetting.

NoteIn the Public Interest

Openness, as this book uses the word, is durable legibility, and the Boulder table shows why the adjective matters. The city did the formal part: a CC0 license, a public API, a portal listing. What it did not do is sustain the series or say that it stopped, so a resident who downloads it today can draw conclusions about “recent” elections from data that ends in 2018. Your README and datasheet supply the interpretability the portal lacks; your versioned citation supplies continuity; your preserved committee_num and object_id columns supply linkability, so a row in your release can be joined back to the source. A city dataset that skips those steps has satisfied the ordinance and failed the public.

6.6 Publishing with Datasette

Datasette, Simon Willison’s open-source tool, publishes a SQLite database as a browsable, queryable website with faceted search, a read-only SQL console, and a JSON API. It is the smallest plausible infrastructure that turns a cleaned CSV into something a council staffer or reporter can explore.

pip install datasette sqlite-utils

sqlite-utils insert boulder.db contributions \
  data/boulder_contributions_clean.csv --csv --detect-types

datasette serve boulder.db --metadata metadata.yml --port 8001

The metadata.yml file is where your README’s column documentation lives inside Datasette:

description_html: |
  <p>City of Boulder election contributions, 2010 to 2018 filings,
  cleaned. See README.md and datasheet.md for provenance.</p>
license: CC0 1.0
license_url: https://creativecommons.org/publicdomain/zero/1.0/
source: City of Boulder Open Data, Election Contributions
source_url: https://open-data.bouldercolorado.gov/
databases:
  boulder:
    tables:
      contributions:
        description: "One row per reported contribution, post-cleaning."
        columns:
          committee_num: "Committee identifier assigned by the City Clerk."
          transaction_date: "Date of the contribution (see date_flag)."
          amount: "USD. Negative values are refunds (parentheses in source)."
          city_is_boulder: "Mailing city matches a Boulder spelling; not city limits."

Open http://localhost:8001/. If a column has no description, the interface tells your reader nothing, and you will see why the YAML matters. datasette publish can deploy the same database to a hosted service when you want a public URL, but a live URL is a front door, not a citation: the layered model is Datasette for exploration, a versioned deposit (Chapter 22) for the permanent record, and the README tying them together.

Before you publish, run a category audit, which Chapter 19 develops: is anyone harmed by how this file names, categorizes, or makes people joinable? The city published donors’ street addresses under CC0. That does not oblige you to republish them. Dropping street from your derived release while documenting that you did so is openness with judgment, a theme Chapter 18 extends.

6.7 Exercises

Exercise 6.1 (Guided). Pull the Boulder Election Contributions layer (or a Denver open data table of your choice) and reproduce the ingest and clean steps. Commit the script, the raw pull, and the cleaned CSV and Parquet files. Write 150 words on one cleaning decision a reasonable person could have made differently.

Exercise 6.2 (Guided). Write a README and a datasheet for the dataset from 6.1. Target one page and two pages. Save them in exercises/ch06_openness/, choose a license, and justify it in one sentence.

Exercise 6.3 (Technical). Build a SQLite database from your cleaned file, write metadata.yml with column descriptions, and run datasette serve. Submit a screenshot of a column description rendering. Optional: deploy it and record one thing that broke.

Exercise 6.4 (Diagnostic). Follow one decision through a city’s records. Using Denver’s council API or Boulder’s agenda portal (respecting its robots.txt; download by hand if crawling is disallowed), pick one agenda item from the last year and trace it from agenda to vote to contract or expenditure. For each link, record the format, whether it has a stable identifier or URL, and whether you could have found it in bulk. In 400 words, name where the chain breaks and which installed-base element is missing.

Exercise 6.5 (Piece 1 component). Package the city data you will use to extend your published finding as a publication, not a file: raw pull, cleaning script, cleaned data, README, short datasheet, license, provenance.jsonl from Chapter 4, and a two-sentence category-audit note on what you excluded and why. This folder is the data half of your Piece 1 technical artifact.

6.8 Looking ahead

You now have a city finding you can stand behind, documented well enough that someone else can check it. A dataset nobody reads is openness in form without openness in effect. Chapter 7 closes the module by turning that finding into the genre journalism uses to reach a city audience: a 750-word argument, pitched to a named local editor, with one chart and one ask.

6.9 Further Reading and Resources

Barari, Soubhik, and Tyler Simko. 2023. “LocalView, a Database of Public Meetings for the Study of Local Politics and Policy-Making in the United States.” Scientific Data 10 (1): 135.
Brown, Eva Maxfield, To Huynh, Isaac Na, et al. 2021. “Council Data Project: Software for Municipal DataCollection, Analysis, and Publication.” Journal of Open Source Software 6 (68): 3904.
Einstein, Katherine Levine, Maxwell Palmer, and David M. Glick. 2019. “Who Participates in Local Government? Evidence from Meeting Minutes.” Perspectives on Politics 17 (1): 28–46.
Gebru, Timnit, Jamie Morgenstern, Briana Vecchione, et al. 2021. “Datasheets for Datasets.” Communications of the ACM 64 (12): 86–92. https://doi.org/10.1145/3458723.
Hutchinson, Ben, Andrew Smart, Alex Hanna, et al. 2021. “Towards Accountability for Machine Learning Datasets: Practices from Software Engineering and Infrastructure.” Proceedings of FAccT 2021, ahead of print. https://doi.org/10.1145/3442188.3445918.
Karasti, Helena, and Jeanette Blomberg. 2018. “Studying Infrastructuring Ethnographically.” Computer Supported Cooperative Work 27: 233–65. https://doi.org/10.1007/s10606-017-9296-7.
Keegan, Brian C. 2026. “Public Interest Data Infrastructuring.” Under Review.
Star, Susan Leigh, and Karen Ruhleder. 1994. “Steps Toward an Ecology of Infrastructure: Complex Problems in Design and Access for Large-Scale Collaborative Systems.” In Proceedings of the 1994 ACM Conference on Computer-Supported Cooperative Work. https://doi.org/10.1145/192844.193021.
Zuiderwijk, Anneke, and Marijn Janssen. 2014. “Open Data Policies, Their Implementation and Impact: A Framework for Comparison.” Government Information Quarterly 31 (1): 17–29.