4  Enclosure: Who Can Observe?

In the sixteenth and seventeenth centuries, across England and Scotland, common pastures and woodlands that had been used collectively for generations were fenced off by landowners and Parliament. The legal instruments were workmanlike. The consequences were not. Villagers who had grazed sheep on the commons, gathered firewood from shared stands, and gleaned fallen grain from open fields found themselves dependent on wage labor, migrating to cities, or both. Enclosure was not a failure of the commons. It was a project, by those with power, to convert shared resources into private ones, often with the language of improvement attached.

James Boyle (2017) argued two decades ago that a “second enclosure movement” was underway, this time over ideas, code, and data. That framing remains useful, but the enclosure most relevant to this book has moved past ideas and into observation itself. The question public interest data science asks is not only what can I license? but what can I see in the first place? The answer, for a researcher in the mid-2020s, is less than it was a decade ago, and the trajectory is not accidental. Following Keegan (2026), you can read the withdrawal of access as an attack on the installed base from Chapter 2: when linkability, interpretability, and safe scrutiny are removed, “public interest” becomes a slogan without an operational grip.

This chapter opens Part II, the first of four modules: journalism, at the level of the city, against enclosure. You will meet enclosure twice: at the scale of a global platform, where it is loud, and at the scale of a city hall, where it is quiet and often more consequential.

4.1 The last ten years, briefly

The easiest way to feel enclosure as a data scientist is to list retrievals you could have done ten years ago and try them today. A decent starter list:

  • Pull one percent of public tweets through the Twitter streaming API, filtered by keyword, for a dissertation on political discourse.
  • Fetch the complete comment history of a subreddit from Pushshift, with timestamps and scores, for a study of moderation.
  • Query Facebook’s Graph API for public-page posts about a municipal election.
  • Retrieve CrowdTangle data on viral content during a disaster.

Each was a routine graduate-student assignment within the last decade. None is routine now. Some require partnerships most researchers cannot get; some are impossible; some have become expensive products whose terms forbid the reuse the original data supported. Bruns (2019) called this shift the APIcalypse: platforms that depended on researcher goodwill in their growth phase have, in their extraction phase, decided that goodwill is an expensive luxury.

Twitter (now X) is the sharpest case. For years its API included a credible research tier, and an enormous empirical literature grew out of it. In February 2023 the company announced that free API access would end; the replacement tiers cost tens of thousands of dollars per month at research volumes. Longitudinal panels built over a decade became unrenewable.

Reddit’s trajectory followed. Pushshift, a volunteer-run archive that had collected and served Reddit data since 2015, lost its public API in 2023 after Reddit revoked its access (lift_ticket83 2023). Reddit’s own API pricing, announced weeks later, pushed third-party clients (including accessibility tools like Apollo) out of the ecosystem and put researcher access behind tiers that, for any nontrivial use, meant large bills. Gaffney and Matias (2018) had already documented silent gaps in Pushshift’s mirror, which compounded the loss: not only is the new data expensive, the old data you thought you had is incomplete, and you cannot fill it in.

The political economy is not mysterious. Wu (2021) describes a new Gilded Age of concentrated platform power: a handful of firms hold the data that describes contemporary public life, and their incentives run toward monetizing it through advertising, model training, and partnerships rather than letting outsiders audit it. “Partnership” is doing a lot of work there. It almost never includes journalists, civil-society researchers, or the public.

4.2 Soft enclosure at city hall

City government is the quiet case. Every council meeting, permit, contract, and vote leaves a record, and sunshine laws say most of those records are public. Keegan (2026) calls what happens next soft enclosure: records that are nominally public but practically fenced off by scanned PDFs, inconsistent metadata, fragile portals, and weak search. Researchers who have tried to build datasets of local meetings describe the same obstacles: materials scattered across departmental sites, agenda packets posted as image scans, and video hosted on third-party services that make systematic analysis a project rather than a query (Brown et al. 2021; Barari and Simko 2023, 2025).

You can see the pattern in the two cities this module uses. When this chapter was revised in October 2026, the City of Boulder’s “City Council Agendas and Materials” page pointed to a meeting portal hosted by a vendor, PrimeGov, and to folders in a Laserfiche document repository. The portal’s robots.txt file contained two lines: User-agent: * and Disallow: /. Every automated agent is asked to stay out of every page. The City and County of Denver runs its council records through Legistar, which exposes a public JSON web API listing meetings, agenda files, and minutes files; the agendas and minutes themselves are PDFs. Neither arrangement is a scandal. Both platforms were chosen to help clerks publish agendas on deadline, not to help residents ask questions across ten years of votes. Soft enclosure is rarely anyone’s decision. It is the accumulated side effect of procurement choices made for other reasons.

This matters for the module’s portfolio piece, which asks you to extend a published finding with city data. Some city data sits behind a documented API and an open license, some behind a portal that asks you not to crawl it, some in PDFs you download one at a time. The responsible move depends on which situation you are in.

4.3 Reading what the platform tells you you can do

Before you write any code, read what the site says about your presence. Two documents matter: robots.txt and the Terms of Service. Neither is a law. Both are signals, and you ignore them at your peril, ethically and practically.

A robots.txt file is served at the root of a domain and tells automated agents which paths are and are not welcome. The Robots Exclusion Protocol dates to 1994 and was standardized as RFC 9309 in 2022. Python’s standard library reads it:

import urllib.robotparser

rp = urllib.robotparser.RobotFileParser()
rp.set_url("https://open-data.bouldercolorado.gov/robots.txt")
rp.read()

rp.can_fetch("PublicInterestDataScience/0.1",
             "https://open-data.bouldercolorado.gov/datasets/")
# => True
rp.crawl_delay("PublicInterestDataScience/0.1")
# => 60   (as of October 2026: one request per minute for crawlers)

The parser returns True or False for a given path and user agent, and surfaces the crawl delay and sitemap if the file declares them. A sixty-second crawl delay is a strong hint that the city does not want its portal pages crawled, and a weak hint about where it does want you to go: the documented data API.

Terms of Service are harder, because they are written by lawyers who assume you will not read them. For years it was unclear whether violating a ToS could be prosecuted as “unauthorized access” under the U.S. Computer Fraud and Abuse Act (CFAA). The Sandvig v. Barr litigation (American Civil Liberties Union 2019) established in 2020 that academic research violating a ToS is not by that fact alone a criminal CFAA violation. That is not a blanket license: you can still be sued, banned, or sent a cease-and-desist. Read the ToS, document what you read, and know which clauses you rely on. Chapter 8 returns to the CFAA.

TipThe Missing Manual

A bare requests.get(url) does not arrive as “anonymous.” It arrives as python-requests/2.x, a default string many content delivery networks block. Set a User-Agent that names your project and a contact email, so an operator who needs you to stop can email you instead of banning your subnet. A second surprise: when a robots.txt request returns HTTP 401 or 403, Python’s RobotFileParser concludes that everything is disallowed. The ArcGIS host that serves Boulder’s open data API returned 403 for /robots.txt in October 2026, so a strict check refuses a CC0 dataset the city publishes precisely for reuse. The parser is not wrong, exactly. It is answering a narrower question than the one you are asking.

4.4 A responsible scraper, from scratch

You will reuse the function below across this module. The design goals are plain: every request identifies itself, respects robots.txt by default, rate-limits, retries on transient errors, and fails loudly on permanent ones.

import time
import logging
from urllib.parse import urlparse
from urllib.robotparser import RobotFileParser

import requests

log = logging.getLogger("responsible_scraper")

DEFAULT_UA = (
    "PublicInterestDataScience/0.1 "
    "(+https://github.com/yourname/pids-book; you@example.edu)"
)


def check_robots(url, user_agent=DEFAULT_UA):
    """Return True if robots.txt allows user_agent to fetch url."""
    parsed = urlparse(url)
    robots_url = f"{parsed.scheme}://{parsed.netloc}/robots.txt"
    rp = RobotFileParser()
    try:
        rp.set_url(robots_url)
        rp.read()
    except Exception as exc:
        log.warning("Could not read %s: %s", robots_url, exc)
        # Conservative default: if we can't read robots.txt, assume we
        # should not fetch. Override explicitly if you know better.
        return False
    return rp.can_fetch(user_agent, url)


def polite_get(
    url,
    headers=None,
    params=None,
    user_agent=DEFAULT_UA,
    min_interval=1.0,
    max_retries=3,
    backoff=2.0,
    timeout=30,
    respect_robots=True,
):
    """Fetch a URL with rate limiting, retries, and a named User-Agent.

    Returns a requests.Response on success.
    Raises requests.HTTPError on non-retryable failure.
    """
    if respect_robots and not check_robots(url, user_agent=user_agent):
        raise PermissionError(f"robots.txt disallows fetching {url}")

    hdrs = {"User-Agent": user_agent, "Accept": "application/json"}
    if headers:
        hdrs.update(headers)

    last_err = None
    for attempt in range(1, max_retries + 1):
        try:
            response = requests.get(
                url, headers=hdrs, params=params, timeout=timeout
            )
        except requests.RequestException as exc:
            last_err = exc
            log.warning("Attempt %d: network error %s", attempt, exc)
        else:
            if response.status_code in (429, 500, 502, 503, 504):
                log.warning(
                    "Attempt %d: status %s from %s",
                    attempt, response.status_code, url,
                )
                last_err = requests.HTTPError(response=response)
            else:
                response.raise_for_status()
                time.sleep(min_interval)
                return response
        time.sleep(backoff ** attempt)

    raise last_err or RuntimeError("polite_get exhausted retries")

The respect_robots flag defaults to True; you should need a reason to flip it, written into the calling code. The function retries 429 and 5xx responses with exponential backoff, treats other 4xx responses as permanent, and sleeps after every success, the crudest form of rate limiting and the one that gets you furthest. A production version would honor Retry-After, cache to disk, and respect the declared crawl delay; the exercises push you there.

4.4.1 A city fetch with a provenance record

Here is the function used on a real city dataset: Boulder’s Election Contributions table, published on the city’s open data portal under CC0 and served through an ArcGIS REST endpoint. Notice the documented override, and notice that every fetch writes a provenance record you will carry into your portfolio piece.

import hashlib
import json
from datetime import datetime, timezone

LAYER = (
    "https://services.arcgis.com/ePKBjXrBZ2vEEgWd/arcgis/rest/services/"
    "Election_Contributions/FeatureServer/0/query"
)

# Override: the API host's robots.txt returns 403, which RobotFileParser
# reads as "disallow all". The city publishes this layer under CC0 through
# a documented query API, and we request it a few times, not crawl it.
resp = polite_get(
    LAYER,
    params={"where": "1=1", "returnCountOnly": "true", "f": "json"},
    respect_robots=False,
)
resp.json()
# => {'count': 9921}   (October 2026)

def log_provenance(resp, path="provenance.jsonl", note=""):
    record = {
        "url": resp.url,
        "retrieved_at": datetime.now(timezone.utc).isoformat(),
        "status": resp.status_code,
        "sha256": hashlib.sha256(resp.content).hexdigest(),
        "user_agent": resp.request.headers["User-Agent"],
        "note": note,
    }
    with open(path, "a") as fh:
        fh.write(json.dumps(record) + "\n")
    return record

log_provenance(resp, note="robots override: CC0 layer via documented API")

The provenance file is boring on purpose. It answers a skeptical editor’s questions (what, when, from where, in what state) without relying on your memory.

4.5 A worked timeline: Reddit access, 2015–2024

Pick one platform and write its access history as a structured timeline. The exercise forces you to treat enclosure as a series of concrete decisions rather than a diffuse climate. The table below is a first pass; Exercise 4.4 asks you to extend and verify it.

Year Event Consequence for researchers
2015 Pushshift begins continuous archiving of Reddit submissions and comments. A reasonably complete corpus becomes queryable through a free API.
2018 Gaffney and Matias (2018) document silent gaps in the Pushshift mirror. Researchers learn the archive is incomplete; many published studies rest on the incomplete slice.
2023 Reddit revokes Pushshift’s access and announces API pricing (Isaac 2023). The public Pushshift API goes dark. Third-party clients exit or retreat.
June 2023 Large subreddits go private in protest (“the blackout”). Moderator labor, the free substrate of Reddit’s data, becomes briefly legible as labor.
Early 2024 Reddit announces a data-licensing deal with Google for model training. Data refused to researchers is sold, as training data, to a firm that can afford it.

Two observations follow. First, the shutdown was a loss of old data, not only new data: the archive lost the credential it needed to keep serving what it had already collected. Second, the data withdrawn from researchers was sold, within a year, to a commercial AI firm. Enclosure is not “the platform closed up.” It is “the platform closed you out while opening further to capital.” Acker and Kreisberg (2019) saw this coming in their work on social-media archives; Suzor and colleagues (2018) place it within a broader problem of platform governance legitimacy.

4.6 A counter-move: Article 40 of the Digital Services Act

Enclosure is a policy choice, which means it can be met with policy. The European Union’s Digital Services Act (European Parliament and Council of the European Union 2022) includes, in Article 40, a legal pathway for vetted researchers to obtain data from very large online platforms and search engines in order to study systemic risks, along with provisions on access to publicly accessible data. In July 2025 the European Commission adopted the delegated act that sets out how researchers apply and how platforms must respond, including a data access portal (European Commission 2025). Reporting in 2024 already described European authorities investigating X’s refusal to grant researchers API access as a possible breach of the Act (Stokel-Walker 2024).

Article 40 tests the American assumption that access is something platforms grant. In the EU it is, in part, something the law requires. The limits are real: access runs through vetting and conditions, covers only the largest services, and does nothing for a city’s agenda portal. Openness that depends on qualifying as a vetted researcher is better than none and narrower than the public.

NoteIn the Public Interest

The value that answers enclosure is openness, which Chapter 6 develops as durable legibility rather than downloadability. Here is the specific claim for this chapter: enclosure does not abolish observation, it reallocates it. When a platform cuts off independent researchers while licensing data to a model developer, observation flows toward capital. When a city’s meeting portal disallows every crawler, observation of city government flows toward the vendor and toward the few residents with time to sit through meetings and click through packets one at a time. Naming either as an openness failure, rather than a business or procurement decision, is the first move that lets you push back.

4.7 The asymmetry of observation

The firms and vendors that hold the records of public life observe them in extraordinary detail. Everyone outside (journalists, academics, civil-society groups, residents) is left with robots.txt files, PDF fragments, and the occasional leak. The asymmetry is between two classes of observers.

That has consequences for what you build. Assume your access is temporary and keep provenance and copies you control (Chapter 22 shows how to deposit them). Be honest that a study conditioned on platform-provided data is a study of what the platform chose to show you. Remember that refusal runs both ways: Zong and Matias (2024) treat refusal by data subjects as a legitimate counter-move, and Chapter 19 argues that sometimes not collecting is the public-interest choice. And diversify your substrate. City open data, council records, and records requests carry their own pressures (Chapter 8 and Chapter 16), but none has an IPO prospectus that treats your access as lost revenue.

4.8 Exercises

Exercise 4.1 (Guided). Fetch and read the robots.txt files for three sites: Wikipedia, Reddit, and one city meeting-records platform (Boulder’s agenda portal or Denver’s Legistar site). Use urllib.robotparser to check five paths per site and record any crawl delay. Write a 200-word comparison: which is most permissive, which most restrictive, and which surprised you.

Exercise 4.2 (Guided). Implement polite_get from scratch, without copying the listing. Add two features the book’s version lacks: honor the Crawl-delay a site declares, and either honor Retry-After on 429 responses or cache responses in a local SQLite database keyed by URL. Commit the function with a docstring and a test that exercises one feature against a public endpoint.

Exercise 4.3 (Analytic). Pick one platform or portal whose data you expect to use this term. Save its current Terms of Service to a dated file (for example, exercises/ch04_tos/boulder_portal_2027-01-19.txt). Annotate each paragraph: R (restricts research), P (permits it), A (ambiguous; needs counsel), or L (unstable definitions such as “public content”). Add a 300-word summary of what a research plan can and cannot do.

Exercise 4.4 (Analytic). Extend the Reddit timeline with three verified rows since 2024. Then write a second timeline, in the same format, for either the DSA Article 40 researcher-access regime or a city’s meeting-records platform (when it was adopted, what it replaced, what changed in access). Cover at least five dated events. Cite a primary document for each date where one exists.

Exercise 4.5 (Piece 1 component). Choose the city dataset you will use to extend a published finding in Piece 1 (see Chapter 7): a Boulder or Denver open data table, or a set of council records. Write a 400-word access memo: who controls the data, under what license, which pathway you will use (bulk download, API, page-by-page download, records request), and your fallback if it closes. Then retrieve the data with polite_get (or a documented manual download) and start provenance.jsonl with one record per retrieval. Commit both to your Piece 1 repository.

4.9 Looking ahead

Enclosure asks who can observe? Chapter 5 turns to the profession that has answered it longest in practice, treating journalism’s watchdog tradition as working instructions: verification, provenance, right of reply, and publication with documentation. You will reproduce ProPublica’s “Machine Bias,” which exists as a public dataset only because a newsroom chose to publish its evidence.

4.10 Further Reading and Resources

Acker, Amelia, and Adam Kreisberg. 2019. “Social Media Data Archives in an API-Driven World.” Archival Science.
American Civil Liberties Union. 2019. Sandvig v. Barr: Researchers’ Right to Investigate Algorithmic Discrimination. ACLU. https://www.aclu.org/cases/sandvig-v-barr-challenge-cfaa-prohibition-uncovering-racial-discrimination-online.
Barari, Soubhik, and Tyler Simko. 2023. “LocalView, a Database of Public Meetings for the Study of Local Politics and Policy-Making in the United States.” Scientific Data 10 (1): 135.
Barari, Soubhik, and Tyler Simko. 2025. “The Promise of Text, Audio, and Video Data for the Study of US Local Politics and Federalism.” Publius: The Journal of Federalism 55 (2): 223–52.
Boyle, James. 2017. “The Public Domain: Enclosing the Commons of the Mind.” Yale Law Journal.
Brown, Eva Maxfield, To Huynh, Isaac Na, et al. 2021. “Council Data Project: Software for Municipal DataCollection, Analysis, and Publication.” Journal of Open Source Software 6 (68): 3904.
Bruns, Axel. 2019. “After the ‘APIcalypse’: Social Media Platforms and Their Fight Against Critical Scholarly Research.” Information, Communication & Society 22 (11): 1544–66. https://doi.org/10.1080/1369118X.2019.1637447.
European Commission. 2025. Commission Adopts Delegated Act on Data Access Under the Digital Services Act. Shaping Europe’s digital future (news release). https://digital-strategy.ec.europa.eu/en/news/commission-adopts-delegated-act-data-access-under-digital-services-act.
European Parliament and Council of the European Union. 2022. Regulation (EU) 2022/2065 on a Single Market for Digital Services (Digital Services Act). Official Journal of the European Union, L 277. https://eur-lex.europa.eu/eli/reg/2022/2065/oj.
Gaffney, Devin, and J. Nathan Matias. 2018. “Caveat Emptor, Computational Social Science: Large-Scale Missing Data in a Widely-Published Reddit Corpus.” PLoS ONE 13 (7): e0200162.
Isaac, Mike. 2023. “Reddit Wants to Get Paid for Helping to Teach Big a.i. Systems.” The New York Times, April.
Keegan, Brian C. 2026. “Public Interest Data Infrastructuring.” Under Review.
lift_ticket83. 2023. “Reddit Data API Update: Changes to Pushshift Access.” Reddit Post. In R/Modnews. https://www.reddit.com/r/modnews/comments/134tjpe/reddit_data_api_update_changes_to_pushshift_access/.
Stokel-Walker, Chris. 2024. “Under Elon Musk, X Is Denying API Access to Academics Who Study Misinformation.” Fast Company, February. https://www.fastcompany.com/91040397/under-elon-musk-x-is-denying-api-access-to-academics-who-study-misinformation.
Suzor, Nicolas P., Tess Van Geelen, and Sarah Myers West. 2018. “Evaluating the Legitimacy of Platform Governance: A Review of Research and a Shared Research Agenda.” International Communication Gazette 80 (4): 385–405.
Wu, Tim. 2021. “The Curse of Bigness: Antitrust in the New Gilded Age.” Columbia Law Review.
Zong, Jonathan, and J. Nathan Matias. 2024. “Data Refusal from Below: A Framework for Resisting Institutional Data Harms.” ACM Journal on Responsible Computing, ahead of print. https://doi.org/10.1145/3630107.