3  The Post-API Age

TipLearning Objectives
  • Describe the historical arc of web data access from open APIs to restricted, monetized, and disappearing endpoints
  • Analyze specific cases of access restriction using the enclosure, exemption, and erosion framework
  • Identify alternative data access strategies including archives, FOIA requests, data donations, and federated platforms
  • Connect contemporary data access challenges to longer traditions of public interest inquiry
  • Articulate the values of openness, oversight, and ownership as guiding principles for web data practice
  • Measure the current state of API access empirically, and distinguish a platform’s refusal from your own network’s
TipCompanion Notebook

Run this chapter’s code as you read: open the companion notebook (see Appendix A for all of them).

3.1 The Golden Age of Web Data Access

For roughly a decade — from about 2008 to 2018 — the web behaved as if it wanted to be studied. Twitter opened its API in 2006 and for years offered researchers a generous view of public conversation. Facebook launched the Graph API in 2010, and academic teams used it to map friendship networks at unprecedented scale. Reddit’s API was free and essentially unrestricted, Google exposed programmatic access to search trends, maps, and books, and a generation of startups assumed that open data access was simply how platforms worked.

Platforms were not being charitable. Open APIs served their strategic interests: third-party developers built apps that made platforms more useful, academic partnerships lent credibility and surfaced insights, and every integration deepened the platform’s position as infrastructure that other people’s software depended on. The unspoken bargain was that platforms got ecosystems and researchers got data.

What that access enabled was real. Computational social science emerged as a field in this window (Lazer et al. 2009), built substantially on platform data: studies of information diffusion, social movements, public health signals, and political communication that would have been unimaginable with surveys alone. Data journalists used APIs to hold institutions accountable. Civic technologists built tools on open government and platform data. When researchers today describe the “post-API age” (Freelon 2018), this is the era they measure against.

My own research career took shape inside this bargain. The Wikipedia studies that formed my dissertation — how volunteer editors self-organize to cover breaking news events — were possible because the MediaWiki API and public database dumps exposed every edit, by every editor, with timestamps, for free (Chapter 10 teaches you the same interfaces). Later work on Reddit communities depended on Pushshift, a volunteer-run archive that made the site’s full comment history queryable for researchers (Baumgartner et al. 2020). Neither project required permission, partnership, or payment. That is what has changed.

The landscape of web data access has undergone a structural transformation since approximately 2018. Three pressures — enclosure, exemption, and erosion — explain not just that access is shrinking but why, and what forces are driving the change.

TipMissing Manual Reference

Technical systems embed political choices — who gets access, on what terms, and who decides. For a broader introduction to this idea, see Missing Manual Chapter 8: Artifacts Have Politics.

3.2 Three Structural Pressures

3.2.1 Enclosure

Enclosure is the privatization of previously accessible data: platforms converting a shared resource into a controlled asset. The name deliberately echoes the enclosure of common land in England between the sixteenth and nineteenth centuries, when fields that villagers had collectively farmed and grazed for generations were fenced off, by force of law, into private holdings. The land did not disappear. What disappeared was the commons — the shared arrangement under which ordinary people could use it. Data enclosure works the same way: the posts, edits, and traces are still there, but the terms of access have been fenced.

The mechanisms are familiar by now. APIs are shut down or repriced beyond reach. Third-party archives are cut off. Bulk exports are restricted. And, increasingly, access is bundled into exclusive licensing deals.

Two cases anchor the pattern:

  • Twitter/X. In January 2021, Twitter launched an Academic Research track offering qualified researchers free access to the full historical archive — arguably the high-water mark of platform openness. Two years later, following the company’s 2022 acquisition, free access was eliminated, the academic track was discontinued, and enterprise access was priced at $42,000 per month and up. Hundreds of research projects and public-interest tools — bot detectors, misinformation trackers, disaster-response monitors — went dark within months.
  • Reddit. Reddit’s API had been free since 2008. In 2023 the company announced pricing of $0.24 per 1,000 API calls, effective that July (Figure 3.1). The developer of Apollo, a popular third-party client, calculated the change would cost roughly $20 million a year and shut the app down; more than seven thousand subreddits went dark in protest. The same enforcement wave cut off Pushshift (Baumgartner et al. 2020), the archive that had underpinned hundreds of published Reddit studies. The research community lost not just future access but its accumulated observational infrastructure.

Facebook’s trajectory is the template both followed: after the Cambridge Analytica scandal broke in 2018, Meta locked down the Graph API, and in August 2024 it shut down CrowdTangle — the tool journalists and researchers used to monitor viral content — replacing it with a Meta Content Library available only through gated application.

Why did enclosure accelerate when it did? Generative AI is a large part of the answer. Once large language models made user-generated text valuable as training data, the calculus flipped: an open API stopped looking like ecosystem-building and started looking like giving away inventory. Reddit’s reported licensing deal with Google — roughly $60 million per year for training access — was announced within months of the API repricing that had priced out researchers and hobbyists. Data that communities produced collectively is now sold privately, which is the enclosure of the commons in nearly literal form.

3.2.2 Exemption

Exemption is the second pressure: platforms operating beyond meaningful outside scrutiny, invoking privacy, intellectual property, or security to resist independent investigation while conducting the same investigations internally at will.

Platforms run experiments on their users continuously — every A/B test is a behavioral study — and employ large internal research teams with total data access. The asymmetry is written into the documents you read in Chapter 2: when the word “research” appears in a Terms of Service or privacy policy, it is generally granting the platform permission to run research on you, and says nothing about your permission to study the platform. When outside researchers attempt that study, the tools of resistance come out: cease-and-desist letters citing the Computer Fraud and Abuse Act, Terms of Service that prohibit the data collection that public-interest research requires (Fiesler et al. 2020), and account bans for researchers who persist. In 2021, Meta disabled the accounts of NYU’s Ad Observatory team, who were studying political ad targeting with data volunteered by consenting users — an action Meta justified as protecting user privacy. Courts have pushed back at the margins (Chapter 2 covers Sandvig v. Barr, Van Buren, and the hiQ litigation), but legal ambiguity itself does deterrence work: a graduate student weighing a scraping study against a potential federal claim usually just picks a different topic.

Outright refusal is only the bluntest instrument. Platforms also offer collaboration on terms that neutralize it: research partnerships under non-disclosure agreements that let the platform preview findings, and data-sharing arrangements that deliver years late, if at all. The cautionary case is Social Science One, launched in 2018 as a landmark industry–academic partnership to study Facebook’s role in elections. The full URL-shares dataset arrived roughly two years behind schedule, after several funders had threatened to withdraw — and in 2021 Facebook disclosed that the data it had distributed to about 110 research teams accidentally excluded U.S. users with no detectable political leanings, roughly half of all U.S. accounts, quietly undercutting dozens of papers already built on it.

The privacy justification deserves particular scrutiny. Platforms routinely cite user protection when closing researcher access — while continuing to sell fine-grained targeting on the same users to advertisers. Privacy is genuinely important, and Chapter 2 takes it seriously. But when “privacy” restricts only accountability research and never revenue, it is functioning as exemption, not protection. Researchers have names for the pattern — privacywashing and ethicswashing: invoking a value in order to block scrutiny of the very conduct the value is supposed to govern.

Regulation is beginning to respond. The European Union’s Digital Services Act includes a provision — Article 40 (Figure 3.2) — that requires very large platforms to provide vetted researchers with data access for studying systemic risks, with implementing rules still maturing as of mid-2026. Whether such mandates produce usable access, or a new bureaucratic bottleneck, is one of the defining open questions of this decade of internet research.

Screenshot of Article 40 of the Digital Services Act on EUR-Lex, Data access and scrutiny. Paragraphs 1 to 3 give the Digital Services Coordinator or the Commission access to data to monitor compliance, and explanations of algorithmic systems. Marker 1, a box around paragraph 4: on a reasoned request, very large online platforms and search engines shall give vetted researchers access to data for research on systemic risks in the Union.
Figure 3.2: Article 40 of the Digital Services Act on EUR-Lex, the EU’s official site for its law. Paragraphs 1 to 3 concern regulators: their access to a platform’s data, and to explanations of its algorithms. Paragraph 4 ① requires very large online platforms and search engines, when the regulator of the country where they are established asks, to give vetted researchers access to data for research on systemic risks in the EU.

3.2.3 Erosion

Erosion is the third pressure, and the quietest: the decay of the infrastructure for sustained observation. Enclosure and exemption are things platforms do; erosion is what happens to the web itself.

Links rot. Roughly half the hyperlinks cited in U.S. Supreme Court opinions no longer point to the content the justices cited, and a 2024 Pew study found that about a third of webpages that existed in 2013 were gone a decade later. Platforms churn: GeoCities vanished with its era of amateur web culture; Parler went offline overnight when its hosting was pulled; Twitter became X and broke a decade of embedded links and stored identifiers. Sites redesign their HTML and silently break every scraper built against the old structure — a fragility you will experience personally in Chapter 6 and Chapter 8.

For researchers, erosion produces a reproducibility crisis. A study built on the Twitter Academic API cannot be replicated at any price today. Papers that cite Pushshift point at an archive that no longer accepts their queries. When the data underlying published findings disappears, the scientific record quietly loses its evidentiary base (Davidson et al. 2023). And migration to federated platforms — Mastodon, Bluesky — solves some access problems (Chapter 12) while complicating basic measurement: there is no single API endpoint for “all of Mastodon,” because there is no single Mastodon. If a finding worked once, for one team, on data no one can collect again, in what sense is it still science?

Erosion is also why two later chapters exist. The Wayback Machine (Chapter 7) is counter-erosion infrastructure — a deliberate institutional bet that the web is worth preserving. And automation (Chapter 14) is counter-erosion practice: if data disappears, systematic collection now is the only insurance for research later.

3.2.4 Why Platforms Say No

The three pressures describe the fence from the researcher’s side. The view from inside the platform deserves its own account, because the reasons platforms give for restricting access are not uniformly cynical:

  • Privacy and safety. User data can expose identities, enable harassment, or be repurposed in ways users never expected — and platforms have been burned before: Cambridge Analytica began as academic-flavored data collection.
  • Infrastructure. Serving bulk access costs real engineering money, and from the receiving end an aggressive scraper is hard to distinguish from an attack.
  • Business. The data is an asset the platform paid to accumulate. Competitors, AI companies, and researchers all want it, and only one of those groups writes papers instead of products.
  • Legal risk. Copyright, privacy law, contracts, and regulatory exposure make broad access dangerous to the platform itself — it can be held responsible for what a third party does with shared data, which is the other lesson of Cambridge Analytica.
  • Control. From the platform’s side of the gate, “researcher” is a self-description. Requests for bulk access arrive from scholars, journalists, marketers, scrapers-for-hire, and intelligence services, often wearing the same clothes.

When a platform evaluates an access request, its questions are concrete: Who counts as a researcher? Access to what data? Under what rules? And — the question that decides budgets — what does the platform get out of it?

Underneath those questions sits an identification problem: neither side can directly observe the other’s intent. Cross the researcher’s actual purpose with the platform’s actual posture and four outcomes appear:

Table 3.1: Four outcomes of an access request, sorted by intent that neither side can directly verify.
Platform acting in bad faith Platform acting in good faith
Researcher acting in good faith Resistance — legitimate research barred Cooperation — responsible access
Researcher acting in bad faith Extraction — abuse, spam, surveillance Refusal — users protected

At the gate, every cell looks the same. The platform reading your request cannot cheaply tell a public-interest audit from commercial extraction, so gates get priced for the worst visitor — public-interest research pays for the existence of spam. And you, reading a refusal, cannot always tell user protection from self-protection, which is why the audits later in this chapter teach you to attribute a refusal before interpreting it. The counter-values in this chapter’s second half are, among other things, how a researcher becomes distinguishable: named user agents, documented methods, institutional accountability, and vetted-access credentials are all signals about which row of the table you occupy.

3.3 Auditing Access Yourself

Enclosure, exemption, and erosion are claims about the world, and claims about the world can be checked. This section builds two small audits. Neither needs an API key, an account, or an application. Both are things you can re-run against any platform you care about, and both produce evidence you could put in a paper.

A warning about what follows: these audits report the web as it is today. The outputs printed below were recorded on 28 August 2026, and yours will differ — a chapter about disappearing access that shipped fixed numbers would be making the mistake it is warning you about.

3.3.1 The Eight Endpoints

Before probing anything, know what you are probing. The eight services below were chosen to span the gradient this chapter describes — from institutions built to share data to platforms built to fence it — and each one carries real research value.

Wikipedia (homepage, API documentation) — the encyclopedia anyone can edit, run by the non-profit Wikimedia Foundation, and the most-studied website in existence. Every edit to every article is public, timestamped, and attributed, and the pageview logs are published too, which makes Wikipedia a uniquely complete record of collective knowledge production. A researcher can reconstruct how coverage of an event assembled itself, editor by editor — Chapter 10 is devoted to exactly this.

Hacker News (homepage, API documentation) — the technology industry’s link-sharing board, run by the startup incubator Y Combinator. Its API exposes every story, comment, score, and timestamp back to 2006 with no key required. For a researcher it is a clean, tractable window into what the people who build technology pay attention to, and a favorite testbed for studying how ranking algorithms shape attention.

arXiv (homepage, API documentation) — the preprint server where physics, mathematics, and computer science papers appear before journal review, operated by Cornell University. Its API serves the metadata of millions of papers: titles, abstracts, authors, categories, and dates. That corpus powers the “science of science” — studies of how research topics rise, spread across fields, and cluster into collaborations.

Mastodon (homepage, API documentation) — the federated social network: thousands of independently run servers speaking a common protocol, of which the audit probes the largest, mastodon.social. Because openness is a design property here rather than a policy, public data is available instance by instance. Researchers use it to study platform migration — where people went after Twitter’s enclosure — and what community-run governance actually looks like at scale.

Bluesky (homepage, API documentation) — a newer decentralized network built on the AT Protocol, where public posts, follows, and profiles are readable by design. It offers something researchers almost never get: the chance to watch a social network grow from nearly its beginning with the full public record in view, rather than reconstructing its history later through a keyhole. Chapter 12 collects data from it.

Reddit (homepage, API documentation) — thousands of self-governing communities, each with its own rules, moderators, and culture, which made it the field’s natural laboratory for studying community norms, moderation, and online discourse. It is also this chapter’s central enclosure case: free and openly scrapable from 2008 until the 2023 repricing, and now the row of the matrix where you can watch the fence.

X (Twitter) (homepage, developer documentation) — for fifteen years the workhorse of computational social science, the platform on which the field’s methods for studying news diffusion, political communication, and public opinion were built. Nearly all of that research is now impossible to repeat at academic budgets. It stays in the audit partly for what it still is, and partly so you can measure what happened to it.

Meta Graph API (homepage, API documentation) — the programmatic interface to the world’s largest social platform, and once, through tools like CrowdTangle, a real window for researchers and journalists tracking what spreads on Facebook and Instagram. Post-2018 lockdowns closed most of that view. Probing it today teaches you less about Facebook’s users than about the shape of restriction itself.

3.3.2 An Access Matrix

The first audit asks one question of each of these endpoints: if I show up with no credentials, what happens?

import requests
import pandas as pd

# Identify yourself. @sec-ethics explains why this header matters; several of these hosts answer anonymous library defaults differently than they answer a named client.
HEADERS = {"User-Agent": "WebDataScience/1.0 (INFO 4617; you@colorado.edu)"}

# Endpoints that ask for no key, no account, and no application.
TARGETS = {
    "Wikipedia":   "https://en.wikipedia.org/w/api.php?action=query&meta=siteinfo&format=json",
    "Hacker News": "https://hacker-news.firebaseio.com/v0/topstories.json",
    "arXiv":       "http://export.arxiv.org/api/query?search_query=all:electron&max_results=1",
    "Mastodon":    "https://mastodon.social/api/v1/instance",
    "Bluesky":     "https://public.api.bsky.app/xrpc/app.bsky.actor.getProfile?actor=bsky.app",
    "Reddit":      "https://www.reddit.com/r/python/hot.json?limit=1",
    "X (Twitter)": "https://api.twitter.com/2/tweets/search/recent?query=test",
    "Meta Graph":  "https://graph.facebook.com/v20.0/me",
}

# What each status code means for an anonymous request; any other code is "other".
VERDICTS = {200: "open", 401: "gated", 403: "gated", 429: "rate-limited"}


def probe(name, url):
    """Ask one endpoint for data without credentials and describe the answer.

    Returns a row rather than printing, so the results compose into a table.
    A request that never completes is a result too, so the exception is
    recorded instead of ending the audit.
    """
    try:
        response = requests.get(url, headers=HEADERS, timeout=20)
    except requests.RequestException as error:
        return {"api": name, "status": None, "verdict": "unreachable",
                "content_type": type(error).__name__, "rate_limit": None}

    verdict = VERDICTS.get(response.status_code, "other")

    return {
        "api": name,
        "status": response.status_code,
        "verdict": verdict,
        # What a host sends back is as informative as the code it sends.
        "content_type": response.headers.get("content-type", "").split(";")[0],
        "rate_limit": response.headers.get("x-ratelimit-limit"),
    }


access = pd.DataFrame([probe(name, url) for name, url in TARGETS.items()])
access["status"] = access["status"].astype("Int64")
print(access.to_string(index=False))
        api  status verdict             content_type rate_limit
  Wikipedia     200    open         application/json        NaN
Hacker News     200    open         application/json        NaN
      arXiv     200    open     application/atom+xml        NaN
   Mastodon     200    open         application/json        300
    Bluesky     200    open         application/json        NaN
     Reddit     403   gated                text/html        NaN
X (Twitter)     401   gated application/problem+json        NaN
 Meta Graph     400   other         application/json        NaN

Five open, two gated, one refusing outright. Read the rows rather than the totals, because the useful information is in the differences.

Wikipedia, Hacker News, and arXiv answer an anonymous request with data, and have done so for over a decade. They are run by a non-profit, a startup incubator, and a university consortium. None of them monetizes attention, and none has an incentive to fence the commons.

Mastodon is the only host that volunteers its rate limit before you hit it. x-ratelimit-limit: 300 states the budget up front, so a well-behaved client never has to discover the ceiling by crashing into it.

Reddit’s row is the one to sit with. The status is 403, but the content type is text/html — that is not an API refusing you, it is a web page served where JSON should be. Reddit’s public JSON endpoint was free from 2008 until the 2023 repricing described above. Today an unauthenticated request from a script gets a block page. In that row, enclosure is a content-type header.

X returns application/problem+json with a 401: a machine-readable “you are not authorized,” which is at least an honest refusal. Meta’s 400 says an access token must be used. Both are gates, but differently built ones, and knowing which you face determines whether applying for access would even help.

3.3.3 Reading the Failure, Not Just the Status Code

Before you write down “this platform blocks unauthenticated access,” rule out a more boring explanation: that something between you and the platform blocked it.

This is not hypothetical. The first time this audit was run while drafting this chapter, GitHub’s API returned 403 — which would have been a striking finding, since GitHub documents 60 unauthenticated requests per hour for anyone. The body explained it:

{"message":"GitHub access to this repository is not enabled for this session..."}

That was the sandbox the code was running in, not GitHub. Reporting it as enclosure would have described the instrument rather than the thing measured.

def diagnose(url):
    """Show who actually answered, so a refusal can be attributed correctly."""
    response = requests.get(url, headers=HEADERS, timeout=20)
    print(f"status       {response.status_code}")
    print(f"server       {response.headers.get('server')}")
    print(f"via          {response.headers.get('via')}")
    print(f"content-type {response.headers.get('content-type')}")
    print(f"body         {response.text[:120]!r}")


diagnose("https://www.reddit.com/r/python/hot.json?limit=1")
status       403
server       snooserv
via          1.1 varnish
content-type text/html
body         '<body class=theme-beta><div><style>.theme-light,:root{--rem360:22.5rem'

snooserv is Reddit’s own server software and varnish is its cache layer, so this refusal did come from Reddit. Had the server header named a proxy you did not choose, or the body described your own network, the correct conclusion would have been about your network instead.

Three habits follow. Check server and via before attributing a block. Read the body, because a refusal usually says who is refusing. And retest from a second network when a result surprises you: campus wifi, a cloud VM, and a home connection are treated differently by the same platform, and a shared or cloud IP is far likelier to be rate-limited than a residential one.

3.3.4 Dead-Endpoint Forensics

The second audit turns to endpoints that are no longer supposed to work. You might expect a dead API to be simply unreachable. The three below fail in three different ways, and the way each fails says whether anything can still be recovered.

RETIRED = {
    "Pushshift (Reddit archive)": "https://api.pushshift.io/reddit/search/submission/?q=test&size=1",
    "CrowdTangle (Meta)":         "https://api.crowdtangle.com/posts",
    "Twitter API v1.1":           "https://api.twitter.com/1.1/statuses/user_timeline.json?screen_name=nasa",
}


def autopsy(name, url):
    """Ask a retired endpoint for data and classify how it refuses.

    How an endpoint dies tells you whether anything can still be recovered
    from it, so the manner of refusal is the finding.
    """
    try:
        response = requests.get(url, headers=HEADERS, timeout=25)
    except requests.RequestException as error:
        # No HTTP response at all: DNS, TLS, or the host is simply gone.
        return {"endpoint": name, "status": None,
                "cause": "host does not answer", "evidence": type(error).__name__}

    causes = {
        200: "still open",
        400: "rejects the request without a token",
        401: "still running, credentials now required",
        403: "still running, credentials now required",
        404: "host alive, endpoint removed",
    }
    return {"endpoint": name, "status": response.status_code,
            "cause": causes.get(response.status_code, f"other ({response.status_code})"),
            "evidence": response.text[:60].replace("\n", " ")}


for name, url in RETIRED.items():
    result = autopsy(name, url)
    print(f"{result['endpoint']:<28} {str(result['status']):<6} {result['cause']}")
    print(f"{'':<28} {'':<6} {result['evidence']}")
Pushshift (Reddit archive)   403    still running, credentials now required
                                    {"detail":"Not authenticated"}
CrowdTangle (Meta)           None   host does not answer
                                    ConnectionError
Twitter API v1.1             400    rejects the request without a token
                                    {"errors":[{"code":215,"message":"Bad Authentication data."}

Pushshift did not disappear. It answers, and it answers in JSON: {"detail":"Not authenticated"}. The service still runs, now restricted to Reddit moderators. The loss is usually described as the archive going away; more precisely, “no longer accepts queries” means “no longer accepts queries from you.” A gated service can in principle be negotiated with. A deleted one cannot.

CrowdTangle is genuinely gone. There is no HTTP response at all, because the host does not answer. Meta shut it down in August 2024, and the shutdown was thorough: the endpoint is not returning 404, it is not resolving. This is enclosure completed, with nothing left to negotiate with.

Twitter v1.1 still parses your request and rejects it with error code 215, “Bad Authentication data.” The endpoint is alive and serving paying customers; you are simply not one of them. Repricing, not removal.

The exact exception name in the evidence column depends on how you reach the network: a direct connection to a host that has vanished usually raises ConnectionError, while a machine behind a proxy raises ProxyError instead. Either way the status is None, and that is the part that matters — no HTTP conversation happened at all. A browser shows the difference (Figure 3.3): Pushshift’s and Twitter’s refusals are pages they send, while the page for CrowdTangle is Chrome’s own, because nothing came back.

For an endpoint in the second category, the only remaining evidence is archival. The Wayback Machine’s availability API reports whether the documentation was captured before it went:

def last_snapshot(url):
    """Find the closest archived capture of a page.

    The archive is infrastructure under constant load and answers 429 when
    it is busy, so treat a rate-limit response as a real outcome rather than
    a bug in your code. A shared campus or cloud IP hits this often.
    """
    response = requests.get("https://archive.org/wayback/available",
                            params={"url": url}, headers=HEADERS, timeout=30)
    if response.status_code == 429:
        return "rate limited — wait and retry"
    snapshot = response.json().get("archived_snapshots", {}).get("closest")
    return snapshot["timestamp"] if snapshot else "no capture found"


for page in ["api.pushshift.io", "www.crowdtangle.com"]:
    print(f"{page:<24} {last_snapshot(page)}")
api.pushshift.io         20230525131046
www.crowdtangle.com      20240814023608

Those timestamps are the shutdowns themselves, written in the archival record. Pushshift’s closest capture is 25 May 2023 — six weeks before Reddit’s July 2023 pricing took effect and the enforcement wave cut the archive off. CrowdTangle’s is 14 August 2024, the day its own banner gave as its last (Figure 3.4). The archive kept visiting the address until December 2024, but every later capture records only a redirect: its copies of a site end when the site stops being a site.

Chapter 7 develops this properly. The point here is narrower: when an endpoint dies, its documentation is often the only surviving description of what it used to offer, and that documentation is itself a web page subject to erosion. Archived docs are how you reconstruct what the dataset behind a five-year-old paper actually contained.

3.4 Three Counter-Values

In response to these pressures, three values can guide how researchers, technologists, and institutions build and sustain access to web data for the public interest — and they pair off against the pressures one to one: openness against enclosure, oversight against exemption, ownership against erosion. They are not a solution to the post-API age; they are a stance for working inside it.

3.4.1 Openness

Openness means conducting your own work so that it resists enclosure. Transparent methods and pipelines let others evaluate and reproduce what you did. Documenting not just what data you collected but how — which endpoints, which parameters, which dates, what the platform’s constraints were at collection time — is what allows a future researcher to understand your dataset after the source has changed or vanished (datasheets are the tool for this (Gebru et al. 2021)). The habit compounds over time: record the date, preserve the data, and repeat the collection — an access matrix run once is a snapshot, while the same matrix run every semester is a record of enclosure happening. Redundancy is openness in practice: depositing data in durable public repositories like Zenodo, Dataverse, or institutional archives, relying on community resources like the Internet Archive and Common Crawl, and preferring platforms whose data carries open licenses. Wikipedia, which anchors Chapter 10, is the standing counter-example to enclosure: open API, open license, full public dumps.

3.4.2 Oversight

Oversight means insisting that consequential systems be inspectable by parties who do not own them — the direct counter to exemption. In practice this includes independent audits of platform algorithms and content moderation; documentation standards that make datasets and collection methods legible to review; impact assessment when access regimes change; and regulatory frameworks — like DSA Article 40 — that convert researcher access from a privilege platforms grant into an obligation they carry. The professional skills this book teaches are oversight skills: someone who can retrieve, parse, and analyze web data independently is someone a platform cannot simply refuse.

3.4.3 Ownership

Ownership means preferring arrangements where communities govern the infrastructure their data lives in. Federated systems built on open protocols distribute control so that no single acquisition or pricing decision can enclose the whole network — and the model is older than the fediverse: HTTP, RSS, and BitTorrent are open protocols no company owns, which is part of why they still work while whole platforms have come and gone. ActivityPub, which underlies Mastodon, and the AT Protocol, which underlies Bluesky, extend that lineage to social data (Chapter 12 works with both). Community-maintained archives, data cooperatives, and volunteer preservation efforts like ArchiveTeam embody the same value: the traces of collective life are held collectively. Ownership is the long-horizon value, and the counter to erosion at its root: community-run infrastructure does not vanish when one company exits, and collectively held records outlive any single host. It is the hardest of the three to practice, and the only one that removes the single point of failure entirely.

3.5 Public Interest Lineages

The tensions in this chapter give a name to what this book is preparing you to do. Public interest data science asks how technical skills can increase the public’s capacity to know, scrutinize, and act on the consequential institutions and systems that shape collective life. Framed that way, retrieving data is never only a technical problem. It is a question of access (who can observe powerful institutions), of methods (what can be made visible and auditable), of infrastructure (whether evidence remains available over time), and of practice (who benefits from the technical work).

Other professions confronted the same underlying problem — powerful actors controlling information the public needs — and built institutions in response: journalism built scrutiny, law built compelled disclosure, accounting built independent audit, urban planning built participation, and engineering built public safety. None of them settles on a single definition of “the public good.” What they share are working traditions for making power observable, reviewable, and contestable.

Journalism built the most direct precedent — an older professional technology for working with contested evidence. Freedom of information laws exist because reporters and reformers spent decades turning watchdog work from a polite request into a procedural right, with deadlines, exemptions, and appeals; the muckraking tradition established that private power, too, is a legitimate target of sustained scrutiny. The working norms transfer almost directly: verification and provenance as default practice (a claim traces back to a source, a document, a row), the right of reply (the target of an investigation is part of the verification process, and holds no veto), and collaborative investigations that scale scrutiny beyond any one newsroom. When you file a public records request to obtain data an agency will not publish (a strategy revisited below), you are working in this lineage: show your work, preserve provenance, invite challenge, publish under pressure.

Law institutionalized compelled disclosure. Public-records statutes give access a procedural geometry — request, deadline, exemption, appeal, litigation — and discovery rules force parties to produce documents, testimony, and answers that no API or records request would ever surface, on the theory that adjudication without evidence is theater. Due process embodies a proceduralist answer to arbitrary power — decisions must be explainable and reviewable — that algorithmic accountability research inherits almost verbatim. The profession also contributes a strategy alongside its procedures: cause lawyering treats a case as a way to change a rule, not merely to resolve one dispute. Sandvig v. Barr (Chapter 2) was exactly that — brought by researchers to shrink the Computer Fraud and Abuse Act’s chill over audit studies before any platform had sued them.

Accounting shows how disclosure becomes infrastructure. After the 1929 crash, mandatory audits and standardized reporting turned corporate finances from private mystery into public record — and the machinery underneath is what data science can borrow. An audit produces an independent record of what was checked, how, and against which standards. Materiality separates the discrepancies that change a public claim from noise. Workpapers make the audit itself inspectable: the evidence, the sampling choices, the judgment calls, the reviewer’s sign-off. Standardized reporting makes one organization’s claims comparable to another’s. Datasheets, model cards, replication packages, pinned environments, and audit logs are early versions of the same machinery for data and models — trustworthy disclosure at scale requires standards, professional norms, and independent examiners, not goodwill.

Urban planning developed participation as method, starting from the recognition that technical decisions are also decisions about land, exposure, mobility, and political voice. Environmental impact statements force agencies to document consequences before irreversible action; public comment periods create a public record of who objected and why. The tradition carries its own warning: consultation without power curdles into participation theater, a record of voices that changed nothing. Its practical instrument, the stakeholder map — who is in the room, who is affected, who holds authority, who has remedy — transfers directly to algorithmic impact assessments: disclose before deployment, map the affected parties, document expected harms, and create a record that supports challenge.

Engineering professionalized public safety. Licensure gives practitioners standing to refuse unsafe work — a refusal with institutional backing, because the license is not the employer’s to revoke. Codes and standards, many written in the aftermath of collapsed bridges and exploded boilers, convert past failures into shared constraints on future designs; ethics canons name the public as a client even when the public signed no contract; and failure analysis treats every collapse as institutional evidence, with required investigation and required disclosure. Data science is still writing its equivalent — no license to lose, no universally enforced canon, no mature failure-reporting regime. You are entering the field while its canon is being drafted.

The argument of this section is not that data science should copy any one of these professions, but that it is developing its own public interest mission — and that mission has precedents, institutions, and hard-won lessons to draw on.

3.6 What This Means for Practice

The rest of this book equips you with a diverse technical toolkit precisely because no single method of access is reliable. When APIs close, you may need to scrape HTML (Chapter 6, Chapter 8). When web pages change, archives provide continuity (Chapter 7). When documents are locked in inaccessible formats, PDF extraction opens them up (Chapter 9). When platforms restrict automated access, federated alternatives offer open protocols (Chapter 12). And when you need sustained observation over time, automation turns one-time data collection into durable infrastructure (Chapter 14).

Beyond the toolkit, researchers push back on the access regime itself, in three registers: through courts, through infrastructure, and through law-making.

Litigate: sue over the gate. Researchers have gone to court to challenge the premise that platform terms alone define the boundary of legitimate inquiry. Sandvig v. Barr was brought by academics running discrimination audits — testing platforms the way civil-rights testers have long tested landlords, lenders, and employers — and won a ruling that Terms of Service violations are not, by themselves, federal crimes; Van Buren then narrowed the Computer Fraud and Abuse Act the rest of the way (Chapter 2 covers both). The strategic goal is chill reduction: every clarified rule converts a category of research from a legal gamble into a methods decision.

Build: assemble access from pieces you control. When the front door closes, several doors remain, each with its own terms:

  • Official research access programs. Some platforms now offer application-gated researcher access — the Meta Content Library, TikTok’s Research API, Reddit’s researcher program — and controlled-access enclaves (“clean rooms”) where analysis runs on the platform’s premises under the platform’s eye. These are real but rationed: eligibility is narrow, terms restrict publication and sharing, and access can be revoked. Use them, and document their constraints in your methods.
  • Data donation. Users have rights to their own data — in many jurisdictions, legally enforceable export rights — and can volunteer it to research. Donation studies recruit participants to contribute their download packages or install instrumented browser extensions, converting individual data rights into collective observation, with consent built in.
  • Independent archives. The Internet Archive and Common Crawl (Chapter 7) hold web history no platform will serve you; the End-of-Term crawls preserve federal websites across changes of administration; and a local WARC collection you build yourself is an archive no one else can retire.
  • FOIA and public records requests. Government-held data is subject to disclosure law. A well-scoped records request can produce datasets no API offers — inspection records, communications, contracts — with Chapter 11 covering the data agencies publish proactively.
  • Crowdsourcing and citizen science. When no one holds the data, communities can create it: distributed observation, structured reporting, and collaborative annotation have produced public-interest datasets from air quality to political advertising.
  • Shared research infrastructure. Replication packages, codebooks, documented missingness, and synthetic examples let others verify and extend work even when the raw data cannot travel.

Each piece trades away something — comprehensiveness, reliability, or resolution — and assembling several is the point: access built from archives plus enclaves plus donations plus audits has no single point of failure.

Legislate: turn access into a policy problem. The third register makes researcher access a public obligation instead of platform generosity. The EU’s DSA Article 40, discussed under oversight above, is the working example. The U.S. picture is contested on every front — live enough that some of what follows will have changed by the time you read it. The Platform Accountability and Transparency Act, a bipartisan Senate bill reintroduced across several Congresses (most recently in 2025), would create a National Science Foundation–run program through which vetted researchers request platform data, with mandated ad libraries and algorithm disclosures; it has not become law. California’s AB 587 required large platforms to report their content-moderation policies and statistics; X sued, the Ninth Circuit held the content-category reporting provisions likely compel speech in violation of the First Amendment (X Corp. v. Bonta, 2024), and a 2025 settlement permanently enjoined them — a live demonstration of how far states can and cannot compel transparency. And where public officials conduct public business on private platforms, records law and constitutional doctrine can convert their accounts into transparency targets: the Supreme Court’s test in Lindke v. Freed (2024) asks whether the official had actual authority to speak for the state and purported to exercise it. Public power exercised on private infrastructure is still public power, and it can be held to public oversight.

The strategic lesson is redundancy and flexibility. Build your skills across the full range of techniques so that the closure of any single data source does not shut down your research program.

3.7 Common Misconceptions

Students often arrive at this material with intuitions that the post-API age has quietly invalidated. Five are worth correcting explicitly:

  • “The post-API age means APIs are gone.” There are more APIs than ever. What changed is the terms: free and open access for outside observers is what is disappearing, replaced by paid tiers, gated programs, and licensing deals.
  • “If the API closes, I’ll just scrape.” Sometimes — this book teaches you how. But scraping carries its own legal ambiguity (Chapter 2), breaks against increasingly adversarial site design, and scales poorly. Treat it as one tool in a portfolio, not an escape hatch that makes enclosure irrelevant. Scraping while logged in to your own account looks like the cheap way around an API’s price, and it adds a contract you agreed to and an account you can lose; Chapter 8 weighs the two.
  • “If data is publicly visible, access must be permitted and ethical.” Visibility, permission, and ethics are three different questions (Fiesler et al. 2020). A post being public does not mean its author expected it in a dataset, and a page being viewable does not mean its ToS allows automated collection.
  • “This only matters for social media research.” Enclosure logic is spreading wherever data gained AI-training value: news archives, code repositories, forums, image libraries. Government data (Chapter 11) is the notable counter-current, held open by legal mandate.
  • “Nothing can be done about it.” The counter-values are not wishful thinking: regulation is creating access obligations, archives are institutionalizing preservation, and federated platforms are growing. The constraint environment is contested, and practitioners’ choices are part of the contest.

3.9 Additional Exercises

These are open-ended extensions — no scaffold, no fixed path. Use them for further practice or deeper exploration.

  1. Extend the autopsy. Find two more retired or restricted endpoints not covered here and classify each with autopsy(). For each, say whether it is gone, gated, or repriced, and what that implies for a researcher who needs its data: is there anyone to ask, and if so, who?

  2. Timeline exercise. Create a timeline of major API access changes from 2008 to the present. Include at least five events across at least three platforms. For each event, identify which of the three pressures (enclosure, exemption, erosion) best characterizes it.

  3. Framework application. Choose a platform API change not discussed in this chapter. Write a 500-word analysis applying the enclosure/exemption/erosion framework.

  4. Data contingency plan. Design a research project that depends on a specific web data source. Then write a contingency plan: if your primary source disappears mid-study, what are your backup strategies? Consider at least three alternatives.

  5. Public interest audit. Identify a web data research project (published paper, journalism investigation, or civic tech tool) that was enabled by API access. Assess how the project would be affected by current access restrictions. Which of the three values — openness, oversight, or ownership — does the project embody?

  6. Negotiate the gate. With a partner, run an access negotiation: one of you represents a research team, the other a platform. Privately decide your own hidden intent (public-interest or extractive), then negotiate a data-access agreement — which records, which fields, how much history, at what rate, under what restrictions. Afterward, reveal intents and place the outcome in Table 3.1 (cooperation, refusal, resistance, or extraction). Debrief in writing: What did the researchers originally ask for? What did the platform actually agree to provide? Which restriction was hardest to negotiate, and what did each side trade away? Could another team reproduce research conducted under your agreement?

  7. Graduate extension (INFO 5617). Read Davidson et al. (2023) and Perriam et al. (2020). Choose one platform and write a two-to-three page “state of research access” brief: what data researchers could obtain five years ago, what is available today and on what terms (cost, eligibility, restrictions), what official research-access programs exist, and which of this chapter’s three pressures best explains the trajectory. Close with a concrete recommendation for a researcher starting a study of that platform this year.

3.10 Common Issues to Debug

  • Every endpoint returns 403, including ones this chapter lists as open: Something between you and the web is intercepting requests — a corporate proxy, a VPN, a locked-down campus network, or a cloud sandbox. Run diagnose() and read the server header and the body. If the refusal names your own network, the finding is about your network.
  • An endpoint works in your browser but returns 403 in Python: The site is treating your script differently from a browser. Check that you sent the User-Agent header; some hosts refuse library defaults outright. Note that this is exactly what Reddit does, and it is a finding rather than a bug.
  • archive.org returns 429: The Wayback Machine rate-limits by IP, and campus and cloud addresses are shared by many users. Wait several minutes and retry. Do not add threads to make it go faster — that makes it worse.
  • OpenAlex returns 429 with “Insufficient budget”: The keyless daily budget is tiny and shared by everyone behind your IP address, so a classroom exhausts it quickly. Record the X-RateLimit headers — they are the measurement — and stop requesting; the budget resets at midnight UTC, and a free API key raises it. Do not retry in a loop.
  • An endpoint returns 200 but the content-type is text/html: You received a web page, not data — often a “missing key” notice served with a success code. Print the first 120 characters of the body. A 200 tells you the server answered; the content-type tells you whether it answered your question.
  • autopsy() hangs instead of failing: A host that has been shut down may drop packets rather than refuse the connection, so the request waits for the full timeout. That is why every request here passes timeout=. Without it, a dead host stops your notebook indefinitely.
  • access["status"].astype("Int64") raises: Use the capital-I Int64 dtype, not int64. Only the capital form holds missing values, and status is None for any endpoint that never answered.
  • Your table disagrees with the one printed above: Expected, and worth writing down. Access changes; that is the chapter’s argument. Record the date alongside the result so the observation stays interpretable later.

3.11 Key Takeaways

The core argument of this chapter is that understanding why access is shrinking — not just that it is — equips you to navigate the constraints strategically rather than being surprised by them.

  • The open-API era (roughly 2008–2018) was a strategic bargain, not a natural state: platforms traded data access for ecosystems, and computational social science was built on the trade.
  • Enclosure converts collectively produced data into privately controlled assets — API shutdowns, repricing, archive removals, and AI-training licensing deals are its mechanisms.
  • Exemption is the asymmetry of scrutiny: platforms study their users freely while deploying legal and technical tools against independent researchers who study the platforms.
  • Erosion is infrastructural decay — link rot, platform churn, and disappearing datasets — and it makes reproducibility a design problem you must solve at collection time.
  • Openness, oversight, and ownership are the counter-values: document and deposit so your work resists enclosure, insist on inspectability, and prefer community-governed infrastructure.
  • The practical response is redundancy: master multiple access techniques (documents, archives, APIs, automation) so no single closure ends your research program.
  • Access is measurable, not just arguable. An unauthenticated probe tells you where a platform sits on the open-to-closed gradient today, and how a retired endpoint refuses — gone, gated, or repriced — tells you whether anyone is left to negotiate with.
  • Attribute refusals carefully. A 403 can come from the platform, from your network, or from something in between, and the server header and response body usually say which. A measurement that describes your instrument instead of your subject is worse than no measurement.

3.12 Further Reading

  • Freelon (2018) — the original articulation of the “post-API age”
  • Bruns (2019) — platforms’ fight against critical scholarly research
  • Davidson et al. (2023) — empirical study of how API restrictions threaten open science
  • Perriam et al. (2020) — digital methods in a post-API environment
  • Lazer et al. (2020) — obstacles and opportunities in computational social science
Baumgartner, Jason, Savvas Zannettou, Brian Keegan, Megan Squire, and Jeremy Blackburn. 2020. “The Pushshift Reddit Dataset.” Proceedings of the International AAAI Conference on Web and Social Media 14: 830–39.
Bruns, Axel. 2019. “After the ’APIcalypse’: Social Media Platforms and Their Fight Against Critical Scholarly Research.” Information, Communication & Society 22 (11): 1544–66. https://doi.org/10.1080/1369118X.2019.1637447.
Davidson, Brittany I., Darja Wischerath, Daniel Racek, et al. 2023. “Platform-Controlled Social Media APIs Threaten Open Science.” Nature Human Behaviour 7: 2054–57. https://doi.org/10.1038/s41562-023-01750-2.
Fiesler, Casey, Nathan Beard, and Brian C. Keegan. 2020. “No Robots, Spiders, or Scrapers: Legal and Ethical Regulation of Data Collection Methods in Social Media Terms of Service.” Proceedings of the International AAAI Conference on Web and Social Media 14: 187–96.
Freelon, Deen. 2018. “Computational Research in the Post-API Age.” Political Communication 35 (4): 665–68. https://doi.org/10.1080/10584609.2018.1477506.
Gebru, Timnit, Jamie Morgenstern, Brenda Vecchione, et al. 2021. “Datasheets for Datasets.” Communications of the ACM 64 (12): 86–92. https://doi.org/10.1145/3458723.
Lazer, David M. J., Alex Pentland, Duncan J. Watts, et al. 2020. “Computational Social Science: Obstacles and Opportunities.” Science 369 (6507): 1060–62. https://doi.org/10.1126/science.aaz8170.
Lazer, David, Alex Pentland, Lada Adamic, et al. 2009. “Computational Social Science.” Science 323 (5915): 721–23. https://doi.org/10.1126/science.1167742.
Perriam, Jessamy, Andreas Birkbak, and Andy Freeman. 2020. “Digital Methods in a Post-API Environment.” International Journal of Social Research Methodology 23 (3): 277–90. https://doi.org/10.1080/13645579.2019.1682840.