22 Archival Deposits
Each of your four portfolio pieces addressed an audience that already exists. The op-ed (Chapter 7) argued to a city reader holding tomorrow’s paper. The testimony (Chapter 11) addressed county commissioners who had a hearing on the calendar. The report (Chapter 15) answered a state body that could act on it. The public comment (Chapter 20) reached a federal agency on a docket open for sixty days. A proposal (Chapter 21) speaks to reviewers convened this fiscal year. Each of those publics was reachable because each is alive.
An archival deposit is different. You are contributing to a conversation with a public that does not yet exist. A graduate student in 2048 reading old EPA Toxics Release Inventory data to reconstruct an environmental-justice claim. A city staffer in 2035 looking for Colorado groundwater trends that a federal agency has since stopped publishing. A journalist in 2060 writing a retrospective on federal climate portals reorganized during this decade. None of those readers can be pitched to, invited, or surveyed. They will find what you deposited if, and only if, it is still there.
That is why the deposit is the required last step of the final project. You take one portfolio piece, deepen it, recast it in a new genre, write a critical-reflection memo (Chapter 23), and then archive the technical artifact and both texts on Zenodo, or another repository that mints a DOI. Only a DOI meets the requirement; an Internet Archive item identifier does not. A submission without the DOI deposit is incomplete. The deposit is where openness, oversight, and ownership (Chapter 6, Chapter 10, Chapter 18) come to rest in a single physical-digital object: a bitstream, a checksum, a resolvable identifier, a provenance record, held by an institution committed to keeping them. This chapter teaches the workflow on a federal dataset at risk, the kind you audited in Chapter 16, and then applies it to your own project.
22.1 Archives and climate, archives and memory
It is tempting to treat archival deposit as a technical afterthought: you did the work, now you put it somewhere. The archival profession has been saying for a decade that this is wrong. Tansey (2015) argues that climate change is an archival problem because the floods, fires, and storms that threaten records are the events that make those records most urgently needed. Mazurczyk and colleagues (2018) extend the argument quantitatively, mapping American archives against projected sea-level rise and showing how much institutional memory sits in places the country’s institutions are not planning to protect. Blanke (2024) takes it digital: the Anthropocene is also a problem for server rooms, for migrations across storage media, and for the economics of keeping bits online for a century.
Bluntly: the federal agencies whose datasets document the climate crisis are the same agencies whose data-retention budgets are most exposed to it. An EPA dataset is at risk from two directions. The weather can take the building. A single administration can take the dataset. The archive hedges both.
Ghaddar (2016) and Robinson-Sweet (2018) describe a second horizon: truth-and-reconciliation futures in which publics will want to examine what present governments recorded, classified, suppressed, or destroyed. What you deposit, and what you decline to deposit, is part of what a later public can say about us. Someone will go looking. Either your work is there or it is not.
22.2 The institutional landscape
A “durable” archive is not a synonym for “a server you rent.” Durability is an institutional property: a mandate to preserve, a budget funded across political cycles, a succession plan, and federated backup so the loss of one building does not destroy the collection. A handful of institutions meet that standard. Several more claim to and do not.
The Internet Archive operates the Wayback Machine and a general items repository (Internet Archive 2024). It accepts almost anything, mints durable URLs (not DOIs), and exposes a Python client. That makes it a fine mirror, not this course’s deposit of record. Its legal status has been repeatedly challenged in court, a sign of how much it matters.
Zenodo, operated by CERN, accepts datasets, code, and papers up to 50GB, mints a Digital Object Identifier (DOI) on publication, and preserves what it holds on CERN’s long-term infrastructure (CERN 2024). A Zenodo DOI is the shortest path from a student laptop to a citable identifier, and it is what the final project requires.
Figshare and OSF play similar roles; OSF integrates tightly with research-project workflows (Center for Open Science 2024). ICPSR at the University of Michigan is the oldest social-science data archive and the standard for survey data and restricted-use materials (Inter-university Consortium for Political and Social Research 2024). Dryad is the preferred repository for the life sciences (Dryad 2024). Dataverse is a federated platform at dozens of universities; DASH at NCAR is the domain archive for atmospheric data.
Software Heritage archives source code at scale (Software Heritage 2024). LOCKSS (“Lots of Copies Keep Stuff Safe”) is the federated-preservation backbone used by libraries (Stanford University Libraries 2024). Perma.cc captures web pages for legal citation (Harvard Library Innovation Lab 2024). EDGI coordinates volunteer archiving of federal environmental data at risk of removal (Environmental Data and Governance Initiative 2024).
Notice what is not on this list. GitHub is not a preservation institution; its retention guarantees are a product roadmap, not a mandate. Hugging Face is the same. Google Drive, Dropbox, S3, and personal websites are the same. They can be links in a preservation chain; they cannot be the last link. If the only copy of your dataset lives on a platform whose survival depends on next quarter’s earnings, it is not archived. It is stored.
22.3 Federated archiving as political strategy
That list is also a political arrangement. LOCKSS’s founding claim is in its name: safety comes from redundancy across independent custodians (Stanford University Libraries 2024). When the EPA’s climate pages were reorganized in 2017, EDGI’s volunteer “data rescue” events moved copies of at-risk datasets into Internet Archive items before the changes took effect. No single institution was going to do that; a federation did.
An archival deposit is therefore political as well as technical. Depositing the same dataset in two independent repositories, in two jurisdictions, hedges against any one failing, being captured, or being compelled. The final project requires a DOI, so Zenodo is the deposit of record. This chapter recommends an Internet Archive mirror as well. That redundancy is not theater; it is the minimum viable federation.
22.4 The running example: EPA Toxics Release Inventory
The Toxics Release Inventory (TRI), mandated by the Emergency Planning and Community Right-to-Know Act of 1986, collects annual self-reports from industrial facilities on releases of listed toxic chemicals to air, water, and land (US Environmental Protection Agency 2024). It is one of the most-cited environmental-justice datasets in the United States and has run continuously since 1987. Like every federal environmental dataset, it is subject to threshold changes, definitional revisions, and occasional quiet removal of historical files when schemas are updated.
Pick a TRI reporting year’s facility-level file (the 2022 Basic Data File is small enough for a laptop), download it, keep the original bytes. The workflow below transfers to any at-risk federal environmental dataset: NOAA Climate Data Online, USGS groundwater, EPA AirData, NOAA GHCN-D.
22.5 Preservation-appropriate formats
Before any upload, convert your artifacts into formats still readable in fifty years. This is the highest-leverage decision in the workflow. The Library of Congress’s “Sustainability of Digital Formats” registry rates formats on openness, adoption, transparency, and self-documentation (Library of Congress 2024). The short version:
- Tabular data: CSV (UTF-8, RFC 4180) as primary; Parquet as companion for large tables. Not Excel; its binary format is semi-proprietary and its date handling is notorious.
- Documentation: plain text or Markdown. Not Word. A
.docxopened in 2045 will render in ways neither you nor your reader can predict. - Images: TIFF or PNG. Not JPEG; every re-save degrades it.
- Geospatial: GeoJSON or GeoPackage primary; Shapefile as companion (its 10-character field-name limit truncates columns).
- Code: plain-text source files, not notebooks, with an HTML export attached. A
.pyis more likely to parse in 2050 than an.ipynb.
You will work harder than you want on this step. Do it anyway. Format choice is the variable most strongly associated with whether your deposit is legible in thirty years.
22.6 Mirroring with the Internet Archive client
The DOI deposit comes next; an optional mirror comes first, so the DOI record can point to it. The internetarchive Python client and its ia command-line companion are the fastest way to make one (Internet Archive 2024). Run ia configure once to store your credentials, then upload. The item identifier is a short unique string that becomes part of the URL (https://archive.org/details/<identifier>). Identifiers are forever; choose one that describes the content.
from internetarchive import upload, get_item
metadata = {
"title": "EPA Toxics Release Inventory, 2022 Basic Data File (mirror)",
"creator": "Your Name, University of Colorado Boulder",
"date": "2027-04-23",
"subject": ["environment", "toxics", "EPA", "TRI",
"environmental justice"],
"language": "eng",
"licenseurl": "https://creativecommons.org/publicdomain/mark/1.0/",
"description": (
"Mirror of the US EPA Toxics Release Inventory 2022 Basic "
"Data File, retrieved from epa.gov on 2027-04-20. See the "
"attached preservation plan."
),
"mediatype": "data",
"collection": "opensource",
}
files = ["tri_2022_us.csv", "tri_2022_us.parquet",
"datasheet.md", "README.md", "LICENSE.txt",
"preservation_plan.md"]
upload("epa-tri-2022-basic-mirror", files=files, metadata=metadata)
# => [<Response [200]>, <Response [200]>, ...] on success
# Verify the deposit. Compare server SHA-1 to what you computed locally.
item = get_item("epa-tri-2022-basic-mirror")
for f in item.files:
print(f["name"], f["size"], f.get("sha1"))HTTP 200 does not guarantee the files are actually there. Fetch the item back and confirm each file has non-zero size and a matching SHA-1. That verification is what most tutorials skip.
The internetarchive client returns success codes on uploads that silently failed: the item is created, the files are listed with size 0, and the HTTP layer looks fine. Always verify by fetching the item back with ia metadata <item_id> or get_item(id) and comparing the server’s SHA-1 to yours. Separately, Zenodo DOIs take up to an hour to resolve after minting; do not panic when a fresh DOI returns “not found” ten seconds after publishing. The other silent killer is format choice. The difference between a file that survives fifty years and one that does not is usually not storage redundancy but format: CSV beats Excel, TIFF beats JPEG, Markdown beats Word, plain source beats a bundled notebook. Storage is easy. Legibility is hard.
22.7 Minting a DOI with Zenodo
Zenodo’s REST API issues a DOI on publication, which is what makes it the deposit of record in the federated pair (CERN 2024). Request a personal access token from zenodo.org/account/settings/applications/, scope it to deposit:write and deposit:actions, and treat it like a password.
import os, requests
ZENODO = "https://zenodo.org/api"
PARAMS = {"access_token": os.environ["ZENODO_TOKEN"]}
HEADERS = {"Content-Type": "application/json"}
# 1. Create an empty deposition.
r = requests.post(f"{ZENODO}/deposit/depositions",
params=PARAMS, headers=HEADERS, json={})
r.raise_for_status()
dep = r.json()
# 2. Upload files into the deposition bucket.
for path in ["tri_2022_us.csv", "tri_2022_us.parquet",
"datasheet.md", "README.md", "preservation_plan.md"]:
with open(path, "rb") as fh:
requests.put(f"{dep['links']['bucket']}/{os.path.basename(path)}",
data=fh, params=PARAMS).raise_for_status()
# 3. Attach Dublin Core metadata.
meta = {"metadata": {
"title": "EPA Toxics Release Inventory, 2022 Basic Data File (mirror)",
"upload_type": "dataset",
"description": "Mirror of the US EPA TRI 2022 Basic Data File...",
"creators": [{"name": "Your Name",
"affiliation": "University of Colorado Boulder"}],
"keywords": ["environment", "toxics", "EPA", "TRI",
"environmental justice", "at-risk data"],
"license": "cc-zero", "language": "eng",
"related_identifiers": [ # drop the first entry if you made no mirror
{"identifier": "https://archive.org/details/epa-tri-2022-basic-mirror",
"relation": "isIdenticalTo", "resource_type": "dataset"},
{"identifier": "https://www.epa.gov/toxics-release-inventory-tri-program",
"relation": "isDerivedFrom", "resource_type": "other"},
],
}}
requests.put(f"{ZENODO}/deposit/depositions/{dep['id']}",
params=PARAMS, headers=HEADERS, json=meta).raise_for_status()
# 4. Publish. This mints the DOI and makes the record immutable.
r = requests.post(
f"{ZENODO}/deposit/depositions/{dep['id']}/actions/publish",
params=PARAMS)
print(r.json()["doi"]) # the API returns your record's DOI here, e.g. 10.5281/zenodo.7654321The related_identifiers field is where federation becomes legible. You are telling Zenodo that its record is identical to the Internet Archive item and derived from the EPA source. A future user landing on any of the three can follow the chain to the others. Before going live, test against sandbox.zenodo.org: the API is identical, the DOIs are not real, the mistakes are free.
22.8 A preservation plan
A deposit without a preservation plan is a bet that someone else will figure out what you meant. A plan makes the bet explicit by answering four questions.
Retention. How long do you intend the deposit to remain accessible, and under what conditions would it be withdrawn? “Indefinite, absent legal compulsion” is acceptable. So is “Ten years, then re-evaluate.” Specify something.
Migration. What happens when the format becomes unreadable? Parquet in 2027 will not be Parquet in 2057. State whether migration is the depositor’s responsibility (usually yes for the first generation), the repository’s (sometimes yes for major archives), or unresolved. Name a review interval; five years is reasonable.
Custodianship. Who is responsible when you move, change institutions, or leave the work? Karasti and Blomberg (2018) call this “maintenance across decades,” and their infrastructure studies (Chapter 2) is the backbone of the answer. Designate a successor, an institution, or a community. “My GitHub account” is not an answer. “The University of Colorado Boulder libraries’ digital preservation program, with EDGI as federated redundancy” is.
Access and refusal. Tie back to the tiered-access matrix from Chapter 18. Open access or tiered (DUA, embargo windows)? Refusal conditions from Chapter 19 under which portions should be withdrawn? Make the rules legible before a future dispute requires them.
A plan that answers those four questions, committed alongside the dataset in the same deposit, is a preservation plan and a governance document. Someone will need it.
The public interest includes people who do not yet exist. That is not a metaphor. The graduate student in 2048, the council staffer in 2035, the journalist in 2060: they are the public your deposit serves, and the archive is how you speak to them. This is the chapter where the book’s three values come to rest. Openness (Chapter 6) lets future readers find the bits, read them, and cite them. Oversight (Chapter 10) lets them hold present institutions accountable using records the institutions did not destroy. Ownership (Chapter 18) binds the deposit to a custodianship chain that does not evaporate with a funding cycle. The installed base (Chapter 2) was never only about the present; it was always about whether the future could read us.
22.9 Assembling the final-project deposit
The TRI mirror is practice. The deposit that counts is your final project, and most of its parts already exist. The technical artifact from your chosen portfolio piece, exported to plain-source .py files with an HTML rendering, is the core. The original public text and the recast text sit beside it, both as Markdown. Your installed-base note becomes the README’s provenance section, and a datasheet in the form you learned in Chapter 16 documents any data you collected. The preservation plan described above carries the tiered-access decisions from Chapter 18 and any refusal conditions from Chapter 19. The critical-reflection memo from Chapter 23 goes in too, so that a future reader meets your account of the work’s limits alongside the work.
You may also deposit the other three portfolio pieces as a collection, with the final project as its centerpiece. Exercise 22.4 asks you to build the deposit.
22.10 Verifying and citing
A deposit that cannot be cited is a backup. A deposit that can be cited is a contribution. After publication, resolve the DOI and confirm it lands. If you made a mirror, paste the Internet Archive URL into a fresh logged-out session and confirm you can download the files. Compute their SHA-256 and compare to what you deposited. Then write the citation, substituting your minted DOI for the XXXXXXX placeholder, in the form your future readers will paste:
Your Name (2027). EPA Toxics Release Inventory, 2022 Basic Data File (mirror) [Dataset]. Zenodo.
https://doi.org/10.5281/zenodo.XXXXXXX. Also at Internet Archive:https://archive.org/details/epa-tri-2022-basic-mirror.
That citation, on your CV and project pages, is the shortest route from “I did the work” to “here is the permanent link.” It is the minimum unit of public-interest infrastructure Keegan (2026) describes. Write it.
22.11 Exercises
Exercise 22.1 (Guided, mirror). Install the internetarchive client and register at archive.org. Pick an at-risk federal environmental dataset; the endangered resource you named at the end of Chapter 16 is a natural choice, as is EPA TRI 2022, a NOAA Climate Data Online file, or a USGS groundwater extract. Download it, convert to CSV and Parquet, write a one-page datasheet following Chapter 16, and deposit the bundle as an Internet Archive item with full Dublin Core metadata. Verify with ia metadata <item_id> that every file has non-zero size and a matching SHA-1. Submit the item URL and verification output.
Exercise 22.2 (Guided). Create a Zenodo account, request an access token, and test the deposit flow against sandbox.zenodo.org before going live. Then deposit to production Zenodo with a related_identifiers entry pointing back to your Internet Archive mirror. Verify the DOI resolves (allow up to an hour). Submit the DOI, the resolved URL, and a note on how long resolution took.
Exercise 22.3 (Reflective). Write a two-to-three-page preservation plan for the dataset. Address retention, migration, custodianship, and access and refusal (tying back to Chapter 18 and Chapter 19). Reuse the sustainability plan from the last exercise of Chapter 3 as the custodianship spine. Commit the plan into the deposit as preservation_plan.md and update the metadata to reference it.
Exercise 22.4 (Final project deposit). Deposit your final project: the deepened technical artifact in preservation formats, the original and recast texts, the README with provenance, a datasheet, the preservation plan, and the critical-reflection memo from Chapter 23. Mint a DOI on Zenodo or another DOI-minting repository; an Internet Archive identifier does not count. A cross-linked Internet Archive mirror is optional. Verify the files, resolve the DOI from a logged-out browser, and put the full citation at the top of your final-project submission. Optionally add your other three portfolio pieces as a collection.
Exercise 22.5 (Open-ended). Propose a public-interest data infrastructure that should exist in 2035 but does not. A federated archive of state benefits-administration audit logs, a repository of municipal algorithmic impact assessments, a climate-resilient mirror of at-risk federal datasets: anything you have wanted to see while working through this book. Three to five pages. Address what it is, what publics it serves, what the first five years of building it look like year by year, what an archival-deposit strategy for its outputs entails, and who bears the maintenance labor across decades. You are not committing to build it.
22.12 Closing
A deposit is a promise to a reader you will never meet, and a promise needs an honest account of its limits. That is why the reflection memo travels inside the deposit. The last chapter, Chapter 23, turns the book’s framework on itself: whose records it prefers, whose refusals it postpones, and which levels of government it assumes will be there to read what you archived.
22.13 Further Reading and Resources
- Internet Archive, Python library documentation: https://archive.org/developers/internetarchive/. The canonical reference for the client used for this chapter’s mirror.
- Zenodo, REST API documentation: https://developers.zenodo.org. Authoritative for the deposit flow; test in the sandbox at https://sandbox.zenodo.org before going live.
- ICPSR, Guide to social science data preparation and archiving: https://www.icpsr.umich.edu/files/deposit/dataprep.pdf. Sixty years of social-science data stewardship, distilled.
- Open Science Framework, OSF help documentation: https://help.osf.io/. The project-scaffolded alternative to Zenodo.
- Software Heritage: https://www.softwareheritage.org/. The Internet Archive of source code; if your deposit includes code, cite its SWHID alongside the DOI.
- LOCKSS, Program documentation: https://www.lockss.org/. The federated-preservation backbone the chapter’s federation claim leans on.
- Environmental Data and Governance Initiative: https://envirodatagov.org/. Volunteer-driven federal environmental archiving; join a data-rescue event.
- Digital Preservation Coalition, Digital Preservation Handbook: https://www.dpconline.org/handbook. The practitioner reference for everything this chapter only sketches.
- Library of Congress, Sustainability of Digital Formats: https://www.loc.gov/preservation/digital/formats/. The registry that tells you, for any format, how scared to be about using it.
- Perma.cc: https://perma.cc/. For preserving web citations in a form that will resolve when the original URL rots.