21 Research Proposals
The op-ed in Chapter 7 argued for a thing that already exists; the report in Chapter 15 documented a thing that had already happened. The research proposal is stranger than either. It argues, in the present tense, for the significance of work that does not yet exist, using evidence you cannot yet produce, on behalf of findings you cannot yet guarantee, to readers who will decide whether the work should exist at all. You are asking a funder to believe in a future that is currently a document.
In this book the proposal does two jobs. First, it is a fifth genre for the final project. Graduate students may take one of their four portfolio pieces (the city op-ed, the county testimony, the state report, or the federal public comment), deepen it, and recast it as a proposal for the larger study the piece points toward. Second, it is the scholarly backbone of the graduate methods memo that accompanies each portfolio piece. The memo’s two required moves, situating the piece in at least five scholarly sources and defending one methodological choice against that literature, are the “gap” and “approach” moves of a proposal in miniature. Undergraduates can read the chapter as a preview of how research gets paid for.
The hard part is translating a public-interest data question into a shape a funder can recognize, and translating your own working life into a budget a program officer can defend to a board. By the end of the chapter you will have a methods memo template and the four documents a final-project proposal needs: a one-page specific-aims document, a budget with a narrative justification, a data management plan (DMP), and a simulated Institutional Review Board (IRB) protocol. None is a submitted grant. All are the documents a submitted grant is made of.
21.1 The grant-proposal genre landscape
The funders you might realistically approach divide into four ecologies, each with its own vocabulary and its own review cycle.
Federal science and humanities agencies include NSF, NIH, IMLS, and NEH. They publish detailed proposal guides (National Science Foundation 2024), use peer-review panels, require formal budgets with indirect-cost recovery, and expect data management plans and broader-impacts sections. As of 2025 they are volatile ground: appropriations uncertainty and agency restructuring mean a proposal submitted under one quarter’s rules may be reviewed under the next’s (Institute of Museum and Library Services 2024; National Endowment for the Humanities 2024; National Institutes of Health 2024).
National private foundations include Knight, MacArthur, Ford, and Robert Wood Johnson. They run on program cycles, usually expect a letter of inquiry first, and keep review standards less public than NSF’s; you are writing for a program officer who is writing for a board (John S. and James L. Knight Foundation 2024).
Regional and community foundations include the Gates Family Foundation in Colorado, the Boettcher Foundation, the Denver Foundation, and analogues in every state (Gates Family Foundation 2024; Boettcher Foundation 2024). Grants run smaller (often $10,000 to $75,000), cycles are shorter, and priorities are locally specific. For a first public-interest research proposal, they are often the right target.
University seed grants are the quietest category and frequently the most accessible. A research development office will fund a $5,000 to $25,000 pilot with a two-page application. A pilot result is the strongest possible leverage for the next proposal up the chain.
21.2 The political economy of what gets funded
The gap between the money available for public-interest research and the projects that could usefully spend it is large and does not close. A typical program officer declines most of the projects they would like to fund.
Two implications follow. Specificity beats scope: a proposal promising to solve algorithmic discrimination in public-benefits systems will lose to one promising to audit the adverse-action notices issued by one county’s SNAP system over eighteen months. The module ladder helps here. A piece already pinned to a city, county, state, or federal body arrives at the proposal with its scope half-written. And if your specific aims are not legible in ninety seconds, your proposal will not get the second read.
Political economy also shapes what counts as “research.” A community-engaged audit whose primary output is a community report is harder to fund through NSF than through a regional foundation. A refusal specification whose output is a decision not to build a dataset is harder to fund than either. Part of the proposal-writing work is finding the funder for whom your project is recognizable.
21.3 The specific-aims page as sales pitch
The specific-aims page is the most important page of any research proposal: the first thing a reviewer reads, often the last, and the basis of an opinion that is then difficult to move. The report’s executive summary from Chapter 15 is its closest sibling, but runs the opposite direction: the executive summary distills work already done; the specific-aims page promises work not yet done and argues that it is worth paying for.
A conventional specific-aims page has four moves. First, a significance paragraph naming the public-interest problem in the vocabulary the funder uses. Second, a gap paragraph naming what is not yet known, or audited, or refused. Third, two or three numbered aims, each with a one-sentence hypothesis or deliverable. Fourth, an impact paragraph naming the actors who will act on the results. If you are recasting a portfolio piece, those actors are usually the body your piece already addressed: the city council, the Board of County Commissioners, the state committee, the federal agency.
First-time authors miss the first move most often. The significance paragraph is not a general statement about why data matters; it is a claim written in the exact vocabulary of the funder’s open call. If the NSF program says “trustworthy AI,” your first paragraph says “trustworthy AI.” If the Knight Foundation call says “local news ecosystems,” your first paragraph says “local news ecosystems.” This is not cynicism. It is translation.
The specific-aims page is not a summary. It is a sales pitch, written in the vocabulary of the field you are selling to, by a person who has read the funder’s last three years of awards and knows what rhymes with them. Rewrite it five times, once as if you were the skeptic on the review panel. Show each draft to a colleague outside your topic and ask what you are proposing and why a funder should care. If they cannot tell you, the page is not done. And: the budget narrative matters more than the budget numbers. A plausible story about personnel time beats a precise calculation that assumes the work will take half as long as it will. Reviewers have seen the half-time calculation. They are not impressed by it.
21.4 Budgets and justifications
A proposal budget has three components. Personnel covers salary and fringe for the people doing the work (PI, co-investigators, graduate students, hourly undergraduates, a contracted community liaison). Direct costs cover everything else accounted for line by line (software licenses, compute credits, travel, honoraria, transcription, modest equipment). Indirect costs, or facilities and administration (F&A), are overhead the institution charges to keep the lights on. Federal indirect rates at research universities run from roughly 25 to 65 percent of modified total direct costs; foundations often cap indirects at 10 to 15 percent, and some refuse them entirely. Know the rate before you start.
A simple one-year, $50,000 project budget, at a regional foundation that caps indirects at 15 percent, might look like this:
import pandas as pd
budget = pd.DataFrame([
{"category": "Personnel", "item": "PI (0.10 FTE, summer)", "amount": 14000},
{"category": "Personnel", "item": "Graduate RA (0.25 FTE, 9 mo)", "amount": 15000},
{"category": "Personnel", "item": "Fringe benefits (22%)", "amount": 6380},
{"category": "Direct", "item": "Community-liaison honorarium", "amount": 3000},
{"category": "Direct", "item": "Participant honoraria (10 x $75)", "amount": 750},
{"category": "Direct", "item": "Compute and storage (AWS / Zenodo)", "amount": 1200},
{"category": "Direct", "item": "Travel (2 site visits)", "amount": 1500},
{"category": "Direct", "item": "Publication and open-access fees", "amount": 1650},
{"category": "Indirect", "item": "F&A at 15% of MTDC", "amount": 6520},
])
totals = budget.groupby("category")["amount"].sum()
print(totals)
# => Direct 8100
# Indirect 6520
# Personnel 35380
# Name: amount, dtype: int64
print("Total:", budget["amount"].sum())
# => Total: 50000The numbers are not the hard part; the narrative justification is. Every line needs a sentence or two in a companion document saying what the money pays for and why that level of effort is necessary. The graduate RA line is not “0.25 FTE for nine months at the university rate.” It is “25 percent effort over nine academic months to conduct the FOIA log coding, build the reproducible pipeline described in Aim 2, and co-author the final community report.” A reviewer should come away knowing what each person will actually do on each Tuesday afternoon.
Write the narrative first and fit the numbers to it; a vague narrative makes reviewers assume vague numbers.
21.5 The data management plan
Federal funders require a data management plan (DMP) in every proposal, and foundations increasingly ask for one (California Digital Library 2024). The DMP is not a formality: it tells the funder, in two pages, what data your project will produce, where it will live, who can access it, for how long, and under what conditions. The DMPTool, operated by the California Digital Library, provides templates for most major funders (California Digital Library 2024).
A serviceable DMP covers six sections.
Types of data. What you will collect or generate. Be specific. “Administrative data from the Colorado Department of Human Services under a data use agreement” is specific. “Public records” is not.
Formats and standards. CSV, Parquet, GeoJSON, JSON-LD, whatever applies. If you follow a community standard (HUD’s HMIS data dictionary, or the FAIR Guiding Principles (Wilkinson et al. 2016)), name it.
Storage during the project. University-managed storage under a DUA, encrypted at rest, with access restricted to named project personnel.
Retention and destruction. How long you keep the data, and what you do with it at the end. A state-agency DUA typically requires destruction of identifiable data within ninety days of project close.
Sharing. What you will make public, and where. A de-identified derived dataset on Zenodo (CERN 2024) with a CC-BY license and a DOI; a code repository on GitHub linked from the dataset record; a datasheet following Gebru and colleagues (Gebru et al. 2021). If something will not be shared, say why.
Preservation. Where the shared outputs live after you stop paying attention. Zenodo and the Internet Archive provide preservation commitments, but the deposit every final project requires must carry a DOI, such as the one Zenodo mints (Chapter 22). A GitHub repo you maintain is not a preservation plan.
The DMP is where you operationalize the installed base for your own project: linkability in formats and standards, interpretability in the datasheet, continuity in preservation, safe scrutiny in sharing.
21.6 The IRB protocol as a sibling genre
The IRB protocol sits beside the proposal but is not part of it. A proposal is a budget document; an IRB protocol is a risk document arguing the work, if funded, will not harm the humans it touches. Different office, different reader, different conventions.
The framing texts are the Belmont Report (National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research 1978) and the revised Common Rule (US Department of Health and Human Services 2018). Belmont’s three principles (respect for persons, beneficence, justice) structure every IRB form you will fill out. Your local IRB (at CU Boulder, the Office of Research Integrity (University of Colorado Boulder Office of Research Integrity 2024)) operationalizes them through a protocol template covering study population, recruitment, consent, data handling, risks, benefits, and conflict-of-interest disclosures.
Public-interest data science frequently works at the edge of the IRB’s jurisdiction. An audit of administrative data under a DUA may be “not human subjects research.” A community survey unambiguously requires full review. A scrape of public social-media posts lands in a contested middle ground whose answer depends on your institution, your IRB chair, and the year. When in doubt, ask. An IRB that tells you a project is exempt is a better friend than an IRB that finds out about your project in the press.
21.7 Community-engaged research as a proposal-design problem
If the work involves a community partner (a neighborhood association, a legal-aid clinic, a union local, a tribal nation’s data office), the proposal has a different shape. The partner is a collaborator, not a subject; their time is an in-kind contribution or a subcontracted line item; their name belongs on the proposal; their data governance sets the terms.
Costanza-Chock’s design-justice framing (Costanza-Chock 2020) and the CARE Principles for Indigenous data governance (Carroll et al. 2020) are the two sources most worth reading before you draft a community-engaged proposal. They change the budget (honoraria become line items), the DMP (data-sovereignty provisions become load-bearing), and the aims (the community’s question is Aim 1).
21.8 The methods memo: a proposal in miniature
Every graduate portfolio piece carries a 750-to-1,000-word methods memo. It answers the question a review panel asks of a proposal’s approach section: why this method, given what the field already knows? The memo has three parts. Situate the piece in at least five scholarly sources, saying what each claims and where your piece agrees or departs. Name one methodological choice you made (a metric, a sampling frame, a unit of analysis, a decision not to collect) and the alternative the literature would favor. Defend your choice, or concede it, in light of that literature.
A literature matrix keeps the memo honest. Here is one for a Piece 3 disaggregated audit whose key choice was to report false positive rate gaps rather than calibration:
import pandas as pd
matrix = pd.DataFrame([
{"key": "buolamwini2018", "claim": "Disaggregation exposes intersectional error gaps",
"stance": "supports"},
{"key": "chouldechova2018", "claim": "A deployed county tool shows metric choices in practice",
"stance": "complicates"},
{"key": "raji2020", "claim": "Audits need an end-to-end accountability frame",
"stance": "supports"},
{"key": "goodman2023", "claim": "Audits can launder systems they claim to check",
"stance": "challenges"},
{"key": "selbst2019fairness", "claim": "Metric choice abstracts away social context",
"stance": "challenges"},
])
print(len(matrix), "sources")
# => 5 sources
print(matrix["stance"].value_counts())
# => stance
# supports 2
# challenges 2
# complicates 1
# Name: count, dtype: int64The check that matters is the last one. A matrix in which every source “supports” your choice is a reading list, not a literature review. If nothing challenges you, you have not read far enough. The memo you write from this matrix is also the seed of a proposal’s gap paragraph: the place where the challenging sources leave a question open is the place your next study lives.
21.9 Broader impacts, taken seriously
NSF requires a broader-impacts section; most private foundations require something analogous. First-time authors treat these as boilerplate, and reviewers read past it.
The strong version asks who will be better off if the project succeeds, in a way that would not have happened otherwise. It names community partners, open datasets, reproducible code, training opportunities for students who would not otherwise have had them, and downstream uses the research makes possible. It ties to the ownership question from Chapter 18: whose problem is this, and who gets to act on the answer? And it anticipates the sustainability question from Chapter 3: when the grant money is gone, does the infrastructure continue, or evaporate with your last invoice?
The broader-impacts section is where “public interest” becomes concrete and accountable. Tie it explicitly to ownership as you built it in Chapter 18: the community whose data you are using is a partner with a stake in the outcomes, not a population to be studied. And tie it to sustainability as you interrogated it in Chapter 3: a funded pilot that dies when the grant ends has not advanced the installed base; it has spent the trust of staff and community members who will be asked to engage the next pilot with less of it. A section that names the partners, the deliverables they will own at the end, and the maintenance plan that survives the grant is one a reviewer will remember. A section that promises mentorship and open access without naming who, by whom, or where reads as filler.
21.10 Exercises
Exercise 21.1 (Structural reading, real public data). Query the NSF Awards API (https://api.nsf.gov/services/v1/awards.json) with requests for a keyword close to your portfolio piece (for example "public interest technology", "open government data", or "algorithmic accountability"), load the award list into a pandas DataFrame, and keep id, title, awardeeName, fundsObligatedAmt, startDate, and abstractText. Pick three abstracts and mark where significance, gap, aims, and impact appear. In 400 words, describe how the aims echo the program’s vocabulary. Save the notebook and exercises/ch21_nsf_aims_diagram.md.
Exercise 21.2 (Methods memo, graduate). For your most recent portfolio piece, build the literature matrix from this chapter with at least five sources, at least one of which challenges your key methodological choice. Write the 750-to-1,000-word methods memo from it. Submit it with the piece.
Exercise 21.3 (Specific-aims page). If you are using the proposal genre for the final project, draft a one-page specific-aims document for a one-year, $50,000 study that grows out of the portfolio piece you are deepening. Target a plausible funder at the right scale: the Gates Family Foundation, the Boettcher Foundation, or a university seed-grant program. Read three of their recent awards first. Use the four-move structure and name the government body from your piece in the impact paragraph. Rewrite five times. Save the fifth draft as proposal/specific_aims.md.
Exercise 21.4 (Budget and DMP). Produce a budget (adapt the pandas example or write proposal/budget.csv) with personnel, direct costs, and indirects at the funder’s actual cap, and a one-page narrative justification. Then draft a two-page DMP with the DMPTool template (California Digital Library 2024), covering all six sections, naming the datasheet standard (Gebru et al. 2021), and naming as the preservation host the Zenodo record and DOI you will create in Chapter 22. Save as proposal/budget_justification.md and proposal/dmp.md.
Exercise 21.5 (Simulated IRB protocol and assembly). Draft a simulated IRB protocol covering consent, risk, data handling, benefits, and conflicts of interest, grounded in the Belmont Report (National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research 1978) and the revised Common Rule (US Department of Health and Human Services 2018). If the project is probably exempt, say so and explain why. Save as proposal/irb_protocol.md. Then assemble the specific-aims page, budget, DMP, and protocol as your recast final-project text, ready to deposit alongside the deepened technical artifact.
21.11 Closing
A proposal is a promise about future work, and a promise is only as good as the record it leaves. The next chapter turns your final project, proposal or otherwise, into that record: a deposit with a DOI (Chapter 22). The chapter after that asks what the whole framework, including the proposal genre’s habit of translating every problem into fundable aims, cannot see (Chapter 23).
21.12 Further Reading and Resources
- NSF Proposal and Award Policies and Procedures Guide (PAPPG) (National Science Foundation 2024), https://www.nsf.gov/publications/pub_summ.jsp?ods_key=pappg. The authoritative US federal reference.
- DMPTool, California Digital Library (California Digital Library 2024), https://dmptool.org. The simplest way to draft a data management plan that fits the conventions of most major funders.
- The Belmont Report (National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research 1978), https://www.hhs.gov/ohrp/regulations-and-policy/belmont-report/index.html. The framing text for every IRB protocol you will ever write. It is short. Read it.
- Knight Foundation grants page (John S. and James L. Knight Foundation 2024), https://knightfoundation.org/grants/. Representative of the national private-foundation proposal ecology for journalism and civic-technology work.
- Foundation Directory Online (Candid 2024), via most research-library subscriptions. Spend an afternoon with it before you start writing.
- Gates Family Foundation (Gates Family Foundation 2024) and Boettcher Foundation (Boettcher Foundation 2024) grant pages. Representative Colorado regional foundations; the closest live targets for the specific-aims page in Exercise 21.3.
- Goldsmith and Kleiman (Goldsmith and Kleiman 2017) and Robinson (Robinson 2017). Two very different books on proposal writing, both useful.
- Revised Common Rule, 45 CFR 46 (US Department of Health and Human Services 2018), https://www.hhs.gov/ohrp/regulations-and-policy/regulations/45-cfr-46/. Skim the subpart A definitions before Exercise 21.5.
- Your university’s research development office and its IRB template. For Colorado Boulder readers, the Office of Research Integrity (University of Colorado Boulder Office of Research Integrity 2024). The most useful local resource, and the one most students discover too late.