Appendix B — AI Coauthorship and Responsible Disclosure
Generative AI tools are increasingly integrated into research, data analysis, and academic writing. This appendix does two things. First, it discloses how AI was used to write the book you are reading. Second, it offers practical guidance on when and how to disclose AI use in your own work.
The order is deliberate. An appendix that asked you to disclose your methods while concealing its own would not be worth reading.
B.1 How This Book Was Made
This book was drafted with substantial assistance from large language models, directed and reviewed by the author. The account below is reconstructed from the book’s public Git history, which records model attribution commit by commit.
B.1.1 The record
The repository contains 39 commits made between 15 April and 10 July 2026. Excluding merge commits, 29 of 36 commits carry a Co-Authored-By: Claude trailer naming the specific model that assisted:
| Model | Commits |
|---|---|
| Claude Opus 4.6 (1M context) | 14 |
| Claude Fable 5 | 13 |
| Claude Opus 4.7 (1M context) | 1 |
| Claude Sonnet 4.6 | 1 |
The seven commits without attribution are the initial repository scaffolding, the first drop of chapter drafts, a committed build of the rendered HTML, and small housekeeping changes. Attribution was not recorded for that first chapter drop, so the Git history cannot establish how much of it was machine-drafted; the sections that followed are documented precisely.
B.1.2 What the AI did, in three sessions
Session 1 — drafting and expansion (15–16 April 2026). An initial commit of 22 files established the chapters, the Quarto configuration, the bibliography, and two specification documents. Fourteen chapters were then expanded in a single sitting, each by its own commit with the message “Expand Chapter N … to ~3,000 words.” The first landed at 11:12 and the last at 11:33 — fourteen chapters in roughly twenty minutes, all assisted by Claude Opus 4.6. The commit bodies describe what each expansion added: new subsections, worked examples, exercises, and the “social history and public interest” material that closes each chapter. This work was merged as pull request #1.
Session 2 — build and publication infrastructure (23–26 April 2026). Claude Sonnet 4.6 and Claude Opus 4.7 assisted in replacing a committed copy of the rendered site with a GitHub Actions workflow that runs quarto render and publishes automatically. This removed roughly 34,000 lines of generated HTML from version control. Merged as pull request #2.
Session 3 — correction and modernization (10 July 2026). Thirteen commits between 14:27 and 14:52, all assisted by Claude Fable 5, revised the April drafts: Chapter 3 was rewritten as full narrative prose, Chapter 2’s legal landscape was updated, Chapter 12 was rebuilt around Spotify’s 2024 API deprecations, code fences were converted to executable Quarto cells, and the companion notebooks and their generator (tools/make_notebooks.py) were added. Merged as pull request #3.
B.1.4 What the correction pass found
The most useful part of this disclosure is not that AI helped write the book. It is what the July session had to fix in the April drafts, because those errors are characteristic of fluent machine-generated technical prose:
- A citation to something that does not exist. Chapter 11 referenced the FRED series
MHIBS08013A052NCEN. There is no such series —BSis not a state code. It was corrected toMHICO08013A052NCEN. - Code that reads plausibly but is wrong. A section combining two API sources reused a
urlvariable still pointing at the previous endpoint. The prose described one thing; the code did another. - Confidently stated figures that were incorrect. The FEC rate limit was given as hundreds of requests per minute; it is approximately 1,000 per hour.
- Examples overtaken by the web. Chapter 12’s platform code assumed API access that Spotify deprecated in 2024.
None of these are typographical slips. Each reads as authoritative and each is wrong — which is precisely why they survived the initial drafting and needed a dedicated pass to catch.
B.1.5 Limitations you should assume are still present
- The code in this book is not executed when the book is built. Chapters are configured
eval: false, so rendering never verifies that any example runs. Examples were checked by human and machine review, not by the build. When something does not work for you, that is a real possibility, not a misreading. - Live targets drift. Much of the book depends on websites and APIs that change without notice. Every example was correct when written; the web makes no promise beyond that.
- Errors of the kind listed above are likely to remain. One correction pass found several. It is unlikely to have found all of them.
- Citations warrant checking. Machine-drafted prose can attach a real-looking source to a claim it does not support. Verify anything you intend to rely on.
If you are reading this as a student in the course, the failure modes above are the target of the weekly textbook revisions. Finding a broken example, an unsupported claim, or an explanation that skips the step where you actually got stuck is a genuine contribution — open an issue or a pull request against the repository.
B.1.6 Verifying this account
This disclosure is reconstructed from public data, and you can check it. From a clone of the repository:
# Every commit and its model attribution
git log --format='%ad %s' --date=short
# Which commits name an AI coauthor, and which model
git log --format='%B' | grep -i 'Co-Authored-By: Claude' | sort | uniq -c
# What the July correction pass changed, in detail
git log --format='%B' --since=2026-07-01 --no-merges
# The expansion prompt that drove the April session (removed from the
# working tree, still in history)
git log --diff-filter=A --format='%H' -- CLAUDE_CODE_EXPANSION_PROMPT.md |
tail -1 | xargs -I{} git show {}:CLAUDE_CODE_EXPANSION_PROMPT.mdThis appendix was itself drafted with AI assistance, from the repository history described above, and reviewed by the author before merging — the same standard it asks of you.
B.2 When Disclosure Is Required
You should disclose AI use when:
- You used an LLM (ChatGPT, Claude, Gemini, etc.) to generate code, text, analysis, or visualizations that appear in your submitted work
- You used LLM APIs (as in Chapter 13) as analytical tools — for sentiment analysis, content coding, structured extraction, or similar tasks
- You used AI-powered code completion tools (GitHub Copilot, Cursor) for significant code generation beyond individual line completions
B.3 When Disclosure Is Typically Not Required
- Using AI for spell-checking, grammar correction, or basic proofreading
- Using AI to understand error messages or debug syntax errors (analogous to searching Stack Overflow)
- Using standard autocomplete features in code editors
B.4 How to Disclose: A Practical Framework
An effective disclosure statement answers four questions:
- What tool? Name the specific model and version (e.g., “GPT-4o via the OpenAI API, accessed in March 2026”)
- What task? Describe what the AI was used for (e.g., “sentiment classification of 200 bill summaries,” “initial draft of the data collection section”)
- What oversight? Describe how you reviewed, verified, and modified the output (e.g., “all generated code was tested independently,” “LLM classifications were validated against a 20% manual sample”)
- What limitations? Note any concerns about the output’s accuracy or appropriateness
The disclosure that opens this appendix is a worked example of that framework: it names the models and counts their commits, describes what each session produced, identifies the human specification and review that governed them, and states plainly what is likely still wrong.
B.5 Example Disclosure Statements
For a notebook assignment:
AI tools used: GitHub Copilot was used for code autocompletion throughout the notebook. ChatGPT (GPT-4o) was consulted to debug a BeautifulSoup parsing error in Section 3. All code was tested and validated independently.
For a research paper:
Text embeddings were computed using OpenAI’s
text-embedding-3-smallmodel. Sentiment classifications were generated by Claude Sonnet 4 (claude-sonnet-4-20250514) using the prompting strategy described in Section 3.2. A random sample of 50 classifications (25%) was manually validated, achieving 88% agreement with human coding.
For a final project:
This project used AI tools in three ways: (1) GitHub Copilot for code autocompletion, (2) GPT-4o-mini for classifying 500 legislative bill summaries into policy categories using the few-shot prompt in Appendix B, and (3) Claude for copy-editing the final report. All classifications were spot-checked against a 50-item manual sample (92% agreement). The scraping code, analysis pipeline, and visualization code were written by the author.
B.6 Publisher and Funder Requirements
Major academic publishers and funding agencies have established disclosure requirements:
- Elsevier, Springer Nature, Wiley, and Taylor & Francis all require disclosure of generative AI use but do not permit AI as a listed author. The consensus: AI tools can assist research, but humans bear full responsibility for accuracy and integrity.
- NIH (NOT-OD-23-149) prohibits peer reviewers from using AI tools to analyze grant applications, citing confidentiality concerns. Researchers are encouraged to disclose AI use in project descriptions.
- NSF suggests disclosure of AI use in proposals and will formalize requirements in future Proposal and Award Policies and Procedures Guides.
The landscape of publisher and funder policies continues to evolve. Always check the current requirements of your target journal or funding agency before submission.
Note what these policies have in common with the disclosure above: AI is credited as a tool in the record, but authorship — and accountability — stays with a person.
B.7 Core Principles
Three principles underlie responsible AI disclosure:
Accountability is unchanged. AI tools do not change the fundamentals of scholarly accountability. You are responsible for the accuracy, originality, and integrity of everything you submit, regardless of what tools you used to produce it.
Disclosure is about transparency, not permission. The goal is not to ask whether you are “allowed” to use AI, but to ensure that readers can evaluate your work knowing how it was produced. Could someone replicate your results knowing what tools you used and how?
The failure to disclose is the problem. Using AI tools responsibly and transparently is generally acceptable. Using them without disclosure, or in ways that misrepresent the origins of your work, compromises research integrity.
For the full AI disclosure statement used in this book’s companion resource, see Missing Manual Appendix B: Disclosure — Use of AI in Authoring.