Appendix B — Disclosure: Use of AI in Authoring
Purpose
A handbook that teaches you to use AI tools responsibly would be hypocritical if its own authors quietly pretended they hadn’t used them. This appendix is an open record of how large language models were used in drafting, editing, and maintaining the Missing Manual for Information Scientists, plus a discussion of the editorial choices we made about what the AI was and was not allowed to do.
We’re writing it in the voice of the human authors, Brian C. Keegan and Abram Handler, not the model’s. If you’re using this book in a class and want a concrete example of what an “AI disclosure” looks like for a piece of academic work, the first half of this appendix is one. The second half takes up the harder editorial and teaching questions the tools forced us to face.
B.1 What AI tools were used
During the drafting and conversion of this handbook we used:
- Claude (Anthropic), accessed through Claude Code, to draft prose, convert LaTeX source files into Quarto, and assist with editing passes. Claude wrote first drafts of several chapters in Chapter 7, Chapter 15, Chapter 20, Chapter 21, Chapter 18, Chapter 24, Chapter 34, Chapter 5, Chapter 19, Chapter 22, Chapter 23, and this appendix. Claude also drafted the chapter-to-meme template assignments now stored in each chapter’s
meme:frontmatter (curated and reviewed by the human authors); see Chapter 35 for related context. - ChatGPT (OpenAI), occasionally, as a second opinion during outlining.
- GitHub Copilot, inside the editor, for small-scale code completion in worked examples.
- A few earlier chapters were drafted by the human authors before AI assistance was introduced; those chapters were later edited with AI assistance but are substantively human-written.
We chose the models, providers, and versions for access and cost, not as an endorsement. A handbook written today would use somewhat different tools than one written a year ago, and we expect the list above to keep shifting.
B.2 What the AI was asked to do
We used the models for roughly five kinds of work, in decreasing order of editorial latitude:
- Mechanical conversion from LaTeX to Quarto. Chapters originally written in LaTeX were converted to
.qmdwith a pandoc + Python cleanup pipeline (the cleanup script was a one-off that never lived in the repository; the conversion itself is in the commit history). The AI wrote the cleanup script after being shown sample input and desired output. This was the safest and most deterministic use of the tools. - Drafting gap chapters. For a handful of chapters the human authors had already scoped but not written, we asked Claude to draft the prose from a detailed outline. We specified the book’s canonical chapter structure (then eight sections), the target audience, the tone (“friendly guide, second-person, empathetic”), the length, and the substantive points each section had to make. The model produced a first draft; the human authors then read, edited, reorganized, and verified the technical content.
- Editing and consistency passes. On existing chapters we asked the models to suggest reorganizations, tighten phrasing, catch inconsistencies in terminology, and flag places where cross-references were stale. These were suggestions, not changes.
- Code example review. Until September 2026, worked examples in Python, SQL, and shell were run or mentally executed by the human authors regardless of where they came from, and when the AI wrote code, the human authors verified behavior and edited for clarity. In the September 2026 rewrite described below, the agents ran the examples and made the output shown match what the code printed.
- Brainstorming and outlining. Before any prose was written for a gap chapter, we sometimes used Claude or ChatGPT as a brainstorming partner: “What are the seven things a novice data scientist gets wrong about CSVs?” The outputs informed our outlines but never replaced editorial judgment about what belonged in the book.
In September 2026 we used AI for a larger job: rewriting every chapter, and then the introduction, the conclusion, and these appendices, in the book’s current voice, with more links to official documentation and Wikipedia, and with the facts re-checked. The work was done by AI coding agents, sessions of Claude Code, working from written instructions: the style rules in the repository’s AGENTS.md and a plan, docs/plans/2026-09-25-voice-rollout.md, that gave every chapter’s agent the same brief and set out how each rewrite was checked before it was committed. The agents ran the code examples and made the output shown match what the code really printed, checked that every link resolved, and checked the facts they touched, softening or cutting claims they couldn’t source. Each part of the book then went to Brian as its own pull request, and he merged each one (#61 through #67, and the pull request that added this paragraph). The pass found and corrected real errors, about thirty in the seven chapters of Part I alone: code examples that didn’t run or didn’t fail the way the text said they would, outdated error messages, a shell command described as safe that wasn’t, advice about imports that didn’t work, and invented references, including a Further reading item whose authors didn’t exist and whose DOI didn’t resolve. The plan and the repository’s commit history record what changed in each chapter and why.
B.3 What the AI was not asked to do
Some decisions we explicitly kept out of AI hands:
- The book’s overall scope and table of contents. The human authors decided what parts the book would have and what the canonical chapter structure would be. Claude proposed candidate gap chapters during a plan-mode session, but the final choice of which to write, which to defer, and how to order them was ours.
- Claims about specific authors, papers, or attributions. Bibliography entries were curated and verified by the human authors. When the models invented a plausible-sounding citation, we removed it. See the section on what we got wrong.
- Factual claims about specific tools’ behavior. We cross-checked any claim of the form “X does Y” against the tool’s documentation before shipping it. This was particularly important for pandas method signatures, pytest features, SQL keyword support, and HTTP status codes, where the model’s training data may reflect an older version of reality.
- Representation of our own experiences as teachers. The anecdotes and observations about “students in our courses” are from the human authors’ actual experience. When the model drafted a chapter that included a “students often…” claim, we either verified the claim against our experience or cut it.
- Judgments about pedagogical tradeoffs. “Should we teach
condaorvenvfirst?” is a judgment call that requires knowing our students. We made those calls ourselves.
B.4 Why we think this is OK
The case for using AI tools in authoring a textbook — especially a textbook about computing practices — comes down to three things:
- Speed matters in a fast-moving field. Computing tooling changes every year. The marginal cost of a well-placed chapter on virtual environments or pre-commit hooks is lower with AI assistance, which means more chapters get written, fewer gaps linger in the book, and students get better coverage of topics that would otherwise be deferred forever.
- We retain editorial responsibility. A draft is not a decision. Every paragraph written before September 2026 was read by a human author with discretion to rewrite, reorganize, or delete. The September 2026 rewrite came to Brian one part at a time, as pull requests he could change or refuse, and he merged each one. The AI is a fast junior writer, not a co-author; the buck stops with us.
- Transparency is the right norm. Many textbooks use AI tools and don’t say so. We’d rather disclose than let you guess. If our disclosure turns out to be more generous than our peers’, so much the better — it gives readers a clear basis to evaluate our choices.
B.5 What we got wrong, and what we watched for
Being honest means admitting the failure modes we hit. Here they are, roughly from most to least common:
- Plausible-sounding but incorrect API details. The models would confidently describe a parameter to
pd.read_csvor a pytest feature that did not exist in the version we targeted, or that had been renamed. Every code example in the book was checked against documentation or executed. - Fabricated citations. Once or twice, we asked for “a paper on X” and received a plausibly-formatted reference that did not exist. Every bibliography entry in
references.bibcorresponds to a real source that one of the human authors verified. - Over-confident generalizations. Draft prose sometimes made sweeping claims like “most data scientists do X” or “this is the standard approach,” where “most” was the model’s guess rather than a substantiated fact. We softened or removed such claims.
- Lost voice and tone drift. A chapter drafted in a single pass would sometimes slip from our “friendly guide” tone into something more formal or generic. Editing passes by the human authors restored the voice.
- Uneven depth. The model sometimes gave equal weight to a trivial topic and a critical one. Re-outlining and shortening or expanding sections was a common human edit.
None of these disqualifies the tools. They’re workflow problems, and we solved them with editing. But they’re worth naming, so that you can judge the risk for yourself.
B.6 What this means for students using the book
If you’re reading this book to learn computing, three things are worth knowing:
- The code examples have been checked. Before September 2026 the human authors executed or traced them; in the September 2026 rewrite the agents ran them and matched the output shown to what the code printed. If you find one that doesn’t work, please open an issue: use the “Report an issue” link on any chapter page, and the Something is wrong form will walk you through what to include.
- The technical claims have been checked against documentation. We expect a few errors to remain (nothing this long is perfect), and we’ll fix them as they’re reported.
- The pedagogical judgments — what to teach, in what order, with what emphasis — are the human authors’ choices, informed by teaching students like you.
Bring the same skeptical habits we describe in Chapter 35 and Chapter 38 to this book and to anything else you read, however it was written.
B.7 Discussion: why we disclose at all
“Does it matter?” is a fair question. A textbook is supposed to be correct and clear, so what does it matter where each sentence came from?
We think it matters for three reasons, which correspond to three different audiences.
For students. One of the skills this book is trying to build is a literate relationship with AI tools: when to use them, how to verify their output, how to read around their mistakes, and how to disclose them when they are part of your own work (see Chapter 35). If we failed to model that ourselves, the advice in those chapters would be hypocritical. Disclosure is the worked example for the standard we are asking students to adopt.
For instructors. An instructor adopting this book for a course has a right to know what kind of artifact they are teaching from. Some instructors will be more comfortable with this than others. We want to make it possible for them to make that choice knowingly, rather than quietly.
For the broader academic community. Norms around AI authorship are still forming. Journals, publishers, and universities are all working out their own disclosure requirements, and the resulting guidelines are inconsistent and sometimes contradictory. By writing down what we did and why, we contribute one data point to that ongoing conversation. We don’t expect our practices to become the standard, but we’d rather be on the record with a defensible position than hope nobody asks.
B.8 A note on the code we ship
A separate but related question is how AI tools were used in any software we publish alongside the book — currently the helper tools in the repository’s tools/ folder (the meme and figure generators, the issue-form script, a screenshot toolkit adapted from a companion book, and browser checks of the page layout) and the _quarto.yml / CI workflow configuration. Those were drafted with AI assistance and then tested: the original LaTeX conversion script was run end-to-end on every chapter, each tool is checked against the files it produces, and the render pipeline was validated by a clean HTML build with zero warnings. The same “draft, verify, own the result” discipline applies.
If this book grows to include executable code cells (see Chapter 16) or tests that ship with the book itself, the same standard will apply to them: AI may draft, humans verify and take responsibility.
B.9 How we will keep this appendix current
AI tools change. The list of tools we used, the ways we used them, and the kinds of mistakes they made are all moving targets. When we revise the book — adding a chapter, updating an example, responding to reader feedback — we will also update this appendix to reflect the new state of practice. The last substantial update was in September 2026, when we added the account of the voice rewrite above.
If you notice a real gap between what this appendix says and what the book obviously contains, that’s a bug, and we’d like to hear about it.
B.10 Further reading
If you’re interested in the broader conversation about AI and scholarly authorship, here are some places to start:
- The chapters Chapter 35, Chapter 36, Chapter 37, and Chapter 38 in this handbook itself.
- Your own institution’s policies on AI assistance in academic work, which are almost certainly stricter in the student-assignment context than ours are in the textbook-authoring context. Don’t assume our disclosure gives you cover for your own assignments.
- The editorial policies of the journals and publishers in your field. Many now require authors to disclose AI use, and they disagree about what counts as enough; compare Nature Portfolio’s policy on AI with the guidance on AI-assisted technology from the International Committee of Medical Journal Editors.
- Prior essays and position papers on the topic, which are voluminous and moving quickly. Use your library rather than a search engine for the authoritative ones.
B.11 Checklist for your own work
If you ever need to write a disclosure like this one for your own course work or research, these are the questions we found useful to answer:
- Which models or tools did you use?
- What tasks did you use them for? Be specific — “drafting,” “editing,” “brainstorming,” “code completion,” “translation.”
- What tasks did you explicitly not use them for?
- What verification did you do on the output?
- What mistakes did the tools make, and how did you catch them?
- Who takes responsibility for the final result? (Usually: you.)
A disclosure of that shape answers the reasonable questions a reader might have and signals that you understand the tools well enough to use them responsibly.