3  Technical Documentation

Prerequisites: none. This chapter stands on its own.

See also: Chapter 2, Chapter 31, Chapter 32, Chapter 25, Chapter 29.

Purpose

Put It Somewhere Else Patrick Meme: Why don’t we take everything we know and write it down somewhere?

Here’s a situation almost everyone lands in sooner or later. You copy a few lines from a tutorial, run them, and get AttributeError: 'DataFrame' object has no attribute 'append'. You search the error, open six tabs, and find six answers that each say something slightly different. An hour later it works, but you couldn’t say why, and the next error starts the whole cycle over. Or you flip it around: a teammate clones your project and messages you, “how do I run this?”, and you realize the answer has only ever lived in your head.

Both problems are about documentation, from opposite ends. Computing courses teach you what to type, but rarely how to find out what to type next, and almost never how to write things down so that someone else (including you, next month) can pick up where you left off. That’s not a gap in your ability. It’s a skill nobody sat you down to teach, and it’s a learnable one.

This chapter covers both ends. It shows you how to find the right docs quickly, how to read them with the right mix of trust and skepticism, how to line them up with the versions on your own computer, and how to write READMEs, runbooks, diagrams, and decision logs that people actually use. Chapter 5 goes deeper on the anatomy of an official docs site, and Chapter 2 picks up where the docs run out and you need to ask a person.

Why read this chapter

  • You “read the docs” and still didn’t get it, and you’ve started to suspect the problem is you (it usually isn’t: you were reading the wrong kind of docs).
  • You copied an example from a blog post, got AttributeError: 'DataFrame' object has no attribute 'append', and only later learned the post was written for an older version of pandas.
  • You searched “how do I …” and ended up with a dozen tabs of answers that disagree with each other, and no way to tell which one is right.
  • A teammate cloned your repository and asked “how do I run this?”, and you realized the answer was nowhere but your own memory.
  • You came back to your own project after a break and couldn’t remember which script to run first, or in which environment.
  • Your instructor wants a README, a data notes section, or a diagram of your pipeline, and you’re not sure what goes in one.
  • Someone showed you a diagram full of crow’s feet and dotted arrows, and you nodded along without knowing how to read it.
  • You’d like an AI tool to draft your docs, without it inventing a command-line flag that doesn’t exist.

Running theme: documentation is an interface

Documentation is the connection between what you meant and what someone else understands, whether that someone is a teammate, a future version of you, or a tool like a build system. Good docs make the right action easy and the wrong one hard to stumble into.

3.1 Documentation is a stack, not a single thing

People talk about “the docs” as if there were one. In practice there’s a whole stack of them, written by different people for different questions: the official reference, a tutorial someone wrote for beginners, the project’s README, a forum answer from 2019, the comments in the code. If you’ve ever felt lost in documentation, part of what’s happening is that you’re asking one layer of that stack a question another layer was built to answer. Once you see the layers, confusion stops being a personal failing and becomes a routing problem: which source should answer this question?

Learning and coordination

Documentation does two jobs, and they pull in different directions. One is learning: teaching you how a tool works and how to use it. The other is coordination: keeping several people agreed on how a project is run, changed, and checked. The same document rarely does both well, and a lot of frustration comes from asking a document to do the job it wasn’t written for. An API reference makes a poor tutorial because it assumes you already know what you’re looking for. A tutorial makes a poor specification because it follows one path through the tool rather than listing every behavior. A README can’t stand in for a changelog, because it describes the project as it is now, not how it got that way. Knowing which job you need done is half the battle.

Which source wins when they disagree

Sooner or later a blog post will tell you one thing and the official docs another. When that happens, rank your sources by how close they are to the thing itself. Primary sources come first: the official documentation, man pages, the tool’s own source code, and published specifications. Next come project-local sources: your repository’s README and CONTRIBUTING guide, comments in the code, and the project’s issues. Secondary sources come last: blog posts, videos, forum answers, and anything an AI tool tells you.

That doesn’t make secondary sources bad. They’re often where a beginner’s explanation lives, written by someone who remembers being confused. But they’re more likely to be out of date or written for a situation that isn’t yours. Use them as navigation aids that point you toward the primary source, not as the last word.

3.2 Documentation genres and the questions they answer

The most useful map of those layers comes from the Diátaxis framework, which sorts documentation into four genres by the need each one serves. You don’t need to memorize the labels. What helps is being able to tell, a few seconds into a page, which kind you’re reading and whether it’s the kind you need.

Reference documentation

Reference docs answer “what exactly are the inputs, outputs, options, and behaviors?” That covers API references for functions, classes, and parameters; command-line help and man pages; and the schemas of configuration files. Reference is dense, structured, and not written as a story. You don’t read it cover to cover; you dip in for one precise answer and leave. It’s the right genre when you already know which function you want and just need to confirm an argument.

Some of the most useful reference is already on your computer:

# Reference docs you can read without a browser
git help merge          # git's full reference page for merge (also online at https://git-scm.com/docs)
ls --help               # GNU ls option reference (Linux; on a Mac, use `man ls`)
python -c "help(dict)"  # Python's built-in reference (https://docs.python.org/3/)

Tutorials

Tutorials answer “how do I learn this from scratch, with someone guiding me?” They’re linear and scaffolded, and they assume you have time to follow a sequence step by step. A good one gets you to something working quickly, even if you don’t understand every piece on the first pass. The pandas “Getting started” guide and the official Python tutorial are both good examples.

How-to guides

How-to guides answer “how do I get this specific thing done?” They assume you already know roughly what you want (“I need to read a CSV that uses semicolons instead of commas”) and hand you the steps without a lecture. The best ones are short, focused, and opinionated about the right way to do it.

Explanations

Explanations answer “why does this work the way it does?” and “how should I think about this?” They cover terminology, architecture, trade-offs, and the reasons behind a design. They’re often where you learn what not to do: why looping over a DataFrame’s rows is slow, why a mutable default argument is a trap in Python, or why an HTTP POST isn’t idempotent the way a GET is. Reading explanations produces nothing you can run right away, which is why students skip them. It’s also why the same conceptual mistake keeps coming back for students who do.

Why the genres matter

When you say “I read the docs and still don’t get it,” the hidden question is almost always which kind of docs you read compared with the kind you needed. A quick diagnosis helps. If you can’t even find the name of the function you need, reach for a tutorial or a how-to guide to point you at it. If you know the name and just need the right parameter, go to the reference. And if you keep making the same kind of mistake (the same join going wrong, the same KeyError in a new place), you need an explanation that fixes your mental model, not another how-to that patches one symptom.

3.3 Finding documentation efficiently

Finding the right page is harder than it sounds, because the web holds far more about any popular tool than you could ever read, and search engines rank pages by popularity, not correctness. The fix is a routine that reliably gets you to the primary source.

Start with “what kind of thing is this?”

The first move is to work out what kind of thing you’re dealing with, because that decides where its official docs live. Is it a language feature, like list slicing in Python? A library, like pandas, NumPy, or scikit-learn? A tool, like Git, conda, or Jupyter? A platform, like GitHub Actions or Windows Task Scheduler? Or a file format, like CSV, JSON, or Parquet? Each category has its own home. Language features live in the language reference, libraries on their own documentation sites, tools in their man pages and project sites, platforms in their vendor’s docs, and file formats in their public specifications. Once you know the category, you can usually guess the URL.

Search in a way that finds the official page

Here’s the trap nearly everyone falls into: you type “how do I do X” into a search engine and land in a sea of secondary sources, blog posts and forum threads and videos of every age and accuracy. A small change to the query routes you to primary sources instead. Rather than “how do I write a requirements.txt”, search for pip documentation requirements.txt, which lands you on pip’s page for the requirements file format. Rather than “how do I use SSH ProxyJump”, search for ssh man page ProxyJump. The pattern is always the same: name the tool, name the topic, and add a word like documentation, API, or man page to tilt the results toward the source.

# Instead of                          search for
# "how to use pandas merge"    →     "pandas API merge"
# "git rebase tutorial"         →     "git documentation rebase"
# "fix requests timeout"        →     "requests timeout parameter site:requests.readthedocs.io"

You’ll still land on a blog or a forum answer sometimes, and that’s fine. Treat it as a map: pull out the function names, flags, and concepts it mentions, then go straight to the official docs to check them against the version you have installed.

Know the home bases

A small set of sites comes up again and again in data science work, and it’s worth recognizing them on sight. For installing Python packages and managing environments, the authorities are the Python Packaging User Guide, pip’s user guide, and conda’s guide to managing environments (Python Packaging Authority, n.d.; pip developers, n.d.; Conda Project, n.d.). For notebooks, it’s the documentation of Project Jupyter and JupyterLab (Project Jupyter, n.d.-b, n.d.-a). For version control, it’s the free book Pro Git (Chacon and Straub 2014). You don’t need to memorize URLs. You do need to recognize which organizations write the authoritative reference for the tools you use, so that a search result from one of them stands out.

Python Packaging Authority. n.d. Python Packaging User Guide. Documentation. https://packaging.python.org/.
pip developers. n.d. User Guide. Pip documentation. https://pip.pypa.io/en/stable/user_guide/.
Conda Project. n.d. Managing Environments. Conda documentation. https://docs.conda.io/docs/user-guide/tasks/manage-environments.html.
Project Jupyter. n.d.-b. Project Jupyter Documentation. Documentation. https://docs.jupyter.org/.
Project Jupyter. n.d.-a. JupyterLab Documentation. Documentation. https://jupyterlab.readthedocs.io/.
Chacon, Scott, and Ben Straub. 2014. Pro Git. 2nd ed. Apress. https://doi.org/10.1007/978-1-4842-0076-6.

Check your own computer before the web

Before you open a browser, look at what’s already installed. Almost every tool ships with its own reference, and that reference has one big advantage over anything online: it matches the version you actually have. From the command line, try tool --help, man tool, or help tool. Git’s subcommands each have a full help page (git help merge, git help commit), and git merge -h prints a short summary of the options instead. In Python, the built-in help() works on any function or class and prints its docstring, the description its author wrote. Inside Jupyter or IPython it’s even quicker: add ? after a name to see its docstring, or ?? to see its source code, as the IPython tutorial shows. R users have ?function and help(function).

git help rebase                  # full reference page for git rebase
ls --help | head -20             # the first 20 lines of ls's options (Linux)
python -c "help(dict.get)"       # built-in help for dict.get

Two small snags to know about. The Mac’s ls doesn’t understand --help and prints a short usage line instead, so use man ls there. And on some Macs, python doesn’t exist and you need python3. Local help is fast, works offline, and is always right about your version, so reach for it before a search engine.

3.4 Reading documentation actively

Reading technical docs isn’t like reading a novel, and if you’ve ever read a docs page top to bottom and come away with nothing, that’s why. You don’t take it in from start to finish. You question it.

Pull out five facts

When you’re reading about a function, command, or tool, you’re hunting for five things. Inputs: which arguments are required, and what types they accept. Outputs: what comes back, what files get written, and what else changes along the way. Defaults: what happens if you don’t pass anything special. Constraints: version requirements, operating-system limits, and permissions. Failure modes: which errors it can raise and what they usually mean. If you can find all five, you know nearly everything you need to use it with confidence. If you can’t (say, a tutorial shows one way to call a function and never lists the parameters), you’re reading the wrong genre, and it’s time to switch to the reference page.

Find the assumptions nobody wrote down

Every set of docs is written for an imagined reader, and that reader is usually in a slightly different situation from you. Docs routinely assume you’re in the right folder, that your environment is already activated, that you’re online, that you’re allowed to write wherever the example writes, and that your data looks like data the docs never quite describe. As you read, keep asking: what has to be true for these steps to work? Then check each answer against your real situation. The first place they differ is usually where the docs and your computer will collide.

Treat examples as promises

The examples in official docs aren’t decoration. They’re the closest thing documentation has to a promise: the maintainers are saying that this code, in this version of the library, gives this result. So take them at their word. Copy the example into a scratch file, run it exactly as written on your installed version, and confirm it works before you change anything. Then change one thing at a time toward what you really need. If the example fails on your machine, that’s useful evidence, not a dead end: either the docs are out of date, you’re on a different version than they assume, or one of their unwritten assumptions doesn’t hold for you. Any of those beats staring at the page.

Look past the happy path

Tutorials almost always show the happy path: clean data, the right permissions, a fresh environment, a steady network, no edge cases. Real work brings the messy versions of all of those: missing files, corrupted data, half-installed packages, two tools that used to get along and now don’t, and network hiccups that look like bugs but aren’t. So go looking for the pages about the unhappy path, usually titled “Troubleshooting,” “Common issues,” “FAQ,” “Compatibility,” or “Upgrading.” Most projects have them, and they’re often the most useful pages on the whole site, because they describe exactly the situations that send people searching for help.

3.5 Reconciling documentation with versions and environments

Here’s something that saves a lot of grief: a big share of “the docs are wrong” moments are really version mismatches. The docs are correct, just for a different version of the tool than the one you have. That DataFrame.append error from the start of the chapter is a perfect case. append worked for years, then pandas 2.0 removed it in favor of pd.concat, and every older tutorial that uses it now fails on a current install.

Write down your versions first

When something the docs say should work doesn’t, the cheapest first step is to record exactly what you’re running. Five facts solve most version mysteries: the version of the tool or library, your operating system, your Python version (if Python is involved), the environment manager you’re using (conda, pip, or venv), and the path of the program that’s actually running. That last one catches more problems than you’d guess: if you have two Pythons installed, sys.executable tells you which one you’re really using. Collect all five before you change anything, because every fix you try will be judged against them.

python --version
python -c "import sys; print(sys.executable)"
python -c "import pandas; print(pandas.__version__)"
uname -a   # macOS / Linux  (Windows: 'systeminfo' or 'ver')

Recording your versions is one of the habits the reproducibility literature recommends (Wilson et al. 2017; The Turing Way Community 2025), and it has a practical bonus: it makes your question answerable when you need to ask someone for help (see Chapter 2). If you need the whole list of installed packages, pip freeze prints them with their exact versions.

Pick the docs for your version

Many documentation sites have a version switcher, usually a dropdown near the top of the page, and the version you land on from a search engine is often not the one you have installed. If your library has had major releases, pick the right version explicitly. When behavior has changed, check the release notes or changelog. You don’t need to read every entry; you only need to find out whether your symptom matches a known change.

Pin or upgrade

When the docs and the behavior disagree, you usually face a choice. You can pin: keep your environment as it is and read the docs for your version. Or you can upgrade: move to the current version and follow the latest docs. For a class project, upgrading is often fine if it doesn’t break the assignment. For a team project, pinning is usually safer, because everyone stays on the same version. Either way, write the choice down in the project’s documentation, so nobody has to rediscover it.

3.6 Project-local documentation: READMEs and runbooks

Every project needs a front door, and that’s the README: the file GitHub shows on your repository’s front page, and the first thing anyone opens. At minimum it has to answer two questions: “what is this?” and “how do I run it?”

A minimum viable README

A good README for a class project has six parts, roughly in this order. It opens with a one-paragraph purpose: what the project does and why anyone should care. Then comes setup, which walks through creating the environment and installing the dependencies. Next is how to run it, with the exact commands to type. After that, expected outputs says which files or results appear when the run succeeds; this is the part that lets a reader confirm it actually worked, and it’s the part people most often leave out. A project structure section explains what each top-level folder holds. And data notes say where the data came from and any limits on how it can be used. Here’s all six in one short Markdown file:

# Q3 Sales Analysis

A short pipeline that loads quarterly sales data, cleans it, and
produces a report.

## Setup
    python -m venv .venv
    source .venv/bin/activate
    pip install -r requirements.txt

## Run
    python src/run_pipeline.py

## Expected outputs
- outputs/figures/quarterly.png
- outputs/tables/top_customers.csv

## Structure
- data/raw/        # original CSVs (read-only)
- src/             # pipeline code
- outputs/         # generated artifacts (gitignored)

## Data
Source: /shared/sales-data/2026-q3.csv (internal use only).

(The source .venv/bin/activate line is for macOS and Linux. On Windows it’s .venv\Scripts\activate; Chapter 15 has the details, and a README that expects Windows readers should say so.)

Writing this can feel like busywork, especially for a project only you have touched. It pays you back the first time you come back after a week away, and again the first time a teammate clones the repository and tries to run it on their own machine.

Runbooks for recurring jobs

Some projects have jobs you do over and over: refresh the data, rebuild the figures, rerun the pipeline. Each of those deserves a short runbook, a page that someone who has never done the job could follow. It says what the task is, when to run it, the exact command, where the logs and outputs go, and what to do when it fails. Runbooks matter even more once a task is automated (see Chapter 33), because the day the automation breaks is the day nobody remembers how to do the job by hand.

3.7 Writing documentation that people actually use

The hard part of writing documentation isn’t typing the words. It’s deciding what to put in and what to leave out, and most bad docs fail on that choice rather than on grammar.

Write for a specific reader

Before you write, decide who you’re writing for (a classmate, a TA, you in six months), what they already know, and what they need to get done. The reason this matters has a name, the curse of knowledge: once you understand something, it’s genuinely hard to remember what it was like not to, so you skip steps that feel obvious to you and aren’t to anyone else. Picturing one specific reader is the best cure. Documentation that tries to serve everyone usually serves no one; a clear page for a named reader beats a universal one.

Structure beats cleverness

Readers skim before they read, so give them predictable headings and a consistent pattern. For anything procedural, a reliable shape is:

  1. Purpose
  2. Prerequisites
  3. Steps (numbered)
  4. Verification (what success looks like)
  5. Troubleshooting (common failures)

If a reader can skim just the headings and understand the shape of the task, you’ve already saved them most of the friction. Template A at the end of the chapter follows this shape.

Make commands copyable

Nothing’s more frustrating than a README that says “run the script” and stops there. Which script? From which folder? In which environment, and with which arguments? When your docs include commands, put each one in a code block so it can be copied exactly, say which folder to run it from, explain any placeholder like <your-file> rather than leaving the reader to guess, and show the expected output when it helps a reader confirm the command worked.

Explain why, not just what

Good docs explain why the important steps exist: why you pin versions, why you never edit the raw data in place, why you run the formatting check before committing. A step without a reason looks like clutter, and a future collaborator trying to tidy things up will “optimize away” exactly the steps that were keeping the results correct. One sentence of rationale protects them.

Anchor ideas with an example

When you’re explaining a concept, include one fully worked example: a real command with its output, a real function call with realistic inputs, or a sample configuration file. An abstract description gives a reader nothing to check; a worked example is something they can run and compare against.

3.8 Diagrams: reading and drawing them

Some things are hard to hold in your head from prose: which script reads which file, how two tables connect, who sends what to whom and in what order. A diagram shows that structure at once, and a reader who knows the three common kinds below can read most of the diagrams in documentation, papers, and design discussions. If diagrams have always looked like a secret code, it’s because they use conventions nobody explains; once you know the conventions, they’re quick to read.

Flowcharts: what happens, in what order

A flowchart (or pipeline diagram) shows steps as boxes and the order between them as arrows. Figure 3.1 is the kind of project described in Chapter 30, drawn as one:

flowchart TD
    accTitle: A data pipeline
    accDescr: A flowchart. The file data/raw/sales.csv flows into the script clean.py, which the file data/dictionary.csv checks (a dotted arrow). clean.py writes data/processed/sales.parquet, which flows into analysis.ipynb, which writes reports/figures/.
    raw[(data/raw/sales.csv)] --> clean[clean.py]
    dict[(data/dictionary.csv)] -. checks .-> clean
    clean --> tidy[(data/processed/sales.parquet)]
    tidy --> nb[analysis.ipynb]
    nb --> figs[(reports/figures/)]
Figure 3.1: A data pipeline as a flowchart. Cylinders are files; rectangles are code; the dotted arrow is a check, not a flow of data.

Read it from the top, along the arrows. The shapes carry meaning (here, cylinders for stored data and rectangles for code), and a good diagram says what its shapes and line styles mean, in a caption or a legend.

Entity-relationship diagrams: how tables connect

An entity-relationship (ER) diagram shows the tables in a database, their columns, and how their rows relate. Figure 3.2 draws the two tables from Chapter 23:

erDiagram
    accTitle: Customers and orders
    accDescr: An ER diagram of two tables. CUSTOMERS has customer_id (primary key), name, and city. ORDERS has order_id (primary key), customer_id (foreign key), amount, and date. A line labeled places joins them, with two bars at the CUSTOMERS end and a circle and crow's foot at the ORDERS end.
    CUSTOMERS ||--o{ ORDERS : places
    CUSTOMERS {
        int customer_id PK
        string name
        string city
    }
    ORDERS {
        int order_id PK
        int customer_id FK
        float amount
        date date
    }
Figure 3.2: The customers and orders tables from the SQL chapter as an ER diagram. The line between them reads: one customer places zero or more orders.

The marks at each end of the line, called crow’s-foot notation, carry the relationship. Two bars (||) mean “exactly one,” a circle means “zero,” and the three-pronged crow’s foot means “many.” So the line reads, from each side: every order belongs to exactly one customer, and a customer places zero or more orders. PK marks a table’s primary key and FK a foreign key, the column that points to another table. When you join two tables, the ER diagram tells you which column to join on and whether a join can duplicate rows (the “many” end is where they multiply).

Sequence diagrams: who talks to whom

A sequence diagram shows messages between participants over time: each participant gets a vertical line, time runs downward, and each arrow is one message. Figure 3.3 is a script calling a web API, as in Chapter 24:

sequenceDiagram
    accTitle: Fetching from a web API
    accDescr: A sequence diagram with two participants, fetch.py and GitHub API. fetch.py sends GET /repos/pandas-dev/pandas with User-Agent and token; the API answers 200 OK with a JSON body. fetch.py parses the JSON and saves it to data/raw/. It then requests the next page, and the API answers 429 Too Many Requests; a note says wait, then retry.
    participant S as fetch.py
    participant A as GitHub API
    S->>A: GET /repos/pandas-dev/pandas (with User-Agent and token)
    A-->>S: 200 OK, JSON body
    S->>S: parse JSON, save to data/raw/
    S->>A: GET the next page
    A-->>S: 429 Too Many Requests
    Note over S: wait, then retry
Figure 3.3: A script fetching data from a web API as a sequence diagram. Solid arrows are requests; dashed arrows are responses.

Sequence diagrams are how API documentation explains authentication and how people debug a conversation between programs: when something fails, you can point at the arrow where it went wrong.

Drawing your own

Most diagrams in a student project are informal architecture sketches: boxes for the pieces (your laptop, a server, a database, cloud storage), arrows for what moves between them, and labels on everything. A photo of a whiteboard is a fine start. For a diagram that will live in documentation, write it as text, the way the three above are written, in Mermaid. A Mermaid diagram is a few lines in a fenced code block that GitHub, Quarto, and many other tools draw for you; because it’s text, it lives in version control, shows up in diffs, and is easy to change when the project does. The Mermaid Live Editor shows the drawing as you type. For diagrams Mermaid can’t lay out well, draw.io (also known as diagrams.net) is a free drawing tool whose files can also be kept in a repository.

Whichever tool you use, a few habits make a diagram readable:

  • One idea per diagram. A diagram that shows the data flow, the database schema, and the deployment at once shows none of them clearly. Draw three.
  • Label the arrows with verbs (“reads,” “writes,” “checks”), and say what the shapes and line styles mean.
  • Keep one direction of flow, left to right or top to bottom, so the reader never has to hunt for where it starts.
  • Describe it in words too. A caption and the text around a diagram should say what it shows, for readers who can’t see it and for the moment the diagram falls out of date. In Mermaid, an accTitle: line and an accDescr: line inside the diagram give screen readers a title and a description (Mermaid’s accessibility guide explains both); the three diagrams above have them.

3.9 Documentation maintenance as a habit

Here’s the uncomfortable truth about docs: the day you write them is the day they start going out of date. Every change to the code is a chance for the README to drift a little further from reality, until one day a newcomer follows it exactly and nothing works. Documentation isn’t a one-time deliverable. It has to move with the code.

Docs are part of “done”

The simplest policy that works: if a change affects how someone uses the project, update the docs in the same pull request. Not “later,” because later rarely comes. When the docs travel with the change, reviewers see both at once, and nobody has to remember what needs updating.

When docs drift, treat it as a bug

When instructions stop working, file an issue, just as you would for broken code. Drift isn’t shameful; it happens to every project. What causes the harm is drift nobody has written down, because then the next person hits the same wall and assumes the problem is them.

Keep a decision log

Some choices can’t be read from the code: why you picked this dataset, why you used this model or that parameter, why you excluded certain records, why you chose one workflow over another. Three months later, someone (often you) will look at one of those choices and wonder whether it was a mistake. A short decision log, a running list of what you decided, when, and why, answers that question before it’s asked and stops the team from having the same debate twice. Software teams call the bigger version an architectural decision record; for a class project, a few lines per decision in a Markdown file is plenty (Template C below is one). Keeping one is standard advice for reproducible projects (The Turing Way Community 2025; Wilson et al. 2017).

The Turing Way Community. 2025. The Turing Way: A Handbook for Reproducible, Ethical and Collaborative Research. Zenodo. https://zenodo.org/records/15213042.
Wilson, Greg, Jennifer Bryan, Karen Cranston, Justin Kitzes, Lex Nederbragt, and Tracy K. Teal. 2017. “Good Enough Practices in Scientific Computing.” PLOS Computational Biology 13 (6): e1005510. https://doi.org/10.1371/journal.pcbi.1005510.

3.10 AI tools in documentation workflows

AI tools are genuinely good at some documentation chores, and genuinely risky at others. They can make things up with total confidence, including command-line flags and function arguments that don’t exist; they can describe a different version of a tool from the one you have; and anything you paste into them may leave your computer. The way through is to use AI as an assistant for drafting and structure, never as the authority on facts.

Where it helps most is turning material you already have into a better shape. Ask it to draft a README skeleton from your project’s folder structure, to rewrite rough notes into a how-to guide, to produce consistent templates such as issue forms or a pull request checklist, to suggest likely troubleshooting steps (which you then check), or to make a draft clearer for a beginner. In each case you’re supplying the facts and the AI is supplying the arrangement.

A few guardrails keep that arrangement safe:

  1. Never paste secrets: tokens, keys, passwords, or private data (see Chapter 34).
  2. Treat every output as a draft: check it against the official docs and by running the commands yourself.
  3. Prefer local evidence: your logs, your versions, and your environment beat the AI’s guess about them.
  4. Cite primary sources: don’t let an AI’s answer stand in for the reference it should have pointed you to.

And one rule to follow every time: if an AI suggests a command that could destroy something (rm, sudo, a change to file permissions), stop and check it in the official documentation or with an instructor before you run it.

3.11 Stakes and politics

Think about the first line of a typical setup guide: “Just run pip install.” Now picture the student reading it on a library computer where installing software is blocked, on a university-managed laptop without administrator rights, on a campus network that blocks the package index, or through a screen reader on a docs site whose examples are images. For each of them that line is a wall, and the tool is effectively unavailable: not because the technology couldn’t serve them, but because the documentation drew a boundary they couldn’t cross. The old shorthand RTFM treats that as the reader’s problem. The habit worth building treats it as the writer’s.

Whose problems the docs anticipate is a design decision: the operating system in the screenshots, the network connection assumed, the language of the prose. Those choices encode an imagined user, usually English-speaking, on a recent laptop, with fast internet, and anyone who doesn’t match pays the difference in time and frustration. GitHub’s 2017 Open Source Survey found that nearly a quarter of the open-source community reads and writes English less than “very well.” Who maintains the docs is a decision too. The same survey found that 93% of respondents had run into incomplete or outdated documentation, yet 60% of contributors rarely or never contribute to it. Docs are only as inclusive as their writers have the time and motivation to make them.

See Chapter 8 for the broader framework. The concrete prompt to carry forward: when you write a README or a how-to, name your imagined reader explicitly, then add one sentence for the reader you assumed away.

3.12 Worked examples

These three examples show the habits above in the kind of project you’re likely working on right now.

Turning a notebook into a runnable project

Most student projects start life as a single Jupyter notebook, and that’s fine. The trouble starts when someone else needs to reproduce your results: they don’t know which environment you used, which data file to put where, or what “working” is supposed to look like. A few additions fix that, in this order:

  1. Create a README with a one-sentence purpose and a “How to run” section.
  2. Add an environment section explaining how to set up with conda or pip.
  3. Add an “Outputs” section describing what files appear after a successful run.
  4. Add a “Project structure” section describing the folders.

That’s maybe half an hour of work, and it turns a private notebook into something a teammate can pick up and run.

Documenting a data intake pipeline

Even a pipeline as simple as “load a CSV and clean it” rests on assumptions nobody wrote down: the file’s encoding, its delimiter, the codes it uses for missing values, the columns it’s expected to have, and what makes a row valid. Each one is a way the pipeline can break silently when a new version of the file arrives.

A good intake note makes those assumptions visible. It records where the dataset came from and when you got it, what each column means (its schema), the quirks you’ve found (numbers stored as text, say), and the checks you run before any analysis. That note is the bridge between a raw file and the claims you make from it. Chapter 21 shows how to turn those checks into code, and Chapter 30 shows how to write the column meanings up as a data dictionary.

Writing troubleshooting notes

Troubleshooting notes are only useful when they’re specific. Compare these two notes about the same problem.

Weak.

“If conda doesn’t work, reinstall it.”

Better.

“If conda activate fails with ‘Your shell has not been properly configured to use conda activate’, your shell hasn’t been set up for conda yet. Run conda init zsh (or conda init bash, to match your shell), then close the terminal and open a new one. If instead the terminal says conda: command not found, conda isn’t on your PATH: if you installed Anaconda or Miniconda after opening this terminal, a new terminal window usually fixes it.”

The better version names the exact symptom, gives a likely cause, and says what to try next and how to tell whether it worked. Someone searching for that error message will find it, and someone following it won’t need to message you.

3.13 Templates

Template A: How-to guide

# Title: How to <do a specific task>

## Purpose
One sentence describing what this accomplishes.

## Prerequisites
- Required software / permissions
- Required files or environment

## Steps
1.
2.
3.

## Verify
What success looks like (expected output / files / screenshots).

## Troubleshooting
- Symptom -> likely cause -> fix

Template B: README skeleton (student project)

# Project Title

## Purpose
What this project does (2-4 sentences).

## Quickstart
Commands to set up and run.

## Environment
- How to create/activate env
- Key dependencies

## Data
- Source
- How to obtain
- Restrictions (if any)

## How to run
Exact command(s) and expected outputs.

## Project structure
- data/
- src/
- notebooks/
- outputs/

## Reproducibility notes
Versions, seeds, and known pitfalls.

Template C: Decision log entry

Decision:
Date:
Context:
Options considered:
Chosen option:
Rationale:
Consequences / follow-ups:

3.14 Exercises

  1. Pick a tool you used this week (Git, conda, pandas, Jupyter). Find one example each of reference docs, a tutorial, and a how-to guide for it, and write one sentence on the question each one answers best.
  2. Write a one-page README for a class assignment repository. Swap with a classmate and try to run each other’s project using only the README. Write down what was missing, then revise.
  3. Take a confusing error you’ve run into and write a troubleshooting note for it that includes the symptom, the likely cause, a way to check, and the fix.
  4. Rewrite a set of rough notes into a how-to guide using Template A. Include a verification step a reader can actually perform.
  5. Find one assumption in your current project that isn’t written down (a path, a version, a parameter, a quirk of the data). Add it to the README with a sentence explaining why.
  6. Draw your project’s pipeline as a Mermaid flowchart in its README, with an accDescr: line, and check that GitHub draws it. If the project uses more than one table, add an ER diagram of how they join.
  7. Run python -c "import pandas; print(pandas.__version__)", then find the documentation for exactly that version of pandas using the site’s version switcher. What’s the first thing on the release notes page that could have affected you?

3.15 One-page checklist

  • I can tell whether I need reference docs, a tutorial, a how-to guide, or an explanation.
  • I use secondary sources to find primary sources, not to replace them.
  • When reading docs, I pull out the inputs, outputs, defaults, constraints, and failure modes.
  • I check local help (--help, man, help(), ?) before searching the web.
  • I record version and environment details when behavior surprises me.
  • I read the docs for the version I actually have installed.
  • My README covers purpose, setup, how to run, expected outputs, structure, and data.
  • Commands in my docs are copyable and say where to run them.
  • I update the docs in the same pull request as the change that affects them.
  • I keep a short decision log for consequential choices.
  • Where structure is hard to follow in prose, I add a diagram as text (Mermaid) with one idea, labeled arrows, and a description.
  • If I use AI tools, I check their output against official docs and my own experiments, and I never paste secrets.
Note📚 Further reading