5  Reading Official Documentation

Prerequisites (read first if unfamiliar): Chapter 2.

See also: Chapter 3, Chapter 6, Chapter 35.

Purpose

Success Kid Meme: Need To Learn About A New Function, The Docs Have A Perfect Example.

You know this evening. Your code breaks, you paste the error into a search engine, and you land on a blog post that almost matches your problem. You copy a line from it, and now something else breaks. You try a second post, then a forum answer from 2017, then a third post that contradicts the first. Two hours later you’re tired, nothing works, and you’ve started to suspect that programming just isn’t for you.

It is for you. What went wrong is where you looked. The answer was very likely a few clicks away in the library’s official documentation, but maybe you didn’t know it existed, or you opened it once, met a wall of fifty parameters, and closed the tab. That’s a normal reaction, and it comes from reading docs the wrong way: top to bottom, like a textbook. Documentation is a reference work, like a dictionary. You open it with one question, find the answer, and leave. Once you read it that way, the official docs become the fastest and most reliable help you have.

This chapter teaches that skill: finding the right kind of page, reading a function’s signature and docstring, pulling a working example out of the docs, and using release notes to explain code that “worked last year.” The examples come from pandas and Python itself, but the habits carry over to any library. It isn’t about writing documentation; that’s Chapter 3.

Why read this chapter

  • You searched an error, copied a fix from a blog post, and two cells later something else broke.
  • You opened the reference page for pd.read_csv, saw dozens of parameters, and closed the tab without finding the one you needed.
  • Code from a tutorial gives you AttributeError: 'DataFrame' object has no attribute 'applymap', even though it worked for a classmate last year.
  • You sorted a DataFrame, and the next line acted as if you never had.
  • You’ve seen *, /, **kwargs, and <no_default> in a function’s signature and had no idea what they were telling you.
  • You’d like to answer your own question in thirty seconds from inside Jupyter, without opening a browser.
  • You want to be able to say “I checked the docs, and here’s what they say” before you ask anyone for help.

Running theme: start from the docs, not from the search results

For any question about how a library function behaves, make the official docs your first stop, not your last: search results are great for discovering what’s possible, and the docs are where you find out what’s true for the version you have.

5.1 The four genres of documentation

One reason docs feel hostile is that people land on the wrong kind of page for their question. You want to know what one argument does, and you’re reading a friendly tutorial that never mentions it. Or you’re brand new, and you’re staring at a dense reference page that assumes you already know everything. Neither page is bad; each is built for a different moment.

A useful way to see this is the Diátaxis framework, written by Daniele Procida. It sorts documentation into four genres by the kind of need each one serves:

Genre What it’s for When you want it
Tutorials learning “I’m new; walk me through the basics hands-on.”
How-to guides getting a task done “I know roughly what I want; tell me the steps.”
Reference looking things up “I know which function I want; tell me the exact arguments.”
Explanation understanding “Why does this work the way it does?”

Most newcomers find a tutorial, get through it, and then keep reaching for tutorials for every question afterwards. The real skill is noticing which of the four questions you’re asking and switching. Diátaxis boils the choice down to two questions in its compass: do you need to do something or to know something, and are you learning a skill or applying one? Suppose you’ve merged two tables and want to see which rows didn’t find a match. A tutorial on merging will take ten minutes to maybe mention it. The reference page for DataFrame.merge lists every parameter, and a quick scan turns up indicator=True, which adds a column saying where each row came from. On the other hand, if you want to know why .apply is slow, no reference page will tell you; that’s a job for an explanation.

The big libraries usually have all four genres, though they don’t always use these names or keep them neatly apart. pandas’ docs have Getting started (installation, the short “10 minutes to pandas” tour, and tutorials), a User Guide that explains each topic with lots of examples, a Cookbook of short recipes inside the User Guide, and the API reference. Python’s own docs have a tutorial, a library reference for the standard library, a language reference for the precise rules of the language, and a set of HOWTOs. For the two or three libraries you use most, spend five minutes learning where each genre lives. It pays for itself the first week.

5.2 Reading a function signature

Every reference page opens with the function’s signature: its name, followed by every parameter it accepts and each parameter’s default. It’s the densest line on the page, which is exactly why it’s worth reading slowly. Here’s the start of the one for pandas.read_csv, as the pandas 3.0 docs show it (the full list goes on for more than forty parameters):

pandas.read_csv(
    filepath_or_buffer,
    *,
    sep=<no_default>,
    delimiter=None,
    header='infer',
    names=<no_default>,
    index_col=None,
    usecols=None,
    dtype=None,
    ...
    na_values=None,
    ...
    encoding=None,
    ...
)

Fifty parameters looks scary until you notice that you’ll pass two or three of them. You don’t read a signature like a paragraph. You scan it for the parameter you care about, and a few bits of punctuation tell you most of what you need.

The lone * means “from here on, use names.” Everything after it is keyword-only: you have to write the parameter’s name, as in sep=";", rather than just putting the value in position. So filepath_or_buffer can be passed by position, but nothing else can. If you guess that the separator is the second argument and write pd.read_csv("data.csv", ";"), Python stops you:

TypeError: read_csv() takes 1 positional argument but 2 were given

That message confuses almost everyone the first time, because you did pass two arguments and you’d like both to count. The fix is to name the second one:

df = pd.read_csv("data.csv", sep=";")

Library authors do this on purpose, so that a line like read_csv("data.csv", ";", 0, None) can’t exist and nobody has to remember which position means what. Python’s tutorial covers the syntax under special parameters, which also explains the *’s sibling, /. You’ll meet it in built-ins like sorted(iterable, /, *, key=None, reverse=False): parameters before a / can only be passed by position, and parameters after the * only by name.

Each = gives a default. header='infer' means that if you don’t pass header, pandas uses 'infer'. That sounds like a detail, but every parameter you skip is a decision you’ve accepted without reading, so defaults are where surprises come from. They’re also where a newcomer’s intuition fails most often, because some defaults aren’t plain values: they’re sentinel values with special meanings you can only learn from the description.

header is a good example. There are three settings that look nearly the same, and the parameter’s description spells them out. header=0 says “the first row holds the column names.” header=None says “this file has no header row, so number the columns 0, 1, 2.” And the default, 'infer', behaves like header=0 unless you’ve also passed your own column names with names=, in which case it behaves like header=None. You’d never guess that from the word “infer,” which is the point: read the description of any parameter you depend on.

The same goes for None and for the odd-looking <no_default>. None in a signature almost never means “nothing.” For usecols=None it means “load every column”; for dtype=None it means “work out each column’s type from the data”; for encoding=None it means “use UTF-8,” as the description tells you. <no_default> is a marker pandas uses internally so it can tell “you didn’t pass this” apart from “you passed None.” Scroll down to sep and the description says what actually happens: sep : str, default ','. When the signature and the description seem to disagree, the description is the one written for you.

Some signatures end in **kwargs. That means “any other keyword arguments,” which the function collects and usually passes along to some other function. It’s the one place a signature can’t tell you what’s allowed, so the docstring normally names the function that receives them. The worked example on requests.get below shows how to follow that trail.

When you read a signature inside Python rather than on the web, you’ll also see type hints, the text after each colon. The web page leaves them out, but help(pd.read_csv) shows them:

header: "int | Sequence[int] | None | Literal['infer']" = 'infer',

Read the | as “or”: header can be a whole number, a list of them, None, or the exact string 'infer'. You don’t need to understand every type name to get value from that line. It has already told you which kinds of value are allowed.

5.3 Reading a return value

Below the parameters, every reference page says what the function gives back. In pandas, that’s usually another DataFrame, a Series, or a single number. The question that trips people up most is simpler, though: does this change my data, or hand me a changed copy?

Here’s the snag in action. You sort a DataFrame, print it on the next line, and it’s still in the original order. Nothing is broken. sort_values returns a new, sorted DataFrame and leaves the original alone, and your code threw the new one away:

df.sort_values("date")                 # returns a sorted copy; df is unchanged
df = df.sort_values("date")            # keep the copy by assigning it back
df.sort_values("date", inplace=True)   # changes df itself and returns None

The signature tells you this too, if you look at the end. inspect.signature(pd.DataFrame.sort_values) ends in -> 'DataFrame | None': you get a DataFrame back normally, and None if you pass inplace=True. That None is behind another classic bug, df = df.sort_values("date", inplace=True), which sorts the data and then replaces df with None. Many pandas methods have an inplace option for historical reasons; the usual advice now is to skip it and assign the result, which is also what the pandas docs’ own examples do.

5.4 Reading docstrings without leaving Python

You don’t always need a browser. Nearly every Python function, method, and class carries a docstring, a block of documentation written into the code itself, and you can read it right where you’re working. It’s also guaranteed to match the version you have installed, which a web page isn’t.

In the plain Python REPL or a script, use the built-in help function:

help(pd.read_csv)

In Jupyter or IPython, put a ? after the name and run the cell:

pd.read_csv?

That shows the signature and the full docstring. A double question mark, pd.read_csv??, goes one step further and shows the function’s source code as well, when the source is written in Python (IPython’s docs describe both under accessing help). In JupyterLab, pressing Shift+Tab with your cursor inside a function’s parentheses pops up the same information in a small tooltip.

Two more tools round this out. inspect.signature gives you just the signature, without the wall of text:

import inspect
inspect.signature(sorted)
# <Signature (iterable, /, *, key=None, reverse=False)>

And dir lists everything an object has, which is handy when you half-remember a method’s name. Filtering out the names that start with an underscore (Python’s internal machinery) leaves the ones meant for you:

import pandas as pd
df = pd.DataFrame({"a": [1, 2, 3]})
[m for m in dir(df) if not m.startswith("_")][:20]
# ['T', 'a', 'abs', 'add', 'add_prefix', 'add_suffix', 'agg', 'aggregate', 'align',
#  'all', 'any', 'apply', 'asfreq', 'asof', 'assign', 'astype', 'at', 'at_time',
#  'attrs', 'axes']

Notice 'a' in the list: the column itself shows up, because pandas lets you reach a column as df.a. All of these work offline, on a plane, and during an exam with the Wi-Fi down.

5.5 Reading a docstring: the usual shape

Docstrings in the scientific Python world (pandas, NumPy, SciPy, scikit-learn) follow a shared layout with the same named sections every time, so once you’ve read one, you can find your way around all of them. (Python’s general conventions for docstrings are in PEP 257.) Here’s the docstring for DataFrame.resample from pandas 3.0, shortened where you see ...:

Resample time-series data.

Convenience method for frequency conversion and resampling of time series.
The object must have a datetime-like index (`DatetimeIndex`, `PeriodIndex`,
or `TimedeltaIndex`), or the caller must pass the label of a datetime-like
series/index to the ``on``/``level`` keyword parameter.

Parameters
----------
rule : DateOffset, Timedelta or str
    The offset string or object representing target conversion.
...

Returns
-------
pandas.api.typing.Resampler
    :class:`~pandas.core.Resampler` object.

See Also
--------
Series.resample : Resample a Series.
DataFrame.resample : Resample a DataFrame.
groupby : Group Series/DataFrame by mapping, function, label, or list of labels.
asfreq : Reindex a Series/DataFrame with the given frequency without grouping.
...

Examples
--------
Start by creating a series with 9 one minute timestamps.

>>> index = pd.date_range("1/1/2000", periods=9, freq="min")
>>> series = pd.Series(range(9), index=index)
...

Downsample the series into 3 minute bins and sum the values
of the timestamps falling into a bin.

>>> series.resample("3min").sum()
2000-01-01 00:00:00     3
2000-01-01 00:03:00    12
2000-01-01 00:06:00    21
Freq: 3min, dtype: int64

That’s a lot of text, and you don’t have to read it in order. Start with the summary line at the very top. It’s one sentence and costs nothing to read, and surprisingly often it answers your question (“oh, this is for time series; I wanted groupby”). If it doesn’t, jump to Parameters and find the one argument you care about. Each entry gives the parameter’s name, its type, and a description, so your editor’s find tool (or your browser’s Ctrl+F on the web page) can take you straight there. Returns tells you what comes back, which is what you need to know to use the result on the next line. See Also is underrated: when a function is almost but not quite what you want, the right one is often listed there. And Examples is usually the fastest way to understand the whole thing, because it shows real calls on small data with the real output underneath. Some docstrings also have a Notes section, with background and links into the User Guide.

If you only have time for two sections, read the summary line and the examples.

5.6 Pulling a working example out of the docs

The examples on a reference page are written to stand on their own: they create their own little dataset, import what they need, and run top to bottom. That makes them the best possible starting point for using a function you’ve never used before, much better than writing a call from scratch and guessing at the arguments.

The workflow goes like this. Copy the example closest to what you want, paste it into a notebook cell or a scratch script, and run it unchanged. If it works, you know the function behaves as described on your machine, with your version. Then change one thing at a time toward your real problem, and run it again after each change. When something breaks, you know exactly which change broke it. Here it is with the first example from DataFrame.melt, heading toward a survey file with columns respondent_id, q1, q2, and q3:

import pandas as pd

# Step 1: paste the official example unchanged, and run it
df = pd.DataFrame(
    {
        "A": {0: "a", 1: "b", 2: "c"},
        "B": {0: 1, 1: 3, 2: 5},
        "C": {0: 2, 1: 4, 2: 6},
    }
)
df.melt(id_vars=["A"], value_vars=["B"])

# Step 2: swap in your real data, and change nothing else
df = pd.read_csv("survey.csv")
df.melt(id_vars=["A"], value_vars=["B"])
# KeyError: "The following id_vars or value_vars are not present in the DataFrame: ['A', 'B']"

# Step 3: change the column names to your own
df.melt(id_vars=["respondent_id"], value_vars=["q1", "q2", "q3"])

The error in step 2 is the useful kind: it names exactly what’s wrong, and you already know it’s the column names because that’s the only thing that changed. This is the same move as building a minimal reproducible example when you ask for help (see Chapter 2), just run in the other direction. It’s slower than pasting a guess, for about five minutes. After that, you also understand why your call works, which is the difference between knowing a function and cargo-culting it.

5.7 Release notes and changelogs

Here’s a situation that makes people feel like they’re losing their minds. A tutorial, or last year’s homework solution, uses df.applymap(...). You run the same line and get:

AttributeError: 'DataFrame' object has no attribute 'applymap'

Nothing is wrong with your installation. The library changed. pandas 2.1 renamed applymap to DataFrame.map and started warning that the old name was deprecated, and pandas 3.0 removed the old name entirely. Code that ran fine in 2023 fails now, and the error message doesn’t tell you any of that history.

This is what release notes are for. Most libraries keep a page, called “What’s new,” “Release notes,” or a changelog, that lists what changed in each version: new features, changed defaults, and things that were deprecated (still working, but with a warning that they’ll go away) or removed. Three symptoms almost always send you there: code that used to work and doesn’t; a tutorial using an argument or method your version doesn’t seem to have; and a FutureWarning or DeprecationWarning you don’t understand. All three are version mismatches, and the release notes are the quickest way to confirm one. For pandas, the notes are at pandas.pydata.org/docs/whatsnew; for Python itself, at docs.python.org/3/whatsnew.

Before you go looking, find out which version you actually have:

import pandas as pd
print(pd.__version__)

(From a terminal, pip show pandas tells you the same thing.) Then open the “What’s new” page for that version and search it for the function you’re using. The same mismatch works in reverse, too: the pandas website shows the newest release by default, so if your environment has an older one, the docs may describe arguments you don’t have yet. The version menu at the top of the pandas docs lets you switch to the docs for the version you’re running.

5.8 When a blog post is wrong

Blog posts, video tutorials, and forum answers are wonderful for discovery, the “I didn’t know you could do that” moments. They’re risky as a reference, and not because their authors are careless. They were right for the version they were written against, and nobody goes back to update them.

Two warning signs should make you slow down before copying. The first is no date, or a date more than a couple of years old. Libraries change; an answer that was right in 2018 can be quietly wrong now. The second, and the more telling, is code that doesn’t match the signature in the current docs: an argument the reference page doesn’t list, or a method that isn’t there. That almost always means the post was written for a different version, and following it further will only take you somewhere more confusing.

So use the two for what each does well. Discover an approach from a blog post or a forum answer, then check every function in the snippet against your installed version before you trust it:

# saw on a blog: pd.read_csv("file.csv", skipfooter=1, engine="python")
# check it locally before trusting it:
import pandas as pd
help(pd.read_csv)             # search the output for "skipfooter"
# or, in Jupyter:
pd.read_csv?

If skipfooter is there and means what the post says (it does: “Number of lines at bottom of file to skip,” and it needs the Python engine, which is why the post passes engine="python"), you’re fine. If your version’s docstring doesn’t mention it, or describes something different, believe the docstring. The same check applies to code an AI assistant writes for you, which often mixes versions (see Chapter 35).

5.9 Stakes and politics

Open the pandas docs and you’ll find tutorials, a long User Guide, and a reference page with worked examples for nearly every method. Then open the README of the small package that came with a research paper you’re trying to reproduce, and you may find an install command, one example, and a note that says “docs coming soon.” The advice “read the official documentation” quietly assumes the first situation.

Good docs cost money and time, and who pays shows up on the page. pandas is a sponsored project of NumFOCUS, a nonprofit that handles its donations, and its “institutional partners” are companies and universities that employ people to work on it. Many smaller libraries have one maintainer writing docs in their spare time. Nadia Eghbal’s 2016 report for the Ford Foundation, Roads and Bridges, described how much of the software everyone depends on rests on this kind of unpaid work. When the docs are thin, the cost of figuring things out anyway lands on readers, and it lands hardest on people without a mentor to ask, or who read English as a second language, since much library documentation exists only in English.

Conventions carry politics too. The Diátaxis genres, the numpydoc section layout, and the Sphinx tooling that builds most Python docs came from particular communities with particular ideas of what a good explanation looks like. Reading docs fluently partly means learning those communities’ habits, and nobody is born knowing them.

See Chapter 8 for the broader framework. The concrete prompt to carry forward: when the docs you need are bad, ask why they’re bad before you blame yourself for not understanding them.

5.10 Worked examples

Looking up a parameter: what does how="outer" do in a merge?

The slow way is to search “pandas merge outer example,” read four blog posts with four different example tables, and come away less sure than when you started.

The fast way is help(pd.merge), or the reference page for DataFrame.merge, and a search for how :. In pandas 3.0, help(pd.merge) shows:

how : {'left', 'right', 'outer', 'inner', 'cross', 'left_anti', 'right_anti},
    default 'inner'
    Type of merge to be performed.

    * left: use only keys from left frame, similar to a SQL left outer join;
      preserve key order.
    * right: use only keys from right frame, similar to a SQL right outer join;
      preserve key order.
    * outer: use union of keys from both frames, similar to a SQL full outer
      join; sort keys lexicographically.
    * inner: use intersection of keys from both frames, similar to a SQL inner
      join; preserve the order of the left keys.
    ...

Thirty seconds, and a definitive answer: "outer" keeps every key from both tables, like a full outer join in SQL (see Chapter 23), and sorts the keys. You also learned that the default is "inner", which is why rows without a match have been disappearing from your merges. (And if you spotted the missing quote after right_anti, yes, that’s a typo in pd.merge’s real docstring; the DataFrame.merge page has it right. Official docs are written by people too.)

Routing to the right genre: why is my .apply so slow?

The slow way is to search “speed up pandas” and try a pile of unrelated tricks, from changing data types to installing new libraries.

“Why is this slow?” is an explanation question, so look for an explanation page rather than a reference entry. In pandas 3.0, the User Guide’s page on user-defined functions answers it directly. Functions you write yourself and hand to .apply run once per row in plain Python, which pandas can’t speed up, while built-in vectorized operations work on whole columns at once in compiled code. The page times one example both ways: about 5.6 seconds with .apply, and 0.004 seconds for the vectorized version. Its advice is to rewrite the calculation with column arithmetic where you can, and it points to the guide on enhancing performance (Numba and friends) for the cases where you can’t.

Following **kwargs: what arguments does requests.get take?

You’re calling a web API with requests and want to know how to set a timeout. In a Jupyter cell:

import requests
requests.get?

With requests 2.34, the docstring is short, and it may not be what you hoped for:

Sends a GET request.

:param url: URL for the new :class:`Request` object.
:param params: (optional) Dictionary, list of tuples or bytes to send
    in the query string for the :class:`Request`.
:param \*\*kwargs: Optional arguments that ``request`` takes.
:return: :class:`Response <Response>` object
:rtype: requests.Response

No timeout anywhere. This is the **kwargs trail from earlier: the last parameter line tells you that everything else goes on to request. So follow it with requests.request?, and there’s the full list: params, data, json, headers, cookies, files, auth, timeout, allow_redirects, proxies, verify, stream, and cert, each with a description. (This docstring uses a different layout from pandas, with :param name: lines instead of a Parameters heading. It’s the older reStructuredText field style, and it holds the same information.) Chapter 24 puts those arguments to work.

Asking a precise question: does in work on a dict?

You want to check whether a value is in a dictionary, and you’re not sure whether x in my_dict looks at the keys or the values. Precise questions about how the language itself behaves belong in the Python Language Reference, which is a different document from the tutorial. The answer is under Expressions, in membership test operations: all the built-in sequences and sets support in, “as well as dictionary, for which in tests whether the dictionary has a given key.” So "a" in {"a": 1} is True, and 1 in {"a": 1} is False. To search the values, use 1 in my_dict.values().

The language reference is dense, because it’s written for people who want exact answers to exact questions. You won’t read it cover to cover. But you’ll find yourself reaching for it more as you get comfortable, because it’s the one place where “what does Python actually do here?” has a single, authoritative answer.

5.11 Templates

A “read the docs before asking for help” checklist:

  1. Can I find the function in the official reference?
  2. Did I read the summary line?
  3. Did I read the description of the parameter I’m using?
  4. Did I read the Returns section?
  5. Did I copy the official example and run it?
  6. Did I check which version I have installed against the version the docs describe?

If the answer to all six is yes and you’re still stuck, it’s time to ask a question (see Chapter 2), and you’ll be able to say exactly what you’ve already tried.

5.12 Exercises

  1. Using help() in a Python REPL, read the docstring for sorted. What does the key parameter do? What does reverse=True do? What does the / in its signature mean for how you can call it?
  2. In a Jupyter notebook, type pd.read_csv? and scroll through every parameter. Pick three you have never used and read their descriptions.
  3. Open the pandas API reference and navigate to pandas.DataFrame.merge. Find the indicator parameter. Write one sentence explaining what it does.
  4. Find a blog post more than three years old that uses a pandas function you know. Find one thing in its code that is now discouraged, deprecated, or removed, and confirm it against the current docs.
  5. Check your installed pandas version with pd.__version__. Open the “What’s new” page for that version, and list two behaviors that changed in that release.
  6. For a library you use often but have never read the docs of (for example, matplotlib or scikit-learn), find its tutorial, a how-to guide, the reference, and an explanation page. Bookmark each.
  7. The next time you hit a confusing error message, before searching online, open the docs for the function that raised it and read its Parameters and Examples sections. Time how long it takes to find the answer.

5.13 One-page checklist

  • Docs first, blog posts second. Official documentation is almost always faster and more reliable.
  • Know the four genres: tutorials (learn), how-tos (do), reference (look up), explanation (understand).
  • Use help(), ?, and ?? to read docstrings inline, and inspect.signature for just the signature.
  • Read signatures carefully: * means name the arguments after it, and defaults like 'infer', None, and <no_default> need their descriptions.
  • Check whether a method returns a new object or changes yours, and assign the result.
  • When a signature ends in **kwargs, follow the docstring to the function that receives them.
  • Copy the reference-page example first, run it unchanged, then change one thing at a time.
  • Always check your library version against the docs version: pd.__version__.
  • Check “What’s new” or the changelog when code that used to work stops working.
  • Docs are reference, not novels: dip in and out, don’t try to read top to bottom.
  • Bookmark the four main entry points of your favorite libraries.
Note📚 Further reading
  • Python documentation — the front door to the standard library and language docs; learn its tutorial, library reference, language reference, and HOWTOs as four different entry points.
  • pandas API reference — every DataFrame and Series method with signatures and examples; a good model of what excellent reference docs look like.
  • Daniele Procida, Diátaxis framework — the model that distinguishes tutorials, how-tos, reference, and explanation; the most useful single idea for finding your way around any documentation site.
  • numpydoc style guide — the docstring conventions used across the scientific Python stack; knowing them speeds up every ? lookup you’ll ever do.
  • Read the Docs — the hosting platform for much open-source Python documentation; knowing how it builds and versions docs explains a lot of the navigation patterns you’ll run into.
  • Stripe, API reference — a widely cited example of commercial reference documentation, with requests and responses side by side; a benchmark for what well-funded docs can look like.