17 Scripting
Prerequisites (read first if unfamiliar): Chapter 14, Chapter 16.
See also: Chapter 33, Chapter 31, Chapter 28.
Purpose

It’s week ten. The notebook that started as a quick look at the data is now 47 cells long, and your instructor asks you to rerun it on the updated file. You click Run All, and cell 31 fails: it uses a variable you defined in a cell you deleted weeks ago. You scroll up, find two nearly identical copies of your cleaning code, and can’t remember which one is the right one. The notebook worked yesterday. It worked because of the order you happened to run things in, and nobody, including you, can say what that order was.
If that sounds familiar, you’re not doing anything unusual. Notebooks are wonderful for thinking, and they make it very easy to build something that only runs once. The way out isn’t to give up notebooks. It’s to move the parts that do the work into plain Python files, where they can be run from the terminal, imported into any notebook, and fixed in one place.
This chapter shows you how: writing a script that also works as a module, organizing your functions into a small src/ package, importing that package into a notebook (and picking up your edits without restarting), passing options from the command line, converting notebooks to scripts and back, and deciding which tool fits which job. It assumes you’ve used Jupyter (Chapter 16) and can make a virtual environment (Chapter 15). Scheduling scripts to run on their own is Chapter 33, and keeping them tidy is Chapter 19.
Why read this chapter
- You ran
python scripts/run_cleaning.pyand gotModuleNotFoundError: No module named 'src', even though thesrcfolder is right there. - You fixed a function in a
.pyfile, re-ran the notebook cell that calls it, and the old version ran anyway. - The same cleaning code lives in four notebook cells, and you just fixed a bug in only three of them.
- Your analysis works for one month of data, and now you need it for twelve months, or every Monday, without clicking through the notebook each time.
- You converted a notebook with
jupyter nbconvertand the script died on its first line withNameError: name 'get_ipython' is not defined. - A teammate’s notebook reads
/Users/alex/Downloads/survey.csvand fails on every computer except Alex’s. - Your notebook’s Git diff is a wall of unreadable text, and nobody on your team can review it.
- You’d like a clear answer to “should this be a notebook or a script?” instead of a vague feeling that notebooks aren’t professional.
Running theme: separate logic from presentation
The code that does the work (loading, cleaning, analyzing) belongs in importable functions; the notebook or script around it just calls those functions in order and explains what they found.
17.1 Scripts, modules, and packages
Python has three words for “a file with Python code in it,” and people use them loosely enough that it’s easy to think they’re the same thing. They’re really three roles. A script is a file you run from the command line, like python analyze.py. A module is a file you import from other code, like import analysis (the Python tutorial’s chapter on modules is a friendly introduction). A package is a folder of modules imported as a unit, usually marked by an __init__.py file inside it.
The part that surprises people is that one .py file can play all three roles. A cleaning.py with a few functions is a module when you write from cleaning import clean_sales, a script when you type python cleaning.py, and part of a package when it sits in a folder with an __init__.py. You don’t have to choose. That’s what lets you build a small library of functions that also works as a command-line tool.
The trick that makes it work is the line you’ve probably seen at the bottom of other people’s files and copied without knowing why:
if __name__ == "__main__":
main()Every module has a variable called __name__. When you run a file, Python sets it to the string "__main__". When you import the file, Python sets it to the module’s name instead. You can watch this happen with a two-line file that prints its own name:
$ python clean.py
__name__ is '__main__'
running main()
$ python -c "from clean import clean"
__name__ is 'clean'
So the if block runs when you use the file as a script and stays quiet when a notebook imports a function from it. Without that guard, importing one function would also run the whole analysis, which is exactly as confusing as it sounds.
That leads to the most useful distinction in this chapter: library code versus entry-point code. Library code is the functions that do the work, like clean_sales(df) or plot_distribution(values). Each one takes inputs as arguments, returns outputs, and knows nothing about where its data came from. Entry-point code is the glue: it reads command-line arguments, opens files, calls library functions in the right order, and writes results somewhere. Programmers call keeping these apart separation of concerns, and it pays off quickly.
# Library code: reusable, testable, knows nothing about the command line
def clean_sales(df):
df = df.copy()
df.columns = df.columns.str.strip().str.lower()
return df.dropna(subset=["customer_id"])
# Entry point: arguments, file paths, and the order of operations
def main():
args = parse_args()
df = pd.read_csv(args.input)
cleaned = clean_sales(df) # library call
cleaned.to_csv(args.output, index=False)A quick test for which kind you’re looking at: if a function depends on sys.argv, on a hardcoded path, or on which folder you launched Python from, it’s entry-point code. If it takes its inputs as arguments and hands back its result, it’s library code. Keep them in separate functions, even in the same file, and you can later put a different front end (a notebook, a scheduled job, a web page) on the same library without rewriting it.
17.2 Writing your first script
A scripting language like Python lets you write a useful program in a dozen lines, and most good scripts share the same shape: imports at the top, then constants, then the functions that do the work, then a main() that calls them in order, and finally the __name__ guard.
"""Clean and summarize the Q3 sales CSV."""
from pathlib import Path
import pandas as pd
INPUT = Path("data/raw/sales.csv")
OUTPUT = Path("data/processed/sales_clean.csv")
def clean(df):
df = df.copy()
df.columns = df.columns.str.strip().str.lower()
df["amount"] = pd.to_numeric(df["amount"], errors="coerce")
return df.dropna(subset=["amount"])
def main():
df = pd.read_csv(INPUT)
cleaned = clean(df)
cleaned.to_csv(OUTPUT, index=False)
print(f"wrote {len(cleaned):,} rows to {OUTPUT}")
if __name__ == "__main__":
main()Run it from the project folder and it does its job:
$ python scripts/clean_q3.py
wrote 6 rows to data/processed/sales_clean.csv
Inside main(), the script’s whole job is to take some inputs and produce some outputs. Send those outputs to predictable places (data/processed/ for cleaned tables, figures/ for plots, reports/ for finished documents) so collaborators, and future you, always know where to look. The pathlib module’s Path objects, used above, make paths work the same way on macOS, Linux, and Windows.
For progress messages, print() is fine (“loaded 42,103 rows”). Once a script runs unattended, or other people depend on it, switch to the standard library’s logging module, which adds timestamps and levels (INFO, WARNING, ERROR) and can write to a file. The logging HOWTO walks through it; the smallest useful setup is three lines:
import logging
logging.basicConfig(level=logging.INFO,
format="%(asctime)s %(levelname)s %(message)s")
log = logging.getLogger(__name__)
log.info("loaded %d rows", len(df))which prints lines like 2026-09-25 18:02:48,068 INFO loaded 7 rows. For course assignments, print is plenty.
17.3 From one long notebook to functions in src/
From copy-paste to functions
Nearly every project starts as straight-line code in one notebook, and that’s the right way to start: the fastest way to understand a new dataset is to poke at it and watch what happens. The trouble begins when the same block shows up in a second cell, then a third. Now every bug has to be fixed in three places, and sooner or later you fix it in two. The copies drift apart, and one afternoon disappears into figuring out why two charts disagree.
The fix is to pull the repeated block out into a function, a small refactoring that programmers sum up as don’t repeat yourself. Six lines that appear in three cells become one function you call three times:
# Before: the same lines appear in three cells and one script
df.columns = df.columns.str.strip().str.lower()
df["date"] = pd.to_datetime(df["date"], errors="coerce")
df["amount"] = pd.to_numeric(df["amount"], errors="coerce")
df = df.dropna(subset=["customer_id", "date", "amount"])
df = df[df["amount"] > 0]
df = df.reset_index(drop=True)
# After: one function, called from anywhere that needs it
def clean_sales(df: pd.DataFrame) -> pd.DataFrame:
df = df.copy() # leave the caller's DataFrame alone
df.columns = df.columns.str.strip().str.lower()
df["date"] = pd.to_datetime(df["date"], errors="coerce")
df["amount"] = pd.to_numeric(df["amount"], errors="coerce")
df = df.dropna(subset=["customer_id", "date", "amount"])
df = df[df["amount"] > 0]
return df.reset_index(drop=True)Notice the first line of the function. Without df.copy(), renaming the columns and converting the dates would also change the DataFrame you passed in: after calling clean_sales(raw), your raw table would suddenly have lowercase column names, and a later cell that expects "Date" would fail. That’s the case for making functions pure where you can. A pure function gets its inputs through parameters, hands back its result, and doesn’t change anything else along the way. It’s easy to test (pass it a small table and check what comes back), easy to reuse, and it never surprises you. Not every function can be pure, since something has to read the file and save the plot, but the closer you get, the fewer mysteries you’ll chase.
A layout with a place for everything
Once you have a handful of functions, they need a home. A common layout for data projects, and the one this chapter uses throughout, looks like this:
sales-project/
├── README.md # what, why, and how to run it
├── pyproject.toml # makes src/sales installable (see below)
├── notebooks/ # narrative exploration, one per question
│ └── 01-explore.ipynb
├── src/
│ └── sales/ # your package: import it as `sales`
│ ├── __init__.py
│ ├── cleaning.py
│ ├── paths.py
│ └── plotting.py
├── scripts/ # command-line entry points that import sales
│ └── run_cleaning.py
├── data/
│ ├── raw/ # the files you were given; never edited
│ └── processed/ # made by your code; safe to delete and rebuild
└── figures/ # generated plots
Each folder has one job. notebooks/ is where you think out loud. src/sales/ holds the functions more than one notebook or script depends on. scripts/ holds the command-line entry points you actually run, on a schedule or against a batch of files. data/raw/ is read-only, data/processed/ can always be rebuilt, and the README says how to run everything (Chapter 30 covers the whole-project view). The payoff is that nobody has to ask where anything is. Where’s the cleaning logic? src/sales/cleaning.py. How do I run it? scripts/run_cleaning.py. Where did the output go? data/processed/.
Why a sales folder inside src/, rather than putting modules straight into src/? Because src is a terrible name to import (every project would have one), and because the extra folder is what lets you install your code as a real package, which, as the next section shows, is what makes imports stop hurting. The Python Packaging Authority calls this the “src layout.”
17.4 Paths and imports: why it works here and not there
The working directory decides what a relative path means
Here’s the single most confusing thing about moving code between scripts and notebooks: a relative path like data/raw/sales.csv is relative to wherever Python is running, not to where the file lives. That place is the working directory. Run the script above from the project folder and it works. Run the same file from inside scripts/ and it fails:
$ cd scripts
$ python clean_q3.py
...
FileNotFoundError: [Errno 2] No such file or directory: 'data/raw/sales.csv'
The code didn’t change; the folder you were standing in did. Notebooks add a twist that catches almost everyone: JupyterLab starts each notebook’s kernel in the notebook’s own folder, not in the folder you launched jupyter lab from. A notebook in notebooks/ is standing in notebooks/, so data/raw/sales.csv isn’t there for it either. Run import os; print(os.getcwd()) in a notebook and in a terminal, and you’ll see two different answers.
The dependable fix is to stop relying on the working directory at all. Pick one anchor, the project root, and build every path from it. A module can find its own location through the special variable __file__, and from there walk up to the root (the same trick Chapter 10 describes). Put that in one file and import it everywhere:
# src/sales/paths.py: one place for every path the project uses
from pathlib import Path
PROJECT_ROOT = Path(__file__).resolve().parents[2] # sales/ -> src/ -> project root
DATA_RAW = PROJECT_ROOT / "data" / "raw"
DATA_CLEAN = PROJECT_ROOT / "data" / "processed"
FIGURES = PROJECT_ROOT / "figures"Now a notebook, a script, and a scheduled job can all write pd.read_csv(DATA_RAW / "sales.csv") and get the same file, whatever folder they started in. The one thing never to do is hardcode an absolute path such as pd.read_csv("/Users/alex/Downloads/survey.csv"). It works on exactly one machine, until the day Alex moves the file, and everyone else gets a FileNotFoundError the moment they try it.
Why from src.cleaning import ... fails, and the fix
Imports have the same problem as paths, and they’re even more confusing because the error doesn’t mention folders at all. Suppose scripts/run_cleaning.py starts with from src.cleaning import clean_sales. You run it from the project root, where src/ is plainly sitting there, and get:
$ python scripts/run_cleaning.py
Traceback (most recent call last):
File "/Users/you/sales-project/scripts/run_cleaning.py", line 6, in <module>
from src.cleaning import clean_sales
ModuleNotFoundError: No module named 'src'
If this has happened to you, it isn’t your fault; lots of tutorials say it should work. Python looks for imports in a list of folders called sys.path, and when you run a script, the first entry is the folder that contains the script, not the folder you’re standing in. So python scripts/run_cleaning.py searches scripts/, finds no src there, and gives up. (The interactive python prompt and python -c do search the current folder, which is why the same import works when you try it by hand.) In a notebook the kernel searches the notebook’s folder, notebooks/, with the same result. Chapter 7 has more on reading errors like this one.
People reach for three fixes, and they’re not equally good.
Install your package in editable mode (recommended). Add a small pyproject.toml at the project root:
[build-system]
requires = ["setuptools>=64"]
build-backend = "setuptools.build_meta"
[project]
name = "sales"
version = "0.1.0"
dependencies = ["pandas"]Then, with your project’s virtual environment active, run this once from the project root. Running pip as python -m pip installs into whichever Python python is, so your package can’t end up in a different Python from the one that runs your code:
python -m pip install -e .The -e stands for editable: instead of copying your code somewhere, pip points the environment at your src/ folder (pip’s guide to local installs explains the details). From then on, from sales.cleaning import clean_sales works in every script, notebook, and terminal that uses that environment, from any folder, and edits to your .py files take effect without reinstalling. You only rerun python -m pip install -e . when you change pyproject.toml itself, say to add a dependency. The Packaging User Guide’s tutorial and its guide to writing pyproject.toml go further when you’re ready.
Patch sys.path by hand (a quick fix for one notebook). At the top of a notebook in notebooks/, you can tell Python where to look:
import sys
from pathlib import Path
sys.path.insert(0, str(Path.cwd().parent / "src")) # notebooks/ -> project root -> src/
from sales.cleaning import clean_salesThis works, and you’ll see it in plenty of course notebooks. Its weakness is that it depends on the working directory again: move the notebook one folder deeper and it silently points at the wrong place. Treat it as a stopgap until you set up the install.
Run scripts as modules with python -m. From the project root, python -m scripts.run_cleaning puts the current folder first on sys.path, so imports from the root work. It helps for scripts but does nothing for notebooks, and you have to remember the unusual command every time.
When imports get confusing, resist piling on more sys.path lines. It’s almost always a project-structure problem: check that you have one project root, that your code is a real package under src/, and that it’s installed in the environment your kernel uses. Fix the structure and the workarounds become unnecessary.
17.5 Using your own code in a notebook
Import it, then tell the story around it
Once your functions live in src/sales/ and the package is installed, a notebook pulls them in with a normal import, the same as pandas:
# notebooks/01-explore.ipynb, first code cell
import pandas as pd
from sales.cleaning import clean_sales
from sales.paths import DATA_RAW
from sales.plotting import plot_monthly_revenue
df = clean_sales(pd.read_csv(DATA_RAW / "sales.csv"))
plot_monthly_revenue(df)Look at what the notebook does and doesn’t do. It loads the data, calls the cleaning function, calls the plotting function, and surrounds those calls with Markdown cells explaining what you’re looking at and why. What it doesn’t contain is the body of clean_sales or plot_monthly_revenue; those live in src/, where there’s exactly one copy. A useful rule of thumb: if a code cell grows past ten or fifteen lines, it’s usually hiding a function that wants to move to src/.
Why your edits don’t show up, and autoreload
The first time you fix a function in src/sales/cleaning.py and re-run the notebook cell that calls it, the old version runs. You haven’t done anything wrong. Python loads a module once per session and keeps it in memory; importing it again just hands back the copy it already has, so your edits on disk don’t reach the running kernel. Running from sales.cleaning import clean_sales a second time changes nothing.
You have three ways out. The sure one is to restart the kernel (Kernel → Restart Kernel…) and run your cells again. It always picks up the latest code, but you lose everything in memory, which hurts if loading the data took five minutes. The targeted one is importlib.reload, which re-reads one module:
import importlib
import sales.cleaning
importlib.reload(sales.cleaning) # re-read src/sales/cleaning.py
from sales.cleaning import clean_sales # grab the new version of the functionThe convenient one, and the right default while you’re going back and forth between a .py file and a notebook, is IPython’s autoreload extension. Put this in the first cell:
%load_ext autoreload
%autoreload 2From then on, before every cell runs, IPython checks whether any imported module changed on disk and reloads it. Edit cleaning.py, save, re-run the cell, and the new code runs. The IPython docs are honest that reloading can’t always be done cleanly (changing a class’s structure, for example, can confuse it), so if something behaves strangely after an edit, restart the kernel. And before you hand a notebook in, restart and run everything from the top once with a fresh kernel, so you know it works without any reloading tricks.
Running a script from a notebook instead
Sometimes you want to run a script rather than import from it, exactly as you would in the terminal. A ! at the start of a line hands it to the shell:
!python scripts/run_cleaning.py --input data/raw/sales.csv --output data/processed/sales_clean.csv
!ls -lh data/processed/That’s handy for showing a whole pipeline run inside a notebook, or for a script written in another language. But the script runs as a separate program, so nothing it computes is available to your notebook afterward: if it builds a DataFrame and exits, that DataFrame is gone, and the only way to see the results is to read the file it wrote. (Remember too that the shell starts in the notebook’s folder, so these relative paths need adjusting if your notebook lives in notebooks/.) For exploratory work, importing the functions and calling them directly is almost always simpler: one program, one memory space, no surprises.
A smoke-test cell
A quick check at the top of a notebook, a smoke test in testing jargon, catches the usual setup problems (wrong kernel, missing data, a package that isn’t installed) with a clear message, before they show up as a baffling error five cells later:
# Smoke test: run this cell first
import os
import sys
import pandas as pd
import sales
from sales.paths import DATA_RAW
print("python: ", sys.executable)
print("cwd: ", os.getcwd())
print("sales: ", sales.__file__)
print("pandas: ", pd.__version__)
for name in ["sales.csv"]:
path = DATA_RAW / name
print("OK " if path.exists() else "MISSING", path)In the scratch project used to check this chapter, it printed:
python: /Users/you/sales-project/.venv/bin/python
cwd: /Users/you/sales-project/notebooks
sales: /Users/you/sales-project/src/sales/__init__.py
pandas: 3.0.6
OK /Users/you/sales-project/data/raw/sales.csv
Each line answers one question. If python isn’t inside your project’s .venv, the notebook is using the wrong kernel (see Chapter 16). If import sales fails, the package isn’t installed in that environment. If a file says MISSING, you know to fetch it before the analysis. Each is a thirty-second fix here and a half-hour detour later.
17.6 Passing parameters from the command line
Why bother with parameters
Sooner or later you’ll want to run the same analysis on a different file: next quarter’s data, another city, a different random seed. The tempting move is to copy the script and change one line, and within a month you have clean_q3.py, clean_q4.py, and clean_q4_fixed.py, each slightly different. Parameters fix that. The same script runs on any input, and it becomes something you can automate: run it from a cron job, loop it over a folder of files, or call it from a pipeline (Chapter 33).
Parameters come in three levels, and it’s fine to start at the first. Constants at the top of the script, like INPUT = Path("data/raw/sales.csv"), are quick and fine for one-off work, but every change means editing code. A configuration file (a small .yml or .toml the script reads at startup) suits settings several teammates adjust without touching code. Command-line arguments are the most flexible: you pass values when you run the script, which is exactly what automation needs.
Command-line arguments with argparse
Python’s built-in argparse module handles the fiddly parts for you. You describe each argument once, with a help message and maybe a default, and argparse reads what the user typed, converts types, rejects bad input, and writes a --help page:
import argparse
from pathlib import Path
def parse_args():
p = argparse.ArgumentParser(description="Clean a sales CSV.")
p.add_argument("--input", type=Path, required=True,
help="path to the raw CSV")
p.add_argument("--output", type=Path, required=True,
help="where to write the cleaned CSV")
p.add_argument("--seed", type=int, default=0,
help="random seed for reproducibility")
return p.parse_args()That’s enough to give your script a real command-line interface. Here’s what Python 3.12 prints for --help:
$ python scripts/clean.py --help
usage: clean.py [-h] --input INPUT --output OUTPUT [--seed SEED]
Clean a sales CSV.
options:
-h, --help show this help message and exit
--input INPUT path to the raw CSV
--output OUTPUT where to write the cleaned CSV
--seed SEED random seed for reproducibility
And here’s what happens when someone forgets an argument, or types a word where a number belongs. argparse stops before any of your code runs:
$ python scripts/clean.py
usage: clean.py [-h] --input INPUT --output OUTPUT [--seed SEED]
clean.py: error: the following arguments are required: --input, --output
$ python scripts/clean.py --input data/raw/sales.csv --output out.csv --seed ten
usage: clean.py [-h] --input INPUT --output OUTPUT [--seed SEED]
clean.py: error: argument --seed: invalid int value: 'ten'
Check inputs early, and say what you ran
argparse can check that --seed is a number, but not that the input file exists or that the seed makes sense for your analysis. Check those yourself at the start of main(), following the fail-fast idea: stop immediately with a clear message rather than running for ten minutes and crashing on a missing file. raise SystemExit("message") prints the message and ends the script with a nonzero exit status, which is how other programs (and your automation) know it failed:
def main():
args = parse_args()
if not args.input.exists():
raise SystemExit(f"input not found: {args.input}")
if args.seed < 0:
raise SystemExit(f"seed must be non-negative, got {args.seed}")
print(f"INPUT={args.input} OUTPUT={args.output} SEED={args.seed}")
...$ python scripts/clean.py --input data/raw/q4.csv --output out.csv
input not found: data/raw/q4.csv
That last print in main() is worth keeping too. When a script announces the settings it’s running with, anyone reading its output later can tell exactly which configuration produced a result. Two more habits finish the job: put the settings that matter (a date, a seed, the input’s name) into output file names so you can tell two runs apart, and write the exact command you ran into the README next to the output, so you can repeat it six months from now without guessing.
17.7 Turning notebooks into scripts (and back)
Why convert
Once an analysis settles down, there are good reasons to move it out of the notebook. The first is automation: a .py file runs from a shell, a scheduled job, or continuous integration without a browser or a running kernel. If it needs to run every Monday, it needs to be a script.
The second is version control. A notebook is saved as JSON, and every output is stored inside it: each plot becomes a single line of Base64 text thousands of characters long, and every rerun changes the execution counts. So a diff of a notebook mixes the one line you changed with screens of noise, and merge conflicts in the raw JSON are miserable to resolve by hand (tools like nbdime help). A script’s diff shows only what a person changed. Chapter 31 has more.
The third reason is the most interesting: converting a notebook is an honest test of whether it works. A script runs top to bottom with no leftover memory to rescue it, so a cell that depends on a variable from a deleted cell, or on being run out of order, fails loudly. The first run of a converted notebook is sometimes the first time anyone learns what it really depends on.
Notebook to script with nbconvert
jupyter nbconvert comes with Jupyter and does the conversion in one command:
jupyter nbconvert --to script notebooks/analysis.ipynb --output-dir scripts/It writes scripts/analysis.py, with every code cell in order and Markdown cells turned into comments. Here’s the start of what it produced for a small notebook that used autoreload:
#!/usr/bin/env python
# coding: utf-8
# # Sales analysis
# Exploration of Q3 sales data.
# In[1]:
get_ipython().run_line_magic('load_ext', 'autoreload')
get_ipython().run_line_magic('autoreload', '2')
# In[2]:
import pandas as pd
from sales.cleaning import clean_sales(The first line is a shebang, which lets macOS and Linux run the file directly once it’s marked executable.) Now try running it with plain Python, and it dies on the first real line:
$ python scripts/analysis.py
Traceback (most recent call last):
File "/Users/you/sales-project/scripts/analysis.py", line 10, in <module>
get_ipython().run_line_magic('load_ext', 'autoreload')
^^^^^^^^^^^
NameError: name 'get_ipython' is not defined
Magics like %autoreload and shell lines like !ls only mean something inside Jupyter, so nbconvert translates them into calls to get_ipython(), a function that exists only there. That’s the clue to how to treat the result: the converted file is a starting point, not a finished script. It still has the df.head() calls you ran to peek at the data, the commented-out experiments, and imports scattered through the file. Tidy it in a few passes: delete the magics and the inspection calls, move repeated blocks into functions in src/, wrap the main flow in main() with the __name__ guard, add argparse if you want options, and finally run it from a fresh terminal to confirm it still does what the notebook did.
Keeping a notebook and a script in sync with Jupytext
Sometimes you want both: a notebook to work in and a plain-text file to review. Jupytext pairs a notebook with a .py file in “percent” format, where special comments mark each cell. Install it (python -m pip install jupytext) and pair a notebook once:
jupytext --set-formats ipynb,py:percent notebooks/analysis.ipynbThe paired notebooks/analysis.py is ordinary Python:
# %% [markdown]
# # Sales analysis
# Exploration of Q3 sales data.
# %%
# %load_ext autoreload
# %autoreload 2
# %%
import pandas as pd
from sales.cleaning import clean_salesNotice that Jupytext comments out the magics, so the file even runs with plain python. When you save the notebook in Jupyter, both files are updated. When you edit the .py in another editor, the notebook catches up the next time you open or reload it in Jupyter, or right away if you run jupytext --sync notebooks/analysis.py. Reviewers read the .py, you keep working in the notebook, and you can commit both files or only the .py and regenerate the other. For a project that uses notebooks heavily and gets code review, it’s a small setup cost for a big payoff.
Running one notebook with different parameters: Papermill
There’s also a middle path: keep the notebook, but run it from the command line with different values each time, to make one report per month or the same analysis for fifty cities from a single template. Papermill does this. First, give one cell the tag parameters (in JupyterLab, select the cell, open the Property Inspector with the gear icon in the right sidebar, and type parameters in the Add Tag box). That cell holds the defaults:
# Cell tagged `parameters`
input_file = "data/raw/sales.csv"
output_file = "data/processed/summary.csv"
start_date = "2026-01-01"Then run it with new values:
papermill notebooks/analysis.ipynb outputs/analysis_q3.ipynb \
-p input_file data/raw/sales_q3.csv \
-p output_file data/processed/summary_q3.csv \
-p start_date 2026-07-01Papermill adds a new cell, tagged injected-parameters, right after yours, so your values override the defaults, then runs the whole notebook and saves a new copy with every output in it. One snag: unlike JupyterLab, Papermill starts the kernel in the folder you ran the command from (here, the project root), unless you pass --cwd; paths built from sales.paths sidestep the question. Use Papermill when the thing you want at the end is a readable notebook and only a few inputs change between runs. When the output is just data, a plain script is simpler.
17.8 Notebook or script? Choosing on purpose
What each is good at
Notebooks shine when the goal is exploring, learning, or explaining: when you’re still finding out what the data looks like, when you want to go back and forth between code and charts, and when the finished product is a story with plots and commentary that someone will read. Mixing code, prose, and output in one document is an old idea called literate programming, and it’s something scripts can’t do.
Scripts shine when the goal is running the same thing again: many inputs, on a schedule, with no one watching, with predictable output in predictable places. Anything that should run at 3 a.m., on every push to GitHub, or over a hundred files should be a script. Scripts also make clean diffs, which makes them easy to review.
In a real project the best answer is usually both, in layers. The reusable logic lives in src/sales/ as plain modules. Command-line entry points in scripts/ import from it. Notebooks in notebooks/ import from it too and tell the story. The notebook is the story, the scripts are the engine, and src/ is the shared library underneath, so the cleaning function you run interactively is the very same one your pipeline runs overnight.
The signs that a notebook should become (or call) a script are easy to spot. You keep rerunning it with small changes to a few values. It takes long enough that you’d like to run it overnight. It needs to run on a schedule or in a pipeline. You wish it had proper logs or consistent output files. The signs that it should stay a notebook are just as clear: the point is interpretation or communication, the analysis is still changing fast, or you’re teaching or documenting your reasoning. Notebooks are the wrong tool for a production pipeline and the right tool for thinking; don’t let “scripts are more professional” push you out of a notebook when a notebook is what you need.
Write down how to run it
Any project someone else might run, including future you, deserves a “How to run” section in its README.md with the exact commands, in order, that a fresh reader on a fresh machine would type. Not “activate the environment and run the analysis,” but the actual commands:
## How to run
1. Create and activate the environment:
```bash
python -m venv .venv
source .venv/bin/activate # on Windows: .venv\Scripts\activate
python -m pip install -r requirements.txt
```
2. Place the raw data at `data/raw/sales.csv` (download link: ...).
3. Run the cleaning pipeline:
```bash
python scripts/run_cleaning.py \
--input data/raw/sales.csv \
--output data/processed/sales_clean.csv
```
4. Open the analysis notebook:
```bash
jupyter lab notebooks/01-explore.ipynb
```
Expected runtime: ~30 seconds for cleaning, ~2 minutes for the notebook.
Expected outputs: `data/processed/sales_clean.csv` (~5 MB).(If your project uses the src/sales package from this chapter, add python -m pip install -e . to step 1.) The exact commands save everyone from “I just have to type that thing I remember from last month.” The expected runtime and output size are a sanity check: if cleaning takes five minutes instead of thirty seconds, or the file is 50 KB instead of 5 MB, something’s wrong. A good “How to run” section gets people unblocked without having to ask you a single question.
Keep notebooks friendly to version control
A few habits keep notebooks from taking over your repository (Chapter 31). Clear outputs before you commit a notebook for others to review (Edit → Clear Outputs of All Cells), or let a tool like nbstripout strip them automatically on every commit; a notebook full of plots is mostly Base64 and bloats the repository. Don’t dump whole datasets into cell outputs: a cell that ends with a bare df on a 50 MB table saves a big chunk of it into the notebook, so show df.head(), df.shape, or df.describe() instead, and keep raw data in data/. And for anything that gets code review, pair with Jupytext or convert to a script, so reviewers read plain text. For a notebook handed in once and never reviewed, that’s overkill; for one in a shared project, it’s worth the few minutes.
17.9 Stakes and politics
In 2019, a team of researchers collected 1.4 million Jupyter notebooks from GitHub and tried to run the Python ones again. Of the notebooks they attempted, about 24% ran without an error, and about 4% produced the same results they had shown when they were saved (Pimentel et al., 2019). The most common reasons were the ones this chapter is about: missing dependencies, hidden state and cells run out of order, and data files that weren’t where the code expected. A notebook that won’t rerun isn’t only its author’s problem. When it backs a published figure, a policy memo, or a class project someone else builds on, the cost lands on whoever tries to check or extend the work, often a student or reviewer with less time and less help than the author had.
The fixes carry their own politics. The habits that make work rerunnable (packages, command-line entry points, a README with exact commands) come from software engineering, and they’re mostly learned informally, from a mentor or a job, not in a methods course. People who picked them up that way find reproducibility easy and can read notebook-only work as unserious; people who didn’t pay for the gap in time, in hiring conversations, and in how their work is judged. The default of “it ran on my laptop” quietly shifts the work of reproduction onto everyone downstream.
See Chapter 8 for the broader framework. The concrete prompt to carry forward: before you share an analysis, ask who will try to rerun it, and whether they could do it from your repository alone, without asking you anything.
17.10 Worked examples
Turning notebook code into importable functions
Your notebook has the same cleaning lines in three cells: lowercase the column names, parse the dates, drop rows with no customer id. Time to move them out. Create src/sales/cleaning.py:
# src/sales/cleaning.py
import pandas as pd
def clean_sales(df: pd.DataFrame) -> pd.DataFrame:
df = df.copy()
df.columns = df.columns.str.strip().str.lower()
df["date"] = pd.to_datetime(df["date"], errors="coerce")
return df.dropna(subset=["customer_id", "date"])With the package installed (python -m pip install -e ., once), replace the three cells in the notebook with one import and one call:
from sales.cleaning import clean_sales
from sales.paths import DATA_RAW
df = clean_sales(pd.read_csv(DATA_RAW / "sales.csv"))Now the cleaning logic lives in exactly one place, and an improvement there reaches every notebook and script that uses it.
Writing a command-line script around your functions
Next, make the same logic something you can run on any file from the terminal:
# scripts/run_cleaning.py
import argparse
from pathlib import Path
import pandas as pd
from sales.cleaning import clean_sales
def main():
p = argparse.ArgumentParser(description="Clean a sales CSV.")
p.add_argument("--input", type=Path, required=True)
p.add_argument("--output", type=Path, required=True)
args = p.parse_args()
df = pd.read_csv(args.input)
cleaned = clean_sales(df)
cleaned.to_csv(args.output, index=False)
print(f"wrote {len(cleaned):,} rows to {args.output}")
if __name__ == "__main__":
main()$ python scripts/run_cleaning.py \
--input data/raw/sales.csv \
--output data/processed/sales_clean.csv
wrote 3 rows to data/processed/sales_clean.csv
The cleaning function didn’t change, the notebook didn’t change, and you now have a third way to use the same logic. Changing the input is a flag, not a code edit.
“It imports in the terminal but not in the notebook”
from sales.cleaning import clean_sales works when you type it at the python prompt, but the same line in your notebook fails with ModuleNotFoundError: No module named 'sales'. One cell tells you which of the two usual causes you have:
import os
import sys
print("python:", sys.executable)
print("cwd: ", os.getcwd())If python isn’t the interpreter inside your project’s .venv, the notebook is running a different kernel, one where your package was never installed. Switch kernels, or register your environment as a kernel (see Chapter 16). If the interpreter is right, check how the terminal import succeeded: if you never ran python -m pip install -e ., it only worked because the terminal was sitting in src/, and the notebook’s folder is notebooks/. Install the package into that environment and the import works everywhere.
Cleaning up a converted notebook
You’ve run jupyter nbconvert --to script notebooks/analysis.ipynb --output-dir scripts/ and have a 120-line scripts/analysis.py. Work through it in order. Delete every get_ipython() line (the old magics and ! commands) and every bare df.head() or df.info() left over from exploring. Move the imports to the top. Move repeated blocks into functions in src/sales/, and replace them with calls. Wrap what’s left in main(), add the if __name__ == "__main__": guard, and swap the hardcoded file names for argparse options. Then run it from a fresh terminal. Whatever breaks now (usually a variable that was only defined in a cell you deleted) is hidden state the notebook had been hiding from you. Fix it here, and the notebook that stays behind in notebooks/ can go back to telling the story.
17.11 Templates
Template A: Minimal script skeleton
"""One-sentence purpose.
Inputs:
Outputs:
How to run:
"""
from pathlib import Path
def main():
...
if __name__ == "__main__":
main()Template B: Minimal command-line script
import argparse
from pathlib import Path
def parse_args():
p = argparse.ArgumentParser(description="What this script does.")
p.add_argument("--input", type=Path, required=True, help="input file")
p.add_argument("--output", type=Path, required=True, help="output file")
return p.parse_args()
def main():
args = parse_args()
if not args.input.exists():
raise SystemExit(f"input not found: {args.input}")
print(f"INPUT={args.input} OUTPUT={args.output}")
...
if __name__ == "__main__":
main()Template C: Minimal pyproject.toml for a src/ package
[build-system]
requires = ["setuptools>=64"]
build-backend = "setuptools.build_meta"
[project]
name = "yourpackage" # matches the folder src/yourpackage/
version = "0.1.0"
dependencies = ["pandas"]Template D: A notebook that calls functions from src/
# 1) Purpose, imports, %load_ext autoreload / %autoreload 2
# 2) Smoke test (interpreter, package location, data files)
# 3) Parameters (as variables)
# 4) Call functions from src/
# 5) Save outputs
# 6) Interpretation (Markdown)
# 7) Restart-and-run-all before sharing
17.12 Exercises
Write a script that loads a CSV and prints a short summary (rows, columns, missing values per column).
Refactor the script so the summary logic is a function and the rest is only an entry point, with an
if __name__ == "__main__":guard.Move that function into a
src/package, install it withpython -m pip install -e ., and use it from a notebook on two different datasets.Add a command-line interface with
--inputand--outputflags, and paste its--helpoutput into your README.Convert one of your notebooks to a script with
nbconvert, get it running from a fresh terminal, and list at least three hidden-state problems you had to fix.Write a short paragraph explaining whether your current project should be a notebook, a script, or a hybrid, and why.
17.13 One-page checklist
- My script has a
main()and anif __name__ == "__main__":guard, so importing it doesn’t run the analysis. - Reusable logic lives in functions in a
src/package, not copied between notebook cells. - My package is installed in the project environment with
python -m pip install -e ., so imports work from any folder. - Paths are built from one anchor (
PROJECT_ROOT), not the working directory, and never hardcoded to my laptop. - Notebooks that import my code start with
%load_ext autoreloadand%autoreload 2and a smoke-test cell. - My script takes its inputs from the command line, checks them early, and prints the settings it ran with.
- Every notebook I share passes Restart Kernel and Run All Cells.
- The README has a “How to run” section with exact commands.
- I choose notebooks for exploring and explaining, and scripts for running things again.
17.14 Quick reference: commands
| Task | Command |
|---|---|
| Run a script | python scripts/run_cleaning.py --input ... --output ... |
| See a script’s options | python scripts/run_cleaning.py --help |
Make src/ importable everywhere |
python -m pip install -e . (once, from the project root) |
| Reload edited modules in a notebook | %load_ext autoreload then %autoreload 2 |
| Find where an import comes from | print(module.__file__) |
| Notebook to script | jupyter nbconvert --to script nb.ipynb --output-dir scripts/ |
Pair a notebook with a .py |
jupytext --set-formats ipynb,py:percent nb.ipynb |
Update the pair after editing the .py |
jupytext --sync nb.py |
| Run a notebook with parameters | papermill in.ipynb out.ipynb -p name value |
- Python docs,
argparsetutorial — the official walk-through for building command-line interfaces with the standard library, from one argument to many. - Python docs,
__main__: top-level code environment — the full story behindif __name__ == "__main__":, including__main__.pyfiles in packages. - Pallets, Click documentation — the most widely used third-party library for command-line tools; worth a look when
argparsestarts to feel verbose. - Sebastián Ramírez, Typer documentation — a modern alternative to Click that reads your options from Python type hints; about the least effort it takes to turn an existing function into a command.
- Real Python, Build Command-Line Interfaces With Python’s argparse — a longer tutorial with subcommands, validation, and worked examples.
- Jupytext, Jupytext documentation — everything about pairing notebooks with text files, including the other text formats and editor integrations.
- Python Packaging Authority, src layout vs flat layout — why the
src/layout used in this chapter prevents a class of import mistakes.