35 Using AI Tools
Prerequisites (read first if unfamiliar): Chapter 2.
See also: Chapter 6, Chapter 36, Chapter 3, Chapter 25, Chapter 26, Chapter 29.
Purpose

It’s late, the assignment is due in the morning, and your code won’t run. You paste the error into a chat assistant, and it hands you a friendly explanation and a fix: just call df.clean_columns() first. You try it and get AttributeError: 'DataFrame' object has no attribute 'clean_columns'. Twenty minutes of wondering what’s wrong with your pandas later, it dawns on you that the method never existed. The assistant made it up, and sounded exactly as sure as it does when it’s right.
If that has happened to you, you’re in good company, and it doesn’t mean you’re bad at using these tools. AI assistants are built on large language models, which produce text that fits the patterns in their training data and your prompt. Fitting and being true overlap most of the time, which is why the tools are so useful for drafting and explaining. They don’t always overlap, and the assistant can’t tell you which case you’re in. So this chapter rests on one rule:
AI can propose. You must verify.
Verifying means checking official documentation, running small experiments, and adding tests, and it stays your job whoever wrote the first draft. This chapter covers how much checking a task needs, how to ask for answers you can check, how to use AI for debugging, documentation, and code, what never to paste into a chat window, and how to stay within your course’s policy. How the models work inside is Chapter 36, evaluating AI systematically is Chapter 38, and agents are Chapter 37. It won’t tell you which product to use, either: tools and prices change every few months, and these habits should outlast them.
Why read this chapter
- An assistant told you to call a function that doesn’t exist, and it took you far too long to suspect the advice instead of your code.
- You asked for sources for a paper, and one of them isn’t in the library, on Google Scholar, or anywhere else.
- Code an assistant wrote for you runs without a single error, and you have no idea whether the numbers it prints are right.
- A chatbot suggested
sudo chmod -R 777to fix a permission error, and you’re not sure whether you should run it. - One syllabus bans AI tools, another asks you to disclose them, and a third encourages them, and you’d like habits that hold up under all three.
- You’re about to paste a dataset (or an API key) into a chat window and have a nagging feeling you shouldn’t.
- You keep getting vague, generic answers and suspect a better question would get a better one.
Running theme: AI can propose; you must verify
An assistant sounds just as confident when it’s wrong as when it’s right, so its confidence tells you nothing; the documentation, a small experiment, and a test are what tell you whether to trust it.
35.1 What an AI assistant is doing
Whether it’s a chat window or a coding helper in your editor, an AI assistant is almost always built on a large language model (LLM). Chapter 36 explains how they work; for using them, one idea is enough. An LLM generates the text that most plausibly comes next, given its training and what you’ve typed, so it’s very good at shaping language and weak at guaranteeing facts. Picture a fast, well-read collaborator who has never seen your computer and is occasionally, fluently wrong. Your job is to turn that collaborator’s drafts into work you can stand behind.
The strengths follow from that picture. Assistants shine when the task is mostly language, structure, or common patterns: drafting an outline, a README, a docstring, or an issue template, where the shape is conventional and you need a starting point to edit. They’re good at rewriting an explanation for a different reader, and at boilerplate you’d otherwise look up every time, like the skeleton of a command-line script with argparse or a standard logging setup. They’re good at suggesting search terms, documentation sections worth reading, likely causes of a problem, and long checklists you can cut down. In each case “roughly right” is a fine start, because your editing finishes the job.
The weaknesses follow from the same picture, and they’re worth knowing by name:
- Invented details. A flag, a function, a parameter, or a citation that sounds plausible and doesn’t exist. This is usually called hallucination, and the
clean_columns()story above is a typical case. - The wrong version. Advice that was right for some version of pandas, scikit-learn, or git, just not the one you have. A model learns from text up to a cutoff date, and libraries keep changing after it.
- Hidden prerequisites. The answer assumes you’ve already activated the right environment, moved into the right folder, or installed a system library, and never says so.
- Silent logic bugs. Code that runs without an error on the example in your prompt and is wrong on the data you actually have.
- Unsafe defaults. A fix that technically works by weakening security or by making a mistake much harder to undo.
Some assistants can now search the web or read files you attach. That helps with versions but not the rest: an assistant can misread a page, or cite one that doesn’t say what it claims. The quieter trap is on your side of the screen. People over-trust suggestions from automated systems even when they have evidence to know better (automation bias), and a fluent answer switches off the part of you that would have checked. So assume any of these errors can turn up: ask where a claim comes from, test on real inputs, and prefer small changes you can reason about.
35.2 A risk-based verification policy
“Can I trust AI?” is the wrong question, because it depends on the task. A better one is what happens if this answer is wrong? A clumsy sentence in a draft README costs you a minute; a wrong rm -rf can cost you a semester. Let the cost of a mistake set how carefully you check. (Organizations do the same at a larger scale: the U.S. NIST AI Risk Management Framework is built around matching care to potential harm.)
Low risk: drafting and formatting
Low-risk work is anything where a mistake is cheap to spot and cheap to fix: rewriting a paragraph, drafting a README skeleton, producing a checklist or a template. Being wrong costs a careful read and an edit, so that’s the check: read the output for accuracy and completeness, and make sure it fits the assignment you’re actually doing, not a generic version of it.
Medium risk: technical guidance you can test quickly
Medium-risk work makes technical claims you can test with small experiments: interpreting a stack trace, suggesting commands to inspect your environment, drafting a test scaffold, proposing a refactor your tests can check. Being wrong usually costs time, not damage. The check is to run the proposed commands, look up each function in the official docs to confirm it exists and does what was claimed, and lock in the result with a test or an assertion.
High risk: destructive commands, security changes, and sensitive data
High-risk work is anything whose mistakes can’t easily be undone, or can hurt someone besides you: commands that delete, overwrite, or recursively move files; anything using sudo or broad permission changes like chmod -R 777; changes to SSH keys, VPNs, firewalls, or other network and login settings; and anything touching confidential, protected, or regulated data.
Here the bar goes way up. Read the primary documentation yourself, not the assistant’s summary of it. Get help from an instructor, a TA, or an experienced colleague rather than running the suggestion alone. Where you can, try the change in a sandbox first: a throwaway virtual machine, a scratch folder, a separate branch. And never run a command you don’t understand, however confident the assistant sounds. Treat these as stop signs until you’ve confirmed the paths, the backups, and what you meant to do:
# Stop signs: do not run these on an assistant's say-so
rm -rf <path> # macOS/Linux: recursive delete, no undo
del /s <path> # Windows: recursive delete
mv <source> <dest> # can silently overwrite <dest>
chmod -R 777 <path> # everyone can read, write, and run everything
sudo <anything> # runs with full administrator rights
git push --force origin main # rewrites history your teammates shareWhen an assistant suggests one of these, ask: What exactly will this change? What evidence says it’s the right change? How would I undo it? If you can’t answer all three, don’t run it.
35.3 The assistive loop: a workflow that forces evidence
AI help usually goes wrong not with one bad answer but with a loop. You paste an error, try the fix, get a new error, paste that, and an hour later your code has changed in six places, still doesn’t work, and you can’t say what the original problem was. The way out is to keep the loop yours: the assistant proposes, and evidence decides.
Start with a short specification: a one-sentence goal; the context the assistant can’t see (operating system, Python version, environment, package versions); the inputs (a snippet, a command, a sample of the data); what happened, with the exact error text; what you expected; and what you’ve tried. That’s the shape of a good question for a person, too (Chapter 2): if you can describe the situation clearly to a human, you can describe it to a model.
Ask for alternatives and checkpoints, not one answer. Asking for the fix invites the model to commit to one path and you to follow it without thinking. Ask for two to four plausible causes, checks that tell them apart, and a plan with a way to confirm each step, so you have somewhere to go when the first guess is wrong.
Check claims against primary documentation. When an answer depends on facts (what a function returns, what a flag does, which version added a feature), look it up in the official docs for your installed version (Chapter 5 shows how). If nobody can find a primary source, treat the claim as a guess.
Run the smallest experiment that could fail. Change one thing at a time so the result means something, and write down what you ran and what happened. The point isn’t to prove the assistant right; it’s to find out whether its explanation matches your computer.
# Change one thing, record the result
python -c "import sys; print(sys.executable)" # which Python is running?
python -c "import pandas; print(pandas.__version__)" # which pandas does it see?When the assistant and your computer disagree, your computer wins. The assistant says a function returns a list, and yours returns a Series: trust what happened when you ran the code. Then confirm your versions and environment, read the docs for that version, and if it’s still unclear, ask a person with a well-structured question.
Lock in what you learned. Once the fix works, add something that will catch the problem next time: a unit test for the broken behavior, an assertion for what you just discovered (a shape, a type, a range), or a line in the README. Without one, the same bug comes back the next time someone touches that code.
35.4 Prompt patterns that produce usable how-to guidance
If you’ve asked “why doesn’t my code work?” and got back a polite, generic essay, the problem was probably the question. Assistants fill the gaps you leave with the most typical answer, which rarely fits your situation. Prompts work best when they say what shape the output should take and what constraints it must respect. The craft is called prompt engineering, and the major vendors publish guides to it (see Further reading), but you don’t need special vocabulary: be specific, and ask for answers you can check. These six patterns cover most everyday needs; fill in the angle brackets.
Pattern A, decision-tree diagnosis:
I am seeing: <symptom or exact error>. Give me a decision tree of 8 checks.
Each check must be a concrete command or observation, and you must say what
each outcome implies. Keep it specific to <OS> and <tool and version>.
Pattern B, shrinking a failing example into a minimal reproducible example:
Here is my current failing snippet. Reduce it further if possible, and replace
the real data with synthetic data. Output: (1) the reduced code, (2) what you
removed and why, (3) how to confirm it still fails.
Pattern C, tests first:
Write 4 pytest tests for the intended behavior below, including 2 edge cases.
Then propose an implementation that passes them. State your assumptions.
Pattern D, a documentation rewrite that can’t invent steps:
Rewrite these notes as a how-to guide with sections: Purpose, Prerequisites,
Steps, Verify, Troubleshooting. Do not invent commands I did not provide.
If prerequisites are missing, list them as questions under Prerequisites.
Pattern E, reviewing a command before you run it:
Explain what this command does, what it changes on disk, and how to undo it:
<command>
Then suggest a safer alternative, or a dry-run option if one exists.
Pattern F, a patch instead of a rewrite:
Here is the current function. Give me a minimal patch, as a diff, that fixes
<specific bug>. Do not refactor unrelated code. Explain how to test the fix.
Each one asks for output you can verify (a check, a test, a diff), and each fences off a way assistants go wrong: inventing steps, rewriting too much, or offering one confident guess.
35.5 Using AI to improve technical questions
Sometimes the best use of an assistant is getting a question into shape before you ask a person (Chapter 2 has the details). Paste in your messy description, minus anything private, and ask for it reorganized as goal, expected behavior, actual behavior, steps to reproduce, context (versions and operating system), and what you tried. The structure makes the gaps obvious. Then audit the result: never accept an “improved” question with details you didn’t give it. If you don’t remember the exact error, the right output is a placeholder, not a plausible-looking message.
If you don’t know what context matters, ask for a checklist. For Python data work it’s nearly always the operating system, the Python version, the path of the interpreter actually running, the environment manager (conda or venv) and active environment, the relevant package versions (pip show pandas, conda list scikit-learn), and the working folder plus the paths the program reads or writes. Write “unknown” rather than guessing: the unknowns are often where the bug is.
The question you post should be yours, even if an assistant helped tidy it. Some communities ban AI-written content outright; Stack Overflow’s policy says generative AI tools “may not be used to generate content” for the site.
35.6 Using AI in debugging: hypotheses, checks, and minimal diffs
Debugging is where assistants are most tempting and where the runaway loop bites hardest. Let the assistant propose hypotheses and checks, and keep the loop in your hands. Chapter 6 has the full method; here’s how AI fits the most common situations.
Import errors are usually environment errors. When import pandas fails in one place and works in another, the question is which Python is running and where its packages live. Find out before you accept any fix:
# Which interpreter, and which packages?
python --version
python -c "import sys; print(sys.executable)"
which python # macOS/Linux
where.exe python # Windows (Command Prompt or PowerShell)
pip show <package>
conda list <package> # if you use condaIf sys.executable isn’t inside the environment you expected, you’re running the wrong interpreter. If pip show says the package lives under a different Python than sys.executable, you installed it into the wrong environment. Don’t accept commands that change your environment until you understand them (Chapter 15 explains the moving parts).
Data that loads wrong needs inspecting, not guessing. When a CSV lands in one column, the likely causes are a different delimiter, odd quoting, or an encoding problem. An assistant can list those, but it can’t see your file (and you shouldn’t paste a sensitive one). Look at the first raw lines yourself, try the delimiter you see, and check the column names after loading, as the worked example “A CSV that loads into one column” does.
Path errors need a look around, not a fix. Many “file not found” errors come from running a program in a different folder than you think. Check where you are before changing any code:
pwd # macOS/Linux: where am I?
ls -la # what's here?
cd # Windows (Command Prompt): where am I?
dir # what's here?If an assistant suggests moving, renaming, or deleting files to solve a path problem, stop. Fix the path or where you ran the program from first.
Prefer small diffs to big rewrites. It’s tempting to accept a large rewrite because the error goes away. So does your ability to say what caused it, and a big change has room for new bugs. Ask for the smallest change that keeps your intent (Pattern F), apply one change at a time, and check each. Reading changes as a diff, lines removed and lines added, makes this much easier.
35.7 Using AI for documentation and project hygiene
Documentation is a good job to hand an assistant, because you can review the output by reading it and running its commands. The trap is that a well-formatted README looks finished even when half its commands were invented (Chapter 3 has more).
A README is only good if someone can follow it. Let the assistant draft the structure, give it the commands you actually use, and let it format them. At minimum a project README needs:
- Purpose
- Setup (creating the environment)
- How to run
- Verify (what success looks like)
- Troubleshooting (common failures)
Then follow it yourself in a fresh environment, from a fresh clone, exactly as written. That’s where the missing steps turn up, like the package you installed months ago and forgot.
For troubleshooting sections, ask for entries in a strict format: symptom, likely cause, check (a command or something to look at), and fix (the smallest action that works). Try every fix on your own system and delete any entry you can’t confirm; readers will trust an untested entry.
For jobs you repeat, such as refreshing a dataset, a runbook lists the prerequisites and inputs, the commands in order, the checks that show each step worked, and how to recover if one fails. An assistant can draft it from your notes; it counts only once you’ve run it end to end.
35.8 Using AI for code: constraints, review, and tests
AI-written code is most useful when you keep the pieces small and check them, and most dangerous when it works on the first try and you stop looking.
Ask for one small, testable unit at a time. Instead of a whole pipeline, ask for one function with a clear contract: input types and constraints, the output type and what it guarantees, a realistic example, and tests including edge cases. A function like that is small enough to read top to bottom and test on its own, and cheap to throw away if it’s wrong. A pipeline drafted in one prompt feels efficient, but its inevitable bug is far harder to find.
Expect code that runs and is still wrong. This catches the most people, because nothing looks broken. Say you ask for year-over-year growth in enrollment, and the assistant writes this:
import pandas as pd
def yearly_growth(df):
"""Percent change in enrollment from the year before."""
return df.set_index("year")["enrollment"].pct_change() * 100On the tidy example in your prompt, it’s perfect:
example = pd.DataFrame({"year": [2020, 2021, 2022], "enrollment": [100, 110, 121]})
print(yearly_growth(example))year
2020 NaN
2021 10.0
2022 10.0
Name: enrollment, dtype: float64
But pct_change compares each row with the row above it, not with the previous year. Give it your real file, where the rows happen to arrive as 2021, 2020, 2022, and it reports a 9% drop in 2020 and 21% growth in 2022, with no error and no warning. A test written before you trust the code catches it:
import pytest
def test_growth_does_not_depend_on_row_order():
df = pd.DataFrame({"year": [2021, 2020, 2022], "enrollment": [110, 100, 121]})
assert yearly_growth(df).loc[2021] == pytest.approx(10.0)Run it with pytest and it fails with Obtained: nan and Expected: 10.0 ± 1.0e-05. Adding .sort_index() after set_index("year") fixes it. That’s why asking for tests first (Pattern C, the idea behind test-driven development) works so well: writing tests forces you to pin down what “right” means. For a function that cleans a text column, you’d have to decide how empty strings are treated, what happens to whitespace, whether case is kept, and what a non-string value does. Once those decisions are tests, the assistant’s code meets them or it doesn’t.
Climb the verification ladder for any AI-written code you plan to keep:
- Read it line by line. You should be able to explain every line.
- Run a smoke test on tiny inputs.
- Add assertions for what must always be true (shape, type, ranges).
- Add unit tests, including edge cases.
- If it replaces existing code, compare its output with the old code’s on the same inputs.
If you can’t explain the code, don’t keep it. “It works” isn’t enough, especially in a course, where understanding it is the point.
Be extra careful with code that writes or deletes files. When an assistant’s code deletes or overwrites anything, ask for a dry-run mode that prints what it would do, a confirmation before anything destructive, output to a dedicated outputs/ folder, and explicit paths rather than assumptions about which folder the code runs from.
35.9 Security and privacy hygiene
Most security mistakes with AI tools start as convenience: the error involves a config file, so you paste the whole file, key and all.
Never paste secrets: passwords, API keys or tokens, private SSH keys, university logins, or confidential datasets. Once a secret is in a service you don’t control, assume it’s no longer secret, whatever the interface promises (Chapter 34 says what to do if it happens). Share the structure instead (which environment variable holds the key, what the call looks like) and the error message, never the value:
import os
import requests
api_key = os.environ["WEATHER_API_KEY"] # the real value never appears in the code
response = requests.get(
"https://api.example.com/v1/forecast",
headers={"Authorization": f"Bearer {api_key}"},
timeout=10,
)Code like this is safe to share: the key lives in your environment, not in the text you paste.
Know where your words go. What you type goes to the provider’s servers, where, depending on the service and your settings, it may be stored, used to train future models, or read by people. Google’s Gemini Apps privacy notice, for one, says human reviewers read some chats and asks users not to enter “confidential information that you wouldn’t want a reviewer to see.” Other providers’ terms differ and change often, so read the ones for your tool, and check which tools your university has approved for which data (a university-licensed version may have different terms from the free one).
Share the shape of your data, not the data: the schema (column names and types), three to five made-up rows with the same structure, and a few summary numbers such as counts, missing-value rates, or ranges. That synthetic version almost always gets you the same advice. Never paste raw records containing personal data, and if data is covered by a course policy, an IRB protocol, or a law such as FERPA for student records, assume you can’t paste it anywhere.
Treat requests for more privileges as high risk. Before running a suggested sudo, recursive chmod, or system-wide setting, confirm in the documentation what it changes, how widely, and how to undo it. Prefer the option with the least power (the principle of least privilege): install into your own environment, write to a folder you own. Unsure whether you need administrator rights? Ask someone who knows that system.
35.10 Academic integrity and collaboration
Here’s a frustration nearly every student has now: one course bans AI tools, another asks you to disclose them, a third expects you to use them, and a fourth doesn’t say. There’s no single rule, so find out the rule for this piece of work. Read the syllabus and the assignment; if they’re silent, ask the instructor before you start, not after you submit. “I assumed it was fine” is a hard position to argue in an academic integrity meeting.
When disclosure is expected, be specific about what the assistant did and how you checked it:
- “Used an AI assistant to draft the README structure and suggest unit test scaffolding; verified every command by running it locally; edited the output for accuracy.”
A vague “AI was used in this assignment” tells the reader nothing about what to trust. This book does the same: Appendix B describes how AI tools were used to write it, what they got wrong, and what the human authors checked.
Use AI to learn, not to skip the learning. When an assistant hands you an answer, ask for a smaller example, ask what it assumes and which edge cases it ignores, then predict what happens if an input changes and test the prediction. A good rule: if you can’t explain it, you don’t own it.
Keep people in the loop on a team. Pull requests should explain their intent and include tests, review comments should point to evidence (a test, a log, the docs), and AI-generated code is reviewed, tested, and justified like any other contribution, by the person who submits it. “The AI wrote it” doesn’t move the responsibility anywhere (Chapter 32 covers review).
35.11 Stakes and politics
In 2023, two New York lawyers filed a brief in a personal-injury suit against the airline Avianca that cited court decisions ChatGPT had invented, complete with made-up quotations. When the other side couldn’t find the cases, one of the lawyers asked ChatGPT whether they were real, and it assured him they were. The judge fined the lawyers $5,000 (Mata v. Avianca). No one fined the tool.
That is the arrangement this chapter’s rule quietly accepts. “AI can propose, you must verify” is the right habit for you, and it is also how an industry ships a product to hundreds of millions of people while leaving the cost of each mistake with the individual who trusted it.
The costs run upstream, too. The text these models learned from (books, news, code, forum answers) was largely collected without asking the people who wrote it, and the legal questions are still being fought over in court. The human judgments that make chat assistants usable, including the ratings behind reinforcement learning from human feedback (RLHF), come from people as well, often contract workers paid little. In 2023 TIME reported that workers in Kenya labeling descriptions of violence and sexual abuse, so that ChatGPT could learn to filter them, took home roughly $1.32 to $2 an hour. The value you get from an assistant is partly their labor.
See Chapter 8 for the broader framework. The concrete prompt to carry forward: when you accept an AI suggestion, ask whose labor produced it and who pays when it’s wrong.
35.12 Worked examples
Each of these starts with an assistant’s suggestion and ends with evidence from your own computer deciding what to do.
A notebook kernel mismatch
You can import pandas in the terminal, but in your Jupyter notebook it fails with ModuleNotFoundError: No module named 'pandas'. An assistant suggests the most common cause: the notebook’s kernel runs a different Python from your terminal. Don’t take its word for it; check inside the notebook:
import sys
print(sys.executable)/Users/you/anaconda3/bin/python
That’s Anaconda’s base Python, but in the terminal which python says /Users/you/project/.venv/bin/python. Now evidence, not the assistant’s confidence, confirms the hypothesis. Switch the notebook to a kernel that uses your project’s environment (registering the environment as a kernel if it isn’t listed), or install pandas into the kernel that’s running. The lasting fix is a line in the README, such as “Use the proj-venv kernel,” so the next person to clone the repository doesn’t repeat your evening (Chapter 16 has more on kernels).
A CSV that loads into one column
You load a file and something’s clearly off:
import pandas as pd
df = pd.read_csv("data.csv")
print(df.shape)
print(df.columns.tolist())(2, 1)
['name\tage\tcity']
One column, named after the whole header line. The assistant offers three suspects: tabs instead of commas, odd quoting, or an unusual encoding. Rather than trying fixes at random, look at the raw lines; repr shows invisible characters like tabs:
with open("data.csv", encoding="utf-8") as f:
for _ in range(3):
print(repr(f.readline()))'name\tage\tcity\n'
'Ana\t21\tBoulder\n'
'Ben\t23\tDenver\n'
Tabs: the cause is no longer a guess. Tell read_csv the separator, and add an assertion so the problem can’t come back silently:
df = pd.read_csv("data.csv", sep="\t")
assert {"name", "age", "city"}.issubset(df.columns), df.columns.tolist()The fix rests on what you saw in the file and a check that runs every time, not on hope.
A merge conflict in a notebook
After a git pull, git reports a merge conflict in analysis.ipynb. An assistant offers three options. You can resolve it by hand in a text editor, which is painful, because a notebook is a JSON file with its outputs (images included, encoded as text) stored inside. You can use a notebook-aware tool like nbdime, which shows the conflict cell by cell. Or you can keep one side and rerun the cells to regenerate the outputs. Which is right depends on your team and on whether the outputs are part of what you hand in. Afterwards, fix the cause: set up nbstripout as a pre-commit hook (see Chapter 33) to strip outputs before each commit, and most of these conflicts disappear, since two people editing different cells no longer both change the outputs. Chapter 31 covers the wider workflow.
A high-risk fix you should not run
Sometimes the right answer to an assistant is no. You get Permission denied writing to a folder, and the assistant suggests sudo chmod -R 777 ~/project, or running your script with sudo. Both are stop signs.
Ask instead why the program can’t write where you told it to. Almost always the path is wrong: the script is writing into a system folder like /usr/local/share/... instead of one you own. Change the path, not the permissions, and write under your home directory (a project-level outputs/ works well). If you really think you need system permissions, ask a TA or instructor before running anything. High-risk suggestions are exactly where the assistant knows least about your situation and being wrong costs most.
35.13 Exercises
- Take a recent error. Ask an assistant for three hypotheses and two checks for each. Run the checks and record which hypotheses you eliminated.
- Draft a help request (Goal, Expected, Actual, Steps to reproduce, Context, What I tried). Ask an assistant to make it clearer without changing any facts. Compare the versions, remove anything it invented, then post it.
- Ask an assistant for three published papers, with DOIs, on a topic from one of your courses. Look each up on Google Scholar or at
https://doi.org/<the DOI>. How many exist, and do they say what the assistant claimed? - Ask for a README skeleton for one of your projects. Follow every command in a fresh environment, then add a Verify section and one troubleshooting entry you’ve confirmed.
- Ask for a small function with unit tests. Add one edge-case test of your own, run everything, and revise until all the tests pass.
- Find the AI policy for each course you’re taking this term. Where one is silent, write down the question you’d ask the instructor.
- Write a personal “do not paste” list (credentials, tokens, private or protected data) and keep it where you work.
35.14 One-page checklist
- I treat AI output as a proposal and look for evidence before I trust it.
- I scale my checking to the cost of being wrong.
- I ask for alternatives and checkpoints, not a single answer.
- I confirm commands, functions, and citations in primary sources.
- I run small experiments and change one thing at a time.
- I add a test or an assertion after every fix.
- I never run a command I can’t explain.
- I don’t paste secrets or sensitive data, and I share the shape of data instead.
- I know my course’s AI policy, and I disclose AI help specifically when it’s required.
- Anthropic, Prompt engineering overview — a vendor’s guide to clear, structured prompts, with techniques you can try in any assistant.
- OpenAI, Prompt engineering — a parallel guide, useful for comparing which advice holds across tools.
- Google, Prompt design strategies — a beginner-friendly walk-through from the Gemini documentation.
- Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell, On the Dangers of Stochastic Parrots (FAccT, 2021) — the widely cited critical paper on large language models, and the background to this chapter’s “Stakes and politics.”
- Arvind Narayanan and Sayash Kapoor, AI Snake Oil — a calm, evidence-driven book on what AI can and can’t do; a good counterweight to both hype and panic.
- Mary L. Gray and Siddharth Suri, Ghost Work — the foundational book on the hidden human labor behind “automated” systems, including the data labeling AI depends on.
- Karen Hao, Empire of AI — long-form journalism on the labor and resource flows behind modern AI; pairs with the labeling-labor framing above.