36 How Language Models Work
Prerequisites (read first if unfamiliar): Chapter 35.
See also: Chapter 37, Chapter 38.
Purpose

You ask a chatbot to fix a bug, and it writes forty clean lines of Python in five seconds. Then you ask how many r’s are in “strawberry,” and it gets it wrong. You paste in a long reading with careful instructions, and ten replies later it’s ignoring them. You run the same prompt twice and get two different answers. And now and then it cites a paper or a function that doesn’t exist, in the same confident tone it uses for everything else.
If that has left you unsure when to trust these tools, you’re asking the right question. None of those behaviors is random, and none means the tool is broken. They follow from how a large language model turns your text into numbers and predicts what comes next. Once you can picture those steps, you can tell a bad prompt from a limit of the model, and you stop being surprised by the same failure twice.
This chapter gives you that picture, with code you can run on a laptop, most of it without an API key: tokens, next-token prediction, context windows, temperature, embeddings, calling a model from code, tool use, and why models state false things fluently. It doesn’t teach the math of training. Everyday habits for using AI assistants are in Chapter 35, models that take actions are in Chapter 37, and testing whether an AI system works is in Chapter 38.
Why read this chapter
- You asked a chatbot how many r’s are in “strawberry,” it got it wrong, and you’d like to know how something that writes fluent essays can’t count letters.
- You ran the exact same prompt twice, maybe even with the temperature set to 0, and got two different answers.
- You pasted a long document into a chat, and a few replies later the model was ignoring the instructions you gave at the start.
- A model handed you a citation or a function name that sounded perfect and turned out not to exist.
- You’re paying for an API by the token and want to know what a token is, and why the same sentence in Hindi or Amharic costs more than in English.
- You keep hearing “embeddings,” “RAG,” and “vector database,” and you’d like to see what they mean in a few lines of code.
- Someone said “the model called a tool,” and you want to know who actually ran the code.
Running theme: the model sees text, not meaning
Almost everything a language model does, the impressive parts and the maddening ones, follows from one fact: it turns your text into tokens and predicts, one token at a time, what’s likely to come next.
36.1 Tokens: what the model actually reads
The first surprise is that the model never sees your letters. Your text is cut into chunks called tokens, each chunk is swapped for a number, and the numbers are what the model works with. You can watch this happen with tiktoken, OpenAI’s tokenizer library (pip install tiktoken):
import tiktoken
enc = tiktoken.get_encoding("o200k_base") # the tokenizer GPT-4o uses
for text in ["the quick brown fox", "unbelievable", "strawberry", " strawberry", "Boulder"]:
ids = enc.encode(text)
pieces = [enc.decode([i]) for i in ids]
print(f"{text!r:22} {pieces} {ids}")'the quick brown fox' ['the', ' quick', ' brown', ' fox'] [3086, 4853, 19705, 68347]
'unbelievable' ['un', 'bel', 'ievable'] [373, 9880, 45794]
'strawberry' ['st', 'raw', 'berry'] [302, 1618, 19772]
' strawberry' [' strawberry'] [101830]
'Boulder' ['B', 'oulder'] [33, 62664]
Common words get a token of their own, leading space included, while rarer words are built from pieces. The pieces come from byte-pair encoding, which starts from single characters and keeps merging the pairs that appear together most often in a pile of training text, until it has a vocabulary of a set size (about 200,000 entries here). Hugging Face’s course has a clear walk-through of how the merges are learned.
Now look at “strawberry” again. With a space in front, as it usually appears mid-sentence, it’s one token, number 101830. That’s why the letter-counting question is hard: the model isn’t looking at s-t-r-a-w-b-e-r-r-y, it’s looking at one number, and it has to have learned how many r’s that number contains. Python gets "strawberry".count("r") right because it works on characters. Newer models often get the strawberry question right too, but when a model stumbles on spelling, letter counts, or reversing a word, tokens are usually why.
Tokens are also the unit you pay in and hit limits in. For ordinary English a token averages about four characters, or three-quarters of a word, but tokenizers differ even within one company: Anthropic’s models overview says a million tokens holds about 555,000 words on its current tokenizer and 750,000 on the one before. To count exactly, use the provider’s own tool: tiktoken for OpenAI models, or Anthropic’s token-counting endpoint for Claude, whose tokenizer isn’t public. Code splits differently from prose (df.groupby('state')['income'].median() is nine tokens), and a technical term the tokenizer never learned whole reaches the model as fragments.
The biggest differences are between languages. Here is Article 1 of the Universal Declaration of Human Rights in nine languages, counted with GPT-2’s tokenizer from 2019 and GPT-4o’s from 2024:
| Language | GPT-2 tokens | GPT-4o tokens |
|---|---|---|
| English | 33 | 33 |
| Spanish | 58 | 38 |
| Swahili | 49 | 36 |
| Chinese (simplified) | 82 | 37 |
| Arabic | 120 | 44 |
| Hindi | 296 | 54 |
| Thai | 260 | 63 |
| Yoruba | 171 | 82 |
| Amharic | 309 | 206 |
GPT-2’s tokenizer was learned mostly from English web pages, so everything else came out in tiny pieces. The newer one is far fairer, but Amharic still takes six times as many tokens as English to say the same thing, which means a bigger bill, a slower reply, and less room in the context window.
36.2 Predicting the next token
Once your text is a list of numbers, the model does one thing with it: for every token in its vocabulary, it computes how likely that token is to come next. One token is picked and added to the end, and the whole thing runs again; a 300-word answer is a few hundred trips around that loop. You can see one trip with GPT-2, a small model OpenAI released in 2019 that runs fine on a laptop, using Hugging Face’s Transformers library (pip install transformers torch; the first run downloads about 550 MB):
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("gpt2")
model = AutoModelForCausalLM.from_pretrained("gpt2")
prompt = "My favorite thing to eat for breakfast is"
ids = tokenizer(prompt, return_tensors="pt").input_ids
with torch.no_grad():
scores = model(ids).logits[0, -1] # one score per token in the vocabulary
probs = torch.softmax(scores, dim=-1) # scores -> probabilities that sum to 1
print(len(probs), "possible next tokens")
top = torch.topk(probs, 5)
for p, i in zip(top.values, top.indices):
print(f"{tokenizer.decode(int(i))!r:12} {p.item():.3f}")50257 possible next tokens
' a' 0.095
' the' 0.060
' bacon' 0.025
' spinach' 0.020
' chicken' 0.018
That’s the whole output: a probability for each of 50,257 possible next tokens. There’s no lookup of facts and no plan for the rest of the sentence, just “given everything so far, what usually comes next?” The rest of this chapter is about what happens to that list.
Inside, each token ID becomes a vector, a long list of numbers (768 in GPT-2, 4,096 in a mid-sized model like Mistral 7B, which is where the meme’s number comes from). The vectors pass through a stack of layers built on the transformer design from the 2017 paper “Attention Is All You Need”. Its key step, attention, updates each token’s vector by looking back at earlier tokens and weighing the ones that matter, which is how “it” gets connected to the noun it stands for. Training sets the millions or billions of numbers in those layers by nudging the model, over enormous amounts of text, toward predicting the real next token. Jay Alammar’s illustrated guide in Further reading is the gentlest way into the details.
A chat assistant is the same kind of model, trained further on example conversations and human ratings of its answers (RLHF) so that the likeliest continuation of a question is a helpful answer. The chat app wraps your message in a template with roles and hidden instructions before the model sees it. It’s still predicting the next token; it has been shaped to predict an assistant’s.
36.3 Context windows: why it “forgets”
Here’s something that surprises almost everyone: the model has no memory between requests. Each time you send a message, the chat app sends the entire conversation so far, hidden instructions included, and the model reads it all from scratch. That bundle of tokens is the context window, and it’s everything the model knows about your task.
Every model limits how big that bundle can be. GPT-2’s limit was 1,024 tokens, about two pages. Anthropic’s context windows guide lists up to a million tokens for its current models. The limit counts everything: instructions, history, pasted documents, and the reply being written.
So why does a model “forget”? There are three different reasons. The window can fill up. An API request that’s too long comes back as an error (Anthropic’s says “prompt is too long”). Chat apps try to keep a long conversation going instead: depending on the app, they drop older turns, summarize them, or ask you to start over, and not all of them tell you. Long contexts get used unevenly. Even when everything fits, models use information in the middle of a long context worse than information at the start or end; a 2023 study called it getting “lost in the middle”, and Anthropic’s guide calls the general decline as contexts grow context rot. A new conversation starts empty, unless the app has a memory feature that pastes notes back in.
The fixes follow from the causes. Restate the constraints that matter when a conversation runs long, or start fresh with a short summary. Paste the relevant part of a 10,000-line log, not the whole thing. Put the most important instructions at the beginning or end. And when you call an API, remember that sending the history is your code’s job.
36.4 Sampling and temperature: why the same prompt gives different answers
Something still has to choose one token from that list, and that choice is called sampling. Always taking the top token tends to produce flat, repetitive text, so usually the program draws at random, weighted by the probabilities. That’s why asking twice gets you two answers.
Temperature controls how adventurous the draw is. Before the scores become probabilities (with the softmax function, the same step as in the GPT-2 code), they’re divided by the temperature. A low temperature stretches the gaps, so the favorite nearly always wins; a high one squashes them, so long shots get a chance. A few lines of NumPy show it:
import numpy as np
words = ["pancakes", "eggs", "toast", "cereal", "broccoli"]
scores = np.array([3.0, 2.5, 2.0, 1.0, -1.0]) # made-up scores for five candidates
def softmax(scores, temperature=1.0):
z = scores / temperature
e = np.exp(z - z.max()) # subtracting the max avoids overflow; same result
return e / e.sum()
for t in [0.2, 0.7, 1.0, 1.5]:
probs = softmax(scores, t)
print(f"T={t}: " + " ".join(f"{w} {p:.2f}" for w, p in zip(words, probs)))
rng = np.random.default_rng(seed=42)
for t in [0.2, 1.0, 1.5]:
print(f"8 draws at T={t}:", " ".join(rng.choice(words, size=8, p=softmax(scores, t))))T=0.2: pancakes 0.92 eggs 0.08 toast 0.01 cereal 0.00 broccoli 0.00
T=0.7: pancakes 0.56 eggs 0.27 toast 0.13 cereal 0.03 broccoli 0.00
T=1.0: pancakes 0.47 eggs 0.29 toast 0.17 cereal 0.06 broccoli 0.01
T=1.5: pancakes 0.39 eggs 0.28 toast 0.20 cereal 0.10 broccoli 0.03
8 draws at T=0.2: pancakes pancakes pancakes pancakes pancakes eggs pancakes pancakes
8 draws at T=1.0: pancakes pancakes pancakes toast eggs toast pancakes pancakes
8 draws at T=1.5: eggs pancakes toast eggs toast pancakes cereal cereal
At 0.2, pancakes wins 92% of the time; at 1.5, even cereal gets picked. Temperature 0 is the limit: always take the top token, called greedy decoding. So keep it low for consistency (extraction, classification, code) and raise it for variety (brainstorming, alternative drafts). Where you can set it, OpenAI’s API takes 0 to 2 and Anthropic’s takes 0 to 1.
Two confusions come up constantly. First, temperature 0 still varies. In theory it’s deterministic; in practice hosted models give slightly different answers to identical requests, as Anthropic’s API reference says outright. One reason is that computers round, so in floating-point arithmetic the order you add numbers in changes the answer:
print((0.1 + 0.2) + 0.3)
print(0.1 + (0.2 + 0.3))0.6000000000000001
0.6
A server batching your request with other people’s may add things up in a different order depending on how busy it is, as a 2025 write-up from Thinking Machines explains. When two tokens are nearly tied, a difference in the last decimal place flips which one is “top,” and the rest of the answer follows a different path. Providers also update the model behind a name. So treat temperature 0 as “much less random,” never as “reproducible.”
Second, the newest models often won’t let you set it. Anthropic’s API reference says its newer models reject any temperature but the default, and other providers restrict it on some models too. If setting a temperature gets you a 400 error, check the model’s documentation.
Three related settings travel with temperature. Top-p, or nucleus sampling, draws only from the smallest set of tokens whose probabilities add up to p; providers suggest adjusting it or temperature, not both. Max tokens caps the reply’s length, and set too low it cuts the answer off mid-sentence. Stop sequences are strings, such as END, that end the reply as soon as the model writes one.
36.5 Embeddings: text as points in space
Embeddings are a different job: turning a piece of text into a fixed-length vector, arranged so that texts with similar meanings land close together. This is the meme at the top of the chapter taken literally, each text a point in a space of hundreds or thousands of dimensions. The idea started with single word embeddings and now covers whole documents. “Close” is usually measured with cosine similarity, which is 1 when two vectors point the same way and near 0 when they’re unrelated:
import numpy as np
def cosine(a, b):
return a @ b / (np.linalg.norm(a) * np.linalg.norm(b))
cat = np.array([0.9, 0.8, 0.1]) # made-up 3-number "embeddings"
kitten = np.array([0.8, 0.9, 0.2])
tax = np.array([0.1, 0.2, 0.9])
print(f"cat vs kitten: {cosine(cat, kitten):.2f}")
print(f"cat vs tax: {cosine(cat, tax):.2f}")cat vs kitten: 0.99
cat vs tax: 0.30
Real embeddings have 384, 1,024, or 3,072 numbers that no one can interpret one by one, but the arithmetic is the same, and it powers a lot. Semantic search embeds a query and your documents and returns the closest, so “reduce memory usage” finds “optimize RAM consumption” with no words in common (the third worked example builds one). Retrieval-augmented generation (RAG) runs that search first and pastes the best passages into the model’s context, which is how “chat with your PDFs” tools handle collections bigger than any context window; the vectors usually live in a vector database. Embeddings also let you cluster documents by topic without labels and find near-duplicate records.
They come from embedding models, which are separate from chat models and usually much cheaper. OpenAI has an embeddings guide; Anthropic doesn’t make an embedding model, and its embeddings page points to Voyage AI; and free models from Sentence Transformers run on a laptop. One rule catches people out: compare vectors only from the same model. Each model has its own space, and mixing them gives meaningless scores with no error to warn you.
36.6 Writing prompts that work
A prompt feels like a question, but to the model it’s the start of a document to continue, so everything in it shapes the answer. Keep that in mind and most prompting advice stops seeming like magic.
When you call a model from code, the input comes in roles. The system prompt holds standing instructions: who the model should be, what format to use, what to assume. User turns carry your requests and data, and assistant turns are the model’s earlier replies, sent back as history (you can also write some yourself, as examples). Some APIs once let you start the assistant’s reply for it (“The answer is:”) to steer the format; newer Claude models reject that, and structured-output features do the job instead. In a chat app the company writes the system prompt and mostly hides it; through an API it’s yours.
The habits that help follow from “it continues what you give it.” Be specific about the task and the reader: “Summarize this in three bullet points for someone with no statistics background,” not “Summarize this.” Name the format, and for machine-readable output use the providers’ structured-output features (Anthropic, OpenAI), which make the reply match a JSON schema. Show two or three examples of input and the output you want (few-shot prompting); for classifying, extracting, and rewriting, examples beat descriptions. Separate instructions from data with clear markers, such as XML-style tags or a line of ---. Give it a role when that helps (“You are a careful code reviewer”), but one clear role beats a paragraph of competing ones.
Let it reason before it answers on multi-step problems. Asking for the working first and the answer last, called chain-of-thought prompting, helps because each token of reasoning becomes context for the next. Many current models have a built-in “thinking” or “reasoning” mode that does this on its own, so check the model’s documentation before adding the instruction. And change one thing at a time: keep a few test inputs, change one line, compare, and keep what helped. The prompt engineering guides in Further reading go further.
36.7 Chatbot or API?
A chat app like ChatGPT, Claude, or Gemini is a website or program you type into. An API, such as the Claude API or the OpenAI API, lets your own code send a prompt and get the reply back as data. The model underneath is the same kind, but you get very different things:
| Chat app | API | |
|---|---|---|
| Getting started | Type and go | An account, a key, some code |
| System prompt | Set by the company, mostly hidden | Yours to write |
| Memory | The app keeps the conversation | None: your code resends the history |
| 500 inputs | Paste 500 times | Write a loop |
| Extras (web search, files) | Usually included | Only what you turn on or build |
| Cost | Free tier or a monthly plan | Per token, itemized on every call |
Use a chat app for one-off work: an explanation, a brainstorm, a first draft. Use the API when the same task has to run over many inputs, inside a script, with settings you control. A minimal call from Python (it needs pip install anthropic and an API key in the ANTHROPIC_API_KEY environment variable; see Chapter 34 for keeping keys out of code):
import anthropic # reads your key from the ANTHROPIC_API_KEY environment variable
client = anthropic.Anthropic()
message = client.messages.create(
model="MODEL-NAME", # copy a current model name from the provider's models page
max_tokens=1024,
system="You are a patient statistics tutor. Answer in two sentences or fewer.",
messages=[
{"role": "user", "content": "What is the median of [3, 7, 2, 9, 5]?"},
],
)
print("".join(block.text for block in message.content if block.type == "text"))
print(message.usage.input_tokens, "tokens in,", message.usage.output_tokens, "tokens out")The reply comes back as a list of content blocks, not a string, and on models with a thinking mode the first block may be the (usually hidden) reasoning, so the code keeps only the text blocks instead of grabbing message.content[0]. The usage numbers are what you’re billed for; printing them while you develop saves surprises.
36.8 Tools and function calling: who runs the code?
On its own, a model can only produce text; it can’t check the weather, run code, or query a database. Function calling, also called tool use, is how it gets those abilities, and the common misunderstanding is about who does the work: the model never runs anything. It writes a request, and your code decides whether to carry it out.
First, you describe the tools you’re willing to run: a name, a plain-language description, and the arguments, written as a JSON Schema. Anthropic’s tool use guide calls the schema input_schema; OpenAI’s function calling guide calls it parameters.
{
"name": "get_weather",
"description": "Returns the current weather for a city. Use it when the user asks about current conditions.",
"input_schema": {
"type": "object",
"properties": {
"city": {
"type": "string",
"description": "City name, e.g. Denver"
}
},
"required": ["city"]
}
}Next, the model decides. If a tool would help, it replies with a structured request instead of an answer, in effect “call get_weather with city set to Boulder.” It writes those arguments the way it writes anything, by predicting likely tokens, so they can be wrong. Then your code runs the function, if you decide it should, and sends back the result. Finally, the model writes its answer from that result.
Because every action passes through your code, you control what the model can touch. That’s also where the risk lives: a tool that deletes files or sends email will do exactly what the model asks, even when the model has misread the situation. Expose only what the task needs, and have a person approve anything hard to undo (see Chapter 35 for matching checks to risk). Running this round trip in a loop until the job is done is what makes an agent, the subject of Chapter 37.
36.9 Why models state false things fluently
When a model confidently says something false, people call it a hallucination. The unsettling part is that it sounds exactly like the true answers, and once you remember what the model does, that stops being mysterious. Here’s GPT-2 asked about something that never happened, with sampling off so it always takes the likeliest token:
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("gpt2")
model = AutoModelForCausalLM.from_pretrained("gpt2")
ids = tokenizer("The first person to walk on Mars was", return_tensors="pt").input_ids
out = model.generate(ids, max_new_tokens=12, do_sample=False, # greedy: always the top token
pad_token_id=tokenizer.eos_token_id)
print(tokenizer.decode(out[0]))The first person to walk on Mars was a man named John Glenn.
The first person to
No one has walked on Mars, and John Glenn orbited Earth. But “a man named” plus a famous astronaut is a likely way for that sentence to go on, and likely is all the model is built to produce. Modern assistants would almost certainly get this one right, but they fail the same way on questions where the mistake is harder to spot.
The causes stack up. The model predicts plausible text, not true text, so a false sentence and a true one come out in the same confident tone. Rare facts are learned poorly: a study of this “long tail” found that models answer worse when the answer appears in few training documents, and in its place they produce something that fits the common pattern, like a plausible author for a paper. Training data stops at a date, the knowledge cutoff, so newer library versions and events are missing and the model fills in what used to be true. Guessing gets rewarded: a 2025 paper, “Why Language Models Hallucinate,” argues that training and evaluation favor a confident guess over “I don’t know,” the way an exam with no penalty for wrong answers rewards guessing. And tokenization adds errors in spelling, counting, and arithmetic, as the strawberry example showed.
The practical responses follow:
- Check facts, citations, and numbers against a primary source before you use them.
- Run code and test it; reading it isn’t enough.
- Give the model the source material (paste the documentation, or use retrieval) instead of relying on its memory.
- Ask what might be wrong; models often flag real uncertainty when asked directly, though not reliably.
- Keep the temperature low for factual tasks, where you can set it.
36.10 Stakes and politics
Article 1 of the Universal Declaration of Human Rights says that all human beings are born free and equal in dignity and rights. In the token table earlier in this chapter, that sentence costs 33 tokens in English and 206 in Amharic. Anyone paying by the token pays about six times as much to send it in Amharic, waits longer for the reply, and fits a sixth as much into the context window. Petrov and colleagues found gaps of up to 15 times between languages. Nobody set out to charge Amharic speakers more; the tokenizer’s vocabulary was learned from text dominated by English and a few other widely published languages. It can change, too: Hindi went from 182 tokens with GPT-4’s tokenizer to 54 with GPT-4o’s, which shows these were design choices all along.
The same thing happens at every layer. The training text over-represents formal, published, English-language writing by people who write a lot online, so the model’s default voice is that corpus’s average. Embeddings carry the corpus’s associations: in 2016, Bolukbasi and colleagues found that word embeddings trained on Google News completed “man is to computer programmer as woman is to ___” with “homemaker.” And only a few very well-funded companies can train frontier models, so a small group decides what those models read, how they’re tuned, and what they refuse; open-weights models like the GPT-2 you ran here spread the using, not the deciding. Chapter 38 covers how to test for gaps like these.
See Chapter 8 for the broader framework. The concrete prompt to carry forward: when a model “knows” something, or charges you more to say something, ask whose text it learned from.
36.11 Worked examples
Diagnosing why a prompt produces inconsistent output
A prompt “usually works” but sometimes returns garbage, and you want to know why before piling on more wording. Start by collecting evidence: save the prompt and a few bad outputs side by side, and count the prompt’s tokens; if you’re near the context window, that alone could be the cause. Next, take sampling out of the picture by setting the temperature to 0, if your model allows it, and running the prompt several times. If the outputs now agree, the inconsistency was sampling; if they still vary a lot, look elsewhere, such as a retrieval step that returns different documents each time. Then name the format (“Reply with JSON only, with the keys summary and tags”); format failures often vanish once the format is spelled out. If it still wanders, add one or two examples of exactly the output you want. After each step, write down which change fixed it; that note is the lesson worth keeping.
import anthropic
client = anthropic.Anthropic()
def call(prompt, temperature=None):
settings = {}
if temperature is not None:
# Recent SDKs have no temperature= argument, and some models don't accept one,
# so pass it through extra_body, to a model whose documentation says it does.
settings["extra_body"] = {"temperature": temperature}
message = client.messages.create(
model="MODEL-NAME", # a model that accepts temperature
max_tokens=512,
messages=[{"role": "user", "content": prompt}],
**settings,
)
return "".join(block.text for block in message.content if block.type == "text")
prompt = "Summarize this abstract in one sentence: ..."
for _ in range(3):
print(call(prompt, temperature=0)) # take sampling out of the picture
print(call(prompt + "\n\nReply with JSON only, with the keys summary and tags."))The comment in call is a lesson of its own: version 1 of Anthropic’s Python SDK has no temperature argument, and not every model accepts one. Settings that were standard a year ago can disappear, so when a call fails with an unexpected-argument error or a 400, check the current documentation before assuming your code is wrong.
Converting a chatbot workflow to an API call
For weeks you’ve pasted a document into a chat window, asked the same question, and copied the answer into a spreadsheet. Time to hand it to a script. Pull apart what you do by hand: the question you retype is standing instructions, and the pasted document is the part that changes. Move the standing part into a system prompt and send the document as the user turn. Make one call and compare it with an answer from the chat window, adjusting the system prompt until they agree, with max_tokens set so long answers aren’t cut off. Then make the input a parameter, a file path from the command line (see Chapter 17), and loop over a folder. Log every input, output, and token count to a CSV so you can audit the results and the bill. Keep the API key in an environment variable, never in the script.
Building a simple semantic search with embeddings
You want to search documents by meaning, and see it work before paying for anything. A small free model from Sentence Transformers runs on a laptop (pip install sentence-transformers; the model is a 90 MB download on first use):
import numpy as np
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("all-MiniLM-L6-v2") # small, free, runs on a laptop
docs = [
"How to reduce memory usage in pandas",
"Optimizing RAM consumption for large DataFrames",
"Baking sourdough bread at home",
"My notebook crashes when I load a big CSV file",
"Choosing colors for a bar chart",
]
doc_vecs = model.encode(docs) # one row of numbers per document
print(doc_vecs.shape)
def search(query, k=3):
q = model.encode(query)
sims = doc_vecs @ q / (np.linalg.norm(doc_vecs, axis=1) * np.linalg.norm(q))
best = np.argsort(sims)[::-1][:k] # indexes of the k highest scores
return [(float(sims[i]), docs[i]) for i in best]
for score, doc in search("my laptop runs out of memory reading a huge spreadsheet"):
print(f"{score:.2f} {doc}")(5, 384)
0.60 My notebook crashes when I load a big CSV file
0.47 How to reduce memory usage in pandas
0.46 Optimizing RAM consumption for large DataFrames
The top result shares none of the query’s key words (“laptop,” “memory,” “huge,” “spreadsheet”), so a keyword search would have missed it. To grow this into something real, embed your documents once and save the matrix (np.save), embed only the query at search time, use the same model for both, and test a few queries where you know the right answer against a plain keyword search to see where each wins.
36.12 Exercises
- Tokenize a paragraph you wrote with tiktoken, then the same paragraph in another language. How do the counts compare, and what would the difference mean for cost and context space?
- Tokenize “strawberry”, ” strawberry”, “Strawberry”, and a long word from your field. Then ask a chatbot to spell each one backwards. Does it do worse on words that split into more pieces?
- Run the softmax code with your own made-up scores. At what temperature does the second-best option win at least one draw in ten? Then give the GPT-2 code a prompt from your field and look at the top five next tokens.
- Write the same request (“Explain what a p-value is”) vaguely, then with a role, an audience, and a format. Compare the answers and describe what changed.
- Write a system prompt and user prompt that produce JSON with the fields
summary,key_terms, andconfidence. Test it on three documents. Did the format hold every time? Before you put thatconfidencenumber in a table, know what it is: text the model wrote, by predicting likely tokens like everything else, not a probability it measured. It isn’t calibrated, so “0.9” doesn’t mean it’s right nine times in ten, and a 2023 study found that models stating their confidence this way tend to be overconfident. If you want to check, run the prompt on ten or so items where you already know the right answer and compare the stated confidence with how often the model was actually right. - Get a model to state something false with confidence, for example about an event after its knowledge cutoff or an obscure paper in your field. Which cause from this chapter explains it, and how would you check the right answer?
- If you have API access, write a script that takes a question from the command line, sends it with your own system prompt, and prints the reply and token counts. Change the system prompt and confirm the answers change.
36.13 One-page checklist
- The model reads tokens, not letters: check its spelling, letter counts, and arithmetic.
- Count tokens with the provider’s own tool before sending anything long.
- Put key instructions at the start or end, and restate them in long conversations.
- Paste targeted excerpts, not whole files.
- Keep the temperature low for consistency where you can set it, and never count on identical outputs.
- Use a system prompt for standing instructions, name the format, and show examples.
- Separate instructions from data with clear markers.
- Compare embeddings only from the same model.
- Use the API, not a chat window, when you need to loop, log, or control settings; keep keys in environment variables.
- Expose only the tools a task needs, and have a person approve anything hard to undo.
- Verify facts, citations, and code from a model against primary sources.
- Anthropic, Prompt engineering overview — practical patterns for structuring prompts, with advice on testing them.
- OpenAI, Prompt engineering — OpenAI’s parallel guide, including how to use system messages and examples.
- Jay Alammar, The Illustrated Transformer — the classic picture-by-picture explanation of attention; about half an hour that demystifies most of what happens inside the model.
- Andrej Karpathy, Let’s build GPT: from scratch, in code, spelled out — a two-hour video that builds a small transformer line by line; the deepest practical understanding most non-specialists will get.
- Hugging Face, Transformers documentation — the library behind the GPT-2 examples in this chapter, with tutorials for running open models yourself.
- Anthropic, Mapping the Mind of a Large Language Model — a readable summary of interpretability research on what the numbers inside a model represent.
- Emily M. Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell, On the Dangers of Stochastic Parrots (FAccT, 2021) — the widely cited critique of ever-larger language models, covering training data, cost, and who bears the risks; pairs with the Stakes section above.