38  Evaluating and Auditing AI Outputs

Prerequisites (read first if unfamiliar): Chapter 36, Chapter 22, Chapter 21.

See also: Chapter 37, Chapter 6.

Purpose

Ants Meme: Do you want hallucinations? Because that’s how you get hallucinations.

Here’s a situation that comes up in almost every class that lets students use AI. You have 2,000 open-ended comments from a course survey, and you ask a language model to sort each one into positive, negative, or mixed. You try five first, it gets all five right, so you run the rest and write that the model “classified the comments accurately.” Then your instructor asks: how accurately, compared to what, and for whom? You can’t answer any of the three.

If that has happened to you, you’re in good company. Five right answers feel like evidence, and that feeling is how most AI mistakes get into finished work. What you need is what you’d want from any measuring instrument: cases where you already know the right answer, a count of how often the system gets them, and an honest sense of how far off that count could be. That’s evaluation. Its partner, auditing, is looking on purpose for the failures you didn’t anticipate: the groups, topics, and odd inputs where the system does worse than its average suggests.

Chapter 35 teaches you to check one output before you use it. This chapter scales that habit up: building a small evaluation set, checking your own labels, scoring a model with a few lines of pandas and scikit-learn, automating checks, watching a system after it ships, and auditing it for who it fails. Every example uses made-up data and simulated model outputs, so you can run it all without an API key. It won’t teach you to train models; it gives you enough to make a claim about an AI system that you can defend.

Why read this chapter

  • You tried your prompt on five examples, it got all five right, and you’re not sure that means anything.
  • You used a model to label a few thousand survey responses, tweets, or abstracts, and now you need to report how accurate it was with a number you can defend.
  • You and a classmate labeled the same 50 items by hand, disagreed on a dozen, and don’t know whose labels count as “right.”
  • Your script parsed the model’s JSON fine yesterday, and today it dies with JSONDecodeError: Expecting value: line 1 column 1 (char 0).
  • You changed one sentence in a prompt and have no idea whether things got better or worse.
  • Someone suggested having one model grade another model’s answers, and you’d like to know when that’s trustworthy.
  • A tool advertises 90% accuracy, and you want to know who’s in the other 10%.

Running theme: measure before you trust

If you can’t say how you’d know the system is working, you don’t yet know that it is. Decide how you’ll measure before you rely on it, and keep measuring afterwards, because AI systems that look fine in testing regularly surprise people once real inputs arrive.

38.1 Why checking AI is harder than testing code

Testing ordinary code is reasonably easy: a function that adds two numbers should return 4 when you give it 2 and 2, and a one-line test will tell you forever whether it does. AI systems break that comfortable picture in several ways at once.

There’s often no single right answer. “Summarize this article” has many good summaries. Even labeling a comment gets fuzzy at the edges: is “Lectures were fine but the exams felt unfair” negative or mixed? Two careful people can disagree.

Outputs vary. A language model picks each word by sampling from probabilities (see Chapter 36), so the same input can give different outputs on different runs. A low temperature helps but doesn’t guarantee repeatability: Anthropic’s glossary warns that even at temperature 0, identical inputs may produce different outputs across API calls. So you evaluate on many inputs, not one lucky example.

Failures look like successes. Broken code usually throws an error. A model that’s wrong gives you fluent, confident text, which is why people call it a hallucination. A casual read won’t catch it; a comparison against a known answer will.

The target moves. Providers update models, you edit prompts, and the inputs change. A 2023 study by Chen, Zaharia, and Zou found that GPT-4’s accuracy at identifying prime numbers fell from 84% in its March 2023 version to 51% in June: same model name, same questions, very different answers.

People are expensive, and they disagree. Careful human review is the most trustworthy evaluation there is, but it doesn’t scale to thousands of outputs. Most real evaluation is a mix: people set the standard on a small set, and code carries it to the rest.

You’ll also see models ranked on public benchmarks such as MMLU (multiple-choice questions across 57 subjects) or Stanford’s HELM, which scores calibration, fairness, bias, toxicity, and efficiency as well as accuracy. They’re useful for comparing models in general, but they measure the benchmark’s task, not yours. Because the questions are public, some can end up in the text models are trained on, and once everyone competes on a number, people tune for the number: Goodhart’s law, “when a measure becomes a target, it ceases to be a good measure.” The only evaluation that tells you whether a model works for your task is one built from your task.

38.2 Decide what “good” means

“Is this output good?” is too vague to score, and it’s where many projects skip ahead too fast. Break “good” into dimensions, each asking a question you could actually answer:

Dimension The question it asks
Correctness Is the information accurate? Does the code run and pass tests?
Completeness Is everything required there? A summary that skips the key finding fails even if every sentence is true.
Consistency Do rephrasings of the same question get the same answer?
Relevance Does it answer the question that was asked?
Safety Does it avoid harmful, offensive, or inappropriate content?
Format Does it follow the structure you asked for? Format failures break the code downstream.
Tone and style Does it suit the audience, at the right length?

You don’t need all seven. A classifier mostly needs correctness and consistency; a public chatbot needs relevance, safety, and tone; an extraction pipeline lives or dies on format. Pick two or three and write them down before you look at outputs. Choose afterwards, and you’ll tend to pick the ones your system already does well on.

38.3 Match your effort to the stakes

Chapter 35 introduced a risk-based policy for single outputs. The same idea decides how much evaluation a whole system needs, and it saves you from both extremes: hand-checking every output of a harmless tool, or trusting a consequential one on vibes.

For low-stakes bulk work (summarizing your own notes, sorting records for a first look), automate cheap checks such as length and format, and read a small random sample. For medium-stakes work (labels that feed a paper, answers other people will read), build an evaluation set with known answers and rerun it on every prompt change or model update. For high-stakes work (anything medical, legal, financial, or affecting people’s access to jobs, housing, or services), plan for a person to review every output and make the final decision, however good the scores look.

For the formal version, see the U.S. National Institute of Standards and Technology’s AI Risk Management Framework (January 2023), which organizes this thinking into four functions: govern, map, measure, and manage.

38.4 Build a small evaluation set

An evaluation set is a collection of inputs where you already know what a good output looks like. Everything else in this chapter depends on it, and it’s less work than it sounds. The catch is that people ask one set to do two different jobs, and no set can do both.

To say how accurate the tool is, you need a random sample of real inputs, with names and personal details removed. Draw them at random, so every comment had the same chance of being picked; that’s what lets the model’s score on them stand in for its score on everything else. Toy examples you write yourself won’t do, since they’re cleaner and easier than the real thing. If inputs come in kinds (short and long, one language and another), you can draw at random within each kind, which is stratified sampling, so rare kinds aren’t left out; if you take extra of a rare kind, weight each kind back to its real share when you report an overall number.

To catch a change that breaks something, you need a hard-case suite. This one you pick by hand, and twenty items that probe where the system might stumble tell you more than fifty easy ones that all pass. Start with the edges, the inputs you suspect will be hard:

  • very short inputs, and very long ones near the context limit
  • unusual formatting: emoji, accented characters, code, tables
  • writing in another language, or in more than one
  • ambiguous inputs that could honestly be read two ways
  • inputs outside what the system is meant to handle

Then add every failure you find, so the same mistake can’t come back unnoticed. Rerun the suite after every change as a pass-or-fail check, but never report its score as the tool’s accuracy. You chose those items because they’re hard, and you chose how many of each kind to include, so the score reflects your choices, not your data: a suite that’s half ambiguous comments scores far below what the tool does on ordinary ones, and a suite of old failures you’ve since fixed scores far above it.

Hold part of the random sample back. While you tune a prompt or a rubric, you look at the items it gets wrong and reword until they come out right. Each fix fits the prompt a little more closely to those items, so its score on them overstates how it will do on new ones, the same way a practice exam you’ve already seen the answers to overstates what you know. So split the random sample before you start: a development part you look at and tune on as much as you like, and a test part you score once, at the end, for the number you report. Machine learning calls these training, validation, and test sets; the idea is the same with a prompt in place of a model. If you go back and tune after seeing the test score, the test part has become development data, and an honest number needs fresh items.

And keep every set like raw data (see Chapter 21): save it as a CSV, commit it, note where each example came from, and never quietly edit an answer to make a new prompt look better.

Here’s the set this chapter uses: sixteen invented course-survey comments. Imagine they were drawn at random from the 2,000 in the survey (a real project would hold some back, but sixteen is too few to split). Two people labeled each one independently (rater_a, rater_b), then talked through their disagreements and settled a final answer (gold). The model column is simulated for this chapter, not a real model’s output. The english column records whether the writer said English was their first language or an additional one, which you’ll need for auditing.

id,english,comment,rater_a,rater_b,gold,model
1,first,"Loved the weekly labs, best class this term.",positive,positive,positive,positive
2,first,"Lectures were fine but the exams felt unfair.",mixed,negative,mixed,negative
3,first,"Too much reading and never enough time to do it.",negative,negative,negative,negative
4,first,"Honestly the group project was the best part.",positive,positive,positive,positive
5,first,"Grading was slow and nobody answered emails.",negative,negative,negative,negative
6,first,"Pretty good overall, although the textbook cost a fortune.",mixed,mixed,mixed,mixed
7,first,"Labs were useful. Lectures, not so much.",mixed,mixed,mixed,mixed
8,first,"Nothing to complain about.",positive,mixed,positive,positive
9,additional,"The TA office hours saved my grade.",positive,positive,positive,positive
10,additional,"Professor is very kind but homework is too much for me.",mixed,mixed,mixed,negative
11,additional,"I not understand the slides many times.",negative,negative,negative,negative
12,additional,"Class is good, projects are hard but I learn a lot.",positive,mixed,positive,negative
13,additional,"The exam was too long, we had no time to finish.",negative,negative,negative,negative
14,additional,"Very helpful course, I recommend to my friends.",positive,positive,positive,mixed
15,additional,"Sometimes the lecture is fast, but the notes help me.",mixed,mixed,mixed,mixed
16,additional,"Teacher explain very clear, I like it.",positive,positive,positive,positive

Save it as eval_set.csv and load it with pandas (see Chapter 22):

import pandas as pd

df = pd.read_csv("eval_set.csv")
print(df.shape)
print(df["gold"].value_counts())
(16, 7)
gold
positive    7
mixed       5
negative    4
Name: count, dtype: int64

Sixteen rows is tiny on purpose: small enough to read in one sitting, and the next sections show exactly how much (and how little) it can tell you.

38.5 Check your human labels first

It’s tempting to skip straight to scoring the model. But your gold labels are the ruler you’ll measure it with, and if two careful people can’t agree on them, the ruler is bent. So measure the people first.

Write a rubric

A rubric replaces “is this good?” with specific questions that have specific answers. For open-ended outputs such as summaries, it might look like this:

Correctness (0-2):
  0 = Contains factual errors
  1 = Mostly correct but missing key details
  2 = Fully accurate and complete

Format (0-1):
  0 = Does not follow the required JSON structure
  1 = Valid JSON with all required fields

Tone (0-1):
  0 = Too technical for the intended audience
  1 = Appropriate for a non-specialist reader

For a labeling task, the rubric is a definition of each label with an example of the hard cases (“mixed means the comment praises one thing and criticizes another”). Rubrics make scoring faster and raters more consistent. Before using one at scale, pilot it on 10 to 20 examples with two raters.

Measure how often the raters agree

The obvious measure is the share of items both raters labeled the same way. The catch is that some agreement happens by luck: two people guessing among three labels would still match now and then. Cohen’s kappa corrects for that, comparing the agreement you saw with what chance alone would give, so that 1 means perfect agreement and 0 means no better than chance. It’s the standard first measure of inter-rater reliability, and scikit-learn computes it (pip install scikit-learn):

from sklearn.metrics import cohen_kappa_score

agree = (df["rater_a"] == df["rater_b"]).mean()
kappa = cohen_kappa_score(df["rater_a"], df["rater_b"])
print(f"Raw agreement: {agree:.2f}")
print(f"Cohen's kappa: {kappa:.2f}")
Raw agreement: 0.81
Cohen's kappa: 0.72

If that feels like a black box, compute it yourself once. Chance agreement comes from how often each rater used each label:

p_a = df["rater_a"].value_counts(normalize=True)
p_b = df["rater_b"].value_counts(normalize=True)
chance = (p_a * p_b).sum()
print(f"Chance agreement: {chance:.2f}")
print(f"Kappa by hand:    {(agree - chance) / (1 - chance):.2f}")
Chance agreement: 0.33
Kappa by hand:    0.72

Is 0.72 good? You’ll often see the Landis and Koch (1977) scale, which calls 0.61 to 0.80 “substantial,” but the Wikipedia article on Cohen’s kappa notes those cutoffs rest on opinion, not evidence. More useful is reading the disagreements. All three (comments 2, 8, and 12) involve “mixed,” so the rubric’s definition of mixed needs work. Every disagreement is a small bug report about your rubric: fix the wording, relabel, and measure again.

Cohen’s kappa handles exactly two raters who both labeled every item. With three or more coders, or some items one coder skipped, content analysts use Krippendorff’s alpha, which allows any number of coders, missing codes, and ordered scales; the krippendorff package computes it, and for this chapter’s two raters it comes out at 0.73, next to kappa’s 0.72.

One thing neither number can tell you is whether the labels are right. Kappa and alpha measure reliability (do the coders give the same answer?), not validity (does the label capture what you mean by it?). Two coders who share a misreading of the rubric, or a coder and a model that share a blind spot, can agree perfectly and both be wrong. Agreement shows your ruler is consistent; whether it measures the right thing is for you to check against the real comments.

Spot-check what you can’t label

When a system produces far more output than you could ever label, read a sample of it, drawn from every kind of input and not only the common ones. Look at each output beside its input so you can see why it went wrong. A spot-check won’t give you a precise number, but it’s how you find failure modes you didn’t know to put in the set.

38.6 Score the model, with honest uncertainty

With gold labels you trust, you can ask how the model did. The scikit-learn user guide on metrics covers dozens of measures; three get you a long way.

Accuracy, and where the errors went

Accuracy is the share of items the model got right:

df["correct"] = df["model"] == df["gold"]
print(f"Accuracy: {df['correct'].mean():.2f}")
Accuracy: 0.75

Seventy-five percent hides which mistakes the model makes. A confusion matrix shows them, with one row per true label, one column per label the model gave, and a count in every cell. pandas makes one with crosstab:

print(pd.crosstab(df["gold"], df["model"]))
model     mixed  negative  positive
gold
mixed         3         2         0
negative      0         4         0
positive      1         1         5

The diagonal holds the correct answers (3 + 4 + 5 = 12); everything else is an error. The pattern is plain once you look: the model never missed a negative comment, but it called three comments negative that weren’t. It’s too quick to see criticism.

Precision and recall

That pattern has names. For one label, precision asks “when the model said negative, how often was it right?” and recall asks “of the truly negative comments, how many did it find?” (Wikipedia’s precision and recall article has good pictures.) scikit-learn’s classification_report prints both for every label:

from sklearn.metrics import classification_report

print(classification_report(df["gold"], df["model"]))
              precision    recall  f1-score   support

       mixed       0.75      0.60      0.67         5
    negative       0.57      1.00      0.73         4
    positive       1.00      0.71      0.83         7

    accuracy                           0.75        16
   macro avg       0.77      0.77      0.74        16
weighted avg       0.81      0.75      0.75        16

Negative has perfect recall and poor precision: the confusion matrix’s story in two numbers. Which matters more depends on your purpose. If you’re flagging complaints for a person to follow up, a miss is worse than a false alarm. If you’re reporting what share of comments were negative, the false alarms inflate your number. F1 balances the two, and “support” is how many true examples of each label there were.

How sure can you be?

Here’s the question almost everyone skips: if the model got 75% of these sixteen right, what would it get on the other 1,984? (It’s a fair question only because the sixteen stand in for a random sample; a hard-case suite’s score says nothing about the rest.) A confidence interval puts a range around the number, and the easiest way to get one is the bootstrap: treat your evaluation set as the population, draw thousands of same-sized sets from it with replacement, score each, and see how far the scores spread. NumPy’s random generator does the drawing:

import numpy as np

correct = df["correct"].to_numpy()
rng = np.random.default_rng(42)
scores = [rng.choice(correct, size=len(correct)).mean() for _ in range(10_000)]

low, high = np.percentile(scores, [2.5, 97.5])
print(f"Accuracy {correct.mean():.2f}, 95% interval {low:.2f} to {high:.2f}")
Accuracy 0.75, 95% interval 0.50 to 0.94

Honest, and humbling: on this evidence the real accuracy could be anything from a coin flip to excellent. For a single proportion like this one there’s also a formula, the Wilson score interval, which behaves well even with small sets and needs no resampling. SciPy (installed along with scikit-learn) computes it with binomtest:

from scipy.stats import binomtest

ci = binomtest(k=12, n=16).proportion_ci(method="wilson")
print(f"Wilson interval: {ci.low:.2f} to {ci.high:.2f}")
Wilson interval: 0.51 to 0.90

It’s the same story, and it gives any score a range: “92% on 50 items” (46 right) is 0.81 to 0.97. The fix is more examples. The same bootstrap on larger sets, each 75% correct, gives:

Items in the set 95% interval around 75%
16 0.50 to 0.94
50 0.62 to 0.86
200 0.69 to 0.81
1,000 0.72 to 0.78

Two hundred labeled examples turns “somewhere between a coin flip and excellent” into “about 75%, give or take six points.” Report the interval (or at least the number of examples) beside every score.

Comparing two prompts on the same items

Sooner or later you’ll have prompt A and prompt B and want to know which is better. The tempting move is to set their two intervals side by side and see whether they overlap. That throws away the most useful fact you have: both prompts saw the same items. Some comments are easy for any prompt and some are hard for any prompt, and that difference between items is much of what makes each interval wide. Compare the prompts item by item and it cancels out.

Here’s a simulated prompt B that fixes three of A’s four mistakes and makes one new one. A paired bootstrap resamples items, not scores, so each draw keeps an item’s two results together:

# Prompt B's labels (simulated): it fixes comments 2, 10, and 14, and breaks comment 6
df["model_b"] = df["model"]
fixed = df["id"].isin([2, 10, 14])
df.loc[fixed, "model_b"] = df.loc[fixed, "gold"]
df.loc[df["id"] == 6, "model_b"] = "positive"

a = df["correct"].to_numpy()
b = (df["model_b"] == df["gold"]).to_numpy()
rng = np.random.default_rng(42)
diffs = []
for _ in range(10_000):
    rows = rng.integers(0, len(a), len(a))    # the same items for both prompts
    diffs.append(b[rows].mean() - a[rows].mean())
print(f"B minus A: {b.mean() - a.mean():+.3f}, 95% interval", np.percentile(diffs, [2.5, 97.5]))
B minus A: +0.125, 95% interval [-0.125  0.375]

For right-or-wrong outcomes there’s also a classic test, McNemar’s test. It looks only at the items where the prompts disagree, since the eleven both got right and the one both got wrong say nothing about which is better. If the prompts were equally good, each disagreement would be a coin flip, so the exact version is a binomial test on those items:

only_b = (b & ~a).sum()    # items B got right and A got wrong
only_a = (a & ~b).sum()
print(f"B alone right: {only_b}, A alone right: {only_a}")
print(f"McNemar exact p-value: {binomtest(only_b, only_b + only_a).pvalue:.3f}")
B alone right: 3, A alone right: 1
McNemar exact p-value: 0.625

Both say the same thing: 87.5% against 75% looks like progress, but three wins to one is well within luck. (statsmodels’ mcnemar with exact=True gives the same p-value.) On a larger random sample the paired comparison can find a real difference that two overlapping intervals would hide, which is why it’s the one to report. And picking the winner is tuning too, so compare prompts on the development part and save the test part for the winner’s final score.

38.7 Automated checks that run on every change

Once people have set the standard, you can hand parts of it to code. Automated checks run in seconds, so you can run them after every prompt edit or model update, which is exactly when regressions sneak in.

Regression checks

AI checks are rarely exact-match, because the wording changes from run to run. Instead, check properties the output must have: it parses, the required fields are there, and nothing forbidden appears. Here’s a check for a model that extracts a title and a date as JSON, using Python’s json module:

import json

FORBIDDEN = ["John Smith", "555-1234"]   # details that must never reach the output

def check_extraction(response: str) -> list[str]:
    """Return the problems with one model response; an empty list means it passed."""
    try:
        parsed = json.loads(response)
    except json.JSONDecodeError as err:
        return [f"not valid JSON ({err})"]
    problems = []
    for field in ["title", "date"]:
        if not parsed.get(field):
            problems.append(f"missing or empty {field!r}")
    for secret in FORBIDDEN:
        if secret in response:
            problems.append(f"leaked {secret!r}")
    return problems

good = '{"title": "Budget hearing", "date": "2026-03-04"}'
fenced = '```json\n{"title": "Budget hearing", "date": "2026-03-04"}\n```'
leaky = '{"title": "Call John Smith at 555-1234", "date": ""}'

for name, response in [("good", good), ("fenced", fenced), ("leaky", leaky)]:
    print(name, check_extraction(response))
good []
fenced ['not valid JSON (Expecting value: line 1 column 1 (char 0))']
leaky ["missing or empty 'date'", "leaked 'John Smith'", "leaked '555-1234'"]

The fenced case is one of the most common snags there is. The model returned good JSON wrapped in a Markdown code fence, and json.loads gives up on the first character. If your script has died with Expecting value: line 1 column 1 (char 0) on output that looked fine, this is almost certainly why: tighten the prompt, or strip the fence before parsing. Run checks like these over the whole evaluation set after every change; pytest can run them for you as tests.

Using a model as the judge

Tone, helpfulness, or whether a summary caught the main point are hard to check with code. A common workaround is LLM-as-judge: give a second model a rubric, the input, and the output, and ask for a score.

JUDGE_PROMPT = """
You are evaluating the quality of an AI assistant response.

Input to the assistant: {input}
Assistant response: {response}

Rate the response on correctness (0-2) and relevance (0-2).
Respond in JSON: {{"correctness": <score>, "relevance": <score>,
  "reasoning": "<brief explanation>"}}
"""

print(JUDGE_PROMPT.format(input="What is 2 + 2?", response="4"))

Notice the doubled braces around the JSON example. Python’s str.format treats every {...} as a slot to fill, so with single braces the call fails with KeyError: '"correctness"', which is baffling the first time. A doubled brace means a literal one, and the printed prompt shows single braces.

Judges are useful, with known weaknesses. A 2023 study by Zheng and colleagues found that GPT-4 as a judge agreed with human preferences over 80% of the time, about as often as humans agreed with each other, and also documented position bias (favoring the answer shown first), verbosity bias (favoring longer answers), and self-enhancement bias (favoring its own answers). A judge is also another model call: it costs money, sometimes fails, and can make up a score.

So validate the judge the way you validated your raters: have a person and the judge score the same sample, and compare. Here are ten simulated pairs of 0 to 2 scores:

scores = pd.DataFrame({
    "human": [2, 1, 2, 0, 2, 1, 1, 2, 0, 2],
    "judge": [2, 2, 2, 1, 2, 1, 2, 2, 0, 2],
})
print(scores.mean())
print(f"Exact agreement: {(scores['human'] == scores['judge']).mean():.2f}")
print(f"Kappa:           {cohen_kappa_score(scores['human'], scores['judge']):.2f}")
print(f"Weighted kappa:  {cohen_kappa_score(scores['human'], scores['judge'], weights='quadratic'):.2f}")
human    1.3
judge    1.6
dtype: float64
Exact agreement: 0.70
Kappa:           0.47
Weighted kappa:  0.74

Because the scores are ordered, weighted kappa gives partial credit when the judge is off by one rather than two, which is why it’s higher. But the most useful line is the first: in all three disagreements, the judge was more generous. A lenient judge makes every version of your prompt look better than it is. Catching that takes ten minutes of scoring by hand. Use a validated judge for fast first-pass scoring, and keep people in charge of anything high-stakes.

Comparing against a reference answer

When you have a correct reference for each input, you can score by comparison. Exact match works for short outputs such as labels, which is what the survey example does. For longer text, overlap scores count shared words and phrases: ROUGE for summaries, BLEU for translation. They’re rough, since a summary can share most of its words with the reference and still miss the point. Semantic similarity compares embeddings of the two texts (see Chapter 36), which copes better with paraphrase. No automatic score captures quality alone; use them to spot changes, then read what they flag.

38.8 Try to break it

Everything so far measures the system on inputs you expect. Red-teaming, named after the red teams that play the attacker in security exercises, is deliberately trying to make it fail. Write inputs designed to:

  • produce harmful or policy-violating output
  • confuse the model about what’s being asked
  • get it to ignore your instructions (prompt injection)
  • exploit a bias to get a wrong answer
  • make it reveal its system prompt or other confidential details

The OWASP Top 10 for LLM applications is a good checklist of what attackers try. Red-teaming works best when someone other than the builder does it, since you’re too close to imagine every way a stranger might misuse your system, so trade with a classmate. If your system is an agent that takes actions, Chapter 37 covers what to add. Every successful attack becomes a new row in your hard-case suite.

38.9 Keep measuring after you ship

Evaluation doesn’t end at launch. Inputs change, models change, and you’ll change the prompt; the only way to notice what that does is to keep records and keep checking.

Log everything. Without logs you can’t explain a failure after the fact, notice a slow decline, or build new test cases from real use. At a minimum, record:

  • the timestamp
  • the full input, including which system prompt version was used
  • the full output
  • the model version and settings such as temperature and maximum tokens
  • token counts, response time, and any error codes

Real inputs often contain personal information, so mask names and identifying details unless you have explicit permission to keep them, and store logs somewhere private.

Watch for drift. The inputs people send can shift away from what you evaluated on (concept drift), and the model underneath can change, as the GPT-4 study showed. Rerun your evaluation on a sample of real traffic on a schedule, and look for falling scores, new kinds of failure, and changes in output length, format compliance, or error rate. Alert when a score drops below a threshold, not only when something crashes: the dangerous failures are the quiet ones.

Rerun the evaluation after every change (prompt, model version, tool definition, retrieval settings) before you redeploy. Where your provider allows it, name a specific dated model version rather than an alias that moves to newer models. And keep a changelog of every change and every score, so that when something gets worse you can line it up with what changed and roll back.

38.10 Auditing: who does the system fail?

An accuracy score is an average, and averages hide things. Auditing asks where the errors land: whether some users, groups, or topics get systematically worse results. This is how algorithmic bias gets found in practice.

The classic example is Joy Buolamwini and Timnit Gebru’s 2018 study Gender Shades (Chapter 8 tells how it started). Three commercial systems that guessed gender from photos had overall accuracies between about 88% and 94%, which sounds respectable. Broken down by skin type and gender, error rates for darker-skinned women reached 34.7%, against at most 0.8% for lighter-skinned men. The overall number was accurate and said almost nothing about who the systems failed.

Break your scores down by group

You can do the same with your evaluation set in one line, using a pandas group-by:

print(df.groupby("english")["correct"].agg(["mean", "size"]))
             mean  size
english
additional  0.625     8
first       0.875     8

The model got 7 of 8 right for writers whose first language is English and 5 of 8 for everyone else, and looking back at the comments you can guess why: “Very helpful course, I recommend to my friends” is plainly positive, and the model misread it. Before announcing a finding, though, bootstrap the gap:

first = df.loc[df["english"] == "first", "correct"].to_numpy()
additional = df.loc[df["english"] == "additional", "correct"].to_numpy()

rng = np.random.default_rng(42)
gaps = [rng.choice(first, len(first)).mean() - rng.choice(additional, len(additional)).mean()
        for _ in range(10_000)]
print(np.percentile(gaps, [2.5, 97.5]))
[-0.125  0.625]

With eight comments per group, the interval runs from slightly below zero to a very large gap, so you can’t yet tell a real disparity from bad luck. That’s not a reason to drop it; it’s a lead. Collect more comments from writers whose first language isn’t English, and measure again. A disparity that shows up in a small audit and gets waved away is how systems like the ones in Gender Shades reach the public.

Group gaps come in several forms, and it helps to know what you’re looking for. Demographic disparities show up when a system does worse for names, dialects, or topics associated with particular groups. Topic sensitivity shows up as inconsistency: refusing harmless questions on some subjects, answering harmful ones on others. Style sensitivity penalizes nonstandard grammar or non-native English, which is what the survey example hints at. And stereotype reinforcement ties roles or traits to particular groups in ways that echo social biases.

Counterfactual tests

To test a disparity directly, write counterfactual pairs: inputs identical except for the one detail that signals a group. It’s the design of the audit studies social scientists use to measure discrimination in hiring and housing, where matched applicants differ in only one characteristic. The names below echo the best-known one, Bertrand and Mullainathan’s “Are Emily and Greg More Employable than Lakisha and Jamal?”, which sent otherwise identical résumés to employers.

A: "Summarize the qualifications of candidate Emily Walsh: five years of
    data analysis experience, SQL and Python, led a team of three..."
B: "Summarize the qualifications of candidate Lakisha Washington: five years of
    data analysis experience, SQL and Python, led a team of three..."

Change one thing at a time: if the swapped name signals both a different gender and a different race, you won’t know which one caused a difference. Run each pair several times (outputs vary), record the label or score, the length, and whether the framing is positive or negative, and compare across many pairs, not one. If outputs differ systematically when only the name changes, write it down and act: change the prompt, change the design, narrow what the system is used for, or don’t use it.

Finally, test the topics where a wrong answer does real damage (politics, medical and legal advice, people in crisis) against what the system should do: a course-help bot shouldn’t give legal advice. Document what you find either way; “we looked, and here’s what we found” makes the next audit cheaper.

38.11 Stakes and politics

In 2016, ProPublica and the company that sold COMPAS, a risk-assessment tool used in some US criminal courts, looked at the same scores for the same defendants and reached opposite conclusions. ProPublica’s analysis found that Black defendants who didn’t go on to reoffend were nearly twice as likely as white defendants to have been labeled higher risk (45 percent against 23 percent). The company answered that a given score meant the same chance of reoffending whatever a defendant’s race. Both were right. They were measuring different things, and deciding which one counts as “working well” is deciding whose errors the system is allowed to make.

It wasn’t a quirk of one dataset. Common definitions of fairness in classification, such as equal false-positive rates, equal false-negative rates, and equal calibration, can’t all hold at once when groups have different base rates, unless predictions are perfect. Every audit has to choose, and a good one makes the choice visible. Audits can also launder more than they expose: one that tests only the metrics the system was tuned for, or only the groups its builders thought of, can leave large disparities in place while “we ran a bias audit” reads as accountability. And who runs it matters. A vendor auditing its own product, an organization checking a tool it bought, and an independent group with reason to dig are different things. Most audits are the first kind, and the independent ecosystem (researchers, groups like AI Now and the Algorithmic Justice League, newsrooms like ProPublica and The Markup) is small next to the scale of deployment.

See Chapter 8 for the broader framework. The concrete prompt to carry forward: when you design or read an audit, ask whose definition of fairness it operationalizes, whose data it measures, and who could have run it differently.

38.12 Worked examples

Building an evaluation suite for a summarization system

You want to know whether your summarization prompt is any good, and whether a change makes it better or worse. Collect 20 real documents at random from the domain it will summarize; toy documents will mislead you. Split them into 10 for development and 10 you set aside. Write a rubric with three dimensions, say completeness, accuracy (nothing made up), and conciseness, each scored 0 to 2. Have two people score the 10 development summaries independently and compute kappa for each dimension; wherever they disagree, reword the rubric until they don’t. Those 10 settled examples are your calibration set. Build an LLM judge from the rubric and tune it on the same 10. Then have the people score the other 10, which neither the rubric nor the judge has seen, and compare the judge with them there, checking both kappa and whether the judge runs high; keep people on any dimension it can’t score. Those held-out 10 give your baseline, with an interval. After that, compare each prompt change with a paired comparison on the development documents, and keep a hard-case suite of the documents that have tripped the system up, rerun after every change.

Auditing a classification system for demographic disparities

You suspect a classifier treats groups differently and want evidence either way. Identify the group signals in typical inputs: names, locations, writing style, dialect. Write counterfactual pairs that differ in only one signal, such as the same résumé with a different name. Run each pair several times and record the label, any confidence score, and anything odd in the wording. Aggregate by group (accuracy, false-positive and false-negative rates), each with a bootstrap interval so you can tell a real gap from noise. Follow up on any gap that holds up with more examples, and document the findings whether or not they show a problem; a clean audit is a useful record too.

Setting up production monitoring for an AI application

You’ve shipped an AI feature and want to know if it’s quietly getting worse. Log every input and output with its metadata (model version, temperature, prompt version) and any feedback users give. Define a few metrics that match your dimensions: the share of outputs that parse as valid JSON, the average rubric score on a sample, the error rate. Set alert thresholds that fit the stakes, such as format compliance below 95% or errors above 1%. Run the evaluation on a stratified sample of real traffic every week and compare with your baseline. Keep a changelog of every model, prompt, and configuration change, and look into every alert before you ship the next change, not after.

38.13 Exercises

  1. Choose a simple AI task (classify sentiment, extract names, summarize a paragraph). Write a rubric with three dimensions scored 0 to 2, and three test cases that would each score differently on at least one dimension.
  2. Save this chapter’s eval_set.csv and run every code block in order. Then edit the model column so it makes the same number of mistakes, all on first-language writers. Which numbers change, and which don’t?
  3. With a classmate, label the same 30 short texts (reviews, headlines, comments) independently with three labels. Compute raw agreement and Cohen’s kappa, rewrite your label definitions to fix the disagreements, then label 30 new texts and measure again.
  4. Write three automated checks for a system that extracts company name, location, and salary range from job postings. Each should test a property of the output, not exact text.
  5. Design a 10-example evaluation set for a question-answering system about Python, with at least two easy cases, two edge cases, and one adversarial case. For each, write the input, the expected behavior, and the dimension it tests.
  6. Choose an AI tool you use regularly. Design one test for inconsistency across paraphrased inputs and one counterfactual test for demographic bias, precisely enough that a classmate could run them.
  7. Two weeks after you deploy a summarization tool, user complaints are up but your evaluation scores haven’t moved. List three possible explanations, and the log data you’d examine to tell them apart.

38.14 One-page checklist

  • Decide which quality dimensions matter before you look at outputs.
  • Write a rubric with specific, answerable questions.
  • Estimate accuracy from a random sample of real inputs, and hold part of it back from all tuning; keep edge cases and every failure you find in a separate hard-case suite. Version both.
  • Measure agreement between human raters (Cohen’s kappa, or Krippendorff’s alpha for more coders) before trusting gold labels, and remember that agreement isn’t correctness.
  • Report accuracy with a confusion matrix, and precision and recall for the labels that matter.
  • Put an interval (bootstrap or Wilson), or at least the number of examples, beside every score.
  • Compare two prompts or models on the same items with a paired comparison, not two separate scores.
  • Validate any automated scorer, including an LLM judge, against human scores, and check whether it runs high.
  • Rerun the evaluation after every prompt, model, or configuration change, before redeploying.
  • Log inputs, outputs, and metadata with personal details removed, and alert on falling scores, not only errors.
  • Break scores down by group, and test with counterfactual pairs that change one thing at a time.
  • Document audit findings whether or not they show a problem, and scale the rigor to the stakes.
Note📚 Further reading
  • Hugging Face, Evaluate library — a toolkit of standard metrics with a helpful guide to choosing one; for evaluating language models, its docs now point to the more actively maintained LightEval.
  • Anthropic, Define success criteria and build evaluations — practical guidance on writing success criteria and building evaluations, from exact-match checks to model-based grading.
  • OpenAI, Evals framework — an open-source framework for writing and running evaluations of language models, with many examples to copy from.
  • Solon Barocas, Moritz Hardt, and Arvind Narayanan, Fairness and Machine Learning: Limitations and Opportunities — the standard free textbook on the technical and political sides of fairness, including the incompatibility result in “Stakes and politics.”
  • Margaret Mitchell et al., Model Cards for Model Reporting (FAT* 2019) — a short documentation standard for a model’s intended use, training data, and evaluation results, including results broken down by group.
  • Inioluwa Deborah Raji and Joy Buolamwini, Actionable Auditing (AIES 2019) — the Gender Shades follow-up: within seven months, all three audited companies released new versions that reduced their disparities, the clearest example of an audit with teeth.
  • AI Now Institute, Publications, and Algorithmic Justice League, Research — ongoing independent research and advocacy on AI accountability; the third-party audit ecosystem named in “Stakes and politics.”