A trained classifier is a function from features to predicted classes, but until it has been measured against held-out data there is no honest answer to the only question that matters: how well will it work tomorrow? This lesson builds the measurement toolkit: the confusion matrix as the universal tabulation of every prediction, accuracy and its well-known paradox, the precision/recall/F1 family, ROC curves and AUC as threshold-free summaries, the threshold-versus-metric relationship that puts the engineer in control, and the class-imbalance toolbox (class weights, resampling, per-class and macro averaging) for problems where the default metrics lie.
Learning Objectives
- Define the four confusion-matrix cells (TP, FP, FN, TN), tabulate a classifier's predictions against the truth, and derive accuracy, error rate, precision, recall, specificity, and F1 from those counts with correct formulas.
- Explain the accuracy paradox with a worked numeric example where a 99%-accurate model is useless, and identify the conditions under which accuracy is the wrong metric.
- Describe how the classifier produces a score, how the threshold converts that score into a class label, and how every threshold produces a different (precision, recall) pair and a different (FPR, TPR) pair.
- Construct a ROC curve by sweeping the threshold from 1 to 0, plot TPR against FPR, compute AUC by the trapezoidal rule, and interpret AUC as the probability that a random positive scores above a random negative.
- Diagnose class imbalance, apply three mitigations (class weighting, undersampling, oversampling), and explain how each changes the confusion matrix and the metrics.
- Compute per-class, macro, and weighted averaging for multiclass classification, and choose the right averaging strategy for the problem.
1. The Confusion Matrix
The confusion matrix is a 2-by-2 (or K-by-K) table that tabulates, for every example in the evaluation set, the predicted class against the true class. For binary classification with the convention that class 1 is "positive" and class 0 is "negative", the four cells have fixed names:
| Predicted positive | Predicted negative | |
|---|---|---|
| Actually positive | True Positive (TP) | False Negative (FN) |
| Actually negative | False Positive (FP) | True Negative (TN) |
The cells are read row-by-row: a "false negative" is something the model said no to that should have been yes; a "false positive" is something the model said yes to that should have been no. There are no other possibilities — every prediction is one of the four cells.
1.1 A concrete worked example
Suppose a classifier is run on a test set of 200 emails and produces these counts:
- 120 spam emails correctly flagged as spam → TP = 120
- 30 spam emails missed (predicted "not spam") → FN = 30
- 10 legitimate emails wrongly flagged as spam → FP = 10
- 40 legitimate emails correctly let through → TN = 40
| Predicted spam | Predicted not spam | Total | |
|---|---|---|---|
| Actually spam | TP = 120 | FN = 30 | 150 |
| Actually not spam | FP = 10 | TN = 40 | 50 |
| Total | 130 | 70 | 200 |
Every metric in the rest of this lesson is a ratio computed from these four counts. The matrix is the single source of truth.
The row totals are the class-conditional base rates (150 spam, 50 not spam — a 3:1 imbalance); the column totals are the model's predictions (130 spam predictions, 70 not-spam predictions). Both views will reappear in the class-imbalance section.
1.2 Implementation
from sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay
import matplotlib.pyplot as plt
y_true = [1] * 150 + [0] * 50 #150 spam, 50 not-spam
y_pred = [1] * 120 + [0] * 30 + [1] * 10 + [0] * 40
#120 spam→spam (TP), 30 spam→not-spam (FN),
#10 not-spam→spam (FP), 40 not-spam→not-spam (TN).
cm = confusion_matrix(y_true, y_pred, labels=[1, 0])
#cm = [[120, 30], [10, 40]]
# row 0 (actually spam): TP=120, FN=30
# row 1 (actually not-spam): FP=10, TN=40
ConfusionMatrixDisplay(cm, display_labels=["spam", "not spam"]).plot()
plt.show()
confusion_matrix orders rows and columns by labels=[1, 0]; pass
labels=[0, 1] if you want the conventional negative/positive ordering.
The ConfusionMatrixDisplay plot draws the same table with a colourmap — it
is the same information, easier to read.
2. Accuracy and the Accuracy Paradox
The most natural summary of a confusion matrix is accuracy, the fraction of all predictions that are correct:
For the spam example:
A score of 0.80 reads "the classifier is right 80% of the time" — a number worth reporting. The complement, , is the error rate.
2.1 The accuracy paradox
The problem with accuracy is that it weights every example equally. If 99% of the test set belongs to the negative class, a trivial classifier that always predicts negative reaches 99% accuracy:
- examples, 9,900 negative and 100 positive.
- Trivial classifier: always predict negative.
- Counts: TP = 0, FN = 100, FP = 0, TN = 9,900.
- .
A 99% accuracy on a model that detects zero positives is not a victory; it is a measurement error. The 99% number is real but it is measuring the prior, not the model. This is the accuracy paradox: when the classes are heavily imbalanced, accuracy can be near-perfect on a model that detects none of the rare class.
The condition under which accuracy misleads is precisely when the marginal class probabilities are very different from each other. The accuracy of the trivial "always-predict-majority" baseline equals the majority-class prior (0.99 here), and any useful model must beat that baseline by enough margin to actually detect the minority class.
2.2 When accuracy is fine
Accuracy is the right metric when:
- The class distribution in the evaluation set reflects the distribution the model will see at deployment.
- The cost of every misclassification is the same (false positives and false negatives cost the same).
- The classes are roughly balanced.
If any of those fails, the lesson's remaining metrics are the answer.
3. Precision, Recall, Specificity, F1
The four cells of the confusion matrix admit six ratios; the four that earn their place are precision, recall, specificity, and the F1 score.
| Metric | Formula | Question it answers |
|---|---|---|
| Precision | Of everything I called positive, how many really were? | |
| Recall (TPR, sensitivity) | Of all the actual positives, how many did I catch? | |
| Specificity (TNR) | Of all the actual negatives, how many did I let through? | |
| F1 score | The harmonic mean of precision and recall |
For the spam example:
- .
- .
- .
- .
Notice that precision and recall are not interchangeable. Precision answers "when I alert, am I right?"; recall answers "did I catch all the cases?". A model can have very high precision by predicting positive only when it is certain (low FP, but high FN — recall collapses); a model can have very high recall by predicting positive almost always (high TP and high FP — but no positives missed, so precision collapses). The F1 score is the harmonic mean of the two, which penalises any extreme imbalance between them.
The harmonic mean is the right choice rather than the arithmetic mean: an arithmetic mean of precision 0 and recall 1 is 0.5, which looks respectable. The harmonic mean is 0.0, which correctly reports the model as worthless. The arithmetic mean rewards one strong axis at the expense of the other; the harmonic mean requires both to be good.
3.1 Why specificity matters
Specificity (the true negative rate) is the recall of the negative class: "of all the actual negatives, how many did the model correctly leave alone?". The complement of specificity is the false positive rate:
The ROC curve, Section 5, plots recall (TPR) against FPR; specificity is the positive-class version of the same question.
3.2 Implementation
from sklearn.metrics import (
accuracy_score, precision_score, recall_score,
f1_score, confusion_matrix,
)
P = precision_score(y_true, y_pred)
R = recall_score(y_true, y_pred)
F = f1_score(y_true, y_pred)
print(f"precision={P:.3f} recall={R:.3f} F1={F:.3f}")
#precision=0.923 recall=0.800 F1=0.857
4. The Threshold Is Yours
Every classifier that produces a probability or a score has a hidden knob: the
threshold that converts the score into a class label. scikit-learn's
LogisticRegression.predict_proba(X)[:, 1] returns a number in ; the
default .predict(X) calls "positive if score > 0.5", but the 0.5 is a choice,
not a property of the model.
The threshold-vs-metric relationship is the central insight of this lesson. The model produces a single score per example; every choice of threshold produces a different confusion matrix and therefore a different (precision, recall, FPR, TPR) tuple. The curve you plot is the set of all those tuples as the threshold varies.
4.1 The precision-recall tradeoff
Lowering the threshold makes the classifier say "positive" more often. More positives mean more TP (good for recall) and more FP (bad for precision). Raising the threshold does the opposite. The two metrics trade against each other by construction; you cannot push both to 1.
A simple illustration with the spam classifier scores:
| Threshold | Predicted positive | TP | FP | FN | Precision | Recall |
|---|---|---|---|---|---|---|
| 0.9 (strict) | 60 | 60 | 0 | 90 | 1.000 | 0.400 |
| 0.7 | 100 | 95 | 5 | 55 | 0.950 | 0.633 |
| 0.5 (default) | 130 | 120 | 10 | 30 | 0.923 | 0.800 |
| 0.3 (loose) | 170 | 140 | 30 | 10 | 0.824 | 0.933 |
| 0.1 (always-yes) | 200 | 150 | 50 | 0 | 0.750 | 1.000 |
At threshold 0.9, precision is perfect (every alert is right) but recall is only 0.4 (60% of spam missed). At threshold 0.1, recall is 1.0 (every spam caught) but precision has collapsed (half of all alerts are wrong). No threshold is universally best; the right one depends on the cost of an FP versus an FN in the application.
4.2 Picking the threshold
Three principled approaches:
- Optimise a single metric on the validation set. Pick the threshold that maximises F1, or a domain-specific weighted score, or a cost matrix .
- Use the ROC operating point. Choose a target FPR (say, "we can absorb 5% false alarms") and read the threshold that achieves it from the ROC curve.
- Use the precision constraint. Pick the highest recall subject to precision (or whatever the application needs).
A useful implementation pattern:
import numpy as np
from sklearn.metrics import precision_recall_curve
#`scores` are the classifier's positive-class probabilities.
#`y_true` are 0/1 labels.
precisions, recalls, thresholds = precision_recall_curve(y_true, scores)
#Pick the threshold that maximises F1.
f1 = 2 * precisions * recalls / (precisions + recalls + 1e-12)
best = np.argmax(f1[:-1]) # last entry has no threshold; drop it.
print(f"best threshold = {thresholds[best]:.3f}, F1 = {f1[best]:.3f}")
The precision-recall curve is the natural place to inspect this tradeoff; the ROC curve is the natural place to inspect the threshold-free summary. Both are tools for the same question, viewed from different angles.
5. ROC Curves and AUC
A receiver operating characteristic (ROC) curve plots the true positive rate (recall) against the false positive rate at every threshold:
As the threshold sweeps from 1 (always predict negative) to 0 (always predict positive), the curve traces a path from to . A useless classifier traces the diagonal — every threshold gives TPR = FPR, because the scores carry no information. A perfect classifier traces a right-angle up the left axis and across the top: TPR = 1 at any threshold where FPR is still 0.
5.1 Building the curve
The algorithm is simple: sort the scores in descending order, walk down the list one example at a time, and update the running TP and FP counts. Each time you cross a positive example, TPR steps up; each time you cross a negative example, FPR steps up. The result is a step function whose "averaged" form is the ROC curve.
5.2 AUC
The area under the ROC curve (AUC) collapses the whole curve into one number. It can be computed as the integral
or equivalently by the trapezoidal rule on the empirical curve.
The interpretation of AUC is the reason it deserves to be reported: AUC is the probability that a randomly chosen positive example scores above a randomly chosen negative example. The proof is the Wilcoxon-Mann-Whitney two-sample statistic:
For the spam classifier, suppose the AUC is 0.94. Read aloud: "if I pick one spam email at random and one legitimate email at random, the spam has a 94% chance of receiving a higher spam score from this classifier". That is the model, summarised.
AUC values:
- 0.5 = no discrimination (random).
- 0.7–0.8 = acceptable.
- 0.8–0.9 = excellent.
- 0.9+ = outstanding, but inspect the curve.
- 1.0 = perfect on the test set (suspect leakage).
5.3 Implementation
from sklearn.metrics import roc_curve, roc_auc_score
import matplotlib.pyplot as plt
#`scores` is the column of positive-class probabilities.
fpr, tpr, thresholds = roc_curve(y_true, scores)
auc = roc_auc_score(y_true, scores)
print(f"AUC = {auc:.3f}")
plt.plot(fpr, tpr, label=f"spam classifier (AUC = {auc:.2f})")
plt.plot([0, 1], [0, 1], "k--", label="random baseline")
plt.xlabel("False Positive Rate")
plt.ylabel("True Positive Rate")
plt.title("ROC curve")
plt.legend()
plt.show()
A practical warning. AUC is invariant to the threshold — it summarises the rank ordering of the scores, not the probabilities. A model with great AUC but uncalibrated probabilities (the scores are wrong as probabilities, only their order is right) will mislead any downstream user who treats the score as a probability. Calibrate (Platt scaling, isotonic regression) if probabilities matter; rely on AUC alone only if ranking is all you need.
5.4 Precision-recall curves vs ROC curves
ROC curves and precision-recall (PR) curves are cousins. Both sweep the threshold; they differ in the y-axis. The rule of thumb:
- Use the PR curve when the positive class is rare or the cost of FP is very different from the cost of FN. PR curves do not condition on the negatives, so a flood of true negatives does not flatter the curve.
- Use the ROC curve when the class distribution is roughly balanced and you want a threshold-free summary that does not depend on the prior.
For the spam problem (a 3:1 imbalance), the PR curve is the more honest picture; the ROC curve will look rosier than it should because the specificity axis has plenty of headroom.
6. Class Imbalance
The accuracy paradox is a symptom of a deeper problem: class imbalance. Most real-world classification problems are imbalanced — fraud, disease, defects, churn, conversion — and the default behaviour of training algorithms is to optimise accuracy, which means favour the majority class.
Three families of mitigation, each with a different effect on the confusion matrix.
6.1 Class weighting
The cleanest fix is to tell the loss function that misclassifying the rare
class is more expensive. scikit-learn's classifiers accept a
class_weight argument:
from sklearn.linear_model import LogisticRegression
#Inverse-frequency weights: positive class gets weight n_neg / n_pos.
clf = LogisticRegression(class_weight="balanced").fit(X_tr, y_tr)
class_weight="balanced" sets each class weight to
where is the number of training examples in class . For the spam
problem (, , ),
the weights would be and
.
The effect on the confusion matrix is to raise TP at the cost of FP: the model now dares to predict positive, catching more spam but also crying wolf on more legitimate mail. The decision boundary shifts, the ROC curve shifts, and the optimal threshold shifts. Class weighting is a modelling-time intervention — it changes what the model learns.
6.2 Resampling
The other fix is to change the training set itself.
| Method | What it does | Effect |
|---|---|---|
| Random undersampling | Drop examples from the majority class until balanced | Less data; cheaper training; risk of discarding useful signal |
| Random oversampling | Replicate examples from the minority class until balanced | No data lost; risk of overfitting on the duplicated minority examples |
| SMOTE | Synthesise new minority examples by interpolating between neighbours of the same class | Smoother decision boundary; risk of generating examples in regions the true distribution does not occupy |
| Tomek links / undersampling hybrids | Remove majority examples that are "ambiguous" — close to minority examples | Cleaner separation; combines undersampling with information about hard cases |
The effect of resampling on the test-set confusion matrix is identical to class weighting — both approaches change the operating point so that the classifier is more willing to predict the minority class. The difference is where the change happens: weighting changes the loss, resampling changes the data.
6.3 Threshold adjustment
The cheapest mitigation of all is to leave the training alone and just lower the threshold. With the spam problem at threshold 0.5, recall is 0.8; lowering the threshold to 0.3 raises recall to 0.93 at the cost of precision (Section 4). No retraining is needed; the model's scores are unchanged. This is also the most honest mitigation — it does not pretend to have trained on a balanced distribution.
6.4 Which fix wins?
There is no universally best fix. The literature and practice converge on a rough hierarchy:
- Pick a metric that does not collapse under imbalance. F1, PR-AUC, recall at fixed FPR, cost-weighted score.
- Lower the threshold. Often enough on its own.
- Apply class weighting or resampling if the threshold cannot move far enough (e.g. the model produces poorly calibrated scores).
Stacking all three mitigations is rarely better than picking the right one and tuning it.
7. Multiclass and Averaging Strategies
For a -class problem the confusion matrix becomes ; row , column counts examples that were truly class but predicted as class . The metrics generalise, but the question "what is the precision?" has two answers: per-class precision and macro-averaged precision.
7.1 Per-class metrics
For each class , treat as the positive class and aggregate the rest into the negative class, then compute precision, recall, and F1. This gives numbers per metric, one per class. scikit-learn reports them by default:
from sklearn.metrics import classification_report
print(classification_report(y_true, y_pred,
target_names=["setosa", "versicolor", "virginica"]))
The per-class report makes minority-class failures visible. A 95% overall accuracy on a 3-class problem can hide a recall of 0.10 on one class; the per-class report surfaces that immediately.
7.2 Averaging strategies
To collapse the per-class scores into a single number, three strategies are common:
| Strategy | Formula | When to use |
|---|---|---|
| Macro | When every class matters equally, regardless of size | |
| Weighted | When classes should be weighted by their frequency | |
| Micro | Compute the metric globally from the sum of all TP, FP, FN | When you want "the metric on the pooled confusion matrix" |
For precision, recall, and F1:
- Macro F1 treats class A and class B as equally important. A classifier that scores 0.95 on a 9,000-example class and 0.20 on a 100-example class has macro F1 = (0.95 + 0.20) / 2 = 0.575 — much worse than 0.95, which is the right behaviour if minority-class detection is important.
- Weighted F1 averages by class size, so the 0.95 dominates and the score reads 0.875. The minority class is still visible, but the metric reflects the deployment distribution.
- Micro F1 (for multiclass, micro F1 equals accuracy) pools the confusion matrix and computes the metric globally. It weights every example equally, so it suffers from the same accuracy-paradox problem on imbalanced data.
7.3 Picking the right average
The right choice depends on the question:
- Are all classes equally important? Macro. It treats every class the same regardless of size.
- Does the deployment distribution favour the majority class? Weighted. It honours the deployment prior while still penalising large per-class failures.
- Do you only care about overall correctness? Micro / accuracy. This collapses to accuracy in the multiclass case and is fine for balanced problems.
For multiclass AUC, scikit-learn uses one-vs-rest by default and averages the AUCs. The same macro/weighted/micro choice applies.
8. Worked Example: Spam Detection End-to-End
A single end-to-end script ties the lesson together: fit a logistic regression on imbalanced spam data, sweep the threshold, and report the metrics at every step.
import numpy as np
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import (
confusion_matrix, precision_score, recall_score, f1_score,
roc_curve, roc_auc_score, precision_recall_curve,
)
#Synthetic imbalanced dataset: 5% positive (rare-fraud-like).
X, y = make_classification(
n_samples=10_000, n_features=20, weights=[0.95, 0.05],
random_state=0,
)
X_tr, X_te, y_tr, y_te = train_test_split(X, y, stratify=y, random_state=0)
clf = LogisticRegression(max_iter=1000).fit(X_tr, y_tr)
scores = clf.predict_proba(X_te)[:, 1]
#Confusion matrix at the default threshold.
y_pred = (scores >= 0.5).astype(int)
cm = confusion_matrix(y_te, y_pred, labels=[1, 0])
print("default threshold (0.5):")
print(cm)
print(f"precision={precision_score(y_te, y_pred):.3f} "
f"recall={recall_score(y_te, y_pred):.3f} "
f"F1={f1_score(y_te, y_pred):.3f}")
#Threshold-free summary.
fpr, tpr, _ = roc_curve(y_te, scores)
auc = roc_auc_score(y_te, scores)
print(f"AUC = {auc:.3f}")
#Optimise the threshold for F1.
prec, rec, thr = precision_recall_curve(y_te, scores)
f1 = 2 * prec * rec / (prec + rec + 1e-12)
best = int(np.argmax(f1[:-1]))
print(f"best F1 threshold = {thr[best]:.3f}, F1 = {f1[best]:.3f}")
A typical run prints something like:
default threshold (0.5):
[[ 14 36] # TP=14, FN=36 → recall = 14/50 = 0.28
[ 12 938]] # FP=12, TN=938
precision=0.538 recall=0.280 F1=0.368
AUC = 0.918
best F1 threshold = 0.082, F1 = 0.673
Read the numbers together:
- The default threshold gives an apparently high accuracy of , but recall is 0.28 — the model only catches 28% of positives. This is the accuracy paradox in action.
- AUC = 0.918 says the ranking is strong: a random positive has a 91.8% chance of scoring above a random negative.
- Lowering the threshold to 0.082 raises F1 from 0.368 to 0.673, with a sharply different confusion matrix (more TP, more FP, fewer FN). The threshold change alone doubled F1.
The fix is not retraining with class weights (though that helps too); the fix is to pick the operating point the deployment actually needs.
9. Putting It All Together
A short, opinionated checklist for evaluating a binary classifier on a serious project:
- Look at the confusion matrix first. Every other metric is a ratio from it. If the matrix has surprises (large FN, large FP, diagonal not where you expect), the model is not what you thought.
- Report at least three numbers: accuracy, F1, and AUC. They answer different questions and the disagreement between them is informative.
- Report the threshold. A metric without a threshold is incomplete; the operating point is a decision the engineer makes.
- Inspect per-class metrics. A single overall accuracy can hide a minority-class collapse; the classification report makes it visible.
- Check calibration. If the scores are used as probabilities, plot a reliability curve and apply Platt or isotonic calibration if needed.
- Validate on held-out data, not training data. The point of every metric is to estimate generalisation. A 100% train accuracy is a warning sign, not a victory.
The order matters: matrix → threshold-free summary (AUC) → thresholded metrics (precision, recall, F1) → threshold choice. Skipping any layer leaves a hole in the picture.
Key Takeaways
- The confusion matrix is the universal tabulation of predictions against truth: every other metric is a ratio of its four cells (TP, FP, FN, TN). Look at it first.
- Accuracy is the right metric only when classes are balanced and the costs are symmetric; under imbalance, the accuracy paradox makes a useless "always-majority" classifier look near-perfect.
- Precision, recall, specificity, and F1 answer different questions — precision is "when I alert, am I right?", recall is "did I catch all the cases?", specificity is "did I leave the negatives alone?", and F1 is the harmonic mean that requires both to be good.
- The threshold is a choice, not a property of the model. Every threshold produces a different confusion matrix and a different (precision, recall) and (FPR, TPR) pair; the model produces scores, the metric is a decision about where to put the threshold.
- The ROC curve plots TPR against FPR at every threshold; AUC is the area under it and equals the probability that a random positive scores above a random negative (Wilcoxon-Mann-Whitney).
- Class imbalance is mitigated by class weighting (change the loss), resampling (change the data), or threshold adjustment (change the operating point). Choose by metric and cost, not by reflex.
- Multiclass metrics need an averaging strategy: macro treats every class equally, weighted honours the deployment distribution, micro collapses to accuracy. Use macro when minority-class detection matters.