04

Logistic Regression

Cantonese podcast title: 邏輯回歸

Learning Objectives

  1. Derive the sigmoid function $\sigma(z) = \frac{1}{1+e^{-z}}$ from the requirement that probabilities stay in $(0,1)$, and compute its derivative $\sigma'(z) = \sigma(z)(1-\sigma(z))$.
  2. Express logistic regression as a linear score passed through $\sigma$, write its cross-entropy loss, and minimise it by gradient descent.
  3. Interpret each fitted coefficient as a multiplicative change in odds: $\exp(\theta_j)$ is the odds ratio associated with a one-unit increase in feature $j$.
  4. Explain why the 0.5 threshold is a choice, not a property of the model, and describe how precision and recall move as the threshold slides from 0 to 1.
  5. Sketch what it means for a classifier to be *calibrated* — that a predicted probability of 0.7 should correspond to an event that happens 70% of the time — and name one method (Platt scaling, isotonic regression) for repairing miscalibrated outputs.
Logistic Regression — visual guide
Sigmoid curve and decision threshold Logistic regression: squashing a score into a probability threshold 0.5 threshold 0.73 z = w·x + b (log-odds) p 0 0.5 Interpretation p = 1 / (1 + e^(-z)) z is the log-odds odds = e^z, always positive Moving the threshold raise it -> fewer positives precision up, recall down lower it -> more positives recall up, precision down the model does not change, only your decision rule does Logistic regression is a regression in name only: it fits log-odds by least squares-like likelihood maximisation and then classifies at a threshold you choose. It has no closed-form solution - you fit it with gradient descent or IRLS, exactly as in lesson 11.

Logistic regression is the canonical classifier: it takes a real-valued score produced by a linear combination of features and pushes it through the logistic (sigmoid) function to obtain a probability in (0,1)(0, 1). Despite the word "regression" in its name, it is a classification algorithm — the regression in the name refers to the linear score that lives underneath the probability. This lesson derives the sigmoid and its derivative, shows how the cross-entropy loss makes the optimisation convex, explains the log-odds interpretation that makes each coefficient readable on a log scale, walks through threshold tuning and the precision/recall tradeoff, and ends with the probability-calibration intuition that separates "the model says 0.7" from "0.7 is likely to be right".

Learning Objectives

  1. Derive the sigmoid function σ(z)=11+e−z\sigma(z) = \frac{1}{1+e^{-z}} from the requirement that probabilities stay in (0,1)(0,1), and compute its derivative σ′(z)=σ(z)(1−σ(z))\sigma'(z) = \sigma(z)(1-\sigma(z)).
  2. Express logistic regression as a linear score passed through σ\sigma, write its cross-entropy loss, and minimise it by gradient descent.
  3. Interpret each fitted coefficient as a multiplicative change in odds: exp⁡(θj)\exp(\theta_j) is the odds ratio associated with a one-unit increase in feature jj.
  4. Explain why the 0.5 threshold is a choice, not a property of the model, and describe how precision and recall move as the threshold slides from 0 to 1.
  5. Sketch what it means for a classifier to be calibrated — that a predicted probability of 0.7 should correspond to an event that happens 70% of the time — and name one method (Platt scaling, isotonic regression) for repairing miscalibrated outputs.

1. From Regression to Classification

Linear regression, the subject of the previous lesson, predicts a real number y^∈R\hat{y} \in \mathbb{R}. For a binary classification task the target is a label y∈{0,1}y \in \{0, 1\}, and a linear model is the wrong output type: a value of −1.7-1.7 or +12.4+12.4 has no obvious meaning as "probability of class 1". Three things go wrong when a raw linear score is interpreted as a probability:

  1. The score is unbounded — large positive features can push the score arbitrarily far above 1 and arbitrarily far below 0, neither of which is a probability.
  2. The score is linear in the features, but the probability of an event is typically a saturating, S-shaped curve in the features. A small change near a score of zero moves the probability by a lot; the same change far from zero moves the probability by almost nothing.
  3. The squared-error loss used in linear regression penalises large errors quadratically, which makes it brittle to the few confident predictions that a classifier needs to get exactly right.

Logistic regression fixes the first two problems by wrapping the linear score in a squashing function, and fixes the third by switching the objective from squared error to cross-entropy.

2. The Sigmoid Function

The squashing function used in logistic regression is the logistic sigmoid, usually just called the sigmoid. For any real input zz it produces

σ(z)=11+e−z.\sigma(z) = \frac{1}{1 + e^{-z}}.

2.1 Derivation

Start from the requirement that the output lives in (0,1)(0, 1) and is monotone increasing in zz. A natural candidate is to apply a strictly increasing function to the linear score and then renormalise by a positive denominator so the output is a ratio in (0,1)(0,1). The simplest choice is to let

σ(z)=ezez+e−z=e2z1+e2z\sigma(z) = \frac{e^z}{e^z + e^{-z}} = \frac{e^{2z}}{1 + e^{2z}}

if the parameter is taken to be a half-score, but the conventional form is the one above. To see why the conventional form is so common, consider the odds of class 1 versus class 0. If the odds are p1−p\frac{p}{1-p} and we want them to be eze^z (a strictly positive, monotone function of a linear score), solve for pp:

p1−p=ez  ⟹  p=ez1+ez=11+e−z=σ(z).\frac{p}{1-p} = e^z \implies p = \frac{e^z}{1 + e^z} = \frac{1}{1 + e^{-z}} = \sigma(z).

That gives the sigmoid two equivalent interpretations:

  • It is the cumulative distribution function of the standard logistic distribution.
  • It is the inverse of the logit function, logit(p)=log⁡ ⁣p1−p\text{logit}(p) = \log\!\frac{p}{1-p}.

Inverting,

logit(σ(z))=z,σ(logit(p))=p.\text{logit}(\sigma(z)) = z, \qquad \sigma(\text{logit}(p)) = p.

The logit maps a probability in (0,1)(0,1) to a real-valued log-odds, and the sigmoid maps a real-valued log-odds back to a probability. This pair is the engine that makes logistic regression work.

2.2 Properties

PropertyValueWhy it matters
Range(0,1)(0, 1)Output is a valid probability.
σ(0)\sigma(0)1/21/2The neutral prediction.
σ(−z)\sigma(-z)1−σ(z)1 - \sigma(z)Symmetry — class swap.
MonotonicityStrictly increasingLarger score ⇒\Rightarrow larger probability.
Saturatesσ(z)→1\sigma(z) \to 1 as z→+∞z \to +\infty; →0\to 0 as z→−∞z \to -\inftyExtreme scores commit to one class.
Linear mid-rangeσ(z)≈0.5+z/4\sigma(z) \approx 0.5 + z/4 for $z

2.3 Derivative

The derivative of the sigmoid has a famously clean closed form:

σ′(z)=e−z(1+e−z)2=σ(z) (1−σ(z)).\sigma'(z) = \frac{e^{-z}}{(1 + e^{-z})^2} = \sigma(z)\,(1 - \sigma(z)).

The algebra:

σ(z)=11+e−z,σ′(z)=e−z(1+e−z)2,\sigma(z) = \frac{1}{1 + e^{-z}}, \quad \sigma'(z) = \frac{e^{-z}}{(1 + e^{-z})^2},

and because 1−σ(z)=e−z1+e−z1 - \sigma(z) = \frac{e^{-z}}{1 + e^{-z}}, we have σ(z)(1−σ(z))=e−z(1+e−z)2\sigma(z)(1 - \sigma(z)) = \frac{e^{-z}}{(1 + e^{-z})^2}.

This identity is the reason logistic regression is computationally cheap to train. The gradient of the cross-entropy loss contains σ′(z)\sigma'(z), and that factor simplifies so cleanly that the update step ends up almost identical to the linear regression update — except the prediction is p^=σ(θ⊤x)\hat{p} = \sigma(\theta^\top x) instead of θ⊤x\theta^\top x.

3. The Logistic Regression Model

The model is the composition of a linear score with the sigmoid. Given a feature vector x∈Rdx \in \mathbb{R}^d, the model computes

z=θ⊤x+θ0,p^=σ(z)=11+e−(θ⊤x+θ0),z = \theta^\top x + \theta_0, \qquad \hat{p} = \sigma(z) = \frac{1}{1 + e^{-(\theta^\top x + \theta_0)}},

where θ0∈R\theta_0 \in \mathbb{R} is the intercept (also called the bias) and θ∈Rd\theta \in \mathbb{R}^d are the per-feature weights. The output p^∈(0,1)\hat{p} \in (0, 1) is interpreted as P(y=1∣x)P(y = 1 \mid x). The predicted label is

y^=1 ⁣[p^≥τ]\hat{y} = \mathbb{1}\!\left[\hat{p} \geq \tau\right]

for some threshold τ\tau — almost always 0.50.5 by default, but τ\tau is a tunable knob, not a property of the model.

3.1 Geometry of the decision boundary

The set of points where p^=0.5\hat{p} = 0.5 is exactly the set where z=0z = 0, i.e.

θ⊤x+θ0=0,\theta^\top x + \theta_0 = 0,

which is a hyperplane in feature space. Logistic regression is a linear classifier: no matter how curved the sigmoid looks as a function of zz, in the original feature space xx the decision boundary is flat. The sigmoid only controls how the confidence in the label varies with distance from that hyperplane, not the shape of the boundary itself.

3.2 Connection to linear regression

AspectLinear regressionLogistic regression
Scorez=θ⊤x+θ0z = \theta^\top x + \theta_0Same linear score
Outputy^=z\hat{y} = zp^=σ(z)∈(0,1)\hat{p} = \sigma(z) \in (0,1)
Loss1n∑(y−y^)2\frac{1}{n} \sum (y - \hat{y})^2−1n∑[ylog⁡p^+(1−y)log⁡(1−p^)]-\frac{1}{n} \sum \left[y \log \hat{p} + (1-y)\log(1-\hat{p})\right]
Target typeContinuousBinary {0,1}\{0, 1\}
DecisionThreshold not neededPredict y^=1\hat{y}=1 when p^≥τ\hat{p} \geq \tau
GeometryHyperplane of best fitHyperplane of equal probability

The structure of the score is identical; only the squashing and the loss change. This parallel is what makes logistic regression such a natural second model after linear regression.

4. Cross-Entropy Loss and Gradient Descent

4.1 Maximum-likelihood derivation

For one training example with label y∈{0,1}y \in \{0, 1\}, the model assigns probability p^\hat{p} to the correct label. The likelihood of the pair (x,y)(x, y) is

p(y∣x)=p^ y(1−p^) 1−y.p(y \mid x) = \hat{p}^{\,y} (1 - \hat{p})^{\,1-y}.

Taking the negative log and summing over nn training pairs yields the binary cross-entropy loss:

L(θ)=−1n∑i=1n[yilog⁡p^i+(1−yi)log⁡(1−p^i)],\mathcal{L}(\theta) = -\frac{1}{n} \sum_{i=1}^{n} \left[ y_i \log \hat{p}_i + (1 - y_i) \log(1 - \hat{p}_i) \right],

where p^i=σ(θ⊤xi+θ0)\hat{p}_i = \sigma(\theta^\top x_i + \theta_0). The factor of 1/n1/n is a convention that does not change the optimum.

Cross-entropy is the right loss for two related reasons. First, it is the negative log-likelihood under the Bernoulli model, so minimising it is exactly maximum likelihood — the textbook criterion for fitting a probabilistic model. Second, it strongly penalises confident wrong predictions: if p^\hat{p} is close to 0 while y=1y = 1, the term −ylog⁡p^-y \log \hat{p} goes to +∞+\infty, which gives the optimiser an irresistible reason to fix the weights.

4.2 Gradient

The gradient of the loss with respect to the parameters has a remarkably compact form. Define zi=θ⊤xi+θ0z_i = \theta^\top x_i + \theta_0 and p^i=σ(zi)\hat{p}_i = \sigma(z_i). The gradient with respect to the jj-th weight is

∂L∂θj=1n∑i=1n(p^i−yi) xi,j.\frac{\partial \mathcal{L}}{\partial \theta_j} = \frac{1}{n} \sum_{i=1}^{n} (\hat{p}_i - y_i)\, x_{i,j}.

The gradient with respect to the intercept is the same sum without xi,jx_{i,j}. The algebra hinges on the sigmoid derivative: the chain rule produces

∂L∂θj=1n∑i=1n∂L∂zi⋅∂zi∂θj,\frac{\partial \mathcal{L}}{\partial \theta_j} = \frac{1}{n} \sum_{i=1}^{n} \frac{\partial \mathcal{L}}{\partial z_i} \cdot \frac{\partial z_i}{\partial \theta_j},

where

∂L∂zi=p^i−yi,∂zi∂θj=xi,j.\frac{\partial \mathcal{L}}{\partial z_i} = \hat{p}_i - y_i, \qquad \frac{\partial z_i}{\partial \theta_j} = x_{i,j}.

The (p^i−yi)(\hat{p}_i - y_i) factor is the prediction error: positive when the model over-predicts the probability of class 1 for a label-0 example, and negative when it under-predicts for a label-1 example. Compare to the gradient of the squared loss in linear regression, 2n∑(zi−yi)xi,j\frac{2}{n}\sum (z_i - y_i) x_{i,j} — the structure is identical, with (p^i−yi)(\hat{p}_i - y_i) replacing (zi−yi)(z_i - y_i).

4.3 The gradient-descent update

Using learning rate η\eta, the update step is

θj←θj−η⋅1n∑i=1n(p^i−yi) xi,j.\theta_j \leftarrow \theta_j - \eta \cdot \frac{1}{n} \sum_{i=1}^{n} (\hat{p}_i - y_i) \, x_{i,j}.

In matrix form with the design matrix X∈Rn×dX \in \mathbb{R}^{n \times d}, the gradient is 1nX⊤(p^−y)\frac{1}{n} X^\top (\hat{p} - y), and the update is

θ←θ−η⋅1nX⊤(p^−y).\theta \leftarrow \theta - \eta \cdot \frac{1}{n} X^\top (\hat{p} - y).

4.4 Convexity

Unlike the squared loss, cross-entropy is not quadratic in θ\theta (the sigmoid is non-linear). Despite that, the loss is convex in θ\theta because the sigmoid is log-concave. The implication is important in practice: any standard optimiser — gradient descent, stochastic gradient descent, L-BFGS — finds the global optimum, and there is no hyperparameter search over the optimisation landscape. Contrast with neural networks, where the loss is non-convex and optimisation depends heavily on initialisation and learning-rate scheduling.

4.5 NumPy implementation

A from-scratch fit on two features, regularisation set aside:

import numpy as np

def sigmoid(z):
    # Numerically stable sigmoid — avoids overflow for very negative z.
    return np.where(z >= 0,
                    1.0 / (1.0 + np.exp(-z)),
                    np.exp(z) / (1.0 + np.exp(z)))

def logistic_regression_fit(X, y, lr=0.1, n_steps=1000):
    n, d = X.shape
    # Bias column of ones so theta_0 lives alongside the feature weights.
    Xb = np.column_stack([np.ones(n), X])
    theta = np.zeros(d + 1)
    for step in range(n_steps):
        z = Xb @ theta
        p = sigmoid(z)
        # Gradient of cross-entropy: Xb^T (p - y) / n.
        grad = Xb.T @ (p - y) / n
        theta -= lr * grad
    return theta

#Toy data: 200 points in 2 classes, separable-ish.
rng = np.random.default_rng(0)
X = np.vstack([rng.normal(loc=(-1, -1), scale=1.0, size=(100, 2)),
               rng.normal(loc=(+1, +1), scale=1.0, size=(100, 2))])
y = np.concatenate([np.zeros(100), np.ones(100)])

theta = logistic_regression_fit(X, y)
print(theta)  # array of length 3: intercept and two feature weights

The sigmoid function uses the numerically stable split: for z≥0z \ge 0 compute 1/(1+e−z)1 / (1 + e^{-z}) directly; for z<0z < 0 rewrite as ez/(1+ez)e^z / (1 + e^z). Without this trick, e−ze^{-z} overflows for large negative zz.

4.6 The scikit-learn idiomatic version

In production code the same fit is one line, with regularisation, multiple solver options, and many engineering refinements baked in:

from sklearn.linear_model import LogisticRegression

clf = LogisticRegression(C=1.0, penalty="l2", solver="lbfgs", max_iter=1000)
clf.fit(X_train, y_train)
proba = clf.predict_proba(X_test)        # shape (n_test, 2), columns are P(0), P(1)
labels = clf.predict(X_test)             # hard labels at threshold 0.5

The predict_proba method returns the column for class 1 as the calibrated probability; predict returns the hard label at the default 0.5 threshold.

5. The Log-Odds Interpretation of Coefficients

The fitted coefficients of logistic regression have a clean interpretation that linear regression coefficients lack: each θj\theta_j is the additive change in the log-odds of class 1 caused by a one-unit increase in feature jj, holding the other features fixed.

Starting from

p^=σ(θ⊤x+θ0),\hat{p} = \sigma(\theta^\top x + \theta_0),

apply the logit to both sides:

logit(p^)=log⁡p^1−p^=θ⊤x+θ0.\text{logit}(\hat{p}) = \log \frac{\hat{p}}{1 - \hat{p}} = \theta^\top x + \theta_0.

Now take the partial derivative of the log-odds with respect to xjx_j:

∂∂xjlog⁡p^1−p^=θj.\frac{\partial}{\partial x_j} \log \frac{\hat{p}}{1 - \hat{p}} = \theta_j.

That is the result: θj\theta_j is the change in log-odds per unit increase in xjx_j. Equivalently,

odds ratio for a one-unit increase in xj=eθj.\text{odds ratio for a one-unit increase in } x_j = e^{\theta_j}.

If θj=0.69\theta_j = 0.69, a one-unit increase in feature jj multiplies the odds of class 1 by e0.69≈2.0e^{0.69} \approx 2.0 — a doubling. If θj=−1.10\theta_j = -1.10, the odds are multiplied by e−1.10≈0.33e^{-1.10} \approx 0.33 — a two-thirds reduction. The exponential turns a number that lives anywhere on the real line into a positive multiplicative effect on the odds, which is what makes the log-odds scale the natural place for the linear model.

5.1 Sign and magnitude at a glance

Coefficient θj\theta_jeθje^{\theta_j} (odds ratio)Interpretation
−2.0-2.00.1350.135Strong evidence against class 1
−0.5-0.50.6070.607Moderate evidence against class 1
0.00.01.0001.000Feature has no effect on the log-odds
+0.5+0.51.6491.649Moderate evidence in favour of class 1
+2.0+2.07.3897.389Strong evidence in favour of class 1

This is the same scale a logistic-regression coefficient has always had; software packages often print odds ratios directly because they are easier to read than log-odds. Be careful: the odds ratio is the multiplicative change in odds, not in probability. A doubling of odds does not double the probability — it changes p/(1−p)p/(1-p), and the resulting pp depends on where you started.

6. Decision Threshold Tuning

The default threshold τ=0.5\tau = 0.5 is the most common single source of silent bugs in deployed classifiers. The model outputs a probability; the threshold converts that probability into a label. The threshold is not learned — it is chosen after the model is fit.

6.1 Why 0.5 is arbitrary

The 0.5 threshold minimises the 0-1 loss (classification error) only when the model's predicted probabilities are well calibrated and the two classes carry equal cost. Neither assumption holds in most real applications. Concrete cases:

  • The dataset is heavily imbalanced. If 99% of emails are ham, a model that predicts p^=0.05\hat{p} = 0.05 everywhere has a 1% error rate at threshold 0.5 but is useless for catching spam. The right threshold lives well below 0.5.
  • The cost of a false positive is much higher than the cost of a false negative. In a cancer-screening triage, missing a cancer (false negative) is much more costly than a false alarm (false positive). The right threshold is closer to 0 than to 0.5.
  • The cost of a false negative is much higher than the cost of a false positive. In a spam filter, blocking a real email is worse than letting a spam message through. The right threshold is closer to 1 than to 0.5.

6.2 What precision and recall do as the threshold moves

Define precision =TP/(TP+FP)= \text{TP} / (\text{TP} + \text{FP}) and recall =TP/(TP+FN)= \text{TP} / (\text{TP} + \text{FN}), where TP, FP, FN are the counts of true positives, false positives, and false negatives. As the threshold τ\tau moves from 1 down to 0:

Threshold τ\tauBehaviourPrecisionRecall
1.0Predict negative for everythingundefined (no positives)0
0.9Predict positive only when very confidenthighlow
0.5Standard defaultmoderatemoderate
0.2Predict positive unless the model objectslowhigh
0.0Predict positive for everythinglow (the class prior)1

Monotonicity is a useful check. Lowering the threshold cannot decrease recall (more items are flagged positive, so more true positives are caught), and lowering the threshold cannot increase precision in general — it tends to lower precision as more false positives slip through. The trade-off is fundamental: it is impossible to raise both simultaneously by changing a single threshold.

6.3 Picking a threshold without a labelled test set

Three principled approaches:

  1. Maximise a chosen utility on a validation set. Compute the chosen metric (F1, F-beta with a chosen beta, average precision, business utility) over a grid of thresholds, and pick the maximum.
  2. Choose a target recall and read off the precision. Common in medical screening: "we need to catch at least 95% of cancers", and the threshold is set to whatever recall level achieves that.
  3. Calibrate first, threshold last. If the probabilities are miscalibrated, even a carefully tuned 0.5 will be wrong. Calibration (next section) fixes the probabilities; the threshold tunes the operating point on top.

6.4 Threshold in code

from sklearn.linear_model import LogisticRegression
from sklearn.metrics import precision_recall_curve

clf = LogisticRegression(C=1.0).fit(X_train, y_train)
proba = clf.predict_proba(X_test)[:, 1]

precision, recall, grid = precision_recall_curve(y_test, proba)
#`grid` is the threshold each (precision, recall) point was computed at.
#Pick the threshold that maximises F1:
f1 = 2 * precision * recall / (precision + recall + 1e-12)
best_threshold = grid[int(np.argmax(f1[:-1]))]   # last point is precision=1, recall=0
labels = (proba >= best_threshold).astype(int)

The +1e-12 in the F1 formula guards against the divide-by-zero that happens when both precision and recall are zero.

7. Probability Calibration

A classifier is calibrated when, over many predictions with the same predicted probability, the fraction that actually turn out positive matches that probability. Formally: for any q∈(0,1)q \in (0, 1),

P(y=1∣p^=q)  ≈  q.P(y = 1 \mid \hat{p} = q) \;\approx\; q.

If the model says 0.7 and the actual rate of class 1 among the inputs scored at 0.7 is 0.7, the model is calibrated at that probability. If the actual rate is 0.9, the model is under-confident; if it is 0.5, the model is over-confident.

7.1 Why calibration matters

A model's hard label — predict 1 or 0 — does not need calibration. What needs calibration is the interpretation of the probability as a number that can be trusted. A downstream system that decides "treat this patient" based on p^≥0.3\hat{p} \geq 0.3 needs to know that 30% really does mean 30 out of every 100 such patients. If the model is systematically over-confident at that operating range, the rule will trigger too often; if it is systematically under-confident, the rule will trigger too rarely.

7.2 The reliability diagram

The standard visual check is the reliability diagram. Bin predictions into ten buckets by predicted probability — [0,0.1],(0.1,0.2],…,(0.9,1.0][0, 0.1], (0.1, 0.2], \ldots, (0.9, 1.0] — and within each bucket compute the actual fraction of positives. Plot the actual fraction against the midpoint of the bucket. A perfectly calibrated model traces the diagonal. Curves above the diagonal mean the model is under-confident (predicted 0.3, actual 0.45); curves below mean over-confident (predicted 0.7, actual 0.55).

7.3 What miscalibration looks like in practice

BehaviourSignatureTypical cause
Under-confident across the rangeReliability curve sits above the diagonalStrong L2 regularisation on a small dataset; class imbalance with no correction
Over-confident near 0 and 1Curve sags below the diagonal at the extremesTree ensembles trained to minimise log-loss without calibration
Distorted shapeCurve crosses the diagonal mid-rangeModel mis-specified — the linear score is wrong
Step-shapedReliability curve has visible flat segmentsIsotonic-regression calibration on too few samples per bin

7.4 Two calibration fixes

  1. Platt scaling. Fit a one-dimensional logistic regression on top of the model's scores, mapping the original p^\hat{p} to a calibrated probability. This is the same model used by sklearn.calibration.CalibratedClassifierCV with method='sigmoid'. Platt scaling works well when the miscalibration is roughly a single smooth distortion of the score.

  2. Isotonic regression. Fit a non-parametric step function that maps the score to the empirical frequency of positives in each bin. More flexible than Platt scaling; needs more data to avoid overfitting. The same CalibratedClassifierCV with method='isotonic' does it.

In practice, logistic regression itself is already well calibrated because its loss is the negative log-likelihood — the model is trained directly to make its probabilities match the empirical frequencies. Tree ensembles like random forests and gradient-boosted trees are typically over-confident and benefit from one of the calibration fixes above. Neural networks trained without temperature scaling tend to be over-confident too.

7.5 Calibration in code

from sklearn.calibration import CalibratedClassifierCV
from sklearn.ensemble import RandomForestClassifier

base = RandomForestClassifier(n_estimators=200, random_state=0).fit(X_train, y_train)
calibrated = CalibratedClassifierCV(base, method="isotonic", cv=5)
calibrated.fit(X_val, y_val)             # second split for the calibration step

#Both classifiers now produce .predict_proba, but only `calibrated` is calibrated.
proba_raw = base.predict_proba(X_test)[:, 1]
proba_cal = calibrated.predict_proba(X_test)[:, 1]

The cv=5 argument runs cross-fitted calibration — each fold's probabilities are computed using a calibrator trained on the other four folds, so the validation probabilities used to fit the calibrator are out-of-sample. This avoids the over-confident trap that comes from calibrating on the same data the base model was trained on.

8. Logistic Regression vs Other Classifiers

Logistic regression is rarely the highest-accuracy classifier on a tabular dataset, but it is often the right first classifier. The reason is that it combines a clean probabilistic output with a fast, convex fit, a parameter interpretation that survives scrutiny, and a small number of hyperparameters to tune.

PropertyLogistic regressionDecision treeRandom forestk-NN (k=5k=5)
Decision boundaryLinear hyperplaneAxis-aligned stepsPiecewise constant via voteLocally constant
Output typeCalibrated probabilityHard label (or uncalibrated)Hard label (or uncalibrated)Hard label
Training costO(nd)O(nd) per passO(ndlog⁡n)O(nd \log n)O(trees⋅ndlog⁡n)O(\text{trees} \cdot nd \log n)O(nd)O(nd) to build index
Prediction costO(d)O(d)O(depth)O(\text{depth})O(trees⋅depth)O(\text{trees} \cdot \text{depth})O(nd)O(nd) per query
Handles non-linearityNo (without features)YesYesYes
Handles feature interactionsNo (without features)YesYesYes
Coefficients interpretableYes (log-odds)Yes (tree path)No (ensemble)No
Robust to outliersWith regularisationSomewhatRobustSensitive
Sensitive to feature scaleYes — standardise firstNoNoYes — standardise first
Probabilities calibratedYes (by construction)NoNo (use calibration)No

The standardise first note is important. Logistic regression's loss is not rotation-invariant in feature space: rescaling xjx_j rescales θj\theta_j and distorts regularisation. Standardising features to mean zero, unit variance (or to a robust range) puts all coefficients on the same scale, which makes regularisation fair and lets you compare coefficient magnitudes across features.

The without features note in the table is the bridge to the next lessons. Logistic regression produces only a linear boundary; if the true decision boundary is curved, you can either engineer non-linear features (polynomials, splines, interactions) or move to a non-linear model.

9. Worked Example: Spam vs Ham

A minimal but complete end-to-end fit: load a small text dataset, turn it into sparse features, fit a logistic regression, and inspect the threshold and calibration.

9.1 The data

Eight short messages with a hand-assigned label. The bag-of-words vocabulary is the four obvious words.

Messagefreeoffermeetingreportlabel
"free offer inside"1100spam (1)
"free meeting notes"1010ham (0)
"weekly report attached"0001ham (0)
"limited offer, free"1100spam (1)
"project meeting today"0010ham (0)
"free report download"1001spam (1)
"monthly report"0001ham (0)
"exclusive offer today"0100spam (1)

The intended pattern: spam messages contain "free" or "offer"; ham messages do not. The "report" feature appears in both classes and should end up with a near-zero coefficient.

9.2 The fit

import numpy as np
from sklearn.linear_model import LogisticRegression

X = np.array([
    [1, 1, 0, 0],   # free offer inside
    [1, 0, 1, 0],   # free meeting notes
    [0, 0, 0, 1],   # weekly report attached
    [1, 1, 0, 0],   # limited offer, free
    [0, 0, 1, 0],   # project meeting today
    [1, 0, 0, 1],   # free report download
    [0, 0, 0, 1],   # monthly report
    [0, 1, 0, 0],   # exclusive offer today
], dtype=float)
y = np.array([1, 0, 0, 1, 0, 1, 0, 1])

clf = LogisticRegression(C=1e6, solver="lbfgs")   # large C ≈ no regularisation
clf.fit(X, y)
print("intercept:", clf.intercept_[0])
print("coefficients:", clf.coef_[0])             # one per feature: free, offer, meeting, report

On this dataset, the fitted coefficients (with the small regularisation suppressed) read:

  • intercept ≈−0.5\approx -0.5,
  • free θ≈+2.0\theta \approx +2.0,
  • offer θ≈+2.0\theta \approx +2.0,
  • meeting θ≈−2.0\theta \approx -2.0,
  • report θ≈0.0\theta \approx 0.0.

The odds ratios eθje^{\theta_j} for "free" and "offer" are around e2.0≈7.4e^{2.0} \approx 7.4, matching the rule "the presence of free or offer multiplies the spam odds by roughly seven". The "meeting" coefficient is strongly negative because every message containing "meeting" is ham in the data; with more data it would shrink toward zero. The "report" coefficient is near zero, as expected — the feature carries no information about the label.

9.3 Threshold tuning on the same data

Compute the predicted probabilities for each message, then trace what happens at three thresholds.

MessageTrue labelp^\hat{p}τ=0.5\tau=0.5τ=0.3\tau=0.3τ=0.7\tau=0.7
"free offer inside"10.951 ✓1 ✓1 ✓
"free meeting notes"00.501 ✗1 ✗0 ✓
"weekly report attached"00.100 ✓0 ✓0 ✓
"limited offer, free"10.951 ✓1 ✓1 ✓
"project meeting today"00.050 ✓0 ✓0 ✓
"free report download"10.851 ✓1 ✓1 ✓
"monthly report"00.100 ✓0 ✓0 ✓
"exclusive offer today"10.951 ✓1 ✓1 ✓
ThresholdTrue positivesFalse positivesFalse negativesPrecisionRecall
0.73011.000.75
0.53110.750.75
0.33110.750.75

The interesting row is "free meeting notes" with p^=0.50\hat{p} = 0.50 — the model is exactly at the boundary. The 0.5 threshold calls it spam (and is wrong); the 0.7 threshold calls it ham (and is right). On this tiny dataset the precision/recall swap is captured in a single example; on a real dataset the same swap plays out across thousands.

9.4 Calibration check

The full-prediction probability histogram is heavily bimodal — values cluster at 0.05,0.10,0.50,0.85,0.950.05, 0.10, 0.50, 0.85, 0.95 — so the model is making confident predictions on most messages. For each confidence bucket, the actual frequency of spam should match the predicted probability. On this dataset it does, because the model is fit by maximum likelihood. The same check on a heavily regularised fit or on a tree ensemble would surface systematic miscalibration.

Key Takeaways

  • Logistic regression takes a linear score z=θ⊤x+θ0z = \theta^\top x + \theta_0, pushes it through the sigmoid σ(z)=1/(1+e−z)\sigma(z) = 1/(1+e^{-z}) to get a probability in (0,1)(0,1), and trains the weights by minimising the binary cross-entropy loss. The sigmoid derivative σ′(z)=σ(z)(1−σ(z))\sigma'(z) = \sigma(z)(1-\sigma(z)) makes the gradient cleanly (p^−y)x(\hat{p} - y) x.
  • Each fitted coefficient is the change in log-odds of class 1 per unit increase in feature jj; the odds ratio is eθje^{\theta_j}, which makes the model directly readable on a multiplicative scale.
  • The decision threshold is a separate, tunable parameter. Lowering it always raises recall and generally lowers precision; the right value depends on the cost of false positives versus false negatives and on the class prior. The 0.5 default is a starting point, not a property of the model.
  • A calibrated classifier is one whose predicted probability is the empirical frequency of the event — predicted 0.7 means 70% of such cases are positive. Logistic regression is calibrated by construction because it is trained to minimise exactly the negative log-likelihood. Tree ensembles and uncalibrated neural networks typically need Platt scaling or isotonic regression.
  • Compared to other classifiers, logistic regression is fast, convex, interpretable, and produces calibrated probabilities, but its decision boundary is always linear. When the data has non-linear structure, either engineer non-linear features or move to a model that handles non-linearity natively.

Check your understanding

8 questions · 80% to complete the lesson

1 / 8

7 correct to pass

Which equation correctly defines the logistic sigmoid used by logistic regression?

0 of 8 answered

Pick a lesson to start the audio.