Logistic regression is the canonical classifier: it takes a real-valued score produced by a linear combination of features and pushes it through the logistic (sigmoid) function to obtain a probability in . Despite the word "regression" in its name, it is a classification algorithm — the regression in the name refers to the linear score that lives underneath the probability. This lesson derives the sigmoid and its derivative, shows how the cross-entropy loss makes the optimisation convex, explains the log-odds interpretation that makes each coefficient readable on a log scale, walks through threshold tuning and the precision/recall tradeoff, and ends with the probability-calibration intuition that separates "the model says 0.7" from "0.7 is likely to be right".
Learning Objectives
- Derive the sigmoid function from the requirement that probabilities stay in , and compute its derivative .
- Express logistic regression as a linear score passed through , write its cross-entropy loss, and minimise it by gradient descent.
- Interpret each fitted coefficient as a multiplicative change in odds: is the odds ratio associated with a one-unit increase in feature .
- Explain why the 0.5 threshold is a choice, not a property of the model, and describe how precision and recall move as the threshold slides from 0 to 1.
- Sketch what it means for a classifier to be calibrated — that a predicted probability of 0.7 should correspond to an event that happens 70% of the time — and name one method (Platt scaling, isotonic regression) for repairing miscalibrated outputs.
1. From Regression to Classification
Linear regression, the subject of the previous lesson, predicts a real number . For a binary classification task the target is a label , and a linear model is the wrong output type: a value of or has no obvious meaning as "probability of class 1". Three things go wrong when a raw linear score is interpreted as a probability:
- The score is unbounded — large positive features can push the score arbitrarily far above 1 and arbitrarily far below 0, neither of which is a probability.
- The score is linear in the features, but the probability of an event is typically a saturating, S-shaped curve in the features. A small change near a score of zero moves the probability by a lot; the same change far from zero moves the probability by almost nothing.
- The squared-error loss used in linear regression penalises large errors quadratically, which makes it brittle to the few confident predictions that a classifier needs to get exactly right.
Logistic regression fixes the first two problems by wrapping the linear score in a squashing function, and fixes the third by switching the objective from squared error to cross-entropy.
2. The Sigmoid Function
The squashing function used in logistic regression is the logistic sigmoid, usually just called the sigmoid. For any real input it produces
2.1 Derivation
Start from the requirement that the output lives in and is monotone increasing in . A natural candidate is to apply a strictly increasing function to the linear score and then renormalise by a positive denominator so the output is a ratio in . The simplest choice is to let
if the parameter is taken to be a half-score, but the conventional form is the one above. To see why the conventional form is so common, consider the odds of class 1 versus class 0. If the odds are and we want them to be (a strictly positive, monotone function of a linear score), solve for :
That gives the sigmoid two equivalent interpretations:
- It is the cumulative distribution function of the standard logistic distribution.
- It is the inverse of the logit function, .
Inverting,
The logit maps a probability in to a real-valued log-odds, and the sigmoid maps a real-valued log-odds back to a probability. This pair is the engine that makes logistic regression work.
2.2 Properties
| Property | Value | Why it matters |
|---|---|---|
| Range | Output is a valid probability. | |
| The neutral prediction. | ||
| Symmetry — class swap. | ||
| Monotonicity | Strictly increasing | Larger score larger probability. |
| Saturates | as ; as | Extreme scores commit to one class. |
| Linear mid-range | for $ | z |
2.3 Derivative
The derivative of the sigmoid has a famously clean closed form:
The algebra:
and because , we have .
This identity is the reason logistic regression is computationally cheap to train. The gradient of the cross-entropy loss contains , and that factor simplifies so cleanly that the update step ends up almost identical to the linear regression update — except the prediction is instead of .
3. The Logistic Regression Model
The model is the composition of a linear score with the sigmoid. Given a feature vector , the model computes
where is the intercept (also called the bias) and are the per-feature weights. The output is interpreted as . The predicted label is
for some threshold — almost always by default, but is a tunable knob, not a property of the model.
3.1 Geometry of the decision boundary
The set of points where is exactly the set where , i.e.
which is a hyperplane in feature space. Logistic regression is a linear classifier: no matter how curved the sigmoid looks as a function of , in the original feature space the decision boundary is flat. The sigmoid only controls how the confidence in the label varies with distance from that hyperplane, not the shape of the boundary itself.
3.2 Connection to linear regression
| Aspect | Linear regression | Logistic regression |
|---|---|---|
| Score | Same linear score | |
| Output | ||
| Loss | ||
| Target type | Continuous | Binary |
| Decision | Threshold not needed | Predict when |
| Geometry | Hyperplane of best fit | Hyperplane of equal probability |
The structure of the score is identical; only the squashing and the loss change. This parallel is what makes logistic regression such a natural second model after linear regression.
4. Cross-Entropy Loss and Gradient Descent
4.1 Maximum-likelihood derivation
For one training example with label , the model assigns probability to the correct label. The likelihood of the pair is
Taking the negative log and summing over training pairs yields the binary cross-entropy loss:
where . The factor of is a convention that does not change the optimum.
Cross-entropy is the right loss for two related reasons. First, it is the negative log-likelihood under the Bernoulli model, so minimising it is exactly maximum likelihood — the textbook criterion for fitting a probabilistic model. Second, it strongly penalises confident wrong predictions: if is close to 0 while , the term goes to , which gives the optimiser an irresistible reason to fix the weights.
4.2 Gradient
The gradient of the loss with respect to the parameters has a remarkably compact form. Define and . The gradient with respect to the -th weight is
The gradient with respect to the intercept is the same sum without . The algebra hinges on the sigmoid derivative: the chain rule produces
where
The factor is the prediction error: positive when the model over-predicts the probability of class 1 for a label-0 example, and negative when it under-predicts for a label-1 example. Compare to the gradient of the squared loss in linear regression, — the structure is identical, with replacing .
4.3 The gradient-descent update
Using learning rate , the update step is
In matrix form with the design matrix , the gradient is , and the update is
4.4 Convexity
Unlike the squared loss, cross-entropy is not quadratic in (the sigmoid is non-linear). Despite that, the loss is convex in because the sigmoid is log-concave. The implication is important in practice: any standard optimiser — gradient descent, stochastic gradient descent, L-BFGS — finds the global optimum, and there is no hyperparameter search over the optimisation landscape. Contrast with neural networks, where the loss is non-convex and optimisation depends heavily on initialisation and learning-rate scheduling.
4.5 NumPy implementation
A from-scratch fit on two features, regularisation set aside:
import numpy as np
def sigmoid(z):
# Numerically stable sigmoid — avoids overflow for very negative z.
return np.where(z >= 0,
1.0 / (1.0 + np.exp(-z)),
np.exp(z) / (1.0 + np.exp(z)))
def logistic_regression_fit(X, y, lr=0.1, n_steps=1000):
n, d = X.shape
# Bias column of ones so theta_0 lives alongside the feature weights.
Xb = np.column_stack([np.ones(n), X])
theta = np.zeros(d + 1)
for step in range(n_steps):
z = Xb @ theta
p = sigmoid(z)
# Gradient of cross-entropy: Xb^T (p - y) / n.
grad = Xb.T @ (p - y) / n
theta -= lr * grad
return theta
#Toy data: 200 points in 2 classes, separable-ish.
rng = np.random.default_rng(0)
X = np.vstack([rng.normal(loc=(-1, -1), scale=1.0, size=(100, 2)),
rng.normal(loc=(+1, +1), scale=1.0, size=(100, 2))])
y = np.concatenate([np.zeros(100), np.ones(100)])
theta = logistic_regression_fit(X, y)
print(theta) # array of length 3: intercept and two feature weights
The sigmoid function uses the numerically stable split: for compute
directly; for rewrite as . Without
this trick, overflows for large negative .
4.6 The scikit-learn idiomatic version
In production code the same fit is one line, with regularisation, multiple solver options, and many engineering refinements baked in:
from sklearn.linear_model import LogisticRegression
clf = LogisticRegression(C=1.0, penalty="l2", solver="lbfgs", max_iter=1000)
clf.fit(X_train, y_train)
proba = clf.predict_proba(X_test) # shape (n_test, 2), columns are P(0), P(1)
labels = clf.predict(X_test) # hard labels at threshold 0.5
The predict_proba method returns the column for class 1 as the calibrated
probability; predict returns the hard label at the default 0.5 threshold.
5. The Log-Odds Interpretation of Coefficients
The fitted coefficients of logistic regression have a clean interpretation that linear regression coefficients lack: each is the additive change in the log-odds of class 1 caused by a one-unit increase in feature , holding the other features fixed.
Starting from
apply the logit to both sides:
Now take the partial derivative of the log-odds with respect to :
That is the result: is the change in log-odds per unit increase in . Equivalently,
If , a one-unit increase in feature multiplies the odds of class 1 by — a doubling. If , the odds are multiplied by — a two-thirds reduction. The exponential turns a number that lives anywhere on the real line into a positive multiplicative effect on the odds, which is what makes the log-odds scale the natural place for the linear model.
5.1 Sign and magnitude at a glance
| Coefficient | (odds ratio) | Interpretation |
|---|---|---|
| Strong evidence against class 1 | ||
| Moderate evidence against class 1 | ||
| Feature has no effect on the log-odds | ||
| Moderate evidence in favour of class 1 | ||
| Strong evidence in favour of class 1 |
This is the same scale a logistic-regression coefficient has always had; software packages often print odds ratios directly because they are easier to read than log-odds. Be careful: the odds ratio is the multiplicative change in odds, not in probability. A doubling of odds does not double the probability — it changes , and the resulting depends on where you started.
6. Decision Threshold Tuning
The default threshold is the most common single source of silent bugs in deployed classifiers. The model outputs a probability; the threshold converts that probability into a label. The threshold is not learned — it is chosen after the model is fit.
6.1 Why 0.5 is arbitrary
The 0.5 threshold minimises the 0-1 loss (classification error) only when the model's predicted probabilities are well calibrated and the two classes carry equal cost. Neither assumption holds in most real applications. Concrete cases:
- The dataset is heavily imbalanced. If 99% of emails are ham, a model that predicts everywhere has a 1% error rate at threshold 0.5 but is useless for catching spam. The right threshold lives well below 0.5.
- The cost of a false positive is much higher than the cost of a false negative. In a cancer-screening triage, missing a cancer (false negative) is much more costly than a false alarm (false positive). The right threshold is closer to 0 than to 0.5.
- The cost of a false negative is much higher than the cost of a false positive. In a spam filter, blocking a real email is worse than letting a spam message through. The right threshold is closer to 1 than to 0.5.
6.2 What precision and recall do as the threshold moves
Define precision and recall , where TP, FP, FN are the counts of true positives, false positives, and false negatives. As the threshold moves from 1 down to 0:
| Threshold | Behaviour | Precision | Recall |
|---|---|---|---|
| 1.0 | Predict negative for everything | undefined (no positives) | 0 |
| 0.9 | Predict positive only when very confident | high | low |
| 0.5 | Standard default | moderate | moderate |
| 0.2 | Predict positive unless the model objects | low | high |
| 0.0 | Predict positive for everything | low (the class prior) | 1 |
Monotonicity is a useful check. Lowering the threshold cannot decrease recall (more items are flagged positive, so more true positives are caught), and lowering the threshold cannot increase precision in general — it tends to lower precision as more false positives slip through. The trade-off is fundamental: it is impossible to raise both simultaneously by changing a single threshold.
6.3 Picking a threshold without a labelled test set
Three principled approaches:
- Maximise a chosen utility on a validation set. Compute the chosen metric (F1, F-beta with a chosen beta, average precision, business utility) over a grid of thresholds, and pick the maximum.
- Choose a target recall and read off the precision. Common in medical screening: "we need to catch at least 95% of cancers", and the threshold is set to whatever recall level achieves that.
- Calibrate first, threshold last. If the probabilities are miscalibrated, even a carefully tuned 0.5 will be wrong. Calibration (next section) fixes the probabilities; the threshold tunes the operating point on top.
6.4 Threshold in code
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import precision_recall_curve
clf = LogisticRegression(C=1.0).fit(X_train, y_train)
proba = clf.predict_proba(X_test)[:, 1]
precision, recall, grid = precision_recall_curve(y_test, proba)
#`grid` is the threshold each (precision, recall) point was computed at.
#Pick the threshold that maximises F1:
f1 = 2 * precision * recall / (precision + recall + 1e-12)
best_threshold = grid[int(np.argmax(f1[:-1]))] # last point is precision=1, recall=0
labels = (proba >= best_threshold).astype(int)
The +1e-12 in the F1 formula guards against the divide-by-zero that happens when both precision and recall are zero.
7. Probability Calibration
A classifier is calibrated when, over many predictions with the same predicted probability, the fraction that actually turn out positive matches that probability. Formally: for any ,
If the model says 0.7 and the actual rate of class 1 among the inputs scored at 0.7 is 0.7, the model is calibrated at that probability. If the actual rate is 0.9, the model is under-confident; if it is 0.5, the model is over-confident.
7.1 Why calibration matters
A model's hard label — predict 1 or 0 — does not need calibration. What needs calibration is the interpretation of the probability as a number that can be trusted. A downstream system that decides "treat this patient" based on needs to know that 30% really does mean 30 out of every 100 such patients. If the model is systematically over-confident at that operating range, the rule will trigger too often; if it is systematically under-confident, the rule will trigger too rarely.
7.2 The reliability diagram
The standard visual check is the reliability diagram. Bin predictions into ten buckets by predicted probability — — and within each bucket compute the actual fraction of positives. Plot the actual fraction against the midpoint of the bucket. A perfectly calibrated model traces the diagonal. Curves above the diagonal mean the model is under-confident (predicted 0.3, actual 0.45); curves below mean over-confident (predicted 0.7, actual 0.55).
7.3 What miscalibration looks like in practice
| Behaviour | Signature | Typical cause |
|---|---|---|
| Under-confident across the range | Reliability curve sits above the diagonal | Strong L2 regularisation on a small dataset; class imbalance with no correction |
| Over-confident near 0 and 1 | Curve sags below the diagonal at the extremes | Tree ensembles trained to minimise log-loss without calibration |
| Distorted shape | Curve crosses the diagonal mid-range | Model mis-specified — the linear score is wrong |
| Step-shaped | Reliability curve has visible flat segments | Isotonic-regression calibration on too few samples per bin |
7.4 Two calibration fixes
-
Platt scaling. Fit a one-dimensional logistic regression on top of the model's scores, mapping the original to a calibrated probability. This is the same model used by
sklearn.calibration.CalibratedClassifierCVwithmethod='sigmoid'. Platt scaling works well when the miscalibration is roughly a single smooth distortion of the score. -
Isotonic regression. Fit a non-parametric step function that maps the score to the empirical frequency of positives in each bin. More flexible than Platt scaling; needs more data to avoid overfitting. The same
CalibratedClassifierCVwithmethod='isotonic'does it.
In practice, logistic regression itself is already well calibrated because its loss is the negative log-likelihood — the model is trained directly to make its probabilities match the empirical frequencies. Tree ensembles like random forests and gradient-boosted trees are typically over-confident and benefit from one of the calibration fixes above. Neural networks trained without temperature scaling tend to be over-confident too.
7.5 Calibration in code
from sklearn.calibration import CalibratedClassifierCV
from sklearn.ensemble import RandomForestClassifier
base = RandomForestClassifier(n_estimators=200, random_state=0).fit(X_train, y_train)
calibrated = CalibratedClassifierCV(base, method="isotonic", cv=5)
calibrated.fit(X_val, y_val) # second split for the calibration step
#Both classifiers now produce .predict_proba, but only `calibrated` is calibrated.
proba_raw = base.predict_proba(X_test)[:, 1]
proba_cal = calibrated.predict_proba(X_test)[:, 1]
The cv=5 argument runs cross-fitted calibration — each fold's probabilities are
computed using a calibrator trained on the other four folds, so the validation
probabilities used to fit the calibrator are out-of-sample. This avoids the
over-confident trap that comes from calibrating on the same data the base model
was trained on.
8. Logistic Regression vs Other Classifiers
Logistic regression is rarely the highest-accuracy classifier on a tabular dataset, but it is often the right first classifier. The reason is that it combines a clean probabilistic output with a fast, convex fit, a parameter interpretation that survives scrutiny, and a small number of hyperparameters to tune.
| Property | Logistic regression | Decision tree | Random forest | k-NN () |
|---|---|---|---|---|
| Decision boundary | Linear hyperplane | Axis-aligned steps | Piecewise constant via vote | Locally constant |
| Output type | Calibrated probability | Hard label (or uncalibrated) | Hard label (or uncalibrated) | Hard label |
| Training cost | per pass | to build index | ||
| Prediction cost | per query | |||
| Handles non-linearity | No (without features) | Yes | Yes | Yes |
| Handles feature interactions | No (without features) | Yes | Yes | Yes |
| Coefficients interpretable | Yes (log-odds) | Yes (tree path) | No (ensemble) | No |
| Robust to outliers | With regularisation | Somewhat | Robust | Sensitive |
| Sensitive to feature scale | Yes — standardise first | No | No | Yes — standardise first |
| Probabilities calibrated | Yes (by construction) | No | No (use calibration) | No |
The standardise first note is important. Logistic regression's loss is not rotation-invariant in feature space: rescaling rescales and distorts regularisation. Standardising features to mean zero, unit variance (or to a robust range) puts all coefficients on the same scale, which makes regularisation fair and lets you compare coefficient magnitudes across features.
The without features note in the table is the bridge to the next lessons. Logistic regression produces only a linear boundary; if the true decision boundary is curved, you can either engineer non-linear features (polynomials, splines, interactions) or move to a non-linear model.
9. Worked Example: Spam vs Ham
A minimal but complete end-to-end fit: load a small text dataset, turn it into sparse features, fit a logistic regression, and inspect the threshold and calibration.
9.1 The data
Eight short messages with a hand-assigned label. The bag-of-words vocabulary is the four obvious words.
| Message | free | offer | meeting | report | label |
|---|---|---|---|---|---|
| "free offer inside" | 1 | 1 | 0 | 0 | spam (1) |
| "free meeting notes" | 1 | 0 | 1 | 0 | ham (0) |
| "weekly report attached" | 0 | 0 | 0 | 1 | ham (0) |
| "limited offer, free" | 1 | 1 | 0 | 0 | spam (1) |
| "project meeting today" | 0 | 0 | 1 | 0 | ham (0) |
| "free report download" | 1 | 0 | 0 | 1 | spam (1) |
| "monthly report" | 0 | 0 | 0 | 1 | ham (0) |
| "exclusive offer today" | 0 | 1 | 0 | 0 | spam (1) |
The intended pattern: spam messages contain "free" or "offer"; ham messages do not. The "report" feature appears in both classes and should end up with a near-zero coefficient.
9.2 The fit
import numpy as np
from sklearn.linear_model import LogisticRegression
X = np.array([
[1, 1, 0, 0], # free offer inside
[1, 0, 1, 0], # free meeting notes
[0, 0, 0, 1], # weekly report attached
[1, 1, 0, 0], # limited offer, free
[0, 0, 1, 0], # project meeting today
[1, 0, 0, 1], # free report download
[0, 0, 0, 1], # monthly report
[0, 1, 0, 0], # exclusive offer today
], dtype=float)
y = np.array([1, 0, 0, 1, 0, 1, 0, 1])
clf = LogisticRegression(C=1e6, solver="lbfgs") # large C ≈ no regularisation
clf.fit(X, y)
print("intercept:", clf.intercept_[0])
print("coefficients:", clf.coef_[0]) # one per feature: free, offer, meeting, report
On this dataset, the fitted coefficients (with the small regularisation suppressed) read:
intercept,free,offer,meeting,report.
The odds ratios for "free" and "offer" are around , matching the rule "the presence of free or offer multiplies the spam odds by roughly seven". The "meeting" coefficient is strongly negative because every message containing "meeting" is ham in the data; with more data it would shrink toward zero. The "report" coefficient is near zero, as expected — the feature carries no information about the label.
9.3 Threshold tuning on the same data
Compute the predicted probabilities for each message, then trace what happens at three thresholds.
| Message | True label | ||||
|---|---|---|---|---|---|
| "free offer inside" | 1 | 0.95 | 1 ✓ | 1 ✓ | 1 ✓ |
| "free meeting notes" | 0 | 0.50 | 1 ✗ | 1 ✗ | 0 ✓ |
| "weekly report attached" | 0 | 0.10 | 0 ✓ | 0 ✓ | 0 ✓ |
| "limited offer, free" | 1 | 0.95 | 1 ✓ | 1 ✓ | 1 ✓ |
| "project meeting today" | 0 | 0.05 | 0 ✓ | 0 ✓ | 0 ✓ |
| "free report download" | 1 | 0.85 | 1 ✓ | 1 ✓ | 1 ✓ |
| "monthly report" | 0 | 0.10 | 0 ✓ | 0 ✓ | 0 ✓ |
| "exclusive offer today" | 1 | 0.95 | 1 ✓ | 1 ✓ | 1 ✓ |
| Threshold | True positives | False positives | False negatives | Precision | Recall |
|---|---|---|---|---|---|
| 0.7 | 3 | 0 | 1 | 1.00 | 0.75 |
| 0.5 | 3 | 1 | 1 | 0.75 | 0.75 |
| 0.3 | 3 | 1 | 1 | 0.75 | 0.75 |
The interesting row is "free meeting notes" with — the model is exactly at the boundary. The 0.5 threshold calls it spam (and is wrong); the 0.7 threshold calls it ham (and is right). On this tiny dataset the precision/recall swap is captured in a single example; on a real dataset the same swap plays out across thousands.
9.4 Calibration check
The full-prediction probability histogram is heavily bimodal — values cluster at — so the model is making confident predictions on most messages. For each confidence bucket, the actual frequency of spam should match the predicted probability. On this dataset it does, because the model is fit by maximum likelihood. The same check on a heavily regularised fit or on a tree ensemble would surface systematic miscalibration.
Key Takeaways
- Logistic regression takes a linear score , pushes it through the sigmoid to get a probability in , and trains the weights by minimising the binary cross-entropy loss. The sigmoid derivative makes the gradient cleanly .
- Each fitted coefficient is the change in log-odds of class 1 per unit increase in feature ; the odds ratio is , which makes the model directly readable on a multiplicative scale.
- The decision threshold is a separate, tunable parameter. Lowering it always raises recall and generally lowers precision; the right value depends on the cost of false positives versus false negatives and on the class prior. The 0.5 default is a starting point, not a property of the model.
- A calibrated classifier is one whose predicted probability is the empirical frequency of the event — predicted 0.7 means 70% of such cases are positive. Logistic regression is calibrated by construction because it is trained to minimise exactly the negative log-likelihood. Tree ensembles and uncalibrated neural networks typically need Platt scaling or isotonic regression.
- Compared to other classifiers, logistic regression is fast, convex, interpretable, and produces calibrated probabilities, but its decision boundary is always linear. When the data has non-linear structure, either engineer non-linear features or move to a model that handles non-linearity natively.