A model that memorises its training set is not a model that generalises. The gap between training behaviour and deployment behaviour is the central problem of supervised learning, and most of the machinery built around it — held-out splits, cross-validation, regularised loss functions, learning curves — is machinery for diagnosing and closing that gap. This lesson develops the diagnosis: underfitting and overfitting as the two failure modes a learning curve can reveal, the bias-variance decomposition that explains why the gap exists at all, the train/validation/test protocol that lets you measure the gap honestly, k-fold cross-validation when a single split is too noisy, the L1/L2 penalty families that constrain how much the model can memorise, and a remedy table for choosing which lever to pull when the curves tell you which problem you actually have.
Learning Objectives
- Define underfitting and overfitting in terms of train and validation error, and read the characteristic learning curves (both errors high and close together vs. training error low and validation error high with a wide gap) that distinguish them.
- State the bias-variance decomposition of expected prediction error, identify which term is reducible and which is noise, and connect the two terms to underfitting and overfitting respectively.
- Explain why a clean experiment needs three disjoint sets (train/validation/test), why the test set is touched only once, and why stratified splitting preserves the class prior in every fold.
- Describe k-fold cross-validation, the bias-versus-variance trade-off it makes relative to a single hold-out split, and the train/val/test + CV hybrid that lets you tune hyperparameters without ever peeking at the test set.
- Derive the L1 (Lasso) and L2 (Ridge) penalties from the regularised objective, prove via the subgradient condition that L1 can drive a coefficient to exactly zero while L2 only shrinks it, and write the elastic-net blend.
- Use learning curves as the diagnostic tool: identify the high-bias and high-variance shapes, choose the right remedy from the comparison table, and stop when the curves say the model is done.
1. Two Failure Modes: Underfitting and Overfitting
Every supervised model can fail in one of two structurally different ways. Both are diagnosed from the same instrument — a learning curve — but they require opposite remedies, so reading the curve correctly is the prerequisite to picking the right lever.
1.1 Underfitting: high bias, both errors high
An underfit model has not learned enough structure from the training data to capture the underlying pattern. Its training error is high and its validation error is high and the two are close together (or even identical). Adding more training data does not help: the model has too little capacity to absorb the pattern that is already there. The signal is that both curves plateau at a high error with a small gap.
Concretely, suppose polynomial regression of degree 1 is fit to a dataset that is actually generated by a cubic relationship. With 1,000 training points the training MSE settles around 0.42 and the validation MSE settles around 0.44. Doubling the dataset to 2,000 points changes those numbers by less than 0.01. The curves are flat and close together, both well above the irreducible-noise floor of about 0.05. That is the signature of underfitting: the model has hit its representational ceiling and more data cannot push it higher.
1.2 Overfitting: high variance, training error low and validation error high
An overfit model has learned the training set — including its noise — so well that it cannot reproduce the pattern on unseen examples. Its training error is low (often near zero on memorisable cases) while its validation error is substantially higher. There is a wide gap between the two curves, and the gap itself is the diagnosis. Adding more training data does help: more data is the only universally effective overfitting remedy because it dilutes any noise pattern the model was latching onto.
Concretely, with the same cubic data but a polynomial of degree 25, the training MSE falls to about 0.003 after enough epochs (the polynomial can interpolate 1,000 points almost exactly) while the validation MSE sits around 0.31. The gap of 0.30 is the overfitting penalty. Adding 10x more training data would shrink that gap: the polynomial needs more anchors before it stops threading the noise.
1.3 The good place: both errors low and close together
A well-fit model has low training error, low validation error, and a small gap between them. The two curves sit close together near the bottom of the plot. Anything below the irreducible noise floor is impossible — that floor is the variance of the label noise itself, and no model can beat it.
1.4 A worked example: the curves in numbers
| Model | Train MSE | Validation MSE | Gap | Diagnosis |
|---|---|---|---|---|
| Linear regression (degree 1) on cubic data | 0.42 | 0.44 | 0.02 | Underfitting — both high, gap small |
| Decision tree, depth 3 | 0.18 | 0.22 | 0.04 | Slight underfitting — both high, gap small |
| Decision tree, depth 12 | 0.05 | 0.19 | 0.14 | Slight overfitting — train low, validation noticeably higher |
| Decision tree, depth 30 (memorising) | 0.00 | 0.31 | 0.31 | Severe overfitting — train near zero, validation much higher |
The shape of the curves, not the magnitude, is the diagnosis. A depth-3 tree with train MSE 0.18 and validation MSE 0.22 is in the same regime as a linear model on cubic data: both are too simple for the task. The gap is small in both cases, but the absolute error is high because the model cannot represent the function.
import numpy as np
from sklearn.model_selection import learning_curve
from sklearn.tree import DecisionTreeRegressor
def plot_lc(model, X, y, train_sizes=np.linspace(0.1, 1.0, 8)):
sizes, train_err, val_err = learning_curve(
model, X, y, train_sizes=train_sizes,
cv=5, scoring="neg_mean_squared_error", shuffle=True, random_state=0,
)
# neg_mean_squared_error is negated so lower is better on the y-axis
return sizes, -train_err, -val_err
sizes, train_err, val_err = plot_lc(
DecisionTreeRegressor(max_depth=3), X, y,
)
# Underfitting example: train_err[-1] is around 0.18, val_err[-1] around 0.22
sizes, train_err, val_err = plot_lc(
DecisionTreeRegressor(max_depth=30), X, y,
)
# Overfitting example: train_err[-1] is around 0.00, val_err[-1] around 0.31
2. The Bias-Variance Decomposition
The underfitting/overfitting dichotomy has a precise mathematical explanation. Decompose the expected prediction error of a learned model at a single test point into three additive terms and the two failure modes fall out as the two reducible terms of the decomposition.
2.1 The decomposition
For a regression problem with target and prediction , the expected squared error at a single test point decomposes as:
The three terms are read as:
- Bias² is the squared distance between the average prediction of the model (averaged over many training sets drawn from the same distribution) and the truth. Bias is the structural error of the model class — its inability to represent the underlying function.
- Variance is how much itself moves around as the training set changes. A high-variance model is one that fits the noise of any particular training set; different training sets give wildly different functions.
- Irreducible noise is the variance of the label noise itself. No model can reduce it because the data itself is noisy.
The total expected error is the sum of three positive terms; you cannot push the sum below regardless of how clever the model is.
2.2 Bias and variance trade off
Bias and variance are not independent: as you add capacity to a model, bias tends to fall (the model can represent more functions, so its average prediction gets closer to the truth) but variance tends to rise (the model fits each training set more closely, so different training sets give different fits). The classic picture is:
The "trade-off" is the observation that the sum has a minimum somewhere in the middle of the capacity axis: at very low capacity both terms are large (underfitting), at very high capacity the bias term is small but the variance term is large (overfitting), and at the sweet spot the sum is closest to .
2.3 Connecting the decomposition to the curves
The training error tracks roughly — how much the model moves around plus the noise it cannot escape. The validation error tracks roughly . Underfitting is a high-bias regime: both errors are high and close together because the bias term dominates both. Overfitting is a high-variance regime: training error is low (variance alone plus noise) but validation error is much higher (the bias term is small, so the gap is almost entirely variance).
2.4 A numeric illustration
Consider three model classes fit to the same cubic data, with irreducible noise :
| Model | Bias² | Variance | Irreducible | Total expected error |
|---|---|---|---|---|
| Linear regression | 0.37 | 0.01 | 0.05 | 0.43 |
| Cubic regression | 0.01 | 0.04 | 0.05 | 0.10 |
| 25-degree polynomial | 0.00 | 0.27 | 0.05 | 0.32 |
Linear regression has high bias (it cannot represent a cubic) but low variance (a linear fit is stable across training sets). The 25-degree polynomial has zero bias (it can represent the cubic exactly) but high variance (the extra degrees of freedom amplify noise). Cubic regression hits the sweet spot: small bias, modest variance, total error near the irreducible floor. These numbers match the MSEs reported in Section 1.4.
3. Train, Validation, and Test Sets
To measure generalisation honestly you must hold out data the model never sees during training or tuning. The standard protocol uses three disjoint sets, each with a different role.
3.1 The three roles
- Training set. Used to fit the model parameters (weights in linear regression, split thresholds in a tree, etc.). The model sees every example here, multiple times across epochs.
- Validation set. Used to compare model variants — different hyperparameters, different feature sets, different regularisation strengths — and to pick the best one. The model never trains on it, but the engineer's choices are informed by it. The validation set is where early stopping, learning-rate schedules, and hyperparameter search look.
- Test set. Used exactly once, at the very end, to report the generalisation error of the chosen model. The test set must not influence any modelling decision, otherwise the reported number is contaminated and the true generalisation error is higher than reported.
The reason three sets are needed is that the validation set, once used to pick a model, becomes part of the fitting process: you picked the model that did best on it, so the validation error of the chosen model is an optimistic estimate of its deployment error. The test set is held out from every decision so its error is unbiased.
3.2 Common split sizes
For a dataset of examples, a typical split is 60/20/20 or 70/15/15. Modern best practice with large datasets pushes toward 98/1/1 because the test set only needs to be large enough to estimate the generalisation error to the precision you care about, and large training sets always help.
from sklearn.model_selection import train_test_split
# First split: carve off the test set (15% of the data)
X_trainval, X_test, y_trainval, y_test = train_test_split(
X, y, test_size=0.15, random_state=42,
)
# Second split: carve validation out of the remaining 85%
X_train, X_val, y_train, y_val = train_test_split(
X_trainval, y_trainval, test_size=0.176, random_state=43,
)
# Result is train ~70%, val ~15%, test ~15% of the original dataset.
The two splits use different random_state values. If they shared a state
the second split would inherit the first's permutation structure and the
sets would not be independent.
3.3 Stratification
When the labels are categorical with a class imbalance, a naive random split can put very few examples of a rare class into the validation or test set, making those sets unreliable for estimating per-class metrics. Stratification preserves the class prior in every split:
X_train, X_val, y_train, y_val = train_test_split(
X, y, test_size=0.2, stratify=y, random_state=42,
)
# If y has 7% positives in the source, every split has 7% positives.
For regression, stratify= is not directly applicable (no classes), but the
equivalent is to bin the target into quantiles and stratify on the bins. This
preserves the marginal distribution of the target in every split, which
matters most when the target itself has heavy tails or rare values.
3.4 Why the test set must be touched once
If you use the test set to pick a model, your reported test error is biased downward because you chose the model that did best on it. If you use it to pick a feature, a threshold, a regularisation strength, or anything else, the bias grows. The test set is the contract with the reader: this number is the expected error of the model in deployment. Breaking the contract by peeking means the number overstates how well the model will do.
A practical pattern: when you need more than a handful of test-set evaluations during development (for example, when the project has multiple subtasks each with its own metric), maintain a separate holdout set that is only used for the final report, and use cross-validation on the remaining data for everything else.
4. k-Fold Cross-Validation
A single train/validation split wastes data: every example used for validation is not used for training, and on small datasets that loss is significant. k-fold cross-validation rotates the validation fold through the dataset so every example is used for training times and for validation exactly once.
4.1 The procedure
- Shuffle the dataset and partition it into equally sized folds.
- For fold : train on the union of the other folds, evaluate on fold .
- Average the validation scores. This average is the cross-validated estimate of the generalisation error.
The result is a number computed from independent training runs. Common choices are and ; is a sensible default for moderate-sized datasets, is faster, and leave-one-out cross-validation () is the most expensive and is reserved for very small datasets.
4.2 Bias of the cross-validated estimate
The cross-validated estimate is biased upward relative to the true generalisation error because each fold trains on only of the data — a smaller training set than a model deployed on the full dataset would use. With each fold trains on 90% of the data, so the CV estimate is approximately the error of a model trained on 90% of the data, which is slightly higher than the error of a model trained on 100%. The bias shrinks as grows: has about 10% of the bias of .
Conversely, the variance of the CV estimate shrinks as grows because each fold shares more data with the others, so the per-fold scores are correlated and the average is less variable. The variance and bias move in opposite directions as increases; is the standard compromise.
4.3 Train/Val/Test + CV: the hybrid protocol
The full protocol when you have enough data is:
- Carve off a test set. Set it aside. Do not look at it.
- Run -fold cross-validation on the remaining (train + val) data. Use the average validation score across folds to compare hyperparameters, regularisation strengths, feature sets.
- Once the best hyperparameters are chosen, refit the model on the entire (train + val) set using those hyperparameters.
- Evaluate the refit model on the test set exactly once. Report that number.
The hybrid gives you the honest test-set estimate from step 4 while still letting you tune on all the non-test data via cross-validation. Every hyperparameter choice was made on data the model never saw at deployment, and the test set has been touched exactly once.
from sklearn.model_selection import StratifiedKFold, cross_val_score
from sklearn.linear_model import Ridge
skf = StratifiedKFold(n_splits=5, shuffle=True, random_state=0)
scores = cross_val_score(
Ridge(alpha=1.0), X_trainval, y_trainval,
cv=skf, scoring="neg_mean_squared_error",
)
# Mean of -scores across 5 folds gives the CV estimate of MSE
print(f"CV MSE = {-scores.mean():.3f} (+/- {scores.std():.3f})")
The cv argument accepts a StratifiedKFold object so each fold preserves
the class prior, even though cross_val_score does not accept stratify=
directly. The scoring parameter uses neg_mean_squared_error because
scikit-learn's convention is "higher is better", so the negated MSE is the
score you maximise; the printed number negates it back for readability.
4.4 When CV breaks
Cross-validation assumes the examples are independent and identically distributed. It breaks when:
- The data is time-series: using future examples to predict the past inflates the score. Use a rolling-origin or blocked time-series split instead.
- The data has group structure (multiple rows per patient, per user, per
session): random folds leak information between groups. Use
GroupKFoldto keep every group in a single fold. - The data is too large: 5-fold CV on 100 million examples is wasteful when a single 99/1 split is already accurate enough.
5. L1 (Lasso) and L2 (Ridge) Regularisation
Regularisation is the practice of adding a penalty term to the loss function that discourages the model from taking on extreme parameter values. The two penalties that earn their place in every practitioner's toolbox are L2 (Ridge) and L1 (Lasso). They are both small, both convex, both penalise magnitude, but they have structurally different effects on the trained coefficients.
5.1 The regularised objective
For a regression problem with parameter vector and a base loss (e.g. mean squared error on the training set), the regularised objective is:
where is the regularisation strength and is the penalty. The optimiser chooses to minimise , so large values are discouraged when .
5.2 L2 (Ridge) penalty
The L2 penalty is the squared norm:
The objective becomes
The gradient of the L2 term with respect to is . At the optimum the gradient of the loss balances this penalty term, giving the closed-form Ridge solution when is the squared-error loss:
The term is exactly the regulariser — it makes the matrix invertible even when is singular or near-singular, which is the L2 regulariser's second role: numerical stabilisation.
The L2 penalty can shrink a coefficient but cannot set it exactly to zero: is a saddle of the penalty, and any non-zero gradient of the loss pulls the coefficient off zero. The closed-form solution above confirms this: every diagonal entry is positive, so the inverse gives a non-zero coefficient whenever is non-zero on that axis.
5.3 L1 (Lasso) penalty
The L1 penalty is the norm:
The objective becomes
The L1 penalty is non-differentiable at , so the optimum is characterised by the subgradient condition: a vector is a subgradient of at if when and otherwise. At the optimum, the gradient of the loss plus must equal zero on every axis:
When the unregularised optimum has , the loss gradient is some non-zero number and the L1 optimum is (the soft-thresholding rule). When the unregularised optimum has at , the subgradient condition is satisfied by — that is, the optimum is exactly . The L1 penalty therefore selects variables: it sets to zero any coefficient whose loss-gradient magnitude is less than , and the number of zero coefficients grows as grows.
5.4 Why L1 can zero a coefficient and L2 cannot
The L2 penalty at has gradient zero, so it exerts no force to keep at zero — the loss gradient alone determines the optimum, which is generically non-zero. The L1 penalty at has a subgradient that spans the entire interval , so any loss gradient in that interval can be balanced by an L1 force, leaving the optimum exactly at zero. This is the geometric intuition: the L2 unit ball is a smooth sphere that touches the axes tangentially, while the L1 unit ball is a diamond with corners sitting on the axes. The first time the loss contours touch the L2 ball they do so off-axis; the first time they touch the L1 ball they do so on a corner, pinning that axis to zero.
5.5 Implementation
from sklearn.linear_model import Ridge, Lasso
ridge = Ridge(alpha=1.0).fit(X_train, y_train)
# alpha is the regularisation strength lambda. All coefficients shrink
# but none are exactly zero.
lasso = Lasso(alpha=0.1).fit(X_train, y_train)
# At alpha=0.1 some coefficients are exactly zero. The model has
# performed implicit feature selection by zeroing unimportant features.
Ridge and Lasso use alpha= for the regularisation strength. Both
support positive=True if the coefficients are constrained to be
non-negative. For logistic regression the penalty='l1' and penalty='l2'
arguments switch the penalty family; C is the inverse of alpha, so
larger C is weaker regularisation.
6. Elastic Net: L1 and L2 Together
L1 is good at producing sparse models (few non-zero coefficients) but unstable when features are correlated — the penalty splits its mass arbitrarily among correlated variables. L2 is good at shrinking all coefficients smoothly but cannot zero any. Elastic net combines the two penalties to inherit sparsity from L1 and stability under correlation from L2.
6.1 The penalty
where sets the overall strength and sets the L1/L2 mix. The objective is
When this is Lasso; when this is Ridge. Intermediate values give elastic net: sparse models that are robust to correlated features.
6.2 Why elastic net handles correlated features better than pure Lasso
Pure Lasso with highly correlated features tends to pick one of them at random and zero the others, because the L1 penalty has no preference among the correlated directions. Elastic net's L2 component keeps correlated coefficients together — if and are nearly collinear, the L2 penalty discourages putting all the mass on and none on , so the solution is more stable across training sets.
from sklearn.linear_model import ElasticNet
enet = ElasticNet(alpha=0.1, l1_ratio=0.5).fit(X_train, y_train)
# l1_ratio is the rho in the formula: 1.0 gives Lasso, 0.0 gives Ridge,
# intermediate values blend the two penalties together.
The convention in scikit-learn is that alpha scales the whole penalty and
l1_ratio is . When l1_ratio=1.0 the model is mathematically
Lasso; when l1_ratio=0.0 it is Ridge. The defaults are alpha=1.0 and
l1_ratio=0.5.
7. Learning Curves as the Diagnostic Tool
A learning curve plots training error and validation error as functions of training-set size (or of training iteration, in the online-learning case). The shape of the two curves together is the single most informative diagnostic for underfitting vs overfitting, and reading it correctly is a prerequisite for choosing the right remedy.
7.1 The high-bias shape (underfitting)
When the model has high bias:
- Both curves plateau at a high error as training-set size grows.
- The two curves are close together because variance is small.
- Adding more training data does not move either curve meaningfully.
- The plateau is well above the irreducible noise floor.
The only way out is to add capacity: more features, a more flexible model, or fewer regularisation constraints.
7.2 The high-variance shape (overfitting)
When the model has high variance:
- The training error is low and grows slightly with more data (the model has more anchors to memorise, so individual fits get slightly worse on training, but still low).
- The validation error is much higher than the training error with a wide gap.
- Adding more training data closes the gap: both curves converge to a similar low value.
The remedies are: more data, regularisation, fewer features, or an ensemble. Capacity should not be increased — the model already has too much.
7.3 The good-fit shape
A well-fit model shows:
- Both curves converging to a low error close to the irreducible floor.
- A small gap between them.
- The gap continuing to shrink as more data is added (or already being small at the data sizes you have).
7.4 Reading a learning curve: a decision procedure
- Compute the validation error at the largest training-set size you can afford. If it is close to the irreducible-noise estimate of the problem, you are done.
- If the validation error is high, look at the gap between the two
curves at the largest training-set size.
- Small gap → high bias. Increase capacity.
- Large gap → high variance. Get more data or regularise.
- If you cannot get more data and you cannot increase capacity (because the model is already at maximum), revisit the feature set: drop noisy features, add features that have signal, transform existing features.
7.5 Implementation
import numpy as np
import matplotlib.pyplot as plt
from sklearn.model_selection import learning_curve
train_sizes, train_scores, val_scores = learning_curve(
estimator, X, y, cv=5,
train_sizes=np.linspace(0.1, 1.0, 10),
scoring="neg_mean_squared_error", shuffle=True, random_state=0,
)
# Negate because scikit-learn scores follow higher-is-better convention
train_err = -train_scores.mean(axis=1)
val_err = -val_scores.mean(axis=1)
plt.plot(train_sizes, train_err, label="training error")
plt.plot(train_sizes, val_err, label="validation error")
plt.xlabel("training set size"); plt.ylabel("MSE")
plt.legend(); plt.show()
The learning_curve helper does the data-size sweep for you, refits the
model at each size, and returns the per-fold scores. Plot the means; the
shaded bands around each line (the standard deviation across folds) tell
you how stable the diagnosis is.
8. Remedies for High Bias vs High Variance
The two failure modes require opposite remedies. Increasing capacity helps high bias but hurts high variance; getting more data helps high variance but does nothing for high bias. The table below is the rule of thumb.
| Symptom in the learning curve | Diagnosis | Effective remedies | Ineffective (or counterproductive) |
|---|---|---|---|
| Train and validation errors both high and close together | High bias (underfitting) | Add features; use a more flexible model; reduce regularisation strength ; cross features; decrease model bias directly (e.g. switch from linear to tree-based) | More training data; more regularisation; dropping features; smaller models; bagging/ensembling |
| Train error low, validation error much higher, large gap | High variance (overfitting) | Collect more training data; increase regularisation strength (especially L2); use L1 or elastic-net for feature selection; reduce the number of features; early stopping; dropout (in neural networks); bagging/ensembling; data augmentation | Adding more capacity; using a more flexible model; reducing regularisation; using fewer training examples |
| Both errors low and close together | Good fit | Stop. More changes risk pushing into one of the other two regimes. | Continued fiddling without a measured signal to chase |
8.1 Why ensembling helps high variance but not high bias
A bagged ensemble (e.g. random forest) trains many models on bootstrap samples and averages their predictions. Averaging reduces variance by a factor of roughly when the models are independent, but does not move the bias term — each individual model has the same bias, and the average of biased models has the same bias. Bagging therefore reduces variance without affecting bias, which is why it is the standard remedy for high-variance models like fully grown decision trees.
8.2 A concrete tuning sequence
A practical sequence when the curves say "high variance":
- First try more data — it is the only universally effective remedy.
- If more data is not available, try regularisation. Start with L2, sweep on the validation set.
- If features are many and you suspect most are noise, try L1 or elastic-net to drop the noise automatically.
- If the model is a deep network, try dropout or early stopping.
- Try an ensemble (bagging, boosting) as a final lever.
The order matters: each remedy is cheaper to try than the next, and each addresses a different source of variance. Start at the top and walk down until the curves say the problem is solved.
9. Worked Example End-to-End
A worked example ties the diagnostic procedure together. The task is a regression on a synthetic cubic dataset with ; the dataset has 1,000 points; we will deliberately underfit and overfit and then fix both.
9.1 The underfit case
Fit Ridge(alpha=1000) (very strong regularisation, near the linear limit)
on a polynomial-feature expansion up to degree 10:
from sklearn.preprocessing import PolynomialFeatures
from sklearn.pipeline import make_pipeline
from sklearn.linear_model import Ridge
underfit = make_pipeline(
PolynomialFeatures(degree=10, include_bias=False),
Ridge(alpha=1000),
).fit(X_train, y_train)
print(f"train MSE = {mse(underfit, X_train, y_train):.3f}")
print(f"val MSE = {mse(underfit, X_val, y_val):.3f}")
# train MSE = 0.41 val MSE = 0.43 gap = 0.02 indicating underfitting
Both errors are above 0.40 — well above the 0.05 noise floor — and the gap is small. The diagnosis is high bias. The remedy is to weaken the regularisation (lower ) or to use a model with less built-in shrinkage.
9.2 The overfit case
Fit Ridge(alpha=0.001) on the same degree-10 polynomial features:
overfit = make_pipeline(
PolynomialFeatures(degree=10, include_bias=False),
Ridge(alpha=0.001),
).fit(X_train, y_train)
print(f"train MSE = {mse(overfit, X_train, y_train):.3f}")
print(f"val MSE = {mse(overfit, X_val, y_val):.3f}")
# train MSE = 0.04 val MSE = 0.18 gap = 0.14 indicating overfitting
Training error is much lower than the underfit case, but validation error has barely moved. The gap is large. The diagnosis is high variance, and the remedy is more data or stronger regularisation.
9.3 Tuning by cross-validation
Use 5-fold CV on the train+val set to pick :
from sklearn.model_selection import GridSearchCV
import numpy as np
param_grid = {"ridge__alpha": np.logspace(-3, 3, 13)}
search = GridSearchCV(
make_pipeline(PolynomialFeatures(degree=10, include_bias=False), Ridge()),
param_grid, cv=5, scoring="neg_mean_squared_error",
).fit(X_trainval, y_trainval)
print(f"best alpha = {search.best_params_}")
# best alpha is around 1.0
The best sits in the middle of the search range — small enough to let the polynomial fit the cubic pattern, large enough to keep the high-degree coefficients from amplifying noise.
9.4 Final evaluation on the test set
Refit on the entire train+val data with the chosen , evaluate on the test set:
final = make_pipeline(
PolynomialFeatures(degree=10, include_bias=False),
Ridge(alpha=search.best_params_["ridge__alpha"]),
).fit(X_trainval, y_trainval)
print(f"test MSE = {mse(final, X_test, y_test):.3f}")
# test MSE is about 0.08, close to the 0.05 irreducible floor
The test MSE is 0.08, close to the 0.05 irreducible floor. The model has recovered most of the available signal: bias has fallen (the polynomial can represent the cubic), variance has fallen (the L2 penalty damps the high-degree coefficients), and the remaining gap to the noise floor is the unavoidable 0.03.
9.5 Diagnostic summary
| Step | Train MSE | Val MSE | Test MSE | Diagnosis |
|---|---|---|---|---|
| (underfit) | 0.41 | 0.43 | — | High bias |
| (overfit) | 0.04 | 0.18 | — | High variance |
| (CV-tuned) | 0.06 | 0.09 | 0.08 | Good fit |
The CV-tuned model's train and validation errors are both close to the irreducible floor and the gap is small. The test MSE, evaluated exactly once, lands between the train and validation numbers, as expected. The contract with the reader has been kept: the test set was touched only at the end, and the reported number is the expected error of the model in deployment.
Key Takeaways
- The two failure modes — underfitting (high bias) and overfitting (high variance) — are diagnosed from the gap between the training curve and the validation curve in a learning curve, and they require opposite remedies.
- The bias-variance decomposition shows that the irreducible noise is a floor no model can pass and that bias and variance move in opposite directions as capacity changes.
- Three disjoint sets (train/validation/test) are needed because the validation set becomes contaminated once it is used to choose a model; the test set is touched exactly once and is the only unbiased generalisation estimate.
- k-fold cross-validation rotates the validation fold through the data; it is biased upward (each fold trains on of the data) but lower-variance than a single split, and pairs with a held-out test set in the train/val/test + CV hybrid protocol.
- L1 (Lasso) drives coefficients to exactly zero through the subgradient condition because the ball has corners on the axes; L2 (Ridge) only shrinks coefficients because the ball touches the axes tangentially; elastic net blends the two to inherit sparsity from L1 and stability under correlated features from L2.
- Learning curves are the diagnostic tool: high bias shows both errors high and close together (remedy: more capacity), high variance shows a wide gap with low training error (remedy: more data, more regularisation, fewer features), and the remedy table maps symptoms to the lever that fixes them.