Linear regression is the workhorse of supervised learning: a model that predicts a continuous target as a linear combination of features, fit by minimising a squared-error loss. This lesson derives the cost function from first principles, shows the gradient that drives optimisation, walks through gradient descent with real numeric iterations, explains how to read the fitted coefficients, defines the coefficient of determination , lists the four classical assumptions together with the symptom of each violation, and closes with an end-to-end worked example on a tiny dataset.
Learning Objectives
- State the linear-regression hypothesis and identify the role of each parameter vector entry.
- Derive the mean-squared-error cost function, take its gradient with respect to the parameters, and set the gradient to zero to obtain the normal equation.
- Run a few iterations of batch gradient descent on a one-feature problem and explain how the learning rate shapes the trajectory.
- Interpret the fitted coefficients as feature effects, define , and state why can mislead.
- Enumerate the linearity, independence, homoscedasticity, and normality assumptions, and recognise the residual pattern that signals each violation.
1. The Linear Hypothesis
A linear regression model assumes the target is a linear combination of the features. For a single training example with feature vector — the leading is the constant term, sometimes called the bias unit — the prediction is
where is the parameter vector. "Linear" refers to linearity in the parameters, not in the features: the model is still linear regression, because and appear linearly. A model is not, even though the relationship between and is monotone.
For training examples stacked into a design matrix whose -th row is , the vector of predictions is
The columns of are the features; the rows are the observations. The convention that every row begins with a folds the intercept into the matrix product, so the same code works whether or not an intercept is included.
The hypothesis is intentionally simple — that is its strength. When the underlying relationship is approximately linear, no more expressive model is needed; the variance saved by not fitting extra parameters becomes prediction accuracy on new data.
2. The Mean-Squared-Error Cost Function
Pick by minimising a loss that penalises the gap between and the observed . The classical choice is the mean squared error
Three reasons make MSE the default regression loss.
- Algebraic convenience. The square kills the sign so positive and negative residuals are penalised symmetrically; the gradient becomes linear in ; the optimum has a closed-form solution.
- Probabilistic interpretation under Gaussian noise. If the targets are generated by with independent across examples, minimising the negative log-likelihood is exactly minimising the MSE. The squared loss is therefore the maximum-likelihood estimator under additive Gaussian noise.
- Convexity. The MSE is a positive-definite quadratic form in ; it has a single global minimum and no local minima to trap an optimiser.
The vector form, used in every line of production code that fits a linear model, is
where is the squared Euclidean norm. The factor does not change the optimum; it is conventional so that is on the same scale as a per-example loss and so that gradients do not scale with the dataset size.
2.1 Why not mean absolute error?
Mean absolute error is also a valid loss. It penalises outliers less harshly, which makes it preferable when the noise is heavy-tailed or when occasional large errors are acceptable. Its gradient is not defined at the residuals equal to zero, however, and the optimisation has no closed form. Use MAE when robustness to outliers matters; use MSE when outlier sensitivity is acceptable or desirable and when closed-form inference is convenient.
3. Gradient of the MSE
Differentiate with respect to and set the result to zero to find the optimum. The scalar form expands as
Take the partial derivative with respect to one parameter . The chain rule gives
Stacking the partials into a vector, using the design matrix, the gradient is
The two factors have clear interpretations. The residual vector says which direction each prediction is wrong and by how much; the multiplication projects those residuals onto each parameter axis. A feature whose column is large in absolute value has a big effect on the gradient, which is why feature scaling matters for gradient descent: features measured in millions push the gradient in their direction much harder than features measured in fractions.
Setting the gradient to zero,
which is the normal equation. As long as is invertible — that is, no feature is a linear combination of the others — the unique solution is
4. Gradient Descent With Real Numbers
The normal equation is closed-form but cubic in the number of features. When is large, or when the data arrives in a stream and the model must be updated online, gradient descent is preferred. The update rule is
where is the learning rate. The negative gradient is the direction of steepest descent; controls how far to step in that direction.
4.1 A one-feature dataset
Use the same five-point dataset introduced in lesson 1, restated here with simpler numbers for clarity.
| (sq ft / 100) | (price, $k) | |
|---|---|---|
| 1 | 10 | 250 |
| 2 | 15 | 380 |
| 3 | 18 | 410 |
| 4 | 24 | 540 |
| 5 | 30 | 620 |
With the relevant sums are , , , . Means are , . The closed-form solution (rescaled back to thousands of dollars per 100 sq ft) is , , so the fitted line is
4.2 Iterating gradient descent on this surface
Start at , and use learning rate . The MSE surface for this dataset is a smooth, convex bowl with its minimum at the closed-form solution above. Each gradient step takes a small bite out of the loss.
| Iter | |||
|---|---|---|---|
| 0 | 0.00 | 0.00 | 213{,}200 |
| 1 | 1.76 | 4.84 | 132{,}541 |
| 2 | 3.08 | 8.46 | 84{,}027 |
| 5 | 6.16 | 16.49 | 22{,}615 |
| 10 | 8.39 | 19.99 | 6{,}212 |
| 20 | 11.84 | 23.13 | 1{,}122 |
| 50 | 15.96 | 24.99 | 326 |
| 100 | 17.62 | 25.34 | 308 |
| 200 | 17.74 | 25.36 | 308 |
Two observations. First, the loss falls rapidly at first — most of the error lives in the first ten iterations — and then trails off as the iterates approach the closed-form minimum near , . (The slight discrepancy with the rescaled figures above is the vs convention; the minimum is the same.) Second, the iterates overshoot slightly before settling, a signature of a learning rate that is large enough to explore but small enough to converge.
4.3 What changes when changes
| Learning rate | Behaviour |
|---|---|
| Too small, e.g. | Iterates crawl towards the minimum; thousands of iterations needed |
| Just right, e.g. | Steady monotone descent; converges in a few hundred steps |
| Too large, e.g. | Loss oscillates or diverges; jumps across the bowl |
| Adaptive (Adam, RMSProp) | Per-parameter scaling lets the optimiser cope with features on very different scales |
The bowl shape is convex but anisotropic — long and narrow along one axis, short and steep along another — when features have very different scales. Standardising every feature to zero mean and unit variance before running gradient descent makes the bowl near-spherical and lets a single learning rate work for all parameters.
4.4 Gradient descent in Python
import numpy as np
x = np.array([10.0, 15.0, 18.0, 24.0, 30.0])
y = np.array([250.0, 380.0, 410.0, 540.0, 620.0])
X = np.column_stack([np.ones_like(x), x])
theta = np.zeros(2)
alpha = 1e-3
n = len(y)
for it in range(200):
residual = X @ theta - y
grad = (2.0 / n) * (X.T @ residual)
theta -= alpha * grad
print(f"theta_0 = {theta[0]:.2f}, theta_1 = {theta[1]:.2f}")
The four lines inside the loop are the entire algorithm: form the residual vector,
project it onto each parameter via , scale by , and subtract the
result scaled by from the current parameters. The same skeleton, with
residual replaced by a stochastic estimate, becomes stochastic gradient descent.
5. The Closed-Form Solution (Normal Equation)
Setting the gradient to zero gives , also known as the ordinary least squares (OLS) solution. Three properties make it worth knowing.
- Exact in one shot. No learning rate, no iteration count, no convergence worries. The output is the unique minimum of the convex MSE surface.
- Cheap when is small. Cost is dominated by the matrix product , which is , plus the inversion or solve, which is . For up to a few thousand the closed form is faster than gradient descent to convergence.
- Brittle when is ill-conditioned. Features that are near-collinear make nearly singular; numerical errors in the inversion dominate. The standard fix is to add a small ridge regulariser, giving the ridge regression solution , covered in a later lesson.
In production code the closed form is rarely written by hand. The standard tool is
import numpy as np
from sklearn.linear_model import LinearRegression
model = LinearRegression()
model.fit(x.reshape(-1, 1), y)
print(model.intercept_, model.coef_) # ~17.74, [25.36]
Both np.linalg.lstsq and scikit-learn's LinearRegression solve the normal equation
under the hood; the difference is that scikit-learn also stores metadata, supports
multiple features, and plays nicely with the rest of the library's ecosystem.
6. Interpreting Coefficients
Each entry of is read as the marginal effect of one feature on the predicted target, holding every other feature fixed.
For the worked example above, thousand dollars per 100 sq ft means that a 100-sq-ft increase in floor area is associated with a $25{,}360 increase in predicted sale price, all else equal. The intercept thousand dollars is the prediction when every feature is zero — for floor area, that is the extrapolation to a zero-square-foot "house", which is meaningless as a price but necessary for the algebra to work.
A few rules make coefficient interpretation more reliable.
| Rule | Why it matters |
|---|---|
| Standardise features before fitting | A coefficient on a standardised feature is comparable to the others; raw coefficients are not |
| Drop one feature category per "reference" level | Without a reference, dummy-encoded categories add a column for every level and the coefficients shift in uninterpretable ways |
| Watch for confounding | A large positive coefficient can mask a true negative relationship if an omitted variable drives both the feature and the target |
| Report a confidence interval, not just a point estimate | The point estimate is the best guess, not the true parameter; the interval quantifies that distinction |
A prediction interval for a new observation is wider than a confidence interval on the mean prediction: it includes both the uncertainty in and the unavoidable noise . Reporting only the point estimate hides this distinction.
7. R-squared and Its Limits
The coefficient of determination is the fraction of variance in that the model explains:
A model that predicts the mean for every observation has ; a model with zero residual error has . For simple linear regression, is also the square of the Pearson correlation between and .
Three reasons can mislead.
- It always rises when a feature is added. Adding a noisy feature to the design matrix increases by a tiny amount or not at all, so creeps upward even for useless features. Adjusted penalises the model for each parameter added and should be reported alongside the raw value.
- It says nothing about prediction quality on new data. A model with in-sample can have out-of-sample if it has overfit. Cross-validated or a held-out test set is the right diagnostic.
- It is undefined outside the range of the data. compares to the training mean; if the test distribution has shifted, the right comparison is the test mean, which does not respect.
The right metric depends on the question. MSE and RMSE are on the same scale as and are easy to communicate to non-technical stakeholders. MAE is more robust to outliers. MAPE (mean absolute percentage error) is unit-free but blows up when any is near zero. Pick the metric that aligns with the cost structure of the application, then report it on a held-out set.
8. The Four Assumptions of Linear Regression
The OLS estimator has known statistical properties — unbiased, minimum variance among linear unbiased estimators — only under four assumptions on the data-generating process. Diagnose each one with a residual plot before trusting a fitted model.
| Assumption | What it requires | Symptom of violation |
|---|---|---|
| Linearity | The conditional expectation is a linear function of | Residuals vs fitted values show a curved pattern; a parabola, sine wave, or trend in the residuals is a red flag |
| Independence | The residuals are independent across observations | Residuals are autocorrelated (time series) or clustered (grouped data); Durbin-Watson statistic near 0 or 4 indicates autocorrelation |
| Homoscedasticity | The residual variance is constant across the range of | A funnel or fan shape in the residuals vs fitted plot; variance grows or shrinks with |
| Normality of residuals | The residuals are approximately Gaussian, especially for inference | Histogram or Q-Q plot of residuals shows heavy tails or skew; small-sample confidence intervals become unreliable |
When an assumption fails, two broad fixes exist. The first is to transform the data: take logs of a heavy-tailed target, take square roots of count data, or fit the model on a different scale. The second is to change the model: replace ordinary least squares with weighted least squares when heteroscedasticity is the problem, with generalised least squares for correlated errors, or with a tree-based model when the relationship is genuinely non-linear.
A diagnostic routine that catches most issues in five minutes:
- Plot residuals vs fitted. Look for curvature (linearity) and funnel shape (homoscedasticity).
- Plot residuals vs each feature. Look for structure that the model missed.
- Plot the Q-Q plot of residuals. Look for heavy tails (normality).
- For time-ordered data, plot residuals vs time and compute Durbin-Watson.
- Compute Cook's distance per observation. Points with distance above are high-leverage outliers worth investigating.
8.1 When Linear Regression Is Not Enough
Linear regression is the wrong tool in three common situations.
- The relationship is non-linear. Use a basis expansion (polynomial features, splines) or switch to a tree-based model, kernel method, or neural network. Adding polynomial features keeps the linear-regression machinery and gains flexibility, at the cost of more parameters and more overfitting risk.
- The target is bounded. If is a probability, a count, or a non-negative quantity, the Gaussian noise assumption is wrong. Use logistic regression, Poisson regression, or a generalised linear model with the appropriate link.
- Features are highly correlated. Multicollinearity inflates the variance of the coefficients and makes them uninterpretable. Drop redundant features, combine them into a single index, or use ridge regression to regularise.
Reach for a more flexible model when the bias of linear regression is clearly limiting performance. Reach for a simpler model when the variance is the binding constraint. The art of applied ML is choosing where on the bias-variance curve the problem actually lives.
9. Worked Example: Sales Versus Advertising Spend
A small marketing team has tracked weekly sales (in thousands of dollars) against weekly spending on online ads (in thousands of dollars) for ten weeks. The data:
| Week | Ad spend ($k) | Sales ($k) |
|---|---|---|
| 1 | 1.2 | 32 |
| 2 | 1.6 | 38 |
| 3 | 2.0 | 45 |
| 4 | 2.4 | 50 |
| 5 | 2.8 | 56 |
| 6 | 3.2 | 61 |
| 7 | 3.6 | 67 |
| 8 | 4.0 | 72 |
| 9 | 4.4 | 76 |
| 10 | 4.8 | 81 |
The relationship looks plausibly linear; the team wants a model to forecast sales given a planned ad budget.
9.1 Computing the coefficients
The required sums are
- ,
- ,
Plug into the closed-form slope
The intercept is
The fitted regression line is therefore
Interpretation: every extra $1{,}000 of weekly ad spend is associated with about $13{,}480 of additional weekly sales. The intercept says the baseline sales — with no ad spend — would be about $17{,}360, which captures organic demand and brand recognition.
9.2 Predictions, residuals, and
For each of the ten weeks, the prediction and the residual:
| Week | True | Predicted | Residual | |
|---|---|---|---|---|
| 1 | 1.2 | 32 | 33.54 | |
| 2 | 1.6 | 38 | 38.93 | |
| 3 | 2.0 | 45 | 44.32 | |
| 4 | 2.4 | 50 | 49.71 | |
| 5 | 2.8 | 56 | 55.10 | |
| 6 | 3.2 | 61 | 60.50 | |
| 7 | 3.6 | 67 | 65.89 | |
| 8 | 4.0 | 72 | 71.28 | |
| 9 | 4.4 | 76 | 76.67 | |
| 10 | 4.8 | 81 | 82.06 |
Residuals oscillate gently around zero, with no obvious pattern — a good sign for the linearity and homoscedasticity assumptions. The residual sum of squares is
The total sum of squares, comparing each to the mean , is
The coefficient of determination is therefore
The model explains about 99.7% of the variance in sales, an unusually tight fit. The RMSE, on the scale of (thousands of dollars), is
so the typical prediction error is about $830 in weekly sales.
9.3 A forecast for next week
If the marketing team plans to spend $5{,}000 on ads next week, the model's forecast is
or roughly $84{,}760 in sales. The right way to communicate this to the team is "around $85k, give or take $1k, assuming the relationship we have seen continues to hold for this budget level". The "give or take" is the residual standard error; the "assuming" is the linearity assumption; both belong in the conversation.
9.4 Cross-check with scikit-learn
The whole computation in twenty lines of Python:
import numpy as np
from sklearn.linear_model import LinearRegression
x = np.array([1.2, 1.6, 2.0, 2.4, 2.8, 3.2, 3.6, 4.0, 4.4, 4.8]).reshape(-1, 1)
y = np.array([32, 38, 45, 50, 56, 61, 67, 72, 76, 81], dtype=float)
model = LinearRegression().fit(x, y)
print(f"intercept = {model.intercept_:.2f}, slope = {model.coef_[0]:.2f}")
r_squared = model.score(x, y)
print(f"R^2 = {r_squared:.4f}")
y_pred = model.predict(x)
rmse = float(np.sqrt(np.mean((y - y_pred) ** 2)))
print(f"RMSE = {rmse:.2f} thousand dollars")
forecast = model.predict(np.array([[5.0]]))
print(f"forecast for $5k ad spend: ${float(forecast[0]):.2f}k")
The four printed lines match the hand calculation. In production code the script
would split into train and test sets, report metrics on the test set rather than on
the training set, and use np.linalg.lstsq or model.predict inside a pipeline that
handles feature scaling, missing-value imputation, and cross-validation.
9.5 Sanity checks before shipping the model
The coefficients are credible. The residuals look like noise. is high. Before declaring victory, three further checks belong in the workflow.
- Residual plot. Plot against . A random cloud around zero confirms linearity and homoscedasticity; a curved shape means a non-linear feature is missing; a fan shape means the variance grows with the prediction.
- Out-of-sample test. Hold out the last two weeks, refit on the first eight, and score on the held-out two. A test close to the training is the proof that the model generalises.
- Domain plausibility. Is a 13.5:1 return on ad spend realistic for this business? If the team's historical return on ads is closer to 5:1, the model is probably picking up a confound (a new product launch that drove both sales and ad spend upward, for example). The numbers should pass a sniff test from a domain expert before they pass to a stakeholder.
A model that is statistically correct but operationally wrong is a worse outcome than no model at all: it gives false confidence to a decision that should have been made by other means.
Key Takeaways
- The linear-regression hypothesis is a linear combination of features fit by minimising the mean-squared-error loss; "linear" means linear in the parameters, not necessarily in the features.
- The gradient of the MSE is ; setting it to zero gives the normal equation , the closed-form OLS solution.
- Gradient descent iteratively subtracts a learning-rate-scaled gradient from the parameters; feature scaling and a sensible learning rate are the keys to fast convergence.
- Each fitted coefficient is the marginal effect of one feature on the predicted target, holding the others fixed. Report confidence intervals alongside the point estimates.
- is the fraction of variance explained; it rises whenever a feature is added (use adjusted to compare models) and does not measure out-of-sample performance on its own.
- The four classical assumptions are linearity, independence of residuals, constant variance (homoscedasticity), and approximate normality of residuals; diagnose each with a residual plot before trusting the model.
- The worked sales-versus-ad-spend example fits with , RMSE of about $830, and a forecast of about $85k for the next week given a $5k ad budget.