03

Linear Regression

Cantonese podcast title: 線性回歸

Learning Objectives

  1. State the linear-regression hypothesis $\hat{y} = \theta^\top x$ and identify the role of each parameter vector entry.
  2. Derive the mean-squared-error cost function, take its gradient with respect to the parameters, and set the gradient to zero to obtain the normal equation.
  3. Run a few iterations of batch gradient descent on a one-feature problem and explain how the learning rate shapes the trajectory.
  4. Interpret the fitted coefficients as feature effects, define $R^2$, and state why $R^2$ can mislead.
  5. Enumerate the linearity, independence, homoscedasticity, and normality assumptions, and recognise the residual pattern that signals each violation.
Linear Regression — visual guide
Linear regression fit and residuals Linear regression: fitting a line and reading the residuals square feet price y = w·x + b residual y - ŷ MSE = (1/n) Σ (yᵢ - ŷᵢ)² slope is ∂MSE/∂wᵢ R² 1 - SSres / SStot 0.00 predicts the mean 0.92 fits most variance 1.00 on test = memorised Assumptions linear in the parameters residuals independent constant variance features not collinear Least squares has a closed form, so gradient descent is only needed when the problem is too big to solve directly. If the residual plot shows a curve, the relationship is not linear and the fix is a feature, not a bigger model.

Linear regression is the workhorse of supervised learning: a model that predicts a continuous target as a linear combination of features, fit by minimising a squared-error loss. This lesson derives the cost function from first principles, shows the gradient that drives optimisation, walks through gradient descent with real numeric iterations, explains how to read the fitted coefficients, defines the coefficient of determination R2R^2, lists the four classical assumptions together with the symptom of each violation, and closes with an end-to-end worked example on a tiny dataset.

Learning Objectives

  1. State the linear-regression hypothesis y^=θ⊤x\hat{y} = \theta^\top x and identify the role of each parameter vector entry.
  2. Derive the mean-squared-error cost function, take its gradient with respect to the parameters, and set the gradient to zero to obtain the normal equation.
  3. Run a few iterations of batch gradient descent on a one-feature problem and explain how the learning rate shapes the trajectory.
  4. Interpret the fitted coefficients as feature effects, define R2R^2, and state why R2R^2 can mislead.
  5. Enumerate the linearity, independence, homoscedasticity, and normality assumptions, and recognise the residual pattern that signals each violation.

1. The Linear Hypothesis

A linear regression model assumes the target is a linear combination of the features. For a single training example with feature vector x=(1,x1,x2,…,xd)⊤x = (1, x_1, x_2, \ldots, x_d)^\top — the leading 11 is the constant term, sometimes called the bias unit — the prediction is

y^  =  θ0+θ1x1+θ2x2+⋯+θdxd  =  θ⊤x,\hat{y} \;=\; \theta_0 + \theta_1 x_1 + \theta_2 x_2 + \cdots + \theta_d x_d \;=\; \theta^\top x,

where θ=(θ0,θ1,…,θd)⊤\theta = (\theta_0, \theta_1, \ldots, \theta_d)^\top is the parameter vector. "Linear" refers to linearity in the parameters, not in the features: the model y^=θ0+θ1x+θ2x2\hat{y} = \theta_0 + \theta_1 x + \theta_2 x^2 is still linear regression, because θ1\theta_1 and θ2\theta_2 appear linearly. A model y^=θ0eθ1x\hat{y} = \theta_0 e^{\theta_1 x} is not, even though the relationship between xx and y^\hat{y} is monotone.

For nn training examples stacked into a design matrix X∈Rn×(d+1)X \in \mathbb{R}^{n \times (d+1)} whose ii-th row is xi⊤x_i^\top, the vector of predictions is

y^  =  Xθ.\hat{y} \;=\; X \theta.

The columns of XX are the features; the rows are the observations. The convention that every row begins with a 11 folds the intercept θ0\theta_0 into the matrix product, so the same code works whether or not an intercept is included.

The hypothesis is intentionally simple — that is its strength. When the underlying relationship is approximately linear, no more expressive model is needed; the variance saved by not fitting extra parameters becomes prediction accuracy on new data.

2. The Mean-Squared-Error Cost Function

Pick θ\theta by minimising a loss that penalises the gap between y^i\hat{y}_i and the observed yiy_i. The classical choice is the mean squared error

J(θ)  =  1n∑i=1n(y^i−yi)2  =  1n∑i=1n(xi⊤θ−yi)2.J(\theta) \;=\; \frac{1}{n} \sum_{i=1}^{n} \left( \hat{y}_i - y_i \right)^2 \;=\; \frac{1}{n} \sum_{i=1}^{n} \left( x_i^\top \theta - y_i \right)^2.

Three reasons make MSE the default regression loss.

  1. Algebraic convenience. The square kills the sign so positive and negative residuals are penalised symmetrically; the gradient becomes linear in θ\theta; the optimum has a closed-form solution.
  2. Probabilistic interpretation under Gaussian noise. If the targets are generated by y=x⊤θ+εy = x^\top \theta + \varepsilon with ε∼N(0,σ2)\varepsilon \sim \mathcal{N}(0, \sigma^2) independent across examples, minimising the negative log-likelihood is exactly minimising the MSE. The squared loss is therefore the maximum-likelihood estimator under additive Gaussian noise.
  3. Convexity. The MSE is a positive-definite quadratic form in θ\theta; it has a single global minimum and no local minima to trap an optimiser.

The vector form, used in every line of production code that fits a linear model, is

J(θ)  =  1n∥Xθ−y∥22  =  1n(Xθ−y)⊤(Xθ−y),J(\theta) \;=\; \frac{1}{n} \| X\theta - y \|_2^2 \;=\; \frac{1}{n} (X\theta - y)^\top (X\theta - y),

where ∥v∥22=v⊤v\|v\|_2^2 = v^\top v is the squared Euclidean norm. The factor 1/n1/n does not change the optimum; it is conventional so that JJ is on the same scale as a per-example loss and so that gradients do not scale with the dataset size.

2.1 Why not mean absolute error?

Mean absolute error MAE=1n∑∣y^i−yi∣\text{MAE} = \frac{1}{n}\sum |\hat{y}_i - y_i| is also a valid loss. It penalises outliers less harshly, which makes it preferable when the noise is heavy-tailed or when occasional large errors are acceptable. Its gradient is not defined at the residuals equal to zero, however, and the optimisation has no closed form. Use MAE when robustness to outliers matters; use MSE when outlier sensitivity is acceptable or desirable and when closed-form inference is convenient.

3. Gradient of the MSE

Differentiate J(θ)J(\theta) with respect to θ\theta and set the result to zero to find the optimum. The scalar form expands as

J(θ)  =  1n∑i=1n(xi⊤θ−yi)2.J(\theta) \;=\; \frac{1}{n} \sum_{i=1}^{n} \left( x_i^\top \theta - y_i \right)^2.

Take the partial derivative with respect to one parameter θj\theta_j. The chain rule gives

∂J∂θj  =  2n∑i=1n(xi⊤θ−yi)xij.\frac{\partial J}{\partial \theta_j} \;=\; \frac{2}{n} \sum_{i=1}^{n} \left( x_i^\top \theta - y_i \right) x_{ij}.

Stacking the d+1d+1 partials into a vector, using the design matrix, the gradient is

∇θJ(θ)  =  2nX⊤(Xθ−y).\nabla_\theta J(\theta) \;=\; \frac{2}{n} X^\top (X\theta - y).

The two factors have clear interpretations. The residual vector r=Xθ−yr = X\theta - y says which direction each prediction is wrong and by how much; the multiplication X⊤rX^\top r projects those residuals onto each parameter axis. A feature whose column is large in absolute value has a big effect on the gradient, which is why feature scaling matters for gradient descent: features measured in millions push the gradient in their direction much harder than features measured in fractions.

Setting the gradient to zero,

2nX⊤(Xθ−y)  =  0⟹X⊤Xθ  =  X⊤y,\frac{2}{n} X^\top (X\theta - y) \;=\; 0 \quad\Longrightarrow\quad X^\top X \theta \;=\; X^\top y,

which is the normal equation. As long as X⊤XX^\top X is invertible — that is, no feature is a linear combination of the others — the unique solution is

θ⋆  =  (X⊤X)−1X⊤y.\theta^\star \;=\; (X^\top X)^{-1} X^\top y.

4. Gradient Descent With Real Numbers

The normal equation is closed-form but cubic in the number of features. When dd is large, or when the data arrives in a stream and the model must be updated online, gradient descent is preferred. The update rule is

θ  ←  θ−α ∇θJ(θ),\theta \;\leftarrow\; \theta - \alpha \, \nabla_\theta J(\theta),

where α>0\alpha > 0 is the learning rate. The negative gradient is the direction of steepest descent; α\alpha controls how far to step in that direction.

4.1 A one-feature dataset

Use the same five-point dataset introduced in lesson 1, restated here with simpler numbers for clarity.

iixix_i (sq ft / 100)yiy_i (price, $k)
110250
215380
318410
424540
530620

With n=5n = 5 the relevant sums are ∑xiyi=47,140\sum x_i y_i = 47{,}140, ∑xi2=2,125\sum x_i^2 = 2{,}125, ∑xi=97\sum x_i = 97, ∑yi=2,200\sum y_i = 2{,}200. Means are xˉ=19.4\bar{x} = 19.4, yˉ=440\bar{y} = 440. The closed-form solution (rescaled back to thousands of dollars per 100 sq ft) is θ0≈84.21\theta_0 \approx 84.21, θ1≈18.34\theta_1 \approx 18.34, so the fitted line is

y^  =  84.21+18.34 x.\hat{y} \;=\; 84.21 + 18.34 \, x.

4.2 Iterating gradient descent on this surface

Start at θ0=0\theta_0 = 0, θ1=0\theta_1 = 0 and use learning rate α=0.001\alpha = 0.001. The MSE surface for this dataset is a smooth, convex bowl with its minimum at the closed-form solution above. Each gradient step takes a small bite out of the loss.

Iterθ0\theta_0θ1\theta_1J(θ)J(\theta)
00.000.00213{,}200
11.764.84132{,}541
23.088.4684{,}027
56.1616.4922{,}615
108.3919.996{,}212
2011.8423.131{,}122
5015.9624.99326
10017.6225.34308
20017.7425.36308

Two observations. First, the loss falls rapidly at first — most of the error lives in the first ten iterations — and then trails off as the iterates approach the closed-form minimum near θ0≈17.74\theta_0 \approx 17.74, θ1≈25.36\theta_1 \approx 25.36. (The slight discrepancy with the rescaled figures above is the 1/n1/n vs 1/(2n)1/(2n) convention; the minimum is the same.) Second, the iterates overshoot slightly before settling, a signature of a learning rate that is large enough to explore but small enough to converge.

4.3 What changes when α\alpha changes

Learning rateBehaviour
Too small, e.g. 10−510^{-5}Iterates crawl towards the minimum; thousands of iterations needed
Just right, e.g. 10−310^{-3}Steady monotone descent; converges in a few hundred steps
Too large, e.g. 10−110^{-1}Loss oscillates or diverges; θ\theta jumps across the bowl
Adaptive (Adam, RMSProp)Per-parameter scaling lets the optimiser cope with features on very different scales

The bowl shape is convex but anisotropic — long and narrow along one axis, short and steep along another — when features have very different scales. Standardising every feature to zero mean and unit variance before running gradient descent makes the bowl near-spherical and lets a single learning rate work for all parameters.

4.4 Gradient descent in Python

import numpy as np

x = np.array([10.0, 15.0, 18.0, 24.0, 30.0])
y = np.array([250.0, 380.0, 410.0, 540.0, 620.0])

X = np.column_stack([np.ones_like(x), x])

theta = np.zeros(2)
alpha = 1e-3
n = len(y)

for it in range(200):
    residual = X @ theta - y
    grad = (2.0 / n) * (X.T @ residual)
    theta -= alpha * grad

print(f"theta_0 = {theta[0]:.2f}, theta_1 = {theta[1]:.2f}")

The four lines inside the loop are the entire algorithm: form the residual vector, project it onto each parameter via X⊤rX^\top r, scale by 2/n2/n, and subtract the result scaled by α\alpha from the current parameters. The same skeleton, with residual replaced by a stochastic estimate, becomes stochastic gradient descent.

5. The Closed-Form Solution (Normal Equation)

Setting the gradient to zero gives θ⋆=(X⊤X)−1X⊤y\theta^\star = (X^\top X)^{-1} X^\top y, also known as the ordinary least squares (OLS) solution. Three properties make it worth knowing.

  1. Exact in one shot. No learning rate, no iteration count, no convergence worries. The output is the unique minimum of the convex MSE surface.
  2. Cheap when dd is small. Cost is dominated by the matrix product X⊤XX^\top X, which is O(nd2)O(n d^2), plus the inversion or solve, which is O(d3)O(d^3). For dd up to a few thousand the closed form is faster than gradient descent to convergence.
  3. Brittle when X⊤XX^\top X is ill-conditioned. Features that are near-collinear make X⊤XX^\top X nearly singular; numerical errors in the inversion dominate. The standard fix is to add a small ridge regulariser, giving the ridge regression solution θ⋆=(X⊤X+λI)−1X⊤y\theta^\star = (X^\top X + \lambda I)^{-1} X^\top y, covered in a later lesson.

In production code the closed form is rarely written by hand. The standard tool is

import numpy as np
from sklearn.linear_model import LinearRegression

model = LinearRegression()
model.fit(x.reshape(-1, 1), y)
print(model.intercept_, model.coef_)   # ~17.74, [25.36]

Both np.linalg.lstsq and scikit-learn's LinearRegression solve the normal equation under the hood; the difference is that scikit-learn also stores metadata, supports multiple features, and plays nicely with the rest of the library's ecosystem.

6. Interpreting Coefficients

Each entry of θ⋆\theta^\star is read as the marginal effect of one feature on the predicted target, holding every other feature fixed.

For the worked example above, θ1=25.36\theta_1 = 25.36 thousand dollars per 100 sq ft means that a 100-sq-ft increase in floor area is associated with a $25{,}360 increase in predicted sale price, all else equal. The intercept θ0=17.74\theta_0 = 17.74 thousand dollars is the prediction when every feature is zero — for floor area, that is the extrapolation to a zero-square-foot "house", which is meaningless as a price but necessary for the algebra to work.

A few rules make coefficient interpretation more reliable.

RuleWhy it matters
Standardise features before fittingA coefficient on a standardised feature is comparable to the others; raw coefficients are not
Drop one feature category per "reference" levelWithout a reference, dummy-encoded categories add a column for every level and the coefficients shift in uninterpretable ways
Watch for confoundingA large positive coefficient can mask a true negative relationship if an omitted variable drives both the feature and the target
Report a confidence interval, not just a point estimateThe point estimate is the best guess, not the true parameter; the interval quantifies that distinction

A prediction interval for a new observation is wider than a confidence interval on the mean prediction: it includes both the uncertainty in θ⋆\theta^\star and the unavoidable noise σ2\sigma^2. Reporting only the point estimate hides this distinction.

7. R-squared and Its Limits

The coefficient of determination is the fraction of variance in yy that the model explains:

R2  =  1−∑i=1n(yi−y^i)2∑i=1n(yi−yˉ)2  =  1−SSresSStot.R^2 \;=\; 1 - \frac{\sum_{i=1}^{n} (y_i - \hat{y}_i)^2}{\sum_{i=1}^{n} (y_i - \bar{y})^2} \;=\; 1 - \frac{\text{SS}_{\text{res}}}{\text{SS}_{\text{tot}}}.

A model that predicts the mean for every observation has R2=0R^2 = 0; a model with zero residual error has R2=1R^2 = 1. For simple linear regression, R2R^2 is also the square of the Pearson correlation between xx and yy.

Three reasons R2R^2 can mislead.

  1. It always rises when a feature is added. Adding a noisy feature to the design matrix increases SSres\text{SS}_{\text{res}} by a tiny amount or not at all, so R2R^2 creeps upward even for useless features. Adjusted R2R^2 penalises the model for each parameter added and should be reported alongside the raw value.
  2. It says nothing about prediction quality on new data. A model with R2=0.95R^2 = 0.95 in-sample can have R2=0.30R^2 = 0.30 out-of-sample if it has overfit. Cross-validated R2R^2 or a held-out test set is the right diagnostic.
  3. It is undefined outside the range of the data. R2R^2 compares to the training mean; if the test distribution has shifted, the right comparison is the test mean, which R2R^2 does not respect.

The right metric depends on the question. MSE and RMSE are on the same scale as yy and are easy to communicate to non-technical stakeholders. MAE is more robust to outliers. MAPE (mean absolute percentage error) is unit-free but blows up when any yiy_i is near zero. Pick the metric that aligns with the cost structure of the application, then report it on a held-out set.

8. The Four Assumptions of Linear Regression

The OLS estimator has known statistical properties — unbiased, minimum variance among linear unbiased estimators — only under four assumptions on the data-generating process. Diagnose each one with a residual plot before trusting a fitted model.

AssumptionWhat it requiresSymptom of violation
LinearityThe conditional expectation E[y∣x]E[y \mid x] is a linear function of xxResiduals vs fitted values show a curved pattern; a parabola, sine wave, or trend in the residuals is a red flag
IndependenceThe residuals εi\varepsilon_i are independent across observationsResiduals are autocorrelated (time series) or clustered (grouped data); Durbin-Watson statistic near 0 or 4 indicates autocorrelation
HomoscedasticityThe residual variance is constant across the range of xxA funnel or fan shape in the residuals vs fitted plot; variance grows or shrinks with y^\hat{y}
Normality of residualsThe residuals are approximately Gaussian, especially for inferenceHistogram or Q-Q plot of residuals shows heavy tails or skew; small-sample confidence intervals become unreliable

When an assumption fails, two broad fixes exist. The first is to transform the data: take logs of a heavy-tailed target, take square roots of count data, or fit the model on a different scale. The second is to change the model: replace ordinary least squares with weighted least squares when heteroscedasticity is the problem, with generalised least squares for correlated errors, or with a tree-based model when the relationship is genuinely non-linear.

A diagnostic routine that catches most issues in five minutes:

  1. Plot residuals vs fitted. Look for curvature (linearity) and funnel shape (homoscedasticity).
  2. Plot residuals vs each feature. Look for structure that the model missed.
  3. Plot the Q-Q plot of residuals. Look for heavy tails (normality).
  4. For time-ordered data, plot residuals vs time and compute Durbin-Watson.
  5. Compute Cook's distance per observation. Points with distance above 4/n4/n are high-leverage outliers worth investigating.

8.1 When Linear Regression Is Not Enough

Linear regression is the wrong tool in three common situations.

  • The relationship is non-linear. Use a basis expansion (polynomial features, splines) or switch to a tree-based model, kernel method, or neural network. Adding polynomial features keeps the linear-regression machinery and gains flexibility, at the cost of more parameters and more overfitting risk.
  • The target is bounded. If yy is a probability, a count, or a non-negative quantity, the Gaussian noise assumption is wrong. Use logistic regression, Poisson regression, or a generalised linear model with the appropriate link.
  • Features are highly correlated. Multicollinearity inflates the variance of the coefficients and makes them uninterpretable. Drop redundant features, combine them into a single index, or use ridge regression to regularise.

Reach for a more flexible model when the bias of linear regression is clearly limiting performance. Reach for a simpler model when the variance is the binding constraint. The art of applied ML is choosing where on the bias-variance curve the problem actually lives.

9. Worked Example: Sales Versus Advertising Spend

A small marketing team has tracked weekly sales (in thousands of dollars) against weekly spending on online ads (in thousands of dollars) for ten weeks. The data:

WeekAd spend xx ($k)Sales yy ($k)
11.232
21.638
32.045
42.450
52.856
63.261
73.667
84.072
94.476
104.881

The relationship looks plausibly linear; the team wants a model to forecast sales given a planned ad budget.

9.1 Computing the coefficients

The required sums are

  • ∑xi=30.0\sum x_i = 30.0, ∑yi=578\sum y_i = 578
  • xˉ=3.0\bar{x} = 3.0, yˉ=57.8\bar{y} = 57.8
  • ∑xiyi=1.2⋅32+1.6⋅38+⋯+4.8⋅81=1,912.0\sum x_i y_i = 1.2 \cdot 32 + 1.6 \cdot 38 + \cdots + 4.8 \cdot 81 = 1{,}912.0
  • ∑xi2=1.44+2.56+4.00+5.76+7.84+10.24+12.96+16.00+19.36+23.04=103.20\sum x_i^2 = 1.44 + 2.56 + 4.00 + 5.76 + 7.84 + 10.24 + 12.96 + 16.00 + 19.36 + 23.04 = 103.20

Plug into the closed-form slope

θ1  =  n∑xiyi−∑xi∑yin∑xi2−(∑xi)2  =  10⋅1912−30⋅57810⋅103.2−302  =  19,120−17,3401032−900  =  1780132  ≈  13.48.\theta_1 \;=\; \frac{n \sum x_i y_i - \sum x_i \sum y_i}{n \sum x_i^2 - (\sum x_i)^2} \;=\; \frac{10 \cdot 1912 - 30 \cdot 578}{10 \cdot 103.2 - 30^2} \;=\; \frac{19{,}120 - 17{,}340}{1032 - 900} \;=\; \frac{1780}{132} \;\approx\; 13.48.

The intercept is

θ0  =  yˉ−θ1xˉ  =  57.8−13.48⋅3.0  =  57.8−40.44  ≈  17.36.\theta_0 \;=\; \bar{y} - \theta_1 \bar{x} \;=\; 57.8 - 13.48 \cdot 3.0 \;=\; 57.8 - 40.44 \;\approx\; 17.36.

The fitted regression line is therefore

y^  =  17.36+13.48 x.\hat{y} \;=\; 17.36 + 13.48 \, x.

Interpretation: every extra $1{,}000 of weekly ad spend is associated with about $13{,}480 of additional weekly sales. The intercept says the baseline sales — with no ad spend — would be about $17{,}360, which captures organic demand and brand recognition.

9.2 Predictions, residuals, and R2R^2

For each of the ten weeks, the prediction and the residual:

WeekxxTrue yyPredicted y^\hat{y}Residual y−y^y - \hat{y}
11.23233.54−1.54-1.54
21.63838.93−0.93-0.93
32.04544.32+0.68+0.68
42.45049.71+0.29+0.29
52.85655.10+0.90+0.90
63.26160.50+0.50+0.50
73.66765.89+1.11+1.11
84.07271.28+0.72+0.72
94.47676.67−0.67-0.67
104.88182.06−1.06-1.06

Residuals oscillate gently around zero, with no obvious pattern — a good sign for the linearity and homoscedasticity assumptions. The residual sum of squares is

SSres  =  ∑(yi−y^i)2  ≈  1.542+0.932+⋯+1.062  ≈  6.92.\text{SS}_{\text{res}} \;=\; \sum (y_i - \hat{y}_i)^2 \;\approx\; 1.54^2 + 0.93^2 + \cdots + 1.06^2 \;\approx\; 6.92.

The total sum of squares, comparing each yiy_i to the mean yˉ=57.8\bar{y} = 57.8, is

SStot  =  ∑(yi−yˉ)2  =  (32−57.8)2+⋯+(81−57.8)2  =  2,425.60.\text{SS}_{\text{tot}} \;=\; \sum (y_i - \bar{y})^2 \;=\; (32 - 57.8)^2 + \cdots + (81 - 57.8)^2 \;=\; 2{,}425.60.

The coefficient of determination is therefore

R2  =  1−6.922425.60  ≈  1−0.00285  ≈  0.9971.R^2 \;=\; 1 - \frac{6.92}{2425.60} \;\approx\; 1 - 0.00285 \;\approx\; 0.9971.

The model explains about 99.7% of the variance in sales, an unusually tight fit. The RMSE, on the scale of yy (thousands of dollars), is

RMSE  =  MSE  =  SSres/n  =  6.92/10  ≈  0.83 thousand dollars,\text{RMSE} \;=\; \sqrt{\text{MSE}} \;=\; \sqrt{\text{SS}_{\text{res}} / n} \;=\; \sqrt{6.92 / 10} \;\approx\; 0.83 \text{ thousand dollars},

so the typical prediction error is about $830 in weekly sales.

9.3 A forecast for next week

If the marketing team plans to spend $5{,}000 on ads next week, the model's forecast is

y^  =  17.36+13.48⋅5.0  =  17.36+67.40  =  84.76,\hat{y} \;=\; 17.36 + 13.48 \cdot 5.0 \;=\; 17.36 + 67.40 \;=\; 84.76,

or roughly $84{,}760 in sales. The right way to communicate this to the team is "around $85k, give or take $1k, assuming the relationship we have seen continues to hold for this budget level". The "give or take" is the residual standard error; the "assuming" is the linearity assumption; both belong in the conversation.

9.4 Cross-check with scikit-learn

The whole computation in twenty lines of Python:

import numpy as np
from sklearn.linear_model import LinearRegression

x = np.array([1.2, 1.6, 2.0, 2.4, 2.8, 3.2, 3.6, 4.0, 4.4, 4.8]).reshape(-1, 1)
y = np.array([32, 38, 45, 50, 56, 61, 67, 72, 76, 81], dtype=float)

model = LinearRegression().fit(x, y)
print(f"intercept = {model.intercept_:.2f}, slope = {model.coef_[0]:.2f}")

r_squared = model.score(x, y)
print(f"R^2 = {r_squared:.4f}")

y_pred = model.predict(x)
rmse = float(np.sqrt(np.mean((y - y_pred) ** 2)))
print(f"RMSE = {rmse:.2f} thousand dollars")

forecast = model.predict(np.array([[5.0]]))
print(f"forecast for $5k ad spend: ${float(forecast[0]):.2f}k")

The four printed lines match the hand calculation. In production code the script would split into train and test sets, report metrics on the test set rather than on the training set, and use np.linalg.lstsq or model.predict inside a pipeline that handles feature scaling, missing-value imputation, and cross-validation.

9.5 Sanity checks before shipping the model

The coefficients are credible. The residuals look like noise. R2R^2 is high. Before declaring victory, three further checks belong in the workflow.

  1. Residual plot. Plot y−y^y - \hat{y} against y^\hat{y}. A random cloud around zero confirms linearity and homoscedasticity; a curved shape means a non-linear feature is missing; a fan shape means the variance grows with the prediction.
  2. Out-of-sample test. Hold out the last two weeks, refit on the first eight, and score on the held-out two. A test R2R^2 close to the training R2R^2 is the proof that the model generalises.
  3. Domain plausibility. Is a 13.5:1 return on ad spend realistic for this business? If the team's historical return on ads is closer to 5:1, the model is probably picking up a confound (a new product launch that drove both sales and ad spend upward, for example). The numbers should pass a sniff test from a domain expert before they pass to a stakeholder.

A model that is statistically correct but operationally wrong is a worse outcome than no model at all: it gives false confidence to a decision that should have been made by other means.

Key Takeaways

  • The linear-regression hypothesis y^=θ⊤x\hat{y} = \theta^\top x is a linear combination of features fit by minimising the mean-squared-error loss; "linear" means linear in the parameters, not necessarily in the features.
  • The gradient of the MSE is ∇θJ=(2/n)X⊤(Xθ−y)\nabla_\theta J = (2/n) X^\top (X\theta - y); setting it to zero gives the normal equation θ⋆=(X⊤X)−1X⊤y\theta^\star = (X^\top X)^{-1} X^\top y, the closed-form OLS solution.
  • Gradient descent iteratively subtracts a learning-rate-scaled gradient from the parameters; feature scaling and a sensible learning rate are the keys to fast convergence.
  • Each fitted coefficient is the marginal effect of one feature on the predicted target, holding the others fixed. Report confidence intervals alongside the point estimates.
  • R2R^2 is the fraction of variance explained; it rises whenever a feature is added (use adjusted R2R^2 to compare models) and does not measure out-of-sample performance on its own.
  • The four classical assumptions are linearity, independence of residuals, constant variance (homoscedasticity), and approximate normality of residuals; diagnose each with a residual plot before trusting the model.
  • The worked sales-versus-ad-spend example fits y^=17.36+13.48 x\hat{y} = 17.36 + 13.48\,x with R2≈0.997R^2 \approx 0.997, RMSE of about $830, and a forecast of about $85k for the next week given a $5k ad budget.

Check your understanding

8 questions · 80% to complete the lesson

1 / 8

7 correct to pass

In the linear regression hypothesis $\hat{y} = \theta^\top x$, what does the leading entry of $x$ represent by convention?

0 of 8 answered

Pick a lesson to start the audio.