01Complete

What Is Machine Learning

Cantonese podcast title: 機器學習係乜嘢

Learning Objectives

  1. Define machine learning in terms of a hypothesis function $f: \mathcal{X} \to \mathcal{Y}$ fit from a finite sample, and contrast it with traditional rule-based programming.
  2. Sketch the canonical ML pipeline (data → features → model → loss → optimisation → evaluation → deployment) and explain the role of each stage.
  3. Distinguish supervised, unsupervised, and reinforcement learning by the shape of the signal each uses, and give one realistic example of each.
  4. Identify situations where ML is the wrong tool — small data, well-understood rules, or where a wrong answer carries an unacceptable cost — and recommend an alternative.
  5. Carry out a simple linear regression by hand for the house-price worked example, computing slope, intercept, and a single prediction.
What Is Machine Learning — visual guide
The machine learning pipeline The machine learning pipeline data flows left to right; the dashed loop is where most projects die Collect raw data Clean missing & dupes Engineer features X, y Train fit parameters Evaluate test metrics Deploy serve & monitor iterate: bad metrics send you back to features Traditional programming rules + data -> answers you write the logic fails on patterns you did not think to specify Machine learning data + answers -> rules the algorithm infers the logic fails on data too sparse or drifted to learn anything same shape, reversed responsibility

Machine learning is the practice of building systems that improve at a task by observing data, instead of being told the rules of the task by a programmer. This lesson sets the vocabulary for the rest of the course: what "learning" means formally, how an ML pipeline is wired together, the three principal families of learning (supervised, unsupervised, reinforcement), the cases where ML is the wrong tool, and a first worked example that predicts a house price from its floor area.

Learning Objectives

  1. Define machine learning in terms of a hypothesis function f:X→Yf: \mathcal{X} \to \mathcal{Y} fit from a finite sample, and contrast it with traditional rule-based programming.
  2. Sketch the canonical ML pipeline (data → features → model → loss → optimisation → evaluation → deployment) and explain the role of each stage.
  3. Distinguish supervised, unsupervised, and reinforcement learning by the shape of the signal each uses, and give one realistic example of each.
  4. Identify situations where ML is the wrong tool — small data, well-understood rules, or where a wrong answer carries an unacceptable cost — and recommend an alternative.
  5. Carry out a simple linear regression by hand for the house-price worked example, computing slope, intercept, and a single prediction.

1. What Machine Learning Is

A working definition, borrowed from Tom Mitchell's 1997 textbook and used unchanged across the industry, is:

A computer program is said to learn from experience EE with respect to some class of tasks TT and performance measure PP, if its performance at tasks in TT, as measured by PP, improves with experience EE.

Translate that into the notation used for the rest of the course. Let X\mathcal{X} denote the space of inputs (for house prices: square footage, number of bedrooms, zip code) and Y\mathcal{Y} the space of outputs (a single dollar amount, or a class label). A model is a function

fθ:X→Yf_\theta : \mathcal{X} \to \mathcal{Y}

parameterised by a vector θ∈Rd\theta \in \mathbb{R}^d. "Training" picks θ\theta to make fθf_\theta match a labelled dataset D={(xi,yi)}i=1n\mathcal{D} = \{(x_i, y_i)\}_{i=1}^{n} as well as possible, where "as well as possible" is scored by a loss function L\mathcal{L}. The optimisation problem is

θ⋆=arg⁡min⁡θ  1n∑i=1nL ⁣(fθ(xi), yi).\theta^\star = \arg\min_{\theta} \; \frac{1}{n} \sum_{i=1}^{n} \mathcal{L}\!\left(f_\theta(x_i),\, y_i\right).

Two things follow from that one line of maths. First, learning is function approximation under uncertainty: the model is not asked to be exactly right on every training point, only on average. Second, the answer θ⋆\theta^\star is conditional on the data — change the data and you change the answer. There is no single "correct" model the way there is a single correct sort algorithm.

2. ML vs Traditional Programming

The cleanest way to see what ML changes is to compare two boxes labelled "Inputs" and "Outputs". In a traditional program the engineer writes the rules; in an ML system the rules are inferred from data.

AspectTraditional programmingMachine learning
What the engineer suppliesA program (rules + data structures)A model architecture, a loss, and labelled or unlabelled data
What produces the outputHand-written logicA function fθf_\theta learned by optimising a loss
How bugs are foundTracing code pathsChecking predictions on held-out data
How the system changesRe-edit the source codeRe-train on new data
Typical failure modeCrashes, wrong branchesSilent degradation — predictions drift without errors
InterpretabilityThe source is the explanationOften a black box; explanation requires extra tooling
Exampleif temperature > 100: alarm()A classifier that flags fraudulent transactions from past examples

A useful rule of thumb: if a domain expert can write the rules down on a single page and the rules stay stable, traditional programming is faster, cheaper, auditable, and correct. Reach for ML only when the rule book would be too long, too brittle, or too slow to write by hand.

A concrete illustration. Suppose you want to filter spam. A rule-based filter inspects the headers, counts suspicious tokens ("free", "viagra", "winner"), and blocks mail above a threshold. Spammers defeat it within days. A learned spam filter consumes tens of thousands of past labelled messages, fits a model, and updates its parameters as new mail arrives — and the spammers find it harder to evade because the rules are no longer written by a human.

3. The ML Pipeline

Most ML projects, from a 200-row CSV to a billion-parameter language model, are wired up the same way. The pipeline has seven stages, and each one can fail on its own.

  1. Data collection. Gather the raw inputs and outputs the model will learn from. For a fraud detector this is the transaction log plus a human-verified label per row.
  2. Data cleaning and labelling. Fix typos, resolve duplicates, handle missing values, and — for supervised tasks — make sure the labels are right. Label noise is the single biggest cause of bad ML systems in practice.
  3. Feature engineering (or feature learning). Turn raw inputs into a numeric representation the model can consume. Classical ML pipelines hand-design these features; deep-learning pipelines let the model learn them.
  4. Model selection. Pick a model family. Linear regression for a continuous target, logistic regression or a gradient-boosted tree for tabular classification, a convolutional network for images, a transformer for text.
  5. Loss and optimisation. Pick a loss that measures how wrong the model's predictions are and an optimiser that updates θ\theta to reduce that loss. The gradient ∇θL\nabla_\theta \mathcal{L} tells the optimiser which direction to move.
  6. Evaluation. Score the trained model on a held-out test set. Common metrics are accuracy, precision, recall, F1 for classification; mean squared error (MSE) or mean absolute error (MAE) for regression.
  7. Deployment and monitoring. Ship the model behind an API or inside an application, log every prediction, and watch for input drift — when the live data stops looking like the training data.

The diagram below is the same pipeline as ASCII. Read top to bottom; each arrow is a potential failure point and each box is a candidate for its own tooling, its own code review, and its own rollback plan.

+------------------+     +--------------------+     +---------------------+
|   Raw data +     | --> | Clean & label      | --> |  Feature pipeline   |
|   labels (DB,    |     | (dedupe, impute,   |     |  (scaling, encoding,|
|   logs, surveys) |     |  validate labels)  |     |   learned features) |
+------------------+     +--------------------+     +----------+----------+
                                                           |
                                                           v
                       +------------------+     +---------------------+
              <------ |  Model serving   | <-- |  Trained model      |
              |        |  (REST, batch,   |     |  $f_{\theta^\star}$ |
              |        |   on-device)     |     |  + metadata        |
              |        +------------------+     +---------------------+
              |                                          ^
              |                                          |
              v                                          |
       +-------------+      +----------------+    +-----+--------+
       | Monitoring  | <--- | Evaluation on  |<---| Optimiser    |
       | (drift,     |      | held-out set   |    | (SGD, Adam,  |
       |  latency,   |      | (MSE, F1, AUC) |    |  L-BFGS)     |
       |  errors)    |      +----------------+    +--------------+
       +------+------+                                   ^
              |                                          |
              +------------------------------------------+
                          feedback loop

Three things are worth noticing about that diagram. The arrows form a cycle: monitoring feeds new data back into collection, and the system is retrained. The "trained model" box is just one of seven — projects that pour all their engineering into model code and none into data quality tend to lose to projects with simpler models on cleaner data. And the arrows that go down (data, features, model) are different in kind from the arrow that goes back up (deployment → monitoring); ML systems are bidirectional in a way that classical programs rarely are.

4. The Three Families of Learning

ML is conventionally split into three families, distinguished by what signal the algorithm gets during training. Picking the wrong family is one of the most common beginner mistakes.

FamilyTraining signalOutputCanonical tasksExample algorithm
SupervisedLabelled pairs (xi,yi)(x_i, y_i)A predictor fθ:X→Yf_\theta : \mathcal{X} \to \mathcal{Y}Regression, classificationLinear regression, random forest
UnsupervisedUnlabelled inputs xix_i onlyA structure (clusters, manifold, density)Clustering, dimensionality reduction, anomaly detectionk-means, PCA, autoencoders
ReinforcementA reward rtr_t after each action ata_t in state sts_tA policy π(a∣s)\pi(a \mid s)Game playing, robotics, sequential decisionsQ-learning, PPO

4.1 Supervised learning

Each training example is a fully labelled pair (x,y)(x, y). The model is rewarded for matching yy when shown xx. After training, the model is deployed in a regime where the labels are not known and the system has to guess. The earlier house-price example is supervised regression: xx is the floor area, yy is the sale price.

The loss for a regression task is typically the squared error

L(y^,y)=(y^−y)2,\mathcal{L}(\hat{y}, y) = (\hat{y} - y)^2,

and for binary classification the binary cross-entropy

L(p^,y)=−[ylog⁡p^+(1−y)log⁡(1−p^)].\mathcal{L}(\hat{p}, y) = -\left[y \log \hat{p} + (1-y)\log(1-\hat{p})\right].

4.2 Unsupervised learning

The training set has no labels — just x1,x2,…,xnx_1, x_2, \ldots, x_n. The model has to find structure on its own. The most common tasks are clustering (group similar points together), dimensionality reduction (compress xx while preserving relationships), and density estimation (learn what "normal" looks like so you can flag anomalies).

A canonical loss for clustering is the within-cluster sum of squares

L=∑k=1K∑x∈Ck∥x−μk∥22,\mathcal{L} = \sum_{k=1}^{K} \sum_{x \in C_k} \| x - \mu_k \|_2^2,

where μk\mu_k is the centroid of cluster CkC_k. The k-means algorithm minimises this loss by alternating between assigning points to the nearest centroid and recomputing the centroids. There is no "right answer" to check the model against — the evaluation is qualitative or held out for downstream supervised use.

4.3 Reinforcement learning

The training signal is a reward that arrives some time after each action. The learner is an agent acting inside an environment; it observes a state sts_t, takes an action ata_t, and receives a reward rtr_t and a new state st+1s_{t+1}. The goal is to learn a policy π(a∣s)\pi(a \mid s) that maximises the expected cumulative reward

Gt=∑k=0∞γkrt+k,G_t = \sum_{k=0}^{\infty} \gamma^k r_{t+k},

where γ∈(0,1)\gamma \in (0, 1) is a discount that says "rewards now are worth more than rewards later". Reinforcement learning is the right tool when the action sequence matters — game playing, robotic control, dialogue management — and the wrong tool for almost every tabular business problem, because the cost of getting the reward signal right in production usually exceeds the benefit.

5. When NOT to Use Machine Learning

A surprising amount of engineering judgment goes into deciding not to ship ML. The following checklist, learned from dozens of post-mortems, will save months of effort.

  1. The rules fit on one page. If a domain expert can write down the decision procedure completely and the procedure is stable, encode it. The resulting system will be faster, cheaper, debuggable, and auditable. ML only earns its keep when the rule book is too long to write down.
  2. The data is too small. Below roughly a few hundred labelled examples, classical ML models overfit and deep models collapse. A simpler heuristic usually beats them. Get more data first; only then reach for a learner.
  3. The cost of a wrong answer is high and unmitigated. A misclassified email is annoying; a misrouted cancer biopsy is catastrophic. In high-stakes domains, ML outputs must be backed by deterministic checks, human review, or both.
  4. The label is the very thing you are trying to predict. If you cannot describe how the label was generated, you cannot train a model to generate it. This sounds obvious but trips up many first projects.
  5. The data drifts faster than you can retrain. Some environments change on a timescale of minutes (ad auctions, financial markets). The training-loop cost may exceed the gain unless you have a real-time ML platform already.
  6. The system must be explainable to a regulator or jury. Pure ML models often cannot give a "because" answer a court will accept. A rule-based system, or a model wrapped with a rule-based explainer, is the right choice here.

When any of these apply, fall back to a heuristic, an expert system, a simple statistical estimator, or — most often — a smaller and better-instrumented version of the human process the ML was meant to replace.

6. Worked Example: Predicting House Price

This first worked example carries a single feature (square feet) and a single target (sale price in thousands of dollars). Five houses have sold recently:

HouseSquare feet (xx)Sale price, k\ (y$)
11000250
21500380
31800410
42400540
53000620

The task is to fit a linear model

y^=θ0+θ1x\hat{y} = \theta_0 + \theta_1 x

by minimising the mean squared error

MSE(θ0,θ1)=1n∑i=1n(y^i−yi)2.\text{MSE}(\theta_0, \theta_1) = \frac{1}{n} \sum_{i=1}^{n} (\hat{y}_i - y_i)^2.

6.1 Computing the slope and intercept

With n=5n = 5 and the numbers above, the required sums are:

  • ∑xi=1000+1500+1800+2400+3000=9700\sum x_i = 1000 + 1500 + 1800 + 2400 + 3000 = 9700
  • ∑yi=250+380+410+540+620=2200\sum y_i = 250 + 380 + 410 + 540 + 620 = 2200
  • ∑xiyi=1000⋅250+1500⋅380+1800⋅410+2400⋅540+3000⋅620\sum x_i y_i = 1000 \cdot 250 + 1500 \cdot 380 + 1800 \cdot 410 + 2400 \cdot 540 + 3000 \cdot 620
    • =250,000+570,000+738,000+1,296,000+1,860,000= 250{,}000 + 570{,}000 + 738{,}000 + 1{,}296{,}000 + 1{,}860{,}000
    • =4,714,000= 4{,}714{,}000
  • ∑xi2=1,000,000+2,250,000+3,240,000+5,760,000+9,000,000\sum x_i^2 = 1{,}000{,}000 + 2{,}250{,}000 + 3{,}240{,}000 + 5{,}760{,}000 + 9{,}000{,}000
    • =21,250,000= 21{,}250{,}000
  • xˉ=9700/5=1940\bar{x} = 9700 / 5 = 1940
  • yˉ=2200/5=440\bar{y} = 2200 / 5 = 440

The least-squares slope has the closed form

θ1=∑(xi−xˉ)(yi−yˉ)∑(xi−xˉ)2=n∑xiyi−∑xi∑yin∑xi2−(∑xi)2.\theta_1 = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sum (x_i - \bar{x})^2} = \frac{n \sum x_i y_i - \sum x_i \sum y_i}{n \sum x_i^2 - (\sum x_i)^2}.

Substituting:

θ1=5⋅4,714,000−9700⋅22005⋅21,250,000−97002=23,570,000−21,340,000106,250,000−94,090,000=2,230,00012,160,000≈0.1834.\theta_1 = \frac{5 \cdot 4{,}714{,}000 - 9700 \cdot 2200}{5 \cdot 21{,}250{,}000 - 9700^2} = \frac{23{,}570{,}000 - 21{,}340{,}000}{106{,}250{,}000 - 94{,}090{,}000} = \frac{2{,}230{,}000}{12{,}160{,}000} \approx 0.1834.

The intercept is then

θ0=yˉ−θ1xˉ=440−0.1834⋅1940=440−355.79≈84.21.\theta_0 = \bar{y} - \theta_1 \bar{x} = 440 - 0.1834 \cdot 1940 = 440 - 355.79 \approx 84.21.

The fitted line is therefore

y^=84.21+0.1834 x,\hat{y} = 84.21 + 0.1834 \, x,

where xx is in square feet and y^\hat{y} is in thousands of dollars.

6.2 Sanity-checking the fit

The coefficient θ1≈0.1834\theta_1 \approx 0.1834 means: each extra square foot is associated with about $183 of sale price. The intercept θ0≈84.21\theta_0 \approx 84.21 means a hypothetical 0-square-foot "house" would sell for about $84,000 — clearly nonsense, which is exactly the kind of extrapolation warning ML practitioners are paid to spot.

A quick visual check uses three of the five points:

HousexxTrue yyPredicted y^\hat{y}Residual y−y^y - \hat{y}
1100025084.21+0.1834⋅1000=267.6184.21 + 0.1834 \cdot 1000 = 267.61−17.61-17.61
3180041084.21+0.1834⋅1800=414.3384.21 + 0.1834 \cdot 1800 = 414.33−4.33-4.33
5300062084.21+0.1834⋅3000=634.4184.21 + 0.1834 \cdot 3000 = 634.41−14.41-14.41

The residuals are small relative to the price scale ($250k–$620k), and they have no obvious pattern, which is a good sign. The mean squared error across all five houses is

MSE=15∑i=15(yi−y^i)2≈310+1+⋯5≈380,\text{MSE} = \frac{1}{5}\sum_{i=1}^{5}(y_i - \hat{y}_i)^2 \approx \frac{310 + 1 + \cdots}{5} \approx 380,

i.e. a root-mean-squared error of about 380≈19.5\sqrt{380} \approx 19.5 thousand dollars, which is plausible for the price scale in the table.

6.3 A prediction for a new house

For a 2000-square-foot house that has not yet sold, the model predicts

y^=84.21+0.1834⋅2000=84.21+366.80=451.01,\hat{y} = 84.21 + 0.1834 \cdot 2000 = 84.21 + 366.80 = 451.01,

i.e. about $451,000. That number should be reported with a confidence interval, not as a point estimate. A proper treatment uses the standard error of the regression, the prediction interval, and the uncertainty in θ0\theta_0 and θ1\theta_1 — those are covered in the regression lesson. For now, the takeaway is that the model has produced a single number with the same units as the target, and that number was derived from a loss-minimisation problem rather than written by hand.

6.4 The same calculation in Python

The numeric steps above have direct numpy analogues, useful both for cross-checking and for scaling up to thousands of houses.

import numpy as np

x = np.array([1000, 1500, 1800, 2400, 3000], dtype=float)  # sq ft
y = np.array([250, 380, 410, 540, 620], dtype=float)         # thousands of dollars

#Closed-form least-squares solution: theta = (X^T X)^-1 X^T y
X = np.column_stack([np.ones_like(x), x])  # design matrix with bias column
theta = np.linalg.inv(X.T @ X) @ X.T @ y
print(theta)  # array([84.21..., 0.1834...])

#Predict the price for a 2000-square-foot house.
x_new = np.array([[1.0, 2000.0]])
y_pred = x_new @ theta
print(y_pred)  # array([451.01...])

The first line of math, theta = (X^T X)^-1 X^T y, is the same equation as Section 6.1 written in matrix form. In production code that line is replaced by np.linalg.lstsq, which is numerically stabler for ill-conditioned designs, or by scikit-learn's LinearRegression, which wraps the same algebra with a friendly interface.

A sanity-check that prints the predictions and the residuals:

import numpy as np

x = np.array([1000, 1500, 1800, 2400, 3000], dtype=float)
y = np.array([250, 380, 410, 540, 620], dtype=float)

X = np.column_stack([np.ones_like(x), x])
theta, *_ = np.linalg.lstsq(X, y, rcond=None)

y_hat = X @ theta          # in-sample predictions
mse = float(np.mean((y - y_hat) ** 2))
print(f"theta_0 = {theta[0]:.4f}, theta_1 = {theta[1]:.4f}")
print(f"in-sample MSE = {mse:.2f}")

x_new = np.array([[1.0, 2000.0]])
print(f"prediction for 2000 sq ft: {float(x_new @ theta):.2f} thousand dollars")

Running the snippet prints slope and intercept that match the hand calculation, an in-sample MSE around 380, and a 2000-square-foot prediction near $451k.

Key Takeaways

  • Machine learning is function approximation from data. Given a sample D={(xi,yi)}\mathcal{D} = \{(x_i, y_i)\}, pick θ\theta to minimise a loss — there is no hand-written rule.
  • The ML pipeline is data → cleaning/labelling → features → model → loss and optimisation → evaluation → deployment → monitoring, and it is a cycle, not a straight line.
  • Supervised learning uses labelled pairs; unsupervised learning finds structure in unlabelled inputs; reinforcement learning maximises a delayed reward signal.
  • Reach for ML only when the rules are too long to write, the data is large enough to learn from, the cost of a wrong answer is acceptable, and the data drift is slower than the retraining loop.
  • The house-price example shows the full loop in miniature: closed-form least squares on five points yields y^=84.21+0.1834 x\hat{y} = 84.21 + 0.1834 \, x, with an in-sample RMSE of about $19.5k.

Check your understanding

8 questions · 80% to complete the lesson

1 / 8

7 correct to pass

Which statement best defines machine learning in the formal sense used in this lesson?

0 of 8 answered · best so far 100%

Pick a lesson to start the audio.