Machine learning is the practice of building systems that improve at a task by observing data, instead of being told the rules of the task by a programmer. This lesson sets the vocabulary for the rest of the course: what "learning" means formally, how an ML pipeline is wired together, the three principal families of learning (supervised, unsupervised, reinforcement), the cases where ML is the wrong tool, and a first worked example that predicts a house price from its floor area.
Learning Objectives
- Define machine learning in terms of a hypothesis function fit from a finite sample, and contrast it with traditional rule-based programming.
- Sketch the canonical ML pipeline (data → features → model → loss → optimisation → evaluation → deployment) and explain the role of each stage.
- Distinguish supervised, unsupervised, and reinforcement learning by the shape of the signal each uses, and give one realistic example of each.
- Identify situations where ML is the wrong tool — small data, well-understood rules, or where a wrong answer carries an unacceptable cost — and recommend an alternative.
- Carry out a simple linear regression by hand for the house-price worked example, computing slope, intercept, and a single prediction.
1. What Machine Learning Is
A working definition, borrowed from Tom Mitchell's 1997 textbook and used unchanged across the industry, is:
A computer program is said to learn from experience with respect to some class of tasks and performance measure , if its performance at tasks in , as measured by , improves with experience .
Translate that into the notation used for the rest of the course. Let denote the space of inputs (for house prices: square footage, number of bedrooms, zip code) and the space of outputs (a single dollar amount, or a class label). A model is a function
parameterised by a vector . "Training" picks to make match a labelled dataset as well as possible, where "as well as possible" is scored by a loss function . The optimisation problem is
Two things follow from that one line of maths. First, learning is function approximation under uncertainty: the model is not asked to be exactly right on every training point, only on average. Second, the answer is conditional on the data — change the data and you change the answer. There is no single "correct" model the way there is a single correct sort algorithm.
2. ML vs Traditional Programming
The cleanest way to see what ML changes is to compare two boxes labelled "Inputs" and "Outputs". In a traditional program the engineer writes the rules; in an ML system the rules are inferred from data.
| Aspect | Traditional programming | Machine learning |
|---|---|---|
| What the engineer supplies | A program (rules + data structures) | A model architecture, a loss, and labelled or unlabelled data |
| What produces the output | Hand-written logic | A function learned by optimising a loss |
| How bugs are found | Tracing code paths | Checking predictions on held-out data |
| How the system changes | Re-edit the source code | Re-train on new data |
| Typical failure mode | Crashes, wrong branches | Silent degradation — predictions drift without errors |
| Interpretability | The source is the explanation | Often a black box; explanation requires extra tooling |
| Example | if temperature > 100: alarm() | A classifier that flags fraudulent transactions from past examples |
A useful rule of thumb: if a domain expert can write the rules down on a single page and the rules stay stable, traditional programming is faster, cheaper, auditable, and correct. Reach for ML only when the rule book would be too long, too brittle, or too slow to write by hand.
A concrete illustration. Suppose you want to filter spam. A rule-based filter inspects the headers, counts suspicious tokens ("free", "viagra", "winner"), and blocks mail above a threshold. Spammers defeat it within days. A learned spam filter consumes tens of thousands of past labelled messages, fits a model, and updates its parameters as new mail arrives — and the spammers find it harder to evade because the rules are no longer written by a human.
3. The ML Pipeline
Most ML projects, from a 200-row CSV to a billion-parameter language model, are wired up the same way. The pipeline has seven stages, and each one can fail on its own.
- Data collection. Gather the raw inputs and outputs the model will learn from. For a fraud detector this is the transaction log plus a human-verified label per row.
- Data cleaning and labelling. Fix typos, resolve duplicates, handle missing values, and — for supervised tasks — make sure the labels are right. Label noise is the single biggest cause of bad ML systems in practice.
- Feature engineering (or feature learning). Turn raw inputs into a numeric representation the model can consume. Classical ML pipelines hand-design these features; deep-learning pipelines let the model learn them.
- Model selection. Pick a model family. Linear regression for a continuous target, logistic regression or a gradient-boosted tree for tabular classification, a convolutional network for images, a transformer for text.
- Loss and optimisation. Pick a loss that measures how wrong the model's predictions are and an optimiser that updates to reduce that loss. The gradient tells the optimiser which direction to move.
- Evaluation. Score the trained model on a held-out test set. Common metrics are accuracy, precision, recall, F1 for classification; mean squared error (MSE) or mean absolute error (MAE) for regression.
- Deployment and monitoring. Ship the model behind an API or inside an application, log every prediction, and watch for input drift — when the live data stops looking like the training data.
The diagram below is the same pipeline as ASCII. Read top to bottom; each arrow is a potential failure point and each box is a candidate for its own tooling, its own code review, and its own rollback plan.
+------------------+ +--------------------+ +---------------------+
| Raw data + | --> | Clean & label | --> | Feature pipeline |
| labels (DB, | | (dedupe, impute, | | (scaling, encoding,|
| logs, surveys) | | validate labels) | | learned features) |
+------------------+ +--------------------+ +----------+----------+
|
v
+------------------+ +---------------------+
<------ | Model serving | <-- | Trained model |
| | (REST, batch, | | $f_{\theta^\star}$ |
| | on-device) | | + metadata |
| +------------------+ +---------------------+
| ^
| |
v |
+-------------+ +----------------+ +-----+--------+
| Monitoring | <--- | Evaluation on |<---| Optimiser |
| (drift, | | held-out set | | (SGD, Adam, |
| latency, | | (MSE, F1, AUC) | | L-BFGS) |
| errors) | +----------------+ +--------------+
+------+------+ ^
| |
+------------------------------------------+
feedback loop
Three things are worth noticing about that diagram. The arrows form a cycle: monitoring feeds new data back into collection, and the system is retrained. The "trained model" box is just one of seven — projects that pour all their engineering into model code and none into data quality tend to lose to projects with simpler models on cleaner data. And the arrows that go down (data, features, model) are different in kind from the arrow that goes back up (deployment → monitoring); ML systems are bidirectional in a way that classical programs rarely are.
4. The Three Families of Learning
ML is conventionally split into three families, distinguished by what signal the algorithm gets during training. Picking the wrong family is one of the most common beginner mistakes.
| Family | Training signal | Output | Canonical tasks | Example algorithm |
|---|---|---|---|---|
| Supervised | Labelled pairs | A predictor | Regression, classification | Linear regression, random forest |
| Unsupervised | Unlabelled inputs only | A structure (clusters, manifold, density) | Clustering, dimensionality reduction, anomaly detection | k-means, PCA, autoencoders |
| Reinforcement | A reward after each action in state | A policy | Game playing, robotics, sequential decisions | Q-learning, PPO |
4.1 Supervised learning
Each training example is a fully labelled pair . The model is rewarded for matching when shown . After training, the model is deployed in a regime where the labels are not known and the system has to guess. The earlier house-price example is supervised regression: is the floor area, is the sale price.
The loss for a regression task is typically the squared error
and for binary classification the binary cross-entropy
4.2 Unsupervised learning
The training set has no labels — just . The model has to find structure on its own. The most common tasks are clustering (group similar points together), dimensionality reduction (compress while preserving relationships), and density estimation (learn what "normal" looks like so you can flag anomalies).
A canonical loss for clustering is the within-cluster sum of squares
where is the centroid of cluster . The k-means algorithm minimises this loss by alternating between assigning points to the nearest centroid and recomputing the centroids. There is no "right answer" to check the model against — the evaluation is qualitative or held out for downstream supervised use.
4.3 Reinforcement learning
The training signal is a reward that arrives some time after each action. The learner is an agent acting inside an environment; it observes a state , takes an action , and receives a reward and a new state . The goal is to learn a policy that maximises the expected cumulative reward
where is a discount that says "rewards now are worth more than rewards later". Reinforcement learning is the right tool when the action sequence matters — game playing, robotic control, dialogue management — and the wrong tool for almost every tabular business problem, because the cost of getting the reward signal right in production usually exceeds the benefit.
5. When NOT to Use Machine Learning
A surprising amount of engineering judgment goes into deciding not to ship ML. The following checklist, learned from dozens of post-mortems, will save months of effort.
- The rules fit on one page. If a domain expert can write down the decision procedure completely and the procedure is stable, encode it. The resulting system will be faster, cheaper, debuggable, and auditable. ML only earns its keep when the rule book is too long to write down.
- The data is too small. Below roughly a few hundred labelled examples, classical ML models overfit and deep models collapse. A simpler heuristic usually beats them. Get more data first; only then reach for a learner.
- The cost of a wrong answer is high and unmitigated. A misclassified email is annoying; a misrouted cancer biopsy is catastrophic. In high-stakes domains, ML outputs must be backed by deterministic checks, human review, or both.
- The label is the very thing you are trying to predict. If you cannot describe how the label was generated, you cannot train a model to generate it. This sounds obvious but trips up many first projects.
- The data drifts faster than you can retrain. Some environments change on a timescale of minutes (ad auctions, financial markets). The training-loop cost may exceed the gain unless you have a real-time ML platform already.
- The system must be explainable to a regulator or jury. Pure ML models often cannot give a "because" answer a court will accept. A rule-based system, or a model wrapped with a rule-based explainer, is the right choice here.
When any of these apply, fall back to a heuristic, an expert system, a simple statistical estimator, or — most often — a smaller and better-instrumented version of the human process the ML was meant to replace.
6. Worked Example: Predicting House Price
This first worked example carries a single feature (square feet) and a single target (sale price in thousands of dollars). Five houses have sold recently:
| House | Square feet () | Sale price, k\ (y$) |
|---|---|---|
| 1 | 1000 | 250 |
| 2 | 1500 | 380 |
| 3 | 1800 | 410 |
| 4 | 2400 | 540 |
| 5 | 3000 | 620 |
The task is to fit a linear model
by minimising the mean squared error
6.1 Computing the slope and intercept
With and the numbers above, the required sums are:
The least-squares slope has the closed form
Substituting:
The intercept is then
The fitted line is therefore
where is in square feet and is in thousands of dollars.
6.2 Sanity-checking the fit
The coefficient means: each extra square foot is associated with about $183 of sale price. The intercept means a hypothetical 0-square-foot "house" would sell for about $84,000 — clearly nonsense, which is exactly the kind of extrapolation warning ML practitioners are paid to spot.
A quick visual check uses three of the five points:
| House | True | Predicted | Residual | |
|---|---|---|---|---|
| 1 | 1000 | 250 | ||
| 3 | 1800 | 410 | ||
| 5 | 3000 | 620 |
The residuals are small relative to the price scale ($250k–$620k), and they have no obvious pattern, which is a good sign. The mean squared error across all five houses is
i.e. a root-mean-squared error of about thousand dollars, which is plausible for the price scale in the table.
6.3 A prediction for a new house
For a 2000-square-foot house that has not yet sold, the model predicts
i.e. about $451,000. That number should be reported with a confidence interval, not as a point estimate. A proper treatment uses the standard error of the regression, the prediction interval, and the uncertainty in and — those are covered in the regression lesson. For now, the takeaway is that the model has produced a single number with the same units as the target, and that number was derived from a loss-minimisation problem rather than written by hand.
6.4 The same calculation in Python
The numeric steps above have direct numpy analogues, useful both for cross-checking and for scaling up to thousands of houses.
import numpy as np
x = np.array([1000, 1500, 1800, 2400, 3000], dtype=float) # sq ft
y = np.array([250, 380, 410, 540, 620], dtype=float) # thousands of dollars
#Closed-form least-squares solution: theta = (X^T X)^-1 X^T y
X = np.column_stack([np.ones_like(x), x]) # design matrix with bias column
theta = np.linalg.inv(X.T @ X) @ X.T @ y
print(theta) # array([84.21..., 0.1834...])
#Predict the price for a 2000-square-foot house.
x_new = np.array([[1.0, 2000.0]])
y_pred = x_new @ theta
print(y_pred) # array([451.01...])
The first line of math, theta = (X^T X)^-1 X^T y, is the same equation as
Section 6.1 written in matrix form. In production code that line is replaced by
np.linalg.lstsq, which is numerically stabler for ill-conditioned designs, or by
scikit-learn's LinearRegression, which wraps the same algebra with a friendly
interface.
A sanity-check that prints the predictions and the residuals:
import numpy as np
x = np.array([1000, 1500, 1800, 2400, 3000], dtype=float)
y = np.array([250, 380, 410, 540, 620], dtype=float)
X = np.column_stack([np.ones_like(x), x])
theta, *_ = np.linalg.lstsq(X, y, rcond=None)
y_hat = X @ theta # in-sample predictions
mse = float(np.mean((y - y_hat) ** 2))
print(f"theta_0 = {theta[0]:.4f}, theta_1 = {theta[1]:.4f}")
print(f"in-sample MSE = {mse:.2f}")
x_new = np.array([[1.0, 2000.0]])
print(f"prediction for 2000 sq ft: {float(x_new @ theta):.2f} thousand dollars")
Running the snippet prints slope and intercept that match the hand calculation, an in-sample MSE around 380, and a 2000-square-foot prediction near $451k.
Key Takeaways
- Machine learning is function approximation from data. Given a sample , pick to minimise a loss — there is no hand-written rule.
- The ML pipeline is data → cleaning/labelling → features → model → loss and optimisation → evaluation → deployment → monitoring, and it is a cycle, not a straight line.
- Supervised learning uses labelled pairs; unsupervised learning finds structure in unlabelled inputs; reinforcement learning maximises a delayed reward signal.
- Reach for ML only when the rules are too long to write, the data is large enough to learn from, the cost of a wrong answer is acceptable, and the data drift is slower than the retraining loop.
- The house-price example shows the full loop in miniature: closed-form least squares on five points yields , with an in-sample RMSE of about $19.5k.