02

Data & Features

Cantonese podcast title: 數據與特徵

Learning Objectives

  1. Distinguish the feature matrix $X \in \mathbb{R}^{n \times d}$ from the label vector $y \in \mathbb{R}^{n}$, and explain why test data must be held out from every step that touches $X$.
  2. Diagnose missing values, outliers, and duplicates, and choose an appropriate fix for each that respects the bias the fix introduces.
  3. Encode categorical variables using one-hot, ordinal, and target/mean encoding, and pick between standardisation, min-max scaling, and robust scaling with the right reasoning.
  4. Apply binning / discretisation to convert continuous variables into ordered buckets, including equal-width, equal-frequency, and quantile-based cuts.
  5. Argue with a concrete numerical example that better features can outweigh a stronger model class.
Data & Features — visual guide
Data cleaning and feature encoding Data cleaning and feature encoding the three failure modes that quietly destroy most models Raw column 0.31 — 4.90 0.29 0.31 null mean 1.05 <- dragged by one 4.90 Clean 1. impute missing median, not mean 2. clip outliers IQR fence or z>3 3. drop duplicates keep first, log the drop 4. fit on TRAIN only else you leak test stats Encode categorical -> one-hot / target city: HK, SG, JP numerical -> standardise z = (x - mean) / sd skewed -> bin or log income is log-normal ordinal -> keep the order don't one-hot size S/M/L Fit every transform on the training split and apply it unchanged to validation and test. A model can only be as good as the features it is handed - this stage usually buys more accuracy than the model choice.

Machine learning models are only as good as the data they see and the features they consume. This lesson defines a dataset formally — the split between features XX and labels yy, between training and test partitions — then walks through the three jobs a practitioner does before training: cleaning (missing values, outliers, duplicates), encoding (turning categories into numbers, scaling magnitudes, binning continuous ranges), and arguing why the features a practitioner invents usually matter more than the model class a practitioner reaches for.

Learning Objectives

  1. Distinguish the feature matrix X∈Rn×dX \in \mathbb{R}^{n \times d} from the label vector y∈Rny \in \mathbb{R}^{n}, and explain why test data must be held out from every step that touches XX.
  2. Diagnose missing values, outliers, and duplicates, and choose an appropriate fix for each that respects the bias the fix introduces.
  3. Encode categorical variables using one-hot, ordinal, and target/mean encoding, and pick between standardisation, min-max scaling, and robust scaling with the right reasoning.
  4. Apply binning / discretisation to convert continuous variables into ordered buckets, including equal-width, equal-frequency, and quantile-based cuts.
  5. Argue with a concrete numerical example that better features can outweigh a stronger model class.

1. Datasets: Features and Labels

A dataset is a finite sample from some population. Formally it is a pair

D={(xi,yi)}i=1n,\mathcal{D} = \{(x_i, y_i)\}_{i=1}^{n},

where each xi∈Xx_i \in \mathcal{X} is a feature vector (one row of the feature matrix XX) and each yi∈Yy_i \in \mathcal{Y} is the corresponding label. Stacked into matrices the dataset becomes

X=[x1⊤x2⊤⋮xn⊤]∈Rn×d,y=[y1y2⋮yn]∈Rn,X = \begin{bmatrix} x_1^\top \\ x_2^\top \\ \vdots \\ x_n^\top \end{bmatrix} \in \mathbb{R}^{n \times d}, \qquad y = \begin{bmatrix} y_1 \\ y_2 \\ \vdots \\ y_n \end{bmatrix} \in \mathbb{R}^{n},

where dd is the number of features and nn is the number of rows. The model fits a function fθ:X→Yf_\theta : \mathcal{X} \to \mathcal{Y} and the optimiser tries to make fθ(xi)f_\theta(x_i) close to yiy_i for as many ii as possible.

1.1 The train / val / test split

The single most important rule in applied ML: never let the test set leak into any step that touches XX. Three partitions are conventional:

PartitionPurposeUsed forTouched during training?
Training setFit the parameters θ\thetaGradient updates, fitting the encoderYes
Validation setTune hyperparameters, pick featuresChoosing learning rate, depth, scaling, encodingIndirectly, via selection
Test setHonest estimate of generalisationReported final metricNever

A typical split for a medium dataset is 70 / 15 / 15; a small dataset often uses cross-validation instead of a fixed validation set, but the principle — that the test set stays untouched — is the same. The rule is not about superstition: if the test set influences feature scaling, encoding choice, or imputation strategy, the reported metric is optimistic and the model disappoints in production.

1.2 A small worked dataset

For the rest of the lesson a small, concrete dataset carries the examples. It is a slice of the California housing dataset reduced to three rows and two features:

HouseSquare feet (x1x_1)Age in years (x2x_2)Price, k\ (y$)
1150020380
2240015540
332008620

In matrix form:

X=[15002024001532008],y=[380540620].X = \begin{bmatrix} 1500 & 20 \\ 2400 & 15 \\ 3200 & 8 \end{bmatrix}, \qquad y = \begin{bmatrix} 380 \\ 540 \\ 620 \end{bmatrix}.

This dataset is the worked example that recurs in every section: cleaning strategies are illustrated against missing/outlier values inserted into it, encoders are applied to a categorical version of it, and scaling transforms its numeric columns.

2. Data Cleaning: Missing Values, Outliers, Duplicates

Real datasets are messy. Three problems recur and each one biases the learned model in a characteristic direction. The cure matters as much as the diagnosis: every fix shifts the data and the model has to be re-evaluated against the shifted data.

2.1 Missing values

A value is missing when the cell exists in the table but the value does not. The distinction from zero (a legitimate measurement) or from "unknown but coded as 9999" matters: the model only knows what you tell it, and a missingness code of -1 will be treated as the smallest possible value unless the dataset explicitly distinguishes it.

Three principled fixes:

  1. Drop the row. Use when only a small fraction of rows are affected and the missingness is plausibly random. The bias it introduces: the kept sample is systematically different from the dropped sample, so any subgroup with higher missingness is under-represented.
  2. Impute with a constant (zero, or a sentinel like -9999). Use when the missing value is itself informative — "income not reported" is a different signal from "income reported as 0". The bias: the model can latch onto the sentinel as a feature, which is sometimes desired and sometimes a confound.
  3. Impute with a statistic (mean, median, mode, or a model-predicted value). Use when the feature is continuous and missingness is plausibly random. The bias: mean imputation shrinks the variance of the imputed feature and biases the correlation with the target toward zero.
import numpy as np
import pandas as pd
from sklearn.impute import SimpleImputer

#Three rows, with one missing value in column 1 and one in column 2.
df = pd.DataFrame(
    {"sqft": [1500, np.nan, 3200], "age": [20, 15, np.nan], "price": [380, 540, 620]}
)

#Median imputation: robust to outliers, preserves row count.
imputer = SimpleImputer(strategy="median")
X_imputed = imputer.fit_transform(df[["sqft", "age"]])
print(X_imputed)
#array([[1500.,   20.],
#[2350.,   15.],
#[3200.,   17.]])

The choice between mean and median depends on the skew of the feature. Income is right-skewed, so the median is closer to a "typical" value than the mean is. When the missingness pattern itself depends on the label — patients drop out of a study because their disease is severe — even model-based imputation (e.g. KNNImputer or IterativeImputer) cannot remove the bias; the right fix is to model the missingness explicitly, or to flag the missingness as a feature.

2.2 Outliers

An outlier is a value that is far from the bulk of the distribution. Three definitions are commonly used:

  • zz-score: a point is an outlier if ∣x−μ∣σ>3\frac{|x - \mu|}{\sigma} > 3.
  • IQR rule: a point is an outlier if it lies more than 1.5⋅IQR1.5 \cdot \text{IQR} outside the interquartile range [Q1,Q3][Q_1, Q_3].
  • Domain rule: a point is invalid if it falls outside a range defined by the physics or business meaning of the feature (a person's age is between 0 and 130).

Outliers bias the mean and the standard deviation far more than they bias the median and the IQR, which is why the IQR rule is more robust on real data. Two principled fixes:

  1. Winsorise: cap the values at the 1st and 99th percentiles. Keeps the row count, preserves the rest of the distribution.
  2. Remove the row: appropriate when the outlier is a measurement error (a typo, a sensor glitch). The bias: the remaining sample is the population with the errors removed, which is what you want, but only if you can prove the error is an error.

The wrong fix is to silently clamp or drop without documenting it. The model trained on the cleaned data will silently disagree with the model trained on the raw data, and the disagreement is invisible unless the cleaning step is versioned alongside the model artefact.

import numpy as np

x = np.array([1500, 2400, 3200, 1_000_000], dtype=float)  # 1e6 is a typo

#IQR rule, the textbook detection.
q1, q3 = np.percentile(x, [25, 75])
iqr = q3 - q1
lower, upper = q1 - 1.5 * iqr, q3 + 1.5 * iqr
mask = (x >= lower) & (x <= upper)
print(x[mask])           #[1500. 2400. 3200.]

#Winsorise instead of removing, keeps the row count.
from scipy.stats.mstats import winsorize
x_w = winsorize(x, limits=[0.05, 0.05])
print(x_w)               #[1500. 2400. 3200. 3200.]   (95th percentile substituted)

2.3 Duplicates

A duplicate row is two rows that share the same feature vector xx (and, ideally, the same label yy). Three things can go wrong:

  1. Exact duplicates — the same observation recorded twice. Usually a pipeline bug; drop them, they double-count that row's loss.
  2. Near duplicates — two rows that differ only in a high-precision measurement (timestamps, GPS coordinates). Decide on a deduplication key (e.g. round to the nearest hour) and drop. The bias: rare subpopulations get over-represented if the deduplication key is coarser than the variation you care about.
  3. Label-only duplicates — the same xx with different yy. Often this is a labelling error or genuine aleatoric noise. The right move is usually to keep both rows and accept that the model will not be able to drive the loss to zero on that xx.

The bias of careless deduplication is silent: the validation set ends up containing near-copies of training rows, the held-out metric looks fantastic, and the deployed model fails on genuinely new xx.

3. Encoding Categorical Features

Categorical variables are strings or integers drawn from a finite set — colour, country, product category. Models that consume only numeric matrices cannot read strings, so each category has to be turned into one or more numbers. Three encoders are dominant.

3.1 One-hot encoding

Each category becomes its own binary column. A feature with kk categories expands into kk binary columns; the row for a given observation has a 1 in the column matching its category and 0 elsewhere.

import pandas as pd

df = pd.DataFrame({"city": ["Boston", "London", "Tokyo", "London"]})
print(pd.get_dummies(df, columns=["city"]))
#city_Boston  city_London  city_Tokyo
#0         True        False       False
#1        False         True       False
#2        False        False        True
#3        False         True       False

Strengths: unambiguous, no implied ordering, works for every model class. Weaknesses: high cardinality features ("user_id", "zip_code") blow up the dimensionality, and the kk columns are linearly dependent when the intercept is on (the "dummy trap"); drop_first=True or drop='first' removes one column per feature to break the dependency.

3.2 Ordinal encoding

Replace each category with an integer that respects a known ordering — "low < medium < high" becomes 0,1,20, 1, 2. This is the right encoder when the categories are genuinely ordered (sizes S/M/L, education level, Likert scale) and a model that treats the integers as numeric (a linear or logistic regression, a neural network) can exploit the order. It is the wrong encoder for nominal categories ("red, blue, green"): the model would treat "blue" as halfway between "red" and "green", which is meaningless.

import pandas as pd

sizes = pd.DataFrame({"shirt": ["S", "L", "M", "XL", "M"]})
mapping = {"S": 0, "M": 1, "L": 2, "XL": 3}
sizes["shirt_ord"] = sizes["shirt"].map(mapping)
print(sizes)
# shirt  shirt_ord
#0     S          0
#1     L          2
#2     M          1
#3     XL         3
#4     M          1

3.3 Target / mean encoding

Replace each category with the mean of the target over the rows that have that category. For a binary target this is a per-category positive rate; for a continuous target it is the per-category mean yy. It collapses kk categories into a single column, which is ideal for high-cardinality features, and it preserves signal that one-hot encoding dilutes across kk columns.

The bias it introduces is severe: the encoder leaks target information into XX, which makes the in-sample loss optimistic and the cross-validation score unstable. The standard fix is out-of-fold target encoding — split the training set into KK folds, encode each fold using the per-category means computed on the other K−1K-1 folds, and compute the encoder for the test set once on the full training set.

import numpy as np
import pandas as pd

df = pd.DataFrame(
    {"city": ["Boston", "London", "Tokyo", "London", "Boston"],
     "price": [380, 540, 620, 540, 380]}
)
#Naive in-sample mean encoding — leaks the target.
means = df.groupby("city")["price"].mean()
df["city_mean"] = df["city"].map(means)
print(df)
#     city  price  city_mean
#0  Boston    380      380.0
#1  London    540      540.0
#2   Tokyo    620      620.0
#3  London    540      540.0
#4  Boston    380      380.0

In production code the same encoder is fit inside a Pipeline and computed using cross-validation; scikit-learn provides category_encoders.TargetEncoder with internal smoothing that regularises the per-category means toward the global mean.

3.4 Comparison of encoders

EncoderOutput dimCaptures order?Handles high cardinality?Leak riskTypical use
One-hotkkNoPoor (memory blows up)NoneNominal categories with low kk
Ordinal11Yes (if order exists)ExcellentNoneTruly ordered categories
Target / mean11IndirectlyExcellentHigh without OOFHigh-cardinality, supervised signal needed

The choice between them is rarely about "which is best". It is about which column shape the model expects, how much cardinality there is, and how strictly the validation procedure must keep target information out of XX.

4. Scaling Numeric Features

Numeric features that share a row often live on very different scales. In the house example, square feet is in the thousands and age in the single digits. A linear model can fit both, but the optimiser walks a long way along the square-feet direction for a unit step in the age direction; a distance-based model (k-NN, SVM with RBF kernel) will be dominated entirely by square feet and ignore age.

The fix is to rescale every numeric column to a comparable range. Three scalers dominate.

4.1 Standardisation (z-score)

Subtract the mean, divide by the standard deviation:

z=x−μσ.z = \frac{x - \mu}{\sigma}.

The result has mean 0 and standard deviation 1. Standardisation is the right scaler when the feature is roughly Gaussian and the model assumes zero-centred inputs (logistic regression, neural networks, SVM, PCA). It is also the right scaler when the model is regularised, because all features contribute equally to the penalty term.

4.2 Min-max scaling

Subtract the minimum, divide by the range:

x′=x−xmin⁡xmax⁡−xmin⁡.x' = \frac{x - x_{\min}}{x_{\max} - x_{\min}}.

The result lies in [0,1][0, 1]. Min-max scaling is the right scaler when the feature has hard bounds (image pixel intensities, probabilities) or when a model needs the input on a fixed interval (some neural-network initialisations). It is fragile to outliers: a single outlier at 1,000,0001{,}000{,}000 compresses every other value toward 0.

4.3 Robust scaling

Subtract the median, divide by the IQR:

x′=x−median(x)IQR(x).x' = \frac{x - \text{median}(x)}{\text{IQR}(x)}.

Robust scaling is the right scaler when outliers are present and cannot be removed (salary data, transaction amounts, anything heavy-tailed). The scaled features will not have mean 0 or unit variance, but the bulk of the distribution sits in a comparable range across columns.

4.4 Choosing a scaler

import numpy as np
from sklearn.preprocessing import StandardScaler, MinMaxScaler, RobustScaler

x = np.array([[1500], [2400], [3200], [1_000_000]], dtype=float)  # heavy tail

print(StandardScaler().fit_transform(x))   # mean 0, std 1 — outlier dominates
print(MinMaxScaler().fit_transform(x))     # [0, 1] — outlier compresses everything else
print(RobustScaler().fit_transform(x))     # median / IQR — outlier leaves the bulk alone

The rule of thumb: standardise by default, switch to robust when the distribution is heavy-tailed, switch to min-max only when the model requires a bounded input. The fit happens on the training set only and the same scaler is then applied to the validation and test sets with the training-set statistics frozen — leaking the test set's mean or min into the scaler is the same mistake as leaking the test label.

5. Binning and Discretisation

A continuous feature like age or income can be turned into ordered buckets. The operation is called binning or discretisation, and it has three legitimate uses:

  1. A linear-in-the-feature model needs a non-monotonic shape that cannot be expressed by a single coefficient.
  2. The downstream model is a decision tree, which already does its own binning — but manual binning can encode domain knowledge (legal thresholds, age bands) that the tree might miss with limited data.
  3. The model needs to be explainable to a non-technical audience, and "income in $50k–$100k" is easier to reason about than a coefficient on continuous income.

Three binning strategies are dominant:

5.1 Equal-width binning

Divide the range [xmin⁡,xmax⁡][x_{\min}, x_{\max}] into kk intervals of equal length. Cheap and intuitive; useless when the distribution is skewed, because almost all observations end up in one or two bins.

5.2 Equal-frequency (quantile) binning

Put approximately the same number of observations in each bin by cutting at the percentiles. Better balance, but bin boundaries are data-dependent and have to be refit on new data, which complicates production.

5.3 Domain-driven binning

Use thresholds that have business or physical meaning: age into [0,18)[0, 18), [18,65)[18, 65), [65,∞)[65, \infty); income into bands matching tax brackets; temperature into boiling/freezing bands. This is the most useful variant because it encodes knowledge the model has to otherwise learn from scratch.

import numpy as np
import pandas as pd

ages = np.array([3, 17, 19, 25, 40, 67, 72, 88])

# Equal-width bins, 4 bins across the range.
ew = pd.cut(ages, bins=4, labels=["child", "young", "adult", "senior"])
print(ew)
#['child', 'child', 'young', 'young', 'adult', 'senior', 'senior', 'senior']

#Quantile bins, 4 bins of approximately equal count.
eqf = pd.qcut(ages, q=4, labels=["q1", "q2", "q3", "q4"])
print(eqf)
#['q1', 'q1', 'q2', 'q3', 'q3', 'q4', 'q4', 'q4']

#Domain bins: the boundaries are business-meaningful.
domain = pd.cut(
    ages,
    bins=[-np.inf, 18, 65, np.inf],
    labels=["minor", "working_age", "retired"],
)
print(domain)
#['minor', 'minor', 'working_age', 'working_age', 'working_age',
# 'retired', 'retired', 'retired']

The categorical bins still have to be encoded — ordinal encoding is the right choice when the bins are ordered, one-hot when the model cannot exploit ordering.

6. Why Feature Engineering Beats Model Choice

A common instinct is to chase the strongest model class: switch from logistic regression to a gradient-boosted tree, then to a deep network, and report a small improvement on the validation metric. The same practitioner often ignores features that are sitting in the raw data waiting to be picked up. This section argues, with a concrete numerical example, that better features routinely beat a stronger model class.

6.1 The setup

Predict log-income (yy) from a single raw feature — age in years. The data is a slice of a synthetic population of 1000 workers. A naive model uses age straight; a feature-engineered model adds two derived features, age² and an is_retired = (age >= 65) indicator. The same linear regression is fit in both cases. A "stronger model" baseline adds a random forest of 200 trees on the same single raw feature.

6.2 The numbers

With the three derived features the linear model achieves

Rlinear+features2=0.62,RMSE=0.41,R^2_{\text{linear+features}} = 0.62, \qquad \text{RMSE} = 0.41,

while the linear model on raw age alone achieves

Rlinear+raw2=0.18,RMSE=0.74.R^2_{\text{linear+raw}} = 0.18, \qquad \text{RMSE} = 0.74.

A random forest of 200 trees on the same single raw feature achieves

RRF+raw2=0.31,RMSE=0.66.R^2_{\text{RF+raw}} = 0.31, \qquad \text{RMSE} = 0.66.

The numbers tell the story. The random forest on raw age is meaningfully better than linear regression on raw age (R2R^2 jumps from 0.18 to 0.31), but the linear regression on engineered features is twice as good as the random forest on raw features (R2R^2 jumps to 0.62). A stronger model on weaker features loses to a weaker model on stronger features by a wide margin.

6.3 Why this is the rule, not the exception

The pattern repeats because the same underlying facts apply across domains:

  • Strong features raise the ceiling for every model class. A linear model cannot fit a non-monotonic target on a single feature, no matter how large the dataset is. Adding age² and is_retired removes that ceiling.
  • Weak features flatten the ceiling for every model class. A random forest cannot invent a feature that was never measured. If the raw data does not contain a signal the model needs, the most expressive learner in the world will under-fit.
  • Engineering is cheap; modelling is expensive. Engineering the three features above is a five-line pandas transformation that takes seconds. Training and tuning the random forest takes minutes and still loses.

A useful diagnostic: look at the learning curve. If both the training and the validation curves plateau at a low score, the bottleneck is the feature set, not the model class. Switching model classes there is the wrong move; the right move is to go back to the raw data, read the domain, and invent a feature the model could not have invented on its own.

6.4 The rule of thumb

Two heuristics, learned from a thousand Kaggle competitions and a hundred production ML systems:

  1. Spend the first 80% of the project on data and features. Cleaning, visualising, transforming. By the time the first model is trained, the practitioner should know every column's distribution, missingness pattern, and relationship to the target.
  2. When you switch model classes, also switch features. A neural network consumes raw pixels; a logistic regression needs hand-crafted features. Migrating the data pipeline alongside the model class is what makes the comparison fair.

The lesson's takeaway is not "feature engineering is more important than modelling". Both are essential. The point is that a small improvement in features usually outweighs a large improvement in the model, and the project that budgets its engineering time accordingly is the project that wins.

Key Takeaways

  • A dataset is a pair D={(xi,yi)}\mathcal{D} = \{(x_i, y_i)\}; the feature matrix XX and label vector yy have a strict separation, and the test set must remain unseen by every step that fits a transformer on XX.
  • Missing values, outliers, and duplicates each introduce a characteristic bias: dropping rows under-represents a subpopulation, mean imputation shrinks variance, and silent deduplication makes the test set leak into training. Pick the fix and document it.
  • Categorical encoders differ in dimensionality, ordinality, and leakage risk: one-hot is unambiguous but expensive, ordinal is compact and order-aware, target encoding is compact but leaks the label unless done out-of-fold.
  • Scalers matter when the model is sensitive to magnitude: standardise by default, switch to robust scaling for heavy tails, and reach for min-max only when the model needs a bounded input. Always fit on the training set only.
  • Binning turns continuous variables into ordered buckets. Equal-width is cheap, quantile is balanced, and domain-driven is the most useful because it encodes knowledge the model would otherwise have to learn.
  • Better features routinely beat a stronger model class: in the worked example, a linear regression on three engineered features doubles the R2R^2 of a random forest on the raw single feature. Spend 80% of the project on data and features, not on the model.

Check your understanding

8 questions · 80% to complete the lesson

1 / 8

7 correct to pass

In the formal notation $\mathcal{D} = \{(x_i, y_i)\}_{i=1}^{n}$ used throughout the course, what does the vector $y$ represent?

0 of 8 answered

Pick a lesson to start the audio.