A model that lives only in a notebook does not exist. The skill that separates a machine-learning practitioner from a machine-learning engineer is the ability to take a fitted estimator, hand it to a process that survives a server reboot, and keep it honest as the world under it moves. This lesson covers the full post-training lifecycle: serialising the artefact, designing an interface that serves it, choosing between batch and online inference, watching the served predictions for drift, deciding when to retrain, and auditing the system for bias. It closes by mapping every earlier lesson in the course to the stage of the lifecycle it feeds, so the reader leaves with one mental diagram of the whole craft.
Learning Objectives
- Describe the four operational stages of a machine-learning system — train, serialise, serve, monitor — and explain why they form a closed loop rather than a one-way pipeline.
- Compare
pickleandjoblibfor serialising fitted estimators, identify the version-skew failure mode when the train-time and serve-time library versions differ, and apply mitigations such as pinning and model registries. - Design a prediction API: specify request and response JSON shapes, define a latency budget, decide where batching belongs, and enforce input validation that prevents schema drift from reaching the model.
- Decide between batch and online inference for a given workload, using a cost-and-decision table that contrasts latency, throughput, freshness, and operational complexity.
- Distinguish data drift (the input distribution changes) from concept drift (the mapping from inputs to labels changes), and use metrics such as PSI, KL divergence, and the Wasserstein distance to detect each.
- Define retraining triggers — scheduled, performance-threshold, and drift-threshold — and select the appropriate trigger for a workload based on how quickly its sensitivity is known to change.
- Recognise disparate-impact bias in a deployed model, state the three-way impossibility result for fairness criteria, and apply mitigations such as reweighting and post-processing audits.
- Map each of the fifteen prior lessons to the stage of the deployment lifecycle it feeds, articulating what a working MLOps system inherits from data engineering, modelling, evaluation, and ethics.
1. The Lifecycle as a Closed Loop
A production machine-learning system is not a pipeline that runs once. It is a loop with four stations:
- Train — fit an estimator on labelled data.
- Serialise — persist the fitted object to durable storage.
- Serve — load the artefact into a process that accepts requests and returns predictions.
- Monitor — compare the served predictions, the inputs that produced them, and the eventual ground truth against the training-time distributions.
The arrow from Monitor points back to Train. When the monitor observes that something has gone stale — a feature distribution that no longer resembles training, a calibration curve that has slipped, a fairness metric that has crossed a threshold — the loop fires a retraining job, and the cycle restarts. Every other concern in this lesson hangs off this loop. Serialisation is what makes the train-to-serve hand-off possible. API design is the shape of the serve station. Drift metrics are the eyes of the monitor. Retraining triggers are the gears that connect monitor back to train. Ethics sits across the loop, auditing each station.
+---------+ +-----------+ +--------+ +----------+
| Train |--->| Serialise |--->| Serve |--->| Monitor |
+---------+ +-----------+ +--------+ +----------+
^ |
| retraining trigger |
+----------------------------------------------+
The cost of skipping any station is concrete. A team that does not serialise re-trains on every cold start. A team that does not serve through a versioned API cannot roll back. A team that does not monitor learns about model decay from a customer complaint. A team that does not retrain watches accuracy decline quarter over quarter until the model is retired under cover of night.
2. Serialising a Fitted Estimator
A fitted estimator in scikit-learn is a Python object whose parameters — split
thresholds in a tree, support vectors in an SVM, coefficients in a linear model
— live in attributes like coef_, tree_, or support_. Python's default
serialisation machinery, pickle, turns any object into a byte stream that can
be written to disk and re-instantiated later in another process.
import pickle
from sklearn.linear_model import LogisticRegression
model = LogisticRegression(max_iter=1000).fit(X_train, y_train)
# write
with open("model.pkl", "wb") as fh:
pickle.dump(model, fh)
# read
with open("model.pkl", "rb") as fh:
loaded = pickle.load(fh)
joblib is a sibling library, also from the scikit-learn ecosystem, that
serialises the same objects but is engineered around NumPy arrays. Because most
fitted estimators carry their state inside large numpy.ndarray buffers,
joblib.dump is typically faster and produces a smaller file:
import joblib
joblib.dump(model, "model.joblib") # write
loaded = joblib.load("model.joblib") # read
The choice between them is rarely principled. Both work; both share the same hazard. The bytes on disk encode the object's full state and the names of the classes and modules that produced it. To unpickle, the loading process must have access to those exact classes and modules at the same version. That is the version-skew trap.
2.1 What goes into the bytes
pickle and joblib do not magic away the implementation; they record the
class name, the module path, and the attribute values. If your training
process wrote a sklearn.linear_model._logistic.LogisticRegression object
with scikit-learn 1.4.2, the load process must be able to import
sklearn.linear_model._logistic.LogisticRegression from a compatible
version. The same rule applies to NumPy, scipy, and any custom transformer in
your Pipeline.
2.2 When joblib is genuinely the better choice
For estimators that carry large arrays — a RandomForestClassifier with
hundreds of trees, a TfidfVectorizer with a million-vocabulary matrix, a
neural-net weight tensor — joblib.dump with compression produces files a
fraction of the size and writes them faster:
joblib.dump(model, "model.joblib", compress=("zlib", 3)) # smaller
For a 50-MB forest, the difference between 250 MB and 50 MB matters at load time; a smaller file means a faster container start and a lower egress bill if the model lives in object storage.
3. The Version-Skew Trap
The most common production failure of a serialised model is not a bug in the model. It is the load step failing — or worse, succeeding silently with the wrong model — because the library versions on the serving host do not match the library versions on the training host.
A concrete failure looks like this:
Training host : sklearn 1.3.0, numpy 1.24.3, python 3.10
Serving host : sklearn 1.5.1, numpy 1.26.4, python 3.11
Error : pickle.UnpicklingError or silent attribute mismatch
Two flavours of failure exist. The loud one is an exception at load time: a class that has been renamed, a module that has been split, a required attribute that no longer exists. The silent one is a model that loads, appears to predict, but produces subtly different outputs because an internal helper changed behaviour. Both are catastrophic.
3.1 The mitigation: pin, sign, and register
Three complementary practices close most of this hole.
-
Pin every dependency. A
requirements.txtor lock file that records the exact scikit-learn, NumPy, scipy, and Python versions used at training time. The serving image installs the same lock file. No exceptions. -
Hash and sign the artefact. Store a SHA-256 of the model file alongside the file in object storage. Verify the hash on load. If the file is corrupted in transit, the hash will catch it; if it is silently replaced with a stale artefact, the hash will catch that too.
-
Use a model registry. Tools such as MLflow, Vertex AI Model Registry, or SageMaker Model Registry attach metadata to each artefact — the training run id, the git commit, the metrics, the library versions, the dataset hash. The serving process asks the registry for "model version 42" and receives a pointer to bytes whose provenance is auditable.
import hashlib
with open("model.joblib", "rb") as fh:
blob = fh.read()
digest = hashlib.sha256(blob).hexdigest()
print(digest) # attach this to the registry entry
A model registry does not prevent skew; it makes skew detectable and reversible. If the monitor fires an alert at 2 a.m., the on-call engineer can ask the registry "which version is serving and which library versions does it depend on?" and roll back to the last known-good artefact in minutes rather than hours.
4. Designing a Prediction API
The serve station is a process that accepts a request, loads the model from storage, runs inference, and returns a response. The shape of that contract is the API, and it deserves as much care as the model itself.
4.1 Request and response shape
A prediction request is a JSON document that contains the features the model expects. A response is the prediction and, when appropriate, a confidence score. Two rules hold across virtually every API that survives contact with production.
First, the request must carry a schema version. A field called
"schema_version": "1.4" tells the consumer — and the server — what to
expect. When the schema changes, the server can refuse requests with an
unrecognised version, returning a clear 400 rather than feeding garbage into
the model.
Second, the response must carry enough metadata to debug failures. A bare
{"prediction": 0.87} is untraceable. A response with model version, request
id, and timestamp is auditable:
{
"request_id": "f3a1...",
"model_version": "churn-v42",
"prediction": 0.87,
"prediction_class": 1,
"served_at": "2026-09-28T14:33:12Z"
}
4.2 Latency budget
Every prediction API has a latency budget: a hard ceiling on how long a request may take. The budget is set by the consumer — a real-time bidding system cannot tolerate more than ~50 ms per request, while a nightly batch job is happy with hours.
The throughput the API can sustain is constrained by that budget. If the model takes seconds per inference and the server runs inferences in parallel, the steady-state throughput is
A 50 ms budget with a single-threaded inference therefore caps a single worker at 20 QPS. To serve 1,000 QPS at that budget, the system needs at least 50 workers, which is rarely free. That arithmetic is what makes the batch-vs-online decision (next section) a financial decision, not a technical one.
4.3 Batching
If latency allows, batching — running the model on multiple inputs at once — is often dramatically faster than serving one request at a time. A neural network that takes 8 ms for a batch of 1 may take 25 ms for a batch of 32, giving 0.78 ms per input versus 8 ms — a 10× throughput improvement. The trade-off is added latency for the request that fills the batch.
A micro-batching server waits for up to milliseconds after the first request arrives, collects every request that lands in that window, runs the model once, and dispatches the responses. The expected per-request latency becomes roughly plus the model forward pass.
4.4 Input validation
The cheapest way to take a model down is to feed it inputs it never saw at training. A prediction API must therefore validate the request before the model sees it.
from pydantic import BaseModel, Field, conlist
class PredictRequest(BaseModel):
schema_version: str = Field(pattern=r"^\d+\.\d+$")
features: conlist(float, min_items=10, max_items=10)
class Config:
extra = "forbid" # reject unknown fields
pydantic is one choice; jsonschema, marshmallow, or hand-written checks
are equally valid. What matters is that an out-of-distribution input — a
categorical that is not in the training vocabulary, a negative age, a NaN
that arrived as a string — is rejected at the door with a 400 and never
reaches model.predict.
5. Batch vs Online Inference
The two dominant serving patterns are batch inference and online inference. They differ on every dimension that matters.
| Dimension | Batch inference | Online inference |
|---|---|---|
| Latency budget | Minutes to hours | Tens to hundreds of ms |
| Freshness of features | Hours to days stale | Sub-second |
| Typical workload | Nightly scoring of all customers | Per-request, on a user action |
| Throughput shape | Bulk, scheduled | Bursty, traffic-driven |
| Compute cost | Lower per prediction (bigger batches) | Higher per prediction |
| Operational complexity | A scheduled job | A stateless web service |
| Failure mode | A failed batch retries the next run | A failed request returns a 5xx to the user |
| Auditability | Easy — one run, one log | Harder — request-level logs must be retained |
The choice is rarely binary. Most mature systems run batch inference for pre-computed features (recommendation candidates, lead scores pushed to a CRM) and online inference for event-time features (a fraud score computed when a transaction lands).
5.1 A worked decision
A team is asked to add a churn-risk score to every active customer. The options:
- Batch, nightly. Score 10 million customers in a 90-minute Spark job. Estimated cost: USD 12 per night. Latency to a sales rep: up to 24 hours.
- Online, on demand. Score each customer when the rep opens the account page. Estimated cost: USD 0.004 per request at 1,000 QPS, USD 35 per day. Latency: <100 ms.
The batch option wins on cost. The online option wins on freshness. The correct answer is determined by the consumer's patience: a rep who opens a record to decide what to pitch today needs freshness more than the USD 23 daily saving.
6. Monitoring: Data Drift and Concept Drift
The monitor station watches the served system and the world around it. Two kinds of change matter, and they are not the same.
6.1 Data drift (a.k.a. covariate shift)
Data drift is the situation where the distribution of inputs the model sees in production, , differs from the distribution it saw at training time, . The mapping has not changed; the world has merely started sending different .
Examples are easy to name. A credit-scoring model trained on a population whose average age was 38 sees a population whose average age is 51. A keyword classifier trained on English product reviews suddenly receives Spanish reviews. A pricing model trained pre-pandemic sees pandemic-shaped demand curves.
Detection is a two-sample problem. Two metrics dominate:
- Population Stability Index (PSI) on each feature:
A PSI above 0.1 is "small change", above 0.2 is "moderate", above 0.5 is "major". The thresholds are industry folklore, not laws of nature; the important thing is to compute and store the baseline so a current PSI can be compared against it.
- Wasserstein distance (a.k.a. earth mover's distance) for continuous features:
The Wasserstein distance has the merit of being interpretable in the units of the feature itself (dollars, years, kilograms), which makes it easier to explain to a non-technical stakeholder than a log-likelihood ratio.
6.2 Concept drift
Concept drift is the situation where the mapping from to has changed: . The inputs may look identical; the rule that relates them to the outcome is now different. A spam classifier trained on last year's spam patterns misses this year's phishing campaigns. A demand forecaster trained on pre-2020 behaviour under-predicts a regime change.
Detection is harder because it requires ground truth that arrives late. The typical approach is to monitor a performance metric that can be computed on labelled data, even if the labels are delayed:
- Calibration drift. Compare predicted probabilities to observed positive rates in decile buckets. If a bucket predicted 0.7 but observed 0.55, calibration has slipped.
- AUC drift on a sliding window. If a labelled window's AUC is materially below the training AUC, concept drift is one possible cause (data drift in is another, and the metric alone cannot distinguish them).
6.3 Metrics that go stale
The same metric can be perfectly stable in the lab and totally stale in production. The categories that go stale fastest are:
- Anything computed on a population that has shifted. Accuracy on a test set drawn from the original population, not the current one.
- Anything measured against human-labelled ground truth. The label-generation process is part of the system; if labellers change, the ground truth moves.
- Calibrated probabilities. Calibration is a property of the joint distribution of and ; drift in either moves it.
6.4 The narrow-waist rule
A useful discipline: monitor the smallest possible set of metrics that would catch every plausible failure mode. A typical minimal set is:
- One drift metric per input feature (PSI or Wasserstein).
- One performance metric on a labelled slice (calibration error, AUC, or RMSE).
- One fairness metric on the protected attributes in scope (see §8).
Anything outside that set is noise that distracts the on-call engineer.
7. Retraining Triggers
The gear between the monitor station and the train station is the retraining trigger. Three families dominate.
| Trigger | Fires when | Strength | Weakness |
|---|---|---|---|
| Scheduled | A calendar condition holds (daily, weekly) | Simple, predictable | Wasteful when nothing has changed; late when something has |
| Performance-threshold | A monitored metric crosses a threshold | Reflects actual model quality | Needs labels, which often arrive late |
| Drift-threshold | PSI or Wasserstein on a feature exceeds a threshold | Detects problems early | Drifted inputs may still be predicted well; false positives waste compute |
A mature production system usually uses two of the three. A common combination is scheduled + drift-threshold: retrain on a weekly cadence, but accelerate to a same-day retrain if PSI on a critical feature exceeds 0.3.
7.1 The trigger is a function, not a button
Retraining on a schedule is the easy case — a cron job. The interesting case is drift-triggered retraining, because the trigger has to be defined as code:
def should_retrain(drift_metrics: dict, perf_metrics: dict,
last_train: datetime) -> bool:
# scheduled: weekly
if datetime.utcnow() - last_train > timedelta(days=7):
return True
# drift threshold on the top-3 most-important features
top3 = ["age", "tenure_months", "monthly_spend"]
if any(drift_metrics[f]["psi"] > 0.3 for f in top3):
return True
# performance threshold on the labelled slice
if perf_metrics["auc_30d"] < 0.78:
return True
return False
This function lives in version control, is unit-tested, and is reviewed by the same people who reviewed the model. A retraining system that runs on ad-hoc operator judgement is one Slack message away from a bad week.
8. Ethics, Fairness, and Bias
A deployed model is a deployed policy. If it scores loan applications, it shapes who gets credit. If it ranks resumes, it shapes who gets interviewed. If it triages patients, it shapes who gets treated first. Every station of the lifecycle inherits an obligation to examine its consequences for the people on the receiving end.
8.1 Disparate impact
Disparate impact is the legal and statistical notion that a seemingly neutral model produces systematically different outcomes for a protected group. The standard measure is the disparate impact ratio:
where is a binary protected attribute (e.g. a gender or race indicator). A value below 0.8 is conventionally treated as evidence of disparate impact under US employment law; the threshold is a rule of thumb, not a physical constant, but it is the one most compliance teams ask about.
8.2 The impossibility result
There is a theorem that every deployed fairness-aware system should know about. Three criteria, individually attractive, cannot be satisfied simultaneously except in degenerate cases. The three are:
- Independence (a.k.a. demographic parity): .
- Separation (a.k.a. equalised odds): .
- Sufficiency: .
Chouldechova (2018) and Kleinberg, Mullainathan & Raghavan (2016) show that, when the base rates and differ, no classifier can satisfy all three. The practical consequence is that "fair" is not a single point in metric space but a triangle whose vertices are mutually exclusive. Choosing where on the triangle to sit is a policy decision, not a technical one.
8.3 Mitigations
Three families of mitigation exist, and they act at different stations of the loop:
- Pre-processing: reweight or resample the training data so the protected attribute is balanced. Easiest to apply; risk of distorting genuine signal.
- In-processing: train with a regulariser that penalises disparity (e.g. equalised-odds constraints in a constrained optimiser). More work to implement; usually less distortion of the rest of the model.
- Post-processing: calibrate predictions per group after the fact. Cheap to deploy; can be a compliance band-aid if the underlying model is badly biased.
8.4 Auditing
A model that is not audited is a model whose fairness is asserted, not demonstrated. An audit answers a concrete question: across the deployed population, what are the values of the fairness metrics we promised to uphold? The audit runs on a schedule (monthly, quarterly, or after every retrain) and produces a report. The report, not the metric, is what external regulators and internal stakeholders actually read.
The most common audit failure mode is the silent fairness regression: a model whose accuracy improved after retraining but whose disparate-impact ratio moved from 0.82 to 0.71 because the new training data was under-represented for one group. The retraining trigger of §7 was correct; it just optimises the wrong objective. The audit is what catches it.
9. Capstone: Mapping the Course to the Lifecycle
This course has covered fifteen lessons before this one. Each one feeds a specific station of the deployment loop, and recognising the mapping is what turns the course from a sequence of techniques into a single craft.
| Lesson | Topic | Feeds which station |
|---|---|---|
| 1 | What ML is (and is not) | Frame the problem at the train station |
| 2 | The modelling workflow | Sequence the train station: split, fit, validate, iterate |
| 3 | Linear and logistic regression | The baseline estimator trained at the train station |
| 4 | Trees and ensemble foundations | The non-linear estimators trained at the train station |
| 5 | k-NN, Naive Bayes, SVM | Alternative estimators at the train station |
| 6 | Regularisation and generalisation | The discipline that keeps the train station honest |
| 7 | Optimisation and gradient descent | The mechanics under the hood at the train station |
| 8 | Neural networks and backprop | The deep estimators trained at the train station |
| 9 | Convolutional and sequence models | Specialised architectures for vision and text |
| 10 | Unsupervised and representation learning | Embeddings, dimensionality reduction, clustering |
| 11 | Probabilistic and Bayesian methods | The uncertainty estimates used in monitoring |
| 12 | Ensemble learning | The combined estimators that often go into production |
| 13 | Evaluation metrics and validation | The metrics watched at the monitor station |
| 14 | Hyperparameter search and AutoML | The systematic tuning done at the train station |
| 15 | Pipelines, leakage, and feature stores | The data plumbing feeding train, serialise, and serve |
The arrow runs from the earlier lessons forward into this one. A team that has internalised those fifteen lessons has the vocabulary to reason about every station of the loop; this lesson is what wires those pieces into a running system.
The reverse arrow runs backward. When the monitor station fires an alert on calibration drift (lesson 13), the on-call engineer reaches for the Bayesian tools of lesson 11 to diagnose whether the shift is data or concept. When the retraining trigger of §7 fires, the engineer rebuilds the pipeline of lesson 15, re-tunes the hyperparameters of lesson 14, and re-trains the ensemble of lesson 12. Every station of the loop is a way of exercising the prior fifteen lessons under operational pressure.
9.1 The end-to-end picture
A single end-to-end view of the loop, with each station labelled, is the mental model to leave the course with:
[Lesson 15: Pipeline] [Lesson 14: Tuning]
| |
v v
+---------+ +-----------+ +--------+
| Train |---->| Serialise |---->| Serve |
+---------+ +-----------+ +--------+
^ ^ |
| | v
| | +----------+
| | | Monitor |
| | +----------+
| | |
| +-------- retrain trigger -------+
| |
| [Lesson 13: Metrics]
| [Lesson 11: Uncertainty]
| [Lesson 8: Ethics & Bias] <-- this lesson
+------------------ audit ---------+
The deployment loop is the smallest unit of machine-learning engineering. Every system a practitioner ships from this point forward — every notebook that graduates to a service, every model that earns its keep in production — is some instantiation of this diagram, instantiated with the tools of the prior fifteen lessons and audited against the fairness discipline of this one.
Key Takeaways
- A production ML system is a closed loop of four stations: train, serialise, serve, and monitor. The arrow from monitor back to train closes the loop and is the source of every long-term operational property the system has.
pickleandjoblibboth serialise fitted estimators, but the bytes encode class names and module paths; serving with a different library version than the one used at training is the most common production failure, mitigated by pinning dependencies, hashing artefacts, and using a model registry.- A prediction API needs a versioned request schema, a clear response with metadata, a latency budget that constrains the throughput, batching where the budget allows, and strict input validation that rejects out-of-distribution requests before they reach the model.
- Batch and online inference are not mutually exclusive; the choice between them is a cost-and-decision problem driven by the consumer's latency tolerance, with batch scoring cheaper per prediction and online scoring fresher.
- Data drift is ; concept drift is . PSI and Wasserstein catch the former; calibration and AUC on labelled slices catch the latter.
- Retraining triggers come in three families — scheduled, performance-threshold, and drift-threshold — and a mature system uses two of them, with the trigger defined as code and reviewed like any other production artefact.
- A deployed model is a deployed policy. Disparate impact is measured by the DI ratio; the three-way impossibility result (independence, separation, sufficiency) means "fair" is a triangle, not a point; and audits, not assertions, demonstrate that the chosen vertex is the one the institution can defend.
- Every one of the fifteen prior lessons feeds a station of this loop, and the craft of machine-learning engineering is the integration of those lessons into a running, monitored, audited system.