Naive Bayes is a family of probabilistic classifiers that apply Bayes' theorem with a strong — and almost always false — assumption that every feature is conditionally independent of every other feature given the class label. Despite that assumption, the classifier remains a stubbornly strong baseline on text problems: it trains in a single pass over the data, predicts in microseconds, and is remarkably tolerant of small training sets and high-dimensional sparse feature vectors. This lesson starts from Bayes' theorem, makes the prior-vs-posterior distinction concrete with real numbers, states the conditional-independence assumption explicitly and explains why it is usually wrong, then specialises to the multinomial variant used for text classification and ties it to TF-IDF features. A comparison table of the Gaussian, multinomial, and Bernoulli variants, a pros-and-cons table, and a worked spam-detection example close the lesson.
Learning Objectives
- State Bayes' theorem in the classification setting and identify prior, likelihood, evidence, and posterior for a concrete example.
- State the conditional-independence assumption of naive Bayes, give a real case where it fails, and explain what the failure does to the predicted probabilities.
- Distinguish prior from posterior with arithmetic on real numbers and explain why the prior matters even when the model is highly accurate.
- Compute the parameters of a multinomial naive Bayes classifier from token counts and predict a class for a new document using the log-sum form.
- Choose between Gaussian, multinomial, and Bernoulli naive Bayes given the type of features (continuous, counts, binary) and link the multinomial variant to TF-IDF weighting.
- Argue why naive Bayes is still a strong baseline for text classification: training cost, prediction cost, and behaviour under sparse data.
1. Bayes' Theorem as a Classifier
The starting point is Bayes' theorem, which expresses a conditional probability in terms of its inverse:
In the classification setting the four pieces play distinct roles:
| Symbol | Name | Meaning in a classifier |
|---|---|---|
| Prior | How common the class is overall, before seeing any features | |
| Likelihood | How probable the observed feature vector is under class | |
| Evidence | How probable the feature vector is overall; a normalising constant that does not depend on | |
| Posterior | The updated belief about after observing — what the classifier actually returns |
A classifier only needs the relative values of the posterior across classes, so the evidence can be dropped at prediction time. The decision rule reduces to
This is the maximum a posteriori (MAP) rule: pick the class whose product of likelihood and prior is largest.
1.1 A toy one-feature example
Imagine a single feature — the colour of a light, red or green — and a binary class {stop, go}. Suppose 70% of lights are stop lights, 30% are go lights. Of all stop lights, 90% show red; of all go lights, 85% show red. The joint probabilities:
Given a red light, the posterior probability that it is a stop light is
A red light is therefore about a 71% chance of being a stop light. The classifier predicts "stop" because . Note that the prior of 0.70 is doing real work: if the priors were balanced (50/50), the same likelihoods would give
The decision still favours stop, but only barely. The lesson: the prior matters most when the likelihoods are similar.
2. Prior, Likelihood, and Posterior with Real Numbers
The prior summarises what the classifier knows about class prevalence before seeing any features. The likelihood encodes how the features behave within each class. The posterior combines the two.
2.1 A worked diagnostic example
A medical test for a rare disease has the following properties:
- Disease prevalence: (1 in 100 people have it).
- Sensitivity: .
- Specificity: , so the false-positive rate is .
A patient tests positive. What is the probability they actually have the disease?
with the denominator computed by expanding over the two classes:
The posterior is
A positive test means only about a 17% chance of actually having the disease. The counter-intuitive result comes entirely from the prior: with a 1% base rate, the pool of false positives from the 99% disease-free population overwhelms the true positives from the 1% sick population. The 99% sensitivity and 95% specificity that sound excellent in isolation become much weaker once the prior is honoured.
2.2 What the classifier actually does
The naive Bayes MAP rule
is mathematically equivalent to saying "multiply the prior by the per-feature likelihoods, pick the largest product." With thousands of features the product underflows numerically, so the classifier is implemented in log-space:
The transform turns the product into a sum and the underflow disappears. The prediction — which class wins — is unchanged.
3. The Conditional-Independence Assumption
The naive Bayes model assumes that, given the class label, the features are mutually independent:
In English: once you know whether a document is spam, the presence of the word "free" tells you nothing further about whether it also contains the word "viagra". Whether a tumour is malignant tells you nothing about the patient's blood-pressure reading. Whether a transaction is fraudulent tells you nothing about the merchant category.
This assumption is almost never true in real data. Words in natural language are strongly correlated with one another; tumour features are correlated through biological pathways; fraud features are correlated through fraud rings. The naive Bayes assumption is therefore not a description of the world. It is a modelling choice that makes the classifier tractable.
3.1 What goes wrong when features are correlated
Two failure modes appear when the assumption is violated.
- Double-counting evidence. If two features carry overlapping information, the classifier multiplies the likelihoods twice and over-weights the evidence. The resulting posterior is too confident in the winning class.
- Calibration breaks down. Probabilities predicted by a naive Bayes model on correlated data are systematically distorted. The argmax is often correct (because the ranking survives mild violations) but the actual numbers — 0.99, 0.97, 0.83 — cannot be trusted as probabilities.
A concrete example: classifying email as spam. Suppose the word "free" appears in 60% of spam and 5% of ham; the word "viagra" appears in 30% of spam and 0.1% of ham. In real spam "free" and "viagra" co-occur far more often than chance, but naive Bayes treats them as independent. The likelihood product for spam with both words is
The true joint probability is higher (because the words are positively correlated in spam) but the model does not know that. The model still tends to predict spam correctly because both likelihoods point the same direction; the failure shows up when the calibration matters — for instance, when the classifier is part of a thresholded system that rejects borderline cases.
3.2 Why naive Bayes still works despite being wrong
Three reasons.
- The ranking survives. Even when probabilities are mis-calibrated, the relative ordering of classes is often preserved under moderate violations of independence. When only the argmax matters, naive Bayes keeps winning.
- The bias is conservative. When features are positively correlated and the classifier multiplies their likelihoods, the result is a stronger vote for the winning class. The argmax does not flip; the confidence just inflates.
- The variance is low. Independence makes the model have many fewer parameters than a model that learns the joint distribution. With a small training set the low-variance estimator beats a higher-variance joint model.
These three points are why naive Bayes is the textbook "surprising baseline" — a model whose assumptions are wrong but whose decisions are still useful.
4. Maximum A Posteriori Prediction
Putting the prior and the likelihood together, the MAP decision rule is
In practice three implementation details matter.
4.1 Log-space
Numerical underflow happens as soon as the product involves more than a few factors smaller than 1. Working in log-space:
The argmax is unchanged. Underflow disappears. Sums are faster than products. Every production implementation uses this form.
4.2 Laplace (additive) smoothing
A feature value that never appears in the training data for a class produces , which then zeroes out the entire product. To prevent that, add a small count to every per-class count:
where is the number of possible values of feature . With this is Laplace smoothing; smaller values are called add-k smoothing or Lidstone smoothing. The smoothing regularises the maximum-likelihood estimate toward the uniform distribution, which dramatically improves performance on rare words in text classification.
4.3 A two-feature numeric walkthrough
import numpy as np
"""Three classes with priors and per-feature likelihoods."""
priors = np.array([0.5, 0.3, 0.2]) # P(C)
P_x1_given_C = np.array([0.4, 0.1, 0.7]) # P(X1 | C)
P_x2_given_C = np.array([0.2, 0.8, 0.3]) # P(X2 | C)
"""New observation: X1 = yes, X2 = yes."""
log_posterior = np.log(priors) + np.log(P_x1_given_C) + np.log(P_x2_given_C)
prediction = int(np.argmax(log_posterior))
print(f"log posteriors: {log_posterior}")
print(f"predicted class: {prediction}")
The output (log values truncated to three decimals) is roughly
log posteriors: [-2.526 -1.172 -1.561]
predicted class: 1
Class 1 wins even though its prior is the smallest of the three, because its likelihoods dominate. The posterior itself (after normalising) is
which, after dividing by the sum , gives . Class 1 is predicted with 51.7% confidence.
5. The Multinomial Variant for Text Classification
The variant used most often for text classification is the multinomial naive Bayes. The features are word counts in a document: each document is a vector of token counts, and the likelihood of the document given class is
where is the count of token in document and is the document length. The multinomial coefficient does not depend on , so it drops out of the MAP rule.
Taking logs,
The per-class parameter is the probability of token under class , estimated as the smoothed count
with the vocabulary size. The estimator is the maximum-likelihood multinomial with Laplace smoothing. A 1,000-document training corpus with a 20,000-word vocabulary produces at most 20,000 parameters per class — about 60,000 floats total for a three-class problem.
5.1 Worked training: a tiny corpus
Consider a four-document training corpus with vocabulary {free, viagra, meeting, project, the}, classified as spam ({C₁}) or ham ({C₂}):
| Doc | Tokens | Class |
|---|---|---|
| 1 | free, viagra, the | spam |
| 2 | free, meeting, the | spam |
| 3 | meeting, project, the | ham |
| 4 | project, the | ham |
Class priors:
Per-class token counts:
| Token | spam count | ham count |
|---|---|---|
| free | 2 | 0 |
| viagra | 1 | 0 |
| meeting | 1 | 1 |
| project | 0 | 2 |
| the | 2 | 2 |
| total | 6 | 5 |
With Laplace smoothing (, vocabulary size ), the per-token probabilities are
| Token | ||
|---|---|---|
| free | ||
| viagra | ||
| meeting | ||
| project | ||
| the |
5.2 Predicting a new document
A new document is ["free", "viagra"]. The MAP scores in log-space:
Spam wins by a margin of in log-space, which corresponds to a posterior ratio of about in favour of spam. Normalising gives
The classifier predicts spam with 83.2% confidence. With more training data the same algorithm produces reliable spam filters; the toy corpus just shows the arithmetic.
5.3 The full training and prediction loop
import numpy as np
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.naive_bayes import MultinomialNB
"""Toy corpus: four documents and their labels."""
corpus = [
"free viagra the", # spam
"free meeting the", # spam
"meeting project the", # ham
"project the", # ham
]
labels = ["spam", "spam", "ham", "ham"]
"""Convert text to a document-term count matrix."""
vectorizer = CountVectorizer()
X = vectorizer.fit_transform(corpus)
"""Fit multinomial naive Bayes with Laplace smoothing (alpha=1.0 by default)."""
model = MultinomialNB(alpha=1.0)
model.fit(X, labels)
"""Predict a new document."""
new_doc = vectorizer.transform(["free viagra"])
print(model.predict(new_doc)) # ['spam']
print(model.predict_proba(new_doc)) # [[0.168 0.832]]
The predict_proba output reproduces the manual calculation. The whole pipeline —
vectoriser, smoothing, log-space MAP — fits inside a handful of lines.
6. Linking Naive Bayes to TF-IDF Features
The previous section used raw term frequency counts. A common refinement is to weight features by TF-IDF (term frequency–inverse document frequency), which down-weights tokens that appear in nearly every document and up-weights rare-but- informative tokens:
Here is the number of documents containing and is the total number of documents.
Naive Bayes assumes features are non-negative counts (or non-negative frequencies), and TF-IDF values are non-negative — but they are not integer counts. The right way to combine the two is to treat TF-IDF values as scaled counts: feed them to multinomial naive Bayes as if they were counts. Empirically this often hurts slightly compared to raw counts, because the multinomial likelihood model prefers integer count data and the TF-IDF transform violates the assumption.
A more principled pairing is:
- Raw counts Multinomial naive Bayes. The textbook choice.
- TF-IDF weights Multinomial naive Bayes as a pragmatic baseline; expected to be slightly weaker than counts but often close.
- Binary presence/absence Bernoulli naive Bayes. The right choice when document length is noisy.
In scikit-learn the swap is one argument:
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.naive_bayes import MultinomialNB
X = TfidfVectorizer().fit_transform(corpus)
model = MultinomialNB(alpha=1.0).fit(X, labels)
Empirically the accuracy difference between count features and TF-IDF features under multinomial naive Bayes is usually within a percentage point on standard benchmarks like 20 Newsgroups or the Enron spam corpus. The lesson: the choice of model (multinomial vs Bernoulli vs Gaussian) usually matters more than the choice of feature weighting (counts vs TF-IDF).
7. Variants: Gaussian, Multinomial, Bernoulli
The name "naive Bayes" covers three standard variants, distinguished by the assumed distribution of .
| Variant | Feature type | Model for | Typical use |
|---|---|---|---|
| Gaussian | Continuous | , fit by per-class mean and variance | Small continuous feature sets, sensor measurements |
| Multinomial | Non-negative integer counts | Multinomial over the vocabulary; parameters are per-token probabilities | Text classification (bag of words), bag of n-grams, TF-IDF as scaled counts |
| Bernoulli | Binary (presence/absence) | Bernoulli with per-class, per-feature probability of presence | Short documents, sentiment with "word present" features, spam filtering on binary word indicators |
7.1 Gaussian naive Bayes
When features are continuous, the natural model is the Gaussian:
The parameters and are the per-class mean and variance of feature , fit by maximum likelihood on the training set. The classifier is fast but assumes each feature is normally distributed within each class — a strong assumption that rarely holds in practice. Gaussian naive Bayes is a useful baseline for small continuous datasets where the alternative is a kernel density estimate.
7.2 Multinomial naive Bayes
Covered in detail in section 5. Each class is a probability distribution over the
vocabulary, the per-token counts form a multinomial, and the document likelihood is
the multinomial probability of the observed counts. This is the variant of choice
for text classification and is the basis of scikit-learn's MultinomialNB.
7.3 Bernoulli naive Bayes
When features are binary, the natural model is the Bernoulli:
where is the probability that feature is present in class . Bernoulli naive Bayes treats "free appeared" and "free did not appear" symmetrically and uses both signals. It often beats multinomial naive Bayes on short documents, where the absence of common words is informative; it loses on longer documents, where the count signal is more useful than the presence signal.
7.4 Choosing between them
The decision is usually forced by the feature type:
| If features are… | Use… |
|---|---|
| Continuous, roughly Gaussian | Gaussian naive Bayes |
| Counts (word counts, n-gram counts) | Multinomial naive Bayes |
| Binary (presence/absence) | Bernoulli naive Bayes |
| TF-IDF values | Multinomial naive Bayes (as a pragmatic baseline) |
When in doubt, fit all three and pick the best on a held-out validation set. The classifier is fast enough that fitting all three takes seconds.
8. Pros and Cons of Naive Bayes
| Naive Bayes | |
|---|---|
| Pros | Trains in a single pass over the data; each parameter is one count. Predictions are per example. Works on small training sets because the number of parameters is , far fewer than logistic regression's with cross-terms or a neural network's much larger parameter count. Robust to irrelevant features: an irrelevant feature just contributes a near-uniform likelihood that cancels between classes. Handles missing features gracefully by ignoring them at prediction time. Probabilistic output (after renormalisation) that can be thresholded. Tends to win on text classification benchmarks despite its wrong assumption. |
| Cons | The conditional-independence assumption is almost always wrong in real data; the resulting probabilities are mis-calibrated, even if the argmax is correct. Cannot learn interactions between features (the words "New" and "York" are independent given the class, which is not true for location-classification). Zero-frequency problem without smoothing; must choose carefully. Continuous features usually do not satisfy the Gaussian assumption, so Gaussian naive Bayes is a weak baseline outside its narrow regime. |
| When it shines | Text classification (spam, sentiment, topic), small datasets, very high-dimensional sparse features, online or streaming settings where training must be one pass, baseline models that have to be ready in minutes. |
| When it struggles | Strong feature interactions (vision, position-sensitive NLP), calibrated probability estimates, regression-style problems where the target is continuous. |
The headline point: naive Bayes is fast, simple, and surprisingly hard to beat on text. The headline caveat: never trust its probabilities as probabilities.
9. Worked Example: Spam Detection on Tiny Data
Pull the threads together by training a multinomial naive Bayes spam filter on a 20-message toy corpus and predicting the class of a new message.
9.1 Training data
| # | Message tokens | Class |
|---|---|---|
| 1 | free viagra the | spam |
| 2 | free meeting the | spam |
| 3 | meeting project the | spam |
| 4 | free the project | spam |
| 5 | project meeting free | spam |
| 6 | meeting project the | ham |
| 7 | project the | ham |
| 8 | meeting the project | ham |
| 9 | project meeting the | ham |
| 10 | meeting the | ham |
| 11 | project meeting the | ham |
| 12 | meeting the | ham |
| 13 | project the meeting | ham |
| 14 | meeting the | ham |
| 15 | project the | ham |
| 16 | meeting the | ham |
| 17 | project the | ham |
| 18 | meeting the project | ham |
| 19 | project the | ham |
| 20 | meeting the | ham |
There are 5 spam and 15 ham messages. Priors:
9.2 Token counts per class
| Token | spam count | ham count |
|---|---|---|
| free | 4 | 0 |
| viagra | 1 | 0 |
| meeting | 4 | 11 |
| project | 4 | 11 |
| the | 5 | 13 |
| total | 18 | 35 |
9.3 Smoothed likelihoods (, )
| Token | ||
|---|---|---|
| free | ||
| viagra | ||
| meeting | ||
| project | ||
| the |
9.4 Predicting a new message
A new message is ["free", "viagra"]. The MAP scores in log-space:
The difference is in log-space, equivalent to a posterior ratio of in favour of spam. Normalising gives
The classifier predicts spam with 91% posterior. The two signals — the rare word "viagra" and the elevated rate of "free" in spam — both push in the same direction; the prior of 25% spam drags the posterior down from near 1.0 but not enough to flip the prediction.
9.5 Comparison with logistic regression
A logistic regression on the same data (one-hot encoded message, regularised) typically lands within one percentage point of the multinomial naive Bayes accuracy on a held-out set. The naive Bayes model has 5 parameters per class (the per-token probabilities plus the prior), while logistic regression has 10 weights per class (an intercept plus one per token). With only 20 training examples the smaller parameter count of naive Bayes gives it a small variance advantage; with 20,000 examples the two are indistinguishable.
9.6 A reusable pipeline
import numpy as np
from sklearn.feature_extraction.text import CountVectorizer
from sklearn.model_selection import cross_val_score
from sklearn.naive_bayes import MultinomialNB
from sklearn.pipeline import make_pipeline
"""Real-world pipeline: text in, class out."""
pipeline = make_pipeline(
CountVectorizer(lowercase=True, ngram_range=(1, 2)),
MultinomialNB(alpha=1.0),
)
"""5-fold cross-validation on a labelled corpus."""
scores = cross_val_score(pipeline, texts, labels, cv=5, scoring="accuracy")
print(f"mean accuracy: {scores.mean():.3f} +/- {scores.std():.3f}")
The pipeline converts raw text to a count matrix, fits multinomial naive Bayes with Laplace smoothing, and reports 5-fold cross-validation accuracy. On the Enron spam corpus this combination reaches 98–99% accuracy with a few seconds of training, outperforming logistic regression trained on the same features at a fraction of the cost. That empirical result is the lesson's headline: naive Bayes is a strong baseline precisely because it is fast, simple, and almost always at least as good as a heavier model on text.
Key Takeaways
- Bayes' theorem rewrites the posterior as the prior times the likelihood , divided by the evidence ; classification only needs the unnormalised product because the evidence is constant across classes.
- The conditional-independence assumption is almost always false in real data; when features are correlated the classifier over-counts evidence and the predicted probabilities are mis-calibrated, although the argmax often survives mild violations.
- The prior matters even when the likelihood signal is strong. A positive medical test with 99% sensitivity and 95% specificity only implies a 17% chance of disease when the base rate is 1%.
- The multinomial variant is the workhorse for text classification; counts (or TF-IDF weights used as scaled counts) feed into per-class multinomial likelihoods and the MAP rule is computed in log-space.
- The three standard variants — Gaussian (continuous), multinomial (counts), Bernoulli (binary) — are picked by feature type; the right choice on text data is usually multinomial with raw counts.
- Naive Bayes is fast, simple, and tolerant of small training sets and high-dimensional sparse features, which is why it remains a strong baseline on text classification. Never trust its output probabilities as probabilities.