Part II · Learning to Classify Text

Helping a Model Work on New Examples

A model can fit its training examples too closely. Regularization discourages overly large weights so it can work better on new data.

Lecture 5 reviews this material on PDF pages 17–19. The next topic, multinomial logistic regression and softmax, extends the classifier to more than two classes.

Doing well on the training set is not enough Lec 4 · Sep 3

Suppose “wow” happens to appear only in positive reviews in our training set. A model may give it a very large weight and rely on that coincidence too heavily. When new reviews use the word differently, or do not use it at all, the model may perform poorly.

This is the risk of overfitting: learning accidental details or noise along with the useful pattern. A perfect fit to the training set does not automatically imply overfitting, but it does not prove the model is good either. What matters is generalization, or how well it handles new examples.

The training objective and the real goal

Given training data \(D=\{(x^{(i)},y^{(i)})\}\), we fit a function \(f(x)\) by minimizing a loss. The difficulty is that the loss we can optimize directly is measured on \(D\), while the performance we care about is on examples the model has not seen.

Too simple, a good fit, or too complicated?

UnderfittingA useful balanceOverfitting
The model is too simple for the pattern. For example, features \([1,x]\) describe a line, which may be unable to fit the training data well.The model captures the underlying pattern and predicts new examples well.The model follows training details too closely. In the lecture's illustration, it achieves 0 training loss but makes erratic predictions between or beyond the observed examples.

Regularization puts a cost on overly large weights Lec 4 · Sep 3

One way to reduce overfitting is to discourage the model from relying too strongly on particular features. Regularization does this by adding a penalty to the training loss. The model must now balance fitting the examples against keeping its parameters reasonably small:

Average prediction loss plus a parameter penalty
\[ \hat\theta = \arg\min_\theta \ \underbrace{\frac{1}{N}\sum_{i=1}^N L\big(f(x^{(i)};\theta), y^{(i)}\big)}_{\text{original loss}} \ +\ \lambda\, R(\theta) \]
  • The first term is the average prediction loss. The second term, \(\lambda R(\theta)\), is the added penalty.
  • \(\lambda\) controls the strength of the penalty. It can be any number \(\ge0\): at \(\lambda=0\), there is no regularization; larger values put more emphasis on the penalty. We choose it as a hyperparameter, rather than learning it as an ordinary model weight.
  • \(R(\theta)\) measures the size of the parameters. Two common choices are the squared L2 norm and the L1 norm, explained below.

L2: make large weights more expensive

L2 regularization adds up the squares of the weights. A large weight therefore contributes much more to the penalty than a small one. Differentiating this penalty adds \(2\lambda w\) to the gradient:

Add up the squared weights: the squared L2 norm of \(w\)
\[ R(w) = \|w\|_2^2 = \sum_{j=1}^{d} w_j^2 \]
Add the penalty gradient to the loss gradient
\[ \nabla_w \big( L + \lambda \|w\|_2^2 \big) = \nabla_w L + 2\lambda w \]
The resulting gradient descent update
\[ w \leftarrow w - \eta\,\nabla_w L - 2\eta\lambda\, w \]

The extra update subtracts a multiple of \(w\) itself. This contribution points toward zero, and a sufficiently small step shrinks the weights. That is why the effect is called weight decay. It discourages unusually large weights and helps keep the model from relying too heavily on one feature.

The displayed equivalence between an L2 penalty and weight decay is for ordinary gradient descent. Other optimizers can implement weight decay differently.

L1: encourage some weights to become zero

L1 regularization adds the absolute values of the weights instead of their squares. Away from zero, its derivative depends only on the sign of a weight, not its size:

Add up absolute weight values: the L1 norm of \(w\)
\[ R(w) = \|w\|_1 = \sum_{j=1}^{d} |w_j| \]
The sign-based derivative and update away from zero
\[ \frac{\partial}{\partial w_j} \lambda \|w\|_1 = \lambda \, \mathrm{sign}(w_j) \qquad\Rightarrow\qquad w_j \leftarrow w_j - \eta\,\frac{\partial L}{\partial w_j} - \eta\lambda\,\mathrm{sign}(w_j) \]

At \(w_j=0\), the absolute-value function has a corner and no ordinary derivative. We use a subgradient: for the penalty, any value in \([-\lambda,\lambda]\) is valid there, including 0. Thus the displayed sign-based expression applies away from zero, with 0 as one valid subgradient choice at zero.

How the two penalties differ

QuestionL1L2
How does the penalty change a nonzero weight?It contributes a step of fixed size \(\eta\lambda\) toward zero, regardless of the weight's magnitude.The contribution is proportional to the weight: larger weights receive a larger adjustment.
What kind of solution does it favor?A sparse solution, with many weights exactly zero. A basic subgradient update can cross zero without landing on it, so exact zeros depend on the solution and the optimization method.Smaller weights throughout the model. It generally does not force them to be exactly zero.
Why does that matter?When \(w_j=0\), feature \(j\) no longer affects the prediction. This is feature selection: L1 can remove features the fitted model does not need.It reduces the influence of large weights more smoothly. Weight decay is a common regularization approach in deep learning.
The connection to smoothing

Regularization plays a similar role to smoothing in n-gram models: both accept a slightly worse fit to the training data in the hope of doing better on new examples. Both also introduce settings to choose, such as \(\lambda\) here, or \(k\) and interpolation weights in smoothing. Tune those settings on held-out validation data, never on the test set.

What to remember
  • The goal is to predict new examples well, not simply to drive training loss to zero.
  • Regularization adds \(\lambda R(\theta)\) to the loss. L2 discourages large weights; L1 can produce sparse weights and select features.
  • Choose the strength \(\lambda\) using held-out validation data.