Part II · Learning to Classify Text

Choosing among Several Classes with Softmax

Give each possible class a score, turn the scores into probabilities, and use the correct answer to train the model.

What if there are more than two answers? Lec 5 · PDF pp. 20-22

A review can be positive, negative, or neutral. A word can be a noun, verb, adjective, or another part of speech. These tasks need more than the two answers provided by binary logistic regression.

Multinomial logistic regression extends the model to \(K\) possible classes. Here, the classes are mutually exclusive: we choose one answer for each input. Their probabilities must therefore add up to 1:

\[ \sum_{c=1}^{K} P(y=c\mid x)=1 \]

Giving each class its own independent sigmoid would not ensure that total. Instead, we use softmax, which converts all the class scores into one shared probability distribution. Raising one class's score increases its share and changes the other classes' probabilities too.

Give each class its own set of weights Lec 5 · PDF pp. 25-27

All classes see the same input features \(x\), but they can interpret those features differently. Class \(c\) has its own weights \(w_c\) and bias \(b_c\), giving a score \(z_c=w_c^\top x+b_c\). This score, before it is turned into a probability, is called a logit.

Put the weights for each class in one row of a matrix \(W\). With \(K\) classes and \(d\) features, we need \(K\) rows and \(d\) columns. Using column vectors for inputs and outputs, the whole calculation becomes:

Use one row of weights for each class
\[ x\in\mathbb{R}^{d},\qquad W=\begin{bmatrix}w_1^\top\\ \vdots\\ w_K^\top\end{bmatrix}\in\mathbb{R}^{K\times d},\qquad b\in\mathbb{R}^{K} \] \[ z=Wx+b\in\mathbb{R}^{K},\qquad \hat y=\operatorname{softmax}(z)\in\mathbb{R}^{K} \]
QuantityShapeWhat it contains
\(x\)\(d\times1\)The \(d\) features of one input.
\(W\)\(K\times d\)One row of weights for each class.
\(b,z,\hat y\)\(K\times1\)Respectively, one bias, score, or probability per class.

This stores \(Kd+K\) parameters: \(Kd\) weights and \(K\) biases. For three sentiment classes and six features, \(W\) has shape \(3\times6\). Reversing it to \(6\times3\) would not match the column-vector convention used here.

An exclamation mark can support more than one kind of sentiment

In the lecture, \(x_5=1\) means that a document contains an exclamation mark. The corresponding weights are \(3.5\) for positive sentiment, \(3.1\) for negative sentiment, and \(-5.3\) for neutral sentiment.

This makes sense: an exclamation mark can express strong approval or strong disapproval, while being less typical of a neutral statement. The weights change the class scores; they are not probabilities themselves.

Turn the scores into probabilities with softmax Lec 5 · PDF pp. 22-25; Lec 6 · PDF pp. 4-6

Softmax has two steps. First, take the exponential of each score, making every value positive. Then divide each value by their total. The result is one probability for each class:

The probability assigned to class c
\[ \hat y_c=P(y=c\mid x;W,b)=\frac{\exp(z_c)}{\sum_{j=1}^{K}\exp(z_j)}=\frac{\exp(w_c^\top x+b_c)}{\sum_{j=1}^{K}\exp(w_j^\top x+b_j)} \]
  • \(\exp(z_c)\), also written \(e^{z_c}\), is positive even if the score \(z_c\) is negative.
  • All classes use the same denominator, \(\sum_j\exp(z_j)\), which makes their probabilities add up to 1.
  • Choose the class with the largest probability: \(\hat c=\arg\max_c\hat y_c=\arg\max_c z_c\). If classes tie, use a consistent tie-breaking rule.

For finite scores and \(K>1\), every probability is strictly between 0 and 1. The order stays the same: the highest score gives the highest probability, but it does not get probability 1 automatically.

Work through the lecture's six-class example Lec 5 · PDF p. 24

Number the classes from 1 to 6. The first line gives their scores. The second takes their exponentials, the third adds them up, and the last divides each exponential by that total:

\[ z=[0.6,\;1.1,\;1.5,\;1.2,\;3.2,\;1.1] \] \[ \exp(z)\approx[1.8221,\;3.0042,\;4.4817,\;3.3201,\;24.5325,\;3.0042] \] \[ \sum_{j=1}^{6}\exp(z_j)\approx40.165 \] \[ \hat y\approx[0.0454,\;0.0748,\;0.1116,\;0.0827,\;0.6108,\;0.0748] \]

Class 5 has the largest probability, \(24.5325/40.165\approx0.6108\), so it is the prediction. Classes 2 and 6 have the same score, so they also have the same probability. The displayed values add up to approximately 1 because they have been rounded; before rounding, the sum is 1.

Keep the computation stable by subtracting the largest score

Very large exponentials can exceed what a computer can represent. We can avoid much of this problem without changing the probabilities: subtract the largest score from every score first.

Set \(m=\max_j z_j\), then compute \(\hat y_c=\exp(z_c-m)/\sum_j\exp(z_j-m)\). The same factor cancels from the numerator and denominator, so the result is unchanged. Now the largest exponential is only \(e^0=1\). This practical step follows directly from the lecture's formula; adding or subtracting any common constant leaves softmax unchanged.

Measure how much probability went to the correct answer Lec 5 · PDF p. 28

For training, we need to mark which class is correct. A one-hot label does this with a vector \(y\in\{0,1\}^{K}\): put 1 at the correct class \(c\), so \(y_c=1\), and 0 everywhere else. In contrast, the prediction \(\hat y\) contains probabilities and has positive entries for finite scores.

We use cross-entropy to penalize the model when it gives the correct class little probability. The one-hot label selects just that class's log probability:

Cross-entropy for one training example
\[ L_{CE}(\hat y,y)=-\sum_{k=1}^{K}y_k\log\hat y_k=-\log\hat y_c \] \[ L_{CE}=-z_c+\log\left(\sum_{j=1}^{K}e^{z_j}\right) \]

The other terms disappear because their labels are zero. Thus minimizing the loss is the same as maximizing the conditional probability of the correct answer, just as in the binary model. With \(N\) training examples, we minimize their average loss and may also add regularization.

Use the probabilities from the six-class example

If class 5 is correct, its label is \(y=[0,0,0,0,1,0]\), and the loss is \(L=-\log(0.6108)\approx0.493\). If class 1 is correct instead, the loss is \(L=-\log(0.0454)\approx3.09\).

The prediction has not changed, but the penalty is much larger when the true answer received little probability. These calculations use natural logarithms.

The gradient still compares prediction with truth

Differentiating the model and loss above gives a familiar pattern. For each class, subtract its label from its predicted probability. Multiply that difference by an input feature to obtain the corresponding weight gradient:

\[ \frac{\partial L}{\partial z_k}=\hat y_k-y_k,\qquad \frac{\partial L}{\partial W_{kj}}=(\hat y_k-y_k)x_j,\qquad \frac{\partial L}{\partial b_k}=\hat y_k-y_k \]

These are consequences of the softmax and cross-entropy formulas. Gradient descent or mini-batch SGD can then use them to update the parameters and reduce the loss.

With two classes, softmax becomes sigmoid

Set \(K=2\), with scores \(z_1\) and \(z_0\). Divide the numerator and denominator of class 1's probability by \(e^{z_1}\). What remains is a sigmoid of the difference between the scores:

\[ P(y=1\mid x)=\frac{e^{z_1}}{e^{z_1}+e^{z_0}}=\frac{1}{1+e^{-(z_1-z_0)}}=\sigma(z_1-z_0) \]

So the binary model only needs the difference \(z_1-z_0=(w_1-w_0)^\top x+(b_1-b_0)\). If we set the reference score \(z_0=0\), we recover \(\sigma(w^\top x+b)\). Softmax extends the same idea to a choice among several classes.

How this leads to word embeddings

A classifier chooses an answer from a fixed set. If that set is a vocabulary, each word becomes a class and \(K=|V|\). Given a learned representation of a word or its context, the model can score the candidate words and use softmax to assign probabilities to them.

The roles are different: an embedding is a vector that represents a word, while the softmax output is a distribution over possible answers. By training a word-prediction task, we can learn useful representations. The next topic, word embeddings and vector semantics, explains why the words that appear nearby can tell us something about meaning.

What to remember
  • Each class gets its own weights and score: \(W\in\mathbb{R}^{K\times d}\) and \(b\in\mathbb{R}^{K}\).
  • Softmax takes exponentials and divides by their total to produce probabilities.
  • With a one-hot label, cross-entropy is \(-\log\) of the probability assigned to the correct class.
  • For two classes, this is the same as applying sigmoid to the score difference.
  • Choosing a word from a vocabulary is also a classification task, which can help us learn word representations.

Sources: Lecture 5, PDF pages 20–28, covers the model, example, class weights, and loss. Lecture 6, PDF pages 4–6, reviews softmax. These page numbers count from the start of the PDF, rather than using the printed slide numbers. The stable computation, gradient formulas, sigmoid derivation, and connection to embeddings are additional explanations derived from the presented model.