Lecture 5 reviews this material on PDF pages 7–13. Once the binary model is clear, multinomial logistic regression and softmax show how to handle more than two classes.
II. Build a model that gives a probability Lec 4 · Sep 3
A review may contain several clues about its sentiment. Logistic regression combines those clues into one score. Each feature \(x_i\) is multiplied by a learned weight \(w_i\), and the products are added together.
We also add a bias \(b\), sometimes written \(w_0\). It lets the model shift the score even when the input features stay the same. Unlike the other weights, it is not attached to a particular feature. Together, the parameters are \(\theta = [w; b]\):
This score \(z\) can be any real number, from \(-\infty\) to \(\infty\). A probability cannot. We still need to turn it into \(P(y=1 \mid x; \theta)\), the probability of class 1, and \(P(y=0 \mid x; \theta)\), the probability of class 0.
Use sigmoid to turn the score into a probability
The sigmoid, also called the logistic function, maps a real-valued score to a number between 0 and 1. Higher scores give higher probabilities:
Three properties are especially useful:
- Its graph is an S-shaped curve, not a straight line. Every finite input maps to \((0,1)\), so the output can represent a probability.
- It has a derivative, which lets us use gradients to train the model. The derivative has a simple form: \(\sigma'(z) = \sigma(z)(1-\sigma(z))\).
- The probabilities on the two sides fit together: \(1 - \sigma(z) = \sigma(-z)\).
Applying sigmoid to the weighted score gives the binary logistic regression model. The second class gets whatever probability is left:
We often write the estimated probability of class 1 as \(\hat y = \sigma(w \cdot x + b)\).
Turn the probability into a class decision
With the usual threshold of 0.5, we choose class 1 when its probability exceeds one half. Otherwise, including a tie, we choose class 0. In the decision rule below, \(\hat y\) denotes the final class label:
The dividing line is \(w \cdot x + b = 0\). With two features it is a line; in more dimensions it is a hyperplane, the higher-dimensional version of a flat boundary. This is why logistic regression is a linear classifier, even though sigmoid itself is curved.
How do we choose the weights and bias?
For a training example, we already know the correct label \(y\), either 0 or 1. The model gives an estimated probability \(\hat y\). We want weights that make these predictions fit the known answers. To do that, we need:
- a loss function to score how poor a prediction is; this is part III, below;
- an optimization algorithm to change \(w\) and \(b\) so the loss goes down; this is part IV, on the next page.
III. Measure errors with cross-entropy Lec 4 · Sep 3
A loss assigns a numerical penalty to a prediction. Here we compare the model's probability \(\hat y = \sigma(w \cdot x + b)\) with the correct label \(y \in \{0, 1\}\), also called the ground truth or gold label. We write the penalty as \(L(\hat y, y)\). The loss, sometimes called the cost, should be small when the model puts high probability on the right answer.
Give the observed answers as much probability as possible
Recall the coin example: after observing (H, H, H, T), choosing \(p = 0.75\) makes that data most likely. We use the same idea here, but now the answer's probability depends on the input.
This is conditional maximum likelihood estimation: choose \(w,b\) to maximize the probability of the correct labels given their inputs. We can maximize log probabilities instead because taking a logarithm preserves their order. For one example, the expression is:
For the full training set, we add this log probability across examples. To write the single-example probability compactly, remember that there are only two possible labels. If \(y=1\), the model assigns the correct answer probability \(\hat y\); if \(y=0\), it assigns \(1-\hat y\). Both cases fit in one expression:
The exponents select the relevant term. When the label is 1, the factor raised to \(1-y\) becomes 1; when the label is 0, the factor raised to \(y\) becomes 1. After taking logs, the unused term is multiplied by zero. The result tells us how much probability the model assigns to the observed answer.
Change the sign to get a loss we can minimize
We want a large log probability but a small loss. Negating the log probability turns one goal into the other. This gives binary cross-entropy, which is also the negative log-likelihood of the correct label:
Suppose \(y=1\). If the model predicts \(\hat y=0.9\), the loss is \(-\log 0.9\approx0.105\). If it predicts \(\hat y=0.1\), the loss is \(-\log 0.1\approx2.30\). Putting little probability on the correct answer is much more costly.
The loss reaches 0 at a probability of 1 for the correct class. With a finite sigmoid score, that probability is approached rather than reached. As the probability of the correct class approaches 0, the loss grows without bound.
- First compute a weighted score, then use \(\hat y = \sigma(w \cdot x + b)\) to get a probability. A threshold of 0.5 turns it into a class decision.
- Cross-entropy is the negative log probability of the correct answer. Its derivation uses the same maximum-likelihood idea we used for n-gram probabilities, now conditioned on an input.
- Training chooses \(\theta = [w; b]\) to reduce average cross-entropy on the training examples. Next, gradient descent explains how to make those updates.