How likely is this sentence? Lec 2 · Aug 27 · recapped Lec 3 · Sep 1
After reading the beginning of a sentence, we expect some words more than others. A language model puts numbers on those expectations. It assigns probabilities to word sequences, which lets us compare complete sentences or predict the next word.
These two views are closely connected. We can ask for the probability of a whole sequence, or ask how likely its last word is after the words before it:
Both are ways to describe a language model. For example, \(P(\text{its water is so transparent that you can see the bottom})\) means the joint probability \(P(\text{its}, \text{water}, \text{is}, \ldots, \text{bottom})\): the probability of those words occurring together in that order.
Why put probabilities on words? Lec 2 · Aug 27
Probabilities help whenever a system must choose among several possible word sequences. They can guide what the system writes, or help it interpret an uncertain input.
- Generating text: the model's next-word probabilities give us a way to choose what comes next.
- Recognizing speech: two sequences may sound alike, but one makes more sense as language. For example, \(P(\text{I saw a van}) \gg P(\text{eyes awe of an})\).
- Correcting writing: the surrounding words help distinguish a typo from the intended word. We expect \(P(\text{I'm about fifteen minutes away}) > P(\text{I'm about fifteen minuets away})\) and \(P(\text{high winds tonight}) > P(\text{large winds tonight})\). The same idea helps compare “I will be back soonish” with “I will be bassoon dish,” or “Your so silly” with “You're so silly.”
- Summarizing and answering questions: these tasks also involve choosing suitable word sequences.
That explains why the probabilities are useful. Next, we need a way to calculate them from text.
Build a sentence probability one word at a time Lec 2 · Aug 27
To get the probability of “its water is so transparent,” first consider how likely its is. Then consider water after its, then is after its water, and continue. Multiplying these conditional probabilities gives the probability of the whole sequence.
\(P(\text{its water is so transparent}) = P(\text{its}) \cdot P(\text{water} \mid \text{its}) \cdot P(\text{is} \mid \text{its water}) \cdot P(\text{so} \mid \text{its water is}) \cdot P(\text{transparent} \mid \text{its water is so})\)
This is the chain rule of probability. The compact formula below does the same multiplication for a sentence of any length. At the first position, there are no previous words, so the first factor is simply the probability of the first word.
The chain rule is an exact identity. So far, we have not assumed that earlier words can be ignored.
Learn probabilities by counting what happened Lec 2 · Aug 27 · reused Lec 4 · Sep 3
Start with a simpler problem. A coin lands heads with some unknown probability \(p\). Four flips give (H, H, H, T). A natural estimate is three heads out of four flips: \(p = 3/4 = 0.75\).
There is a mathematical reason for that choice. The probability of observing this particular sequence is \(p \cdot p \cdot p \cdot (1-p)\). Different choices of the parameter give different probabilities to the data. Differentiate this expression and set the derivative to zero: the maximum occurs at \(p = 0.75\).
This is maximum likelihood estimation (MLE): choose the parameter values that make the observed data most probable. It is an estimate, not a guarantee about the coin; even a fair coin with heads probability 0.5 could have produced those four flips.
Predicting which word follows a word \(w\) uses the same idea, with one possible outcome for each word in the vocabulary instead of just heads and tails. Count how often each outcome followed that context, then divide by the total number of observations of the context. This idea appears again as conditional maximum likelihood when we derive the loss for logistic regression.
Why counting whole sentence prefixes runs out of data
In principle, we could estimate a next-word probability by counting an entire prefix and the word that followed it:
The ratio makes sense when the prefix occurs in our data. The difficulty is that there are far too many possible prefixes. A long one may occur only once, or never at all. Its count then gives us either very little evidence or a zero denominator. We need to share evidence across sentences by making the context shorter.
Keep only the most recent words Lec 2 · Aug 27 · recapped Lec 3 · Sep 1
Instead of keeping “its water is so transparent,” suppose we keep only “transparent,” or perhaps “so transparent.” We then have more chances to see that shorter context elsewhere in the training text.
This simplification is the Markov assumption, named after Andrei Markov. We approximate the next-word probability using only the last few words:
For our example, \(P(\text{that} \mid \text{its water is so transparent}) \approx P(\text{that} \mid \text{transparent})\), or \(\approx P(\text{that} \mid \text{so transparent})\). Unlike the chain rule, this is an approximation: it deliberately ignores some earlier words.
A first-order model keeps just the previous word. Watch the counting convention in the slides: a \(k\)-gram model uses the previous \(k-1\) words, because its window includes the word being predicted. “Order” counts the context words; “gram” counts the whole window.
- A language model assigns probabilities to sentences and next words.
- Those probabilities help with generation, ambiguous input, and writing tools.
- Maximum likelihood estimates probabilities from observed counts.
- The Markov assumption shortens the context so those counts are easier to collect.
- The next page puts these ideas together in an n-gram language model.