Part I · Predicting Words

Generating Text and Handling Unseen Words

Follow a model as it writes one word at a time, then see why unfamiliar words and word pairs cause trouble.

Generate a sentence by drawing one word at a time Lec 2 · Aug 27 · recapped Lec 3 · Sep 1

A bigram probability table can do more than score a sentence. It can also generate one. Start at the beginning marker, draw a next word according to its probability, and use that word to decide which distribution to draw from next.

This is called sampling. A more probable word is more likely to be drawn, but it does not have to win every time. With a table such as the one learned from the Berkeley Restaurant Corpus, the steps are:

  1. Start with <s>. Draw the first word using the probabilities of bigrams \((\text{<s>}, w)\).
  2. After choosing that word, draw its successor using the probabilities of \((w, x)\).
  3. Repeat, each time using the latest word as the context, until the draw produces </s>.
  4. Join the generated words to form the sentence.

One possible path is <s> I → I want → want to → to eat → eat Chinese → Chinese food → food </s>. It gives “I want to eat Chinese food”. To practice sampling, try drawing a few steps from a probability table yourself; a new run can produce a different sentence.

Why the generated text starts to sound like Shakespeare Lec 2 · Aug 27

The slides train n-gram models on Shakespeare and compare their generated text. As \(n\) grows, the samples look more like the source: unigrams produce disconnected words, bigrams produce local phrases, and trigrams and 4-grams can produce lines that resemble Shakespeare.

Before concluding that the model has learned to write well, look at how much of the space of possible word combinations it has actually seen:

What is counted?Value
Word occurrences, or tokens \(N\)884,647
Vocabulary size \(V\)29,066
Possible bigrams \(V^2\)≈ 844 million
Distinct bigrams that actually appear300,000
Share of possible bigrams never seen99.96%

Even with nearly 900,000 word occurrences, almost every possible pair is missing. Four-word groups, also called quadrigrams, are rarer still. When a model uses a long context from this small corpus, there may be very few observed ways to continue it. The generated passage can therefore repeat chunks of the original text.

That is the point of the slides' remark that it looks like Shakespeare because it is Shakespeare: much of the apparent success comes from reproducing familiar passages. Most possible n-grams remain unseen.

Doing well on familiar text is only part of the job Lec 2 · Aug 27 · Lec 3 · Sep 1

If a very large n-gram window produces convincing samples, why does that not settle the language-modeling problem or remove the need for models such as GPT? The problem is overfitting: the model can depend too heavily on the particular text it was trained on.

  • A change of domain can break familiar patterns. A model trained on Shakespeare does not have the same vocabulary and usage as one trained on the Wall Street Journal, and vice versa. As the slides put it, “The WSJ is no Shakespeare.” A fixed vocabulary also cannot produce words it does not contain.
  • We want generalization: useful predictions on new text. Doing well on one similar test set is narrower evidence than doing well across the new situations we care about. It does not mean any model can promise success on every arbitrary test set.
  • Missing counts expose this problem directly. A perfectly reasonable word or phrase may first appear when the model is tested.

What happens when a reasonable phrase has count zero? Lec 2 · Aug 27 · Lec 3 · Sep 1

The model has seen “denied the,” but not this continuation

Training text: … denied the allegations / … denied the reports / … denied the claims / … denied the request.
Test text: … denied the offer / … denied the loan.

With raw count estimates, \(P(\text{offer} \mid \text{denied the}) = 0\). Multiplying this factor into the other probabilities makes the whole test sequence's probability zero. Perplexity then has no finite value: the inverse probability diverges to infinity.

Lecture 3 separates two reasons a count can be missing:

  1. The word itself is unknown. It has zero unigram count and is absent from the training vocabulary: an out-of-vocabulary (OOV) word. New terms, names, dialect forms, and changes in language make this common. A token here means one occurrence of a word in the text.
  2. The words are known, but the combination is new. The bigram or longer n-gram has zero count even though the individual words appeared in training.

If the context has a positive count but the continuation never occurred, the raw estimate is zero. If the context itself never occurred, the count ratio has a zero denominator and is undefined. Either situation needs handling before we can assign useful probabilities to new text.

The lecture introduces two complementary tools: use <UNK> to represent unknown words, and use smoothing to give probability to unseen combinations.

Give unknown words a shared representation

The special symbol <UNK> stands for an unknown word. Rather than waiting until testing to introduce it, we include examples of it during training:

  • Choose a small frequency threshold \(n\). Replace each training word that occurs fewer than that many times with <UNK>, then recount the text and estimate the probabilities again.
  • At test time, map any word outside the retained vocabulary to <UNK>.
  • Compare models carefully. Making the vocabulary tiny puts many different words into one large <UNK> category and can artificially lower perplexity. The prediction problem has become easier, even if the model has not become better at distinguishing actual words.

This lets the model assign probabilities to an unknown-word category. It does not recover the identity of each new word, and combinations involving <UNK> may still need smoothing.

Decide how the model should handle words outside its vocabulary

Closed vocabularyOpen vocabulary
Choose a vocabulary in advance, for example from a dictionary, and assume the words to be modeled belong to it. New names, spelling errors, and dialect forms violate that assumption unless an extra handling rule is provided.Allow for words not seen in training. A word-level model can map them to <UNK>; modern models often split new words into smaller known pieces through subword tokenization. The whole-word vocabulary can be open even when the set of pieces is fixed.
What to remember
  • Generate text by repeatedly sampling the next word until the model chooses </s>.
  • A long context can make samples sound convincing by reproducing training passages. That alone does not show that the model will predict new text well.
  • A zero-probability event makes perplexity infinite. Use <UNK> for unknown words and smoothing for unseen combinations.