Index

Glossary & Formula Sheet

A quick reminder of what each term means and when to use each formula. Follow the links for examples and a fuller explanation.

Each formula below has a short reminder of what it does. Jump to the term definitions if you need a word explained, or follow a formula’s title for the full lesson.

What each formula does

Each factor asks for the chance of the next word, given all the preceding words. Multiplying the factors gives the joint probability.

\[ P(w_1 \ldots w_n) = \prod_{k=1}^{n} P(w_k \mid w_1 \ldots w_{k-1}) \]

The approximation limits the context to the most recent k words. Here k counts previous words, so an n-gram model uses k=n−1.

\[ P(w_i \mid w_1 \ldots w_{i-1}) \approx P(w_i \mid w_{i-k} \ldots w_{i-1}) \]

Divide the count of the pair by the count of its first word as a context. The denominator must be nonzero, and the counting convention must be consistent.

\[ P(w_i \mid w_{i-1}) = \frac{c(w_{i-1} w_i)}{c(w_{i-1})} \]

W is the test sequence and N is the number of predicted tokens. The middle form uses the bigram assumption. The exp form uses natural logarithms. Compare scores only under the same evaluation setup.

\[ \mathrm{PP}(W) = P(w_1 \ldots w_N)^{-1/N} = \sqrt[N]{\prod_{i=1}^{N} \frac{1}{P(w_i \mid w_{i-1})}} = \exp\Big(-\tfrac{1}{N}\sum_i \log P(w_i \mid w_{i-1})\Big) \]

r is the frequency rank and s controls how quickly frequency falls. This helps explain why even a large text collection contains many rare words.

\[ \text{freq}(r) \propto r^{-s} \]

Add k to every possible next-word count. There are V vocabulary choices, so the denominator increases by kV. Setting k=1 gives add-one smoothing.

\[ P_{\text{Add-1}}(w_i \mid w_{i-1}) = \frac{c(w_{i-1}w_i) + 1}{c(w_{i-1}) + V} \qquad P_{\text{Add-}k}(w_i \mid w_{i-1}) = \frac{c(w_{i-1}w_i) + k}{c(w_{i-1}) + kV} \]

Multiply the smoothed probability by the original context count. The result makes it easier to see how much the smoothing changed the original evidence.

\[ c^*(w_{i-1}w_i) = \frac{[c(w_{i-1}w_i)+1]\, c(w_{i-1})}{c(w_{i-1}) + V} \]

The lambdas weight the unigram, bigram, and trigram predictions. Each weight must be nonnegative and the weights must sum to 1; choose them using held-out data.

\[ \hat P(w_n \mid w_{n-2}w_{n-1}) = \lambda_1 P(w_n) + \lambda_2 P(w_n \mid w_{n-1}) + \lambda_3 P(w_n \mid w_{n-2}w_{n-1}), \quad \textstyle\sum_i \lambda_i = 1 \]

The score is the weighted feature sum plus a bias. Sigmoid maps it between 0 and 1; the other class gets the remaining probability.

\[ \sigma(z) = \frac{1}{1+e^{-z}}, \qquad P(y=1 \mid x) = \sigma(w \cdot x + b), \qquad P(y=0 \mid x) = \sigma(-(w \cdot x + b)) \]

The gold label y is 0 or 1. That leaves only the term for the correct class, so a low probability for the correct answer causes a large loss.

\[ L_{CE}(\hat y, y) = -\big[ y \log \hat y + (1-y) \log (1-\hat y) \big] \]

First subtract the true label, 0 or 1, from the predicted probability. For a weight, multiply this difference by its feature value. For the bias, the difference itself is the gradient. These are gradients for one example.

\[ \frac{\partial L_{CE}}{\partial w_j} = \big[\sigma(w \cdot x + b) - y\big]\, x_j, \qquad \frac{\partial L_{CE}}{\partial b} = \sigma(w \cdot x + b) - y \]

Subtract the gradient, scaled by learning rate η, from the current parameters θ. A suitable step size matters: taking a bigger step does not always help.

\[ \theta \leftarrow \theta - \eta \nabla_\theta L \]

λ controls regularization strength. L2 adds a shrinkage term proportional to the weight. For L1, sign(w) applies away from zero; at zero use a subgradient. These expressions omit the bias penalty.

\[ L + \lambda \|w\|_2^2 \ \Rightarrow\ w \leftarrow w - \eta \nabla L - 2\eta\lambda w \qquad\qquad L + \lambda \|w\|_1 \ \Rightarrow\ w_j \leftarrow w_j - \eta \tfrac{\partial L}{\partial w_j} - \eta\lambda\,\mathrm{sign}(w_j) \]

Each class c has its own weights and bias. Exponentiate its score and divide by the sum over all K classes, so the resulting probabilities sum to 1.

\[ \hat y_c = \frac{\exp(w_c^\top x+b_c)}{\sum_{j=1}^K \exp(w_j^\top x+b_j)} \]

The label is one-hot: only the correct class c* has value 1. The sum therefore reduces to the negative log probability of that class.

\[ L = -\sum_{c=1}^K y_c\log\hat y_c = -\log\hat y_{c^*} \quad (y_{c^*}=1) \]

Divide the dot product by the two vector lengths. This removes the effect of overall length. Both vectors must be nonzero.

\[ \cos(u,v) = \frac{u^\top v}{\|u\|_2\|v\|_2}, \qquad u\ne 0,\ v\ne 0 \]

For a:b::c:d, take the step from a to b and add it to c. Search near that result for d. This is an approximate pattern, not a rule that every relation follows.

\[ v_d \approx v_b-v_a+v_c \]

For a positive count, take its base-10 logarithm and add 1; an absent term gets zero. This uses the TF convention in Lecture 6, PDF page 54.

\[ \mathrm{tf}_{t,d}=\begin{cases}1+\log_{10} c(t,d)&c(t,d)>0\\0&c(t,d)=0\end{cases} \]

N counts all documents; df counts documents containing the term. A term in every document gets IDF zero. Multiply IDF by TF to get the final weight.

\[ \mathrm{idf}_t=\log_{10}\frac{N}{\mathrm{df}_t},\qquad \mathrm{tfidf}_{t,d}=\mathrm{tf}_{t,d}\,\mathrm{idf}_t \]

Compare the observed joint probability with the product of the two marginals. Base-2 PMI measures the result in bits. PPMI keeps positive values and replaces negative ones with zero.

\[ \operatorname{PMI}(w,c)=\log_2\frac{P(w,c)}{P(w)P(c)},\qquad \operatorname{PPMI}(w,c)=\max(0,\operatorname{PMI}(w,c)) \]

T is the total number of target-context pairs, not necessarily the number of words in the corpus. Use row and column totals from the same table. With positive marginals, a zero cell gives PPMI zero.

\[ \operatorname{PMI}(w,c)=\log_2\frac{f_{wc}\,T}{(\sum_j f_{wj})(\sum_i f_{ic})},\qquad T=\sum_{i,j}f_{ij} \]

The dot product measures how similar the target vector w and the context vector c are; the sigmoid turns that score into a probability. Both vectors are learned. The “not a neighbor” probability is the rest.

\[ P(+\mid w,c)=\sigma(\mathbf{c}\cdot\mathbf{w})=\frac{1}{1+\exp(-\mathbf{c}\cdot\mathbf{w})},\qquad P(-\mid w,c)=\sigma(-\mathbf{c}\cdot\mathbf{w}) \]

The loss is small when the real neighbor gets a high dot product and every sampled word gets a low one. It is the cross-entropy of k+1 yes-or-no decisions, treated as independent.

\[ L_{CE}=-\Big[\log\sigma(\mathbf{c}_{pos}\cdot\mathbf{w})+\sum_{i=1}^{k}\log\sigma(-\mathbf{c}_{neg_i}\cdot\mathbf{w})\Big] \]

Each gradient is (predicted probability − label) times the other vector in the pair. The positive context vector moves toward w, each negative context vector moves away, and w moves in response to all of them. η is the learning rate.

\[ \mathbf{c}_{pos}\leftarrow\mathbf{c}_{pos}-\eta\big[\sigma(\mathbf{c}_{pos}\cdot\mathbf{w})-1\big]\mathbf{w},\qquad \mathbf{c}_{neg}\leftarrow\mathbf{c}_{neg}-\eta\big[\sigma(\mathbf{c}_{neg}\cdot\mathbf{w})\big]\mathbf{w} \] \[ \mathbf{w}\leftarrow\mathbf{w}-\eta\Big[\big[\sigma(\mathbf{c}_{pos}\cdot\mathbf{w})-1\big]\mathbf{c}_{pos}+\sum_{i=1}^{k}\big[\sigma(\mathbf{c}_{neg_i}\cdot\mathbf{w})\big]\mathbf{c}_{neg_i}\Big] \]

M holds the co-occurrence counts. One matrix has a vector per target word and the other a vector per context word, chosen so their dot products approximate log M with only d dimensions.

\[ V^{\top}U\approx\log M \]

Apply the same rule, such as the average or the maximum, to each coordinate across all the word vectors in the text. Word order is lost.

\[ \mathbf{v}=\operatorname{pool}(\mathbf{v}_1,\dots,\mathbf{v}_n) \]

Each maps the weighted sum z to an output. Sigmoid gives (0, 1), tanh gives (−1, 1), and ReLU keeps positive values and zeroes negative ones. ReLU is the most common.

\[ \sigma(z)=\frac{1}{1+e^{-z}},\qquad \tanh(z)=\frac{e^{z}-e^{-z}}{e^{z}+e^{-z}},\qquad \operatorname{ReLU}(z)=\max(0,z) \]

The hidden layer applies an activation function g to weighted sums of the inputs. The output layer is multinomial logistic regression on the hidden vector. The equations follow Jurafsky & Martin, Chapter 7; the slide shows the diagram.

\[ \mathbf{h}=g(W\mathbf{x}+\mathbf{b}),\qquad \hat{\mathbf{y}}=\operatorname{softmax}(U\mathbf{h}) \]

The Lecture 6 TF example uses a different convention from its preceding formula, and the PMI example contains a denominator typo. The linked notes explain both and work through consistent calculations.

Terms in plain language

Listed by English name so you can look up terms from the slides.

<s> / </s>

Special markers for the start and end of a sentence. The first word gets a preceding context, and the model has a way to signal that the sentence is finished.

n-gram Models · Lec 2, Lec 3
Activation function

The non-linear function applied to a unit’s weighted sum. Sigmoid, tanh, and ReLU are the standard choices. Without one, stacked layers collapse into a single linear model.

Add-k / add-one (Laplace) smoothing

Add the same small count to every possible outcome before turning counts into probabilities. This reserves probability for unseen combinations. Add-one is usually too crude for language models, though this family of methods is useful in models such as Naïve Bayes.

Smoothing · Lec 3, Lec 4
Allocational vs. representational harm

Two ways biased embeddings cause damage. Allocational harm is an unfair distribution of resources or opportunities, such as hiring searches skewed by a “programmer is to man” association. Representational harm is a demeaning association carried by the representation itself, whether or not any decision is made with it.

Analogy

Try to carry a relation from one pair of words to another. If a is to b as c is to d, look for a word vector near b − a + c. This works for some relations, not every analogy.

Antonymy

Words such as hot and cold have opposite meanings in a particular respect. They may still appear in very similar sentences, so their word vectors can be close.

Word Meaning & Vectors · Lec 5, Lec 6
Bag of words (BoW)

Represent a document by counting how often each vocabulary word appears. The result keeps word counts but loses word order and the surrounding context.

Data & Features · Lec 3, Lec 4
Bias term \(b\) (or \(w_0\))

A number added to the weighted feature total. It can shift the prediction even when every input feature is zero, allowing the decision boundary to move.

Chain rule of probability

Find the probability of a sequence by working from left to right: take the first word’s probability, then multiply by each next word’s probability given everything before it.

Sentence Probabilities · Lec 2, Lec 3
Closed vs. open vocabulary

A closed-vocabulary setup assumes every test word is already in the vocabulary. An open-vocabulary setup allows unfamiliar words and needs a way to handle them, such as an unknown-word marker.

Conditional maximum likelihood

Choose parameters that give the known correct labels high probability for their inputs. Taking the negative logarithm gives the cross-entropy loss used to train a classifier.

Connotation / VAD

A word can carry an emotional tone as well as a literal meaning. VAD describes pleasantness (valence), emotional intensity (arousal), and a sense of control (dominance).

Word Meaning & Vectors · Lec 5, Lec 6
Convex function

For a convex loss, a local minimum, if it exists, is also a global minimum. There may be more than one minimizer, and gradient descent still needs suitable step sizes and conditions. Logistic regression has a convex loss; neural-network losses are generally not convex.

Cosine similarity

Compare the directions of two nonzero vectors, after accounting for their lengths. The score ranges from −1 to 1; a larger score means more similar directions. It is undefined for a zero vector. Lecture 7 adds why it beats the raw dot product: long vectors, which belong to frequent words, would otherwise win.

Cross-entropy loss

For an example with one correct class, take the negative log of the probability assigned to that class. A confident correct prediction has a small loss; a confident wrong prediction has a large one.

Discriminative classifier

Learn which label to predict from an input. The model estimates the probability of a label given the input; logistic regression is one example.

Distributional hypothesis

Words that appear in similar surroundings often have similar meanings. This lets us learn something about a word by studying where it is used.

Word Meaning & Vectors · Lec 5, Lec 6
Extrinsic vs. intrinsic evaluation

An intrinsic evaluation checks a model on its own, using a score such as perplexity. An extrinsic evaluation checks whether it helps with a real task, such as speech recognition. The first is usually easier; the second is closer to the intended use.

Feed-forward neural network / multilayer perceptron

Layers of units where information flows from the inputs, through one or more hidden layers, to the outputs. Each hidden unit is a weighted sum followed by an activation function. The output layer is multinomial logistic regression on the hidden vector.

Feature / weight

A feature is a measurable clue in the input, such as the number of positive words in a review. Its weight tells the model how strongly, and in which direction, that clue should affect a class score.

Data & Features · Lec 3, Lec 4
Garden-path sentence

A sentence whose opening encourages a reading that later words force us to revise. In “The old man the boat,” man is a verb. These examples show why a short local context can be misleading.

n-gram Models · Lec 3
GloVe

Global Vectors: a static embedding model that starts from the whole word-word co-occurrence matrix and fits target and context vectors whose dot products approximate its logarithm, with a fixed small number of dimensions. It builds on PPMI intuitions and matrix factorization.

Gradient

A collection of derivatives, one for each parameter. It tells us how the loss changes as the parameters move; its direction gives the steepest local increase, so learning usually steps the other way.

Held-out data

Set aside examples that are not used to fit the weights. Use them to choose settings such as smoothing strength or regularization, while keeping a separate test set for the final evaluation.

Smoothing · Lec 3
Hyperparameter

A setting chosen outside the weight-learning process, such as the learning rate, batch size, or strength of smoothing and regularization. Use validation data to compare choices.

Interpolation

Mix predictions from models using different amounts of context, such as a unigram, bigram, and trigram model. Nonnegative weights that sum to 1 keep the result a probability distribution; the weights may depend on the context.

Smoothing · Lec 3
L1 / L2 regularization

Add a cost for large weights. L2 charges for squared weights and tends to shrink them; L1 charges for absolute values and can make some weights exactly zero, effectively selecting features.

Language model

A model that assigns probabilities to sequences of words, or predicts the probabilities of the next word from the words already seen.

Sentence Probabilities · Lec 1, Lec 2, Lec 3
Learning rate \(\eta\)

The number that scales each parameter update. A larger rate takes bigger steps; a rate that is too large can overshoot rather than improve the model.

Lemma / sense / polysemy

A lemma is the dictionary form shared by variants such as break and broke. A sense is one meaning of a word. Polysemy means a word has several, usually related, meanings.

Word Meaning & Vectors · Lec 5, Lec 6
Log space

Store log probabilities and add them instead of multiplying many tiny probabilities. This makes the arithmetic more stable and helps avoid numbers becoming too small for the computer to represent.

n-gram Models · Lec 2, Lec 3
Markov assumption

Approximate the next word’s probability using only the last few words, rather than the whole sentence so far. This makes the amount of context manageable.

Sentence Probabilities · Lec 2, Lec 3
Maximum likelihood estimate (MLE)

Choose the model parameters that make the observed data most likely. For a simple count-based word model, this means dividing the relevant count by the total count of possible outcomes.

Sentence Probabilities · Lec 2, Lec 4
Mini-batch

Use a small group of examples for one update. Compute their gradients, usually average them, and adjust the parameters once for the whole group.

Multinomial logistic regression

Extend logistic regression to several mutually exclusive labels. Calculate one score for each class, then use softmax to turn the scores into probabilities that sum to 1.

Multiclass Predictions · Lec 5, Lec 6
Negative sampling

For each real (target, context) pair, draw k words at random to serve as “not a neighbor” examples. Words are sampled by their unigram frequency raised to a power, commonly 0.75, which gives rare words a slightly better chance.

word2vec · Lec 7
n-gram

A sequence of n neighboring tokens. An n-gram language model predicts a token from the previous n−1 tokens: unigram uses none, bigram uses one, and trigram uses two.

n-gram Models · Lec 2, Lec 3
One-hot vector

A vector with one 1 and zeros everywhere else. It can mark a class or a vocabulary word. Different word identities have no shared nonzero coordinate, so the representation alone does not express meaning similarity.

Word Vectors from Counts · Lec 5, Lec 6
OOV / <UNK>

An out-of-vocabulary word is missing from the model’s vocabulary. One approach maps rare training words and unfamiliar test words to the same <UNK> marker, so the model can learn a probability for unknown words.

Overfitting

A model fits the training examples so closely that it performs worse on new examples. It may rely on accidental patterns or noise instead of learning relationships that carry over.

Parametric model

A model whose behavior is controlled by a specified set of parameters, such as weights and a bias. The structure is fixed while training changes the parameter values.

Pooling

Combine several word vectors into one vector by applying the same rule, such as the average or maximum, to each coordinate. It gives a classifier one input per text but discards word order.

Perplexity

A summary of how much probability a model gives to the test text, adjusted for its length. Lower is better under the same evaluation setup. For equally likely choices, it equals the number of choices.

Measuring Predictions · Lec 2, Lec 3
PMI / PPMI

PMI compares how often two items occur together with how often they would co-occur by independence. It takes the log of that ratio. PPMI replaces negative results with zero; estimates from rare counts can be unreliable.

ReLU / tanh

Two activation functions. ReLU, max(0, z), passes positive values and zeroes negative ones and is the most common choice. Tanh is an S-shaped curve centered at zero with outputs between −1 and 1.

Self-supervision

Training on labels that the data supplies for itself, with no human annotation. In word2vec, a word that really appears near the target is the “yes” answer.

word2vec · Lec 7
Sigmoid

A smooth S-shaped function that turns any finite real-valued score into a number between 0 and 1. Binary logistic regression uses it to produce a probability.

Skip-gram with negative sampling (SGNS)

The word2vec variant developed in lecture. A logistic regression classifier decides whether a context word belongs near a target word, using the sigmoid of their dot product. Real neighbors are positive examples and randomly sampled words are negative ones.

word2vec · Lec 7
Softmax

Turn several class scores into one probability distribution: exponentiate each score, then divide by the sum of those exponentials. For finite scores, the probabilities are positive and sum to 1.

Multiclass Predictions · Lec 5, Lec 6
Sparse vs. dense vectors

A sparse vector has many zero entries, often because most possible contexts were never observed. A dense word vector usually has fewer coordinates, most of them nonzero, and packs information into that smaller space. Lecture 7 gives typical lengths: 20,000–50,000 for sparse, 50–1000 for dense.

Static vs. contextual embeddings

A static embedding gives a vocabulary word the same vector wherever it appears. A contextual embedding changes with the sentence, so different occurrences of the same word can have different vectors.

Stochastic gradient descent (SGD)

Update the parameters after one example, or after a small batch, rather than waiting to process the whole training set. Training usually visits examples in a shuffled order.

Stochastic vs. deterministic system

With the same input and settings, a deterministic system gives the same output. A stochastic system uses randomness, so the output can vary. Sampling from a language model is one example; fixed decoding can instead be deterministic.

What Is NLP? · Lec 1
Stop words

Very common words such as the, is, and of. Some preprocessing pipelines remove them when they add little useful signal, but whether to remove them depends on the task.

Synonymy vs. similarity / relatedness

Synonyms can share a meaning in some contexts. Similar words share important properties, as coffee and tea do. Related words belong together in a situation, as coffee and cup do, without meaning the same kind of thing.

Word Meaning & Vectors · Lec 5, Lec 6
Target vs. context embeddings

Word2vec keeps two vectors per word, one for its role as the target and one for its role as a context word, in matrices W and C. A common final representation adds the two.

word2vec · Lec 7
Term-document matrix

A table with one row per word and one column per document. Each entry records that word’s count in that document. A row describes a word across documents; a column describes one document.

TF-IDF

Give a term more weight if it appears often in one document but in relatively few documents overall. TF measures within-document frequency; IDF discounts terms that appear almost everywhere.

Word Vectors from Counts · Lec 3, Lec 6
Token

One unit produced when text is split for a model. In these word-level examples it is usually a word occurrence, though other tokenizers may use punctuation, word pieces, or characters.

Tokenization / stemming / lemmatization

Tokenization splits text into units. Stemming applies rules that shorten word forms, such as the Porter stemmer. Lemmatization finds a dictionary form. These steps organize the input but do not all do the same job.

Type vs. token

A type is a distinct vocabulary item; a token is one occurrence of it. After lowercasing and ignoring punctuation, “the cat chased the cat” has three types and five tokens.

Word Meaning & Vectors · Lec 5, Lec 6
Underflow

A number becomes too small for the computer’s number format to represent properly. Multiplying many small probabilities can cause this; adding log probabilities helps avoid it.

n-gram Models · Lec 2
Weight decay

Shrink weights toward zero during learning. With ordinary gradient descent, adding an L2 penalty produces this effect alongside the update from the prediction loss.

Word embeddings

Represent a word with a list of numbers. When the vectors are learned from how words are used, similar words can share useful information instead of being treated as unrelated names.

Word Meaning & Vectors · Lec 5, Lec 6
Window size

How many words on each side of the target count as context. Small windows (about ±2) find words that behave alike; large windows (about ±5) find words about the same topic. It shapes both sparse and dense vectors.

Word-context matrix

A table with target words as rows and nearby context words as columns. An entry counts how often that context word appears within the chosen window around the target.

Word2vec / GloVe / SVD

Three routes to dense vectors. Word2vec trains a classifier on word pairs and keeps its weights; it includes skip-gram and CBOW. GloVe factorizes the log co-occurrence matrix. SVD reduces a matrix directly, and LSA applies it to text. Lecture 6 named all three; Lecture 7 developed the first two.

WordNet / synset

WordNet organizes word senses into groups that express the same concept, called synsets. It also records relations such as one thing being a kind of another, or a part of another.

Word Meaning & Vectors · Lec 5, Lec 6
Zeros

An unsmoothed count model may give zero probability to a word combination missing from training. If that combination appears in the test text, its log probability is not finite and perplexity becomes infinite.

Zipf's law

A few words occur very often, while a great many words are rare. Roughly, frequency falls as a power of a word’s rank in the frequency list.

Smoothing · Lec 3