Why compare directions rather than raw dot products? Lec 7 · PDF pp. 46–50
Lecture 7 returns to the similarity measure from Lecture 6 and adds a reason for preferring cosine. The dot product of two vectors adds up the products of matching coordinates:
This number is high when both vectors have large values in the same coordinates. It is also high simply because a vector is long, meaning it has large values in many coordinates. Frequent words such as the, you, and of occur alongside many other words, so their count vectors are long. A raw dot product would therefore favor frequent words over words that are genuinely similar.
Cosine removes that effect. The dot product equals the product of the two lengths times the cosine of the angle between the vectors, so dividing by the lengths leaves only the angle:
The result is \(+1\) when the vectors point the same way, \(0\) when they are at right angles, and \(-1\) when they point in opposite directions. We do not care about the length of the embeddings, only the angle between them.
The slide uses three words and three context columns from a count table:
| pie | data | computer | |
|---|---|---|---|
| cherry | 442 | 8 | 2 |
| digital | 5 | 1683 | 1670 |
| information | 5 | 3982 | 3325 |
Digital and information point almost the same way; cherry points somewhere else. The slide's two-dimensional plot of the pie and computer coordinates shows the same picture: cherry rises along the pie axis while the other two lie nearly flat along the computer axis.
A note on the number: the first ratio evaluates to \(0.01775\), which rounds to \(0.018\). PDF p. 49 prints \(.017\), which is the value cut off rather than rounded.
GloVe: fit the vectors to global co-occurrence statistics Lec 7 · PDF pp. 51–55
Word2vec is hard to interpret. It slides one window at a time across the text, and it is not obvious what quantity the finished vectors approximate. The lecture asks whether there is a more theoretical way to describe what is being learned.
GloVe, short for Global Vectors (Pennington et al., 2014), is another widely used static embedding model. Instead of a narrow window, it starts from the global word-word co-occurrence matrix for the whole corpus, the same object used to compute PPMI. The slide summarizes its idea as using word vectors to approximate the pointwise mutual information of two words:
The slide lists what GloVe builds on: ratios of probabilities taken from the co-occurrence matrix, the intuitions behind count-based models like PPMI, and matrix factorization. The goal is to store most of the important information in a fixed, small number of dimensions, creating a low-dimensional matrix while keeping the reconstruction loss small. That loss is the error made when going back from the low-dimensional vectors to the full matrix. Training is fast and scales to very large corpora.
Here \(M\) collects the co-occurrence counts for all target and context words. One matrix holds a vector per target word and the other a vector per context word, so that the dot product of a target vector and a context vector approximates the log of their co-occurrence entry. The slide describes the solution as a factorization of \(M\) with rank \(d\), where \(d\) is the embedding length.
The slide gives only the approximation above. Pennington et al. fit it with a weighted least-squares objective. Each word \(i\) and context \(j\) also receive bias terms \(b_i\) and \(\tilde b_j\), and a weighting function \(f\) reduces the influence of very rare and very frequent pairs:
\[ J=\sum_{i,j}f(X_{ij})\big(\mathbf{w}_i\cdot\tilde{\mathbf{w}}_j+b_i+\tilde b_j-\log X_{ij}\big)^2 \]Pairs with \(X_{ij}=0\) receive weight zero, so the log of zero never has to be computed. This detail is not in the lecture.
Which method is better? Lec 7 · PDF p. 56
Neither. The slide states that no single embedding method is the best for every application. What matters more is the amount and quality of training data and the choice of hyperparameters: the vector dimension \(d\); for word2vec, the negative-sampling weight \(\alpha\) and the number of negatives \(K\); for GloVe, the matrix factorization algorithm; and others.
Use word vectors as features for a classifier Lec 7 · PDF p. 57
A classifier such as multinomial logistic regression needs one feature vector per document. A sentence gives us one vector per word. The slide's example is I am a student at USC, six words and six vectors \(\mathbf{v}_1,\dots,\mathbf{v}_n\).
The simplest way to combine them is element-wise pooling: apply the same rule to each coordinate across all the word vectors.
Averaging or taking the maximum throws away the order of the words and the way they modify one another. The slide flags this directly: no word-order or context information survives. The result is still a useful, compact input, and it improves on a bag of words because similar words now produce similar features.
What the vectors pick up depends on the window Lec 7 · PDF p. 59
The context window size changes which words end up close together. This holds for sparse and dense vectors alike.
| Window | Nearest neighbors tend to be | Example for Hogwarts |
|---|---|---|
| Small, about \(\pm2\) | Words that behave alike grammatically and belong to the same category | Other fictional schools: Sunnydale, Evernight, Blandings |
| Large, about \(\pm5\) | Words related by topic | Words from the same stories: Dumbledore, half-blood, Malfoy |
So the window size is a modeling choice, not just a speed setting. A small window captures how a word is used; a large window captures what it is about.
How do we know whether the vectors are any good? Lec 7 · PDF p. 60
One check is the analogy test from Lecture 6: in a GloVe space, \(\mathbf{w}_{king}-\mathbf{w}_{man}+\mathbf{w}_{woman}\) lands near \(\mathbf{w}_{queen}\), and a two-dimensional projection shows parallel offsets for related pairs. The lecture repeats the caveats: this works only for frequent words, small distances, and certain relations such as countries and capitals or parts of speech, and not for others. Understanding analogy remains an open research area.
The better test is extrinsic evaluation: plug the vectors into a real task and measure the task's performance. This is the same distinction drawn for language models on the perplexity page.
Embeddings reflect cultural bias Lec 7 · PDF pp. 61–62
Because the vectors are learned from text people wrote, they absorb the associations in that text, including harmful ones. The slide gives three analogy completions:
- “Paris is to France as Tokyo is to ___” → Japan. This one is fine.
- “father is to doctor as mother is to ___” → nurse.
- “man is to computer programmer as woman is to ___” → homemaker.
The lecture calls these allocational harms. An algorithm that uses such embeddings while, for example, searching for programmers to hire could distribute opportunities unfairly. The example is from Bolukbasi et al. (2016), whose title asks exactly this question and proposes ways to reduce the bias.
The same property makes embeddings a tool for studying bias. Garg et al. (2018) trained embeddings on text from different decades and measured, for each adjective, how much closer it sat to words for women than to words for men, or to names associated with particular ethnic groups. Two of their findings:
- Competence adjectives such as smart, wise, brilliant, and logical were biased toward men, with the bias slowly decreasing from 1960 to 1990.
- Dehumanizing adjectives such as barbaric, monstrous, and bizarre were biased toward Asians in the 1930s, with the bias decreasing over the twentieth century.
These measurements match attitude surveys from the 1930s. The lecture labels this kind of effect a representational harm: the representation itself carries a demeaning association, whether or not any decision is made with it.
- Raw dot products favor frequent words because their vectors are long. Cosine divides the length out and compares only direction.
- GloVe fits target and context vectors so their dot products approximate the log co-occurrence matrix. It is a global, factorization-based cousin of word2vec.
- No embedding method wins everywhere. Data and hyperparameters, including \(d\), \(\alpha\), and \(K\), matter more.
- Pooling word vectors gives a classifier one input vector per text, at the cost of word order.
- Small windows find words used alike; large windows find words about the same topic. Judge vectors by the tasks they help with.
- Embeddings inherit the biases of their training text, which can cause harm and can also be measured.