Lecture 7 restates the distributional idea on PDF pp. 9–10 and returns to cosine similarity on pp. 46–50, adding a worked example with real counts and a reason to prefer cosine over the raw dot product. That material is on the GloVe and embedding properties page.
How can a computer tell that two words mean similar things? Lec 5–6
In an n-gram model, rancor and hatred are simply two different symbols. Their names do not tell the model that their meanings are close. One way to express that connection is to give each word a list of numbers, treat the list as a point in a space, and let positions in that space reflect relationships between words. This idea is called vector semantics.
We are choosing how to represent the input. A word vector could, for example, become the input features for multinomial logistic regression. The classifier would still calculate scores and use softmax to turn them into class probabilities. Our question comes first: what should the numbers in a word vector tell us?
What do we mean by “a word”? Lec 5 · pp. 33–35
Counting appearances, counting different spellings, and identifying meanings are different tasks. These terms help keep them apart.
| Name | What it describes | Example |
|---|---|---|
| Token | A word at one particular position in a sequence. | “The cat chased the cat” contains 5 word tokens. |
| Type | A distinct word form, under the rules we choose for processing the text. | After lowercasing and ignoring punctuation, there are 3 types: the, cat, chased. |
| Lemma | The dictionary form that groups a word's grammatical variations. | Break, breaks, broke, broken, breaking all have the lemma break. |
| Sense | One particular meaning of a word. | Bank can mean a financial institution or the side of a river. |
| Polysemy | A word having several meanings, often connected to each other. | Head can mean a body part or the person leading an organization. |
A stem is another way to shorten a word: a procedure removes parts of its form, and the result need not be a dictionary word. A lemma is the accepted base form. Finding a stem or a lemma does not, by itself, tell us which meaning a word has in a particular sentence.
In “The cat chased the cat.”, the two appearances of cat are separate tokens of the same type. The counts can change if we distinguish uppercase from lowercase or count punctuation separately, so state those rules first. A static type embedding stores one vector for the type and looks up that same vector for both appearances.
What kinds of relationship can words have? Lec 5 · pp. 36–40
Two words can be connected in several ways. We need these distinctions when deciding what a similarity score really tells us.
| Relationship | Examples | How to understand it |
|---|---|---|
| Synonymy | couch / sofa, automobile / car | The words mean the same or nearly the same thing in some settings. Their formality, politeness, and usage can still differ. |
| Similarity | car / bicycle, cow / horse | The things share meaningful properties, but the words are not interchangeable names for the same thing. |
| Antonymy | hot / cold, in / out, rise / fall | The meanings oppose each other in one respect while sharing other properties. |
| Relatedness | coffee / cup | The words belong to a shared situation or topic. A cup goes with coffee, but it is not another drink. |
| Connotation | Happy suggests positive feelings; sad suggests negative feelings. | The feelings and judgments a word brings to mind are part of how people understand it. |
Very few synonyms work as replacements everywhere. “My big sister” normally means my older sister, while “my large sister” describes size. Changing one word changes the message. Likewise, coffee / tea are similar drinks, while coffee / cup are mainly related through use.
Opposites can lie at different ends of a scale, as with long / short; express a binary choice, as with in / out; or name opposing processes, as with rise / fall. They often appear in similar sentences, so vectors built from their contexts can be close. Vector similarity alone does not tell us which relationship two words have.
How WordNet records relationships
WordNet is an English lexical database. It groups word senses that express the same concept into a synset. Different senses of one word can belong to different groups. It also records explicit relationships:
- Hypernym and hyponym, or “is-a”: armchair → chair → furniture moves from a narrower category to a broader one. Chair is a hypernym of armchair, and armchair is a hyponym of chair.
- Meronymy, or “part-of”: a leg is part of a chair. A chair leg is not a kind of chair.
- Antonymy: two particular word senses are opposites.
These labels explain the connection. A vector similarity score gives a flexible numerical comparison, but it does not automatically say whether one thing is a kind of another or a part of it.
Learn about a word by looking at where it appears Lec 5 · pp. 43–49; Lec 6 · p. 33
Words used in similar contexts tend to have similar meanings. This is the distributional hypothesis. A context might be the neighboring words, a whole document, or a grammatical pattern. We look for clues in how people use the word.
The lecture opens with examples of predicting a word's neighborhood. Doth fits a Shakespeare passage better than a modern newspaper. Garnish fits a recipe that lists ingredients and serving instructions. The surrounding language gives clues about topic and style as well as meaning.
Suppose ongchoi appears near leaves, garlic, sautéed, rice, and sauce. You have seen spinach, chard, and collard greens in similar sentences. Those shared settings suggest a leafy vegetable. The slides reveal that ongchoi is water spinach (空心菜).
You already had evidence about the meaning before looking up a dictionary definition.
If tesgüino comes in a bottle, is made from corn, and can make someone drunk, an alcoholic drink is a sensible guess. Each clue rules out some possibilities. Together, the clues tell us more than a single neighboring word would.
Car and automobile can fit into similar sentences without often appearing next to each other. We compare the contexts each word tends to have. Their direct co-occurrence is only one possible piece of evidence.
How can numbers represent meaning?
The lecture starts with an example where each number has a clear meaning. Describe a word's emotional associations along three dimensions: valence, or pleasantness; arousal, or emotional intensity; and dominance, or sense of control. Those three values place the word at a point in three-dimensional space.
Vectors built or learned from text extend this idea to many dimensions. The coordinates usually do not have such simple names. A two-dimensional plot helps us inspect the result, but it does not mean that the original representation, or word meaning itself, has only two dimensions.
How word vectors help a model handle new examples Lec 5 · pp. 50–52; Lec 6 · pp. 26–35
A static embedding stores one fixed vector for each type in the vocabulary. The notation below says that a word from the vocabulary \(V\) is mapped to a list of \(d\) numbers.
Suppose a classifier learns that terrible signals negative sentiment. If it only recognizes word identities, it does not automatically transfer that knowledge to awful. If the two words have similar vectors, the classifier has a chance to recognize their shared numerical pattern and respond similarly.
- An identity feature says, “the previous word was terrible.” A different word activates a different feature.
- A vector feature says, “the previous word had these coordinates.” A similar word can provide a similar set of numbers.
- This can help in practice: word similarity can connect a question to an answer phrased differently, support comparisons between sentences or documents, and help track changes in word use over time.
A word might be absent from the training examples for your particular task but already have a vector learned from a larger collection of text. If the word is missing from the embedding table itself, a fixed lookup cannot provide its vector. Success also depends on the representation and the task: nearby vectors do not guarantee the same label.
Where do these vectors come from? The count-based embeddings page starts by counting which documents a word appears in and which words surround it. It then explains term-document and word-context matrices, TF-IDF, PMI/PPMI, and why we might want shorter, denser vectors.
Cosine similarity: how closely do the directions match? Lec 6 · pp. 36–37
A dot product depends both on how vectors line up and on how long they are. To focus on direction, divide the dot product by both lengths. The result is cosine similarity. For nonzero vectors \(\mathbf{u},\mathbf{v}\in\mathbb{R}^{d}\):
| Cosine value | Directions | What this tells us |
|---|---|---|
| \(+1\) | The same | The directions match completely, even if the lengths differ. |
| \(0\) | At right angles | The vectors have no directional alignment in this representation. |
| \(-1\) | Opposite | The vectors point opposite ways. Their words are not necessarily antonyms. |
Cosine generally ranges from \([-1,1]\). If all coordinates are nonnegative, as with raw counts or PPMI, the range is \([0,1]\). It is a similarity score, not a probability. Whether it tells us something useful about meaning depends on what the vectors contain.
Work through a small example
Take \(\mathbf{u}=[1,2,0]\) and \(\mathbf{v}=[2,1,0]\). First calculate the dot product, then both lengths, and finally divide:
Multiplying \(\mathbf{u}\) by three gives \(3\mathbf{u}=[3,6,0]\). Its direction stays the same, so \(\operatorname{cos}(\mathbf{u},3\mathbf{u})=1\). The vector \([0,0,5]\), on the other hand, is at right angles to \(\mathbf{u}\), so their cosine is \(0\).
If either vector has length zero, the denominator is zero and cosine is undefined. A program must skip that case or explicitly decide how to handle it. The formula itself does not give it a score of zero.
If both vectors are first scaled to length one, giving \(\hat{\mathbf{u}},\hat{\mathbf{v}}\), then \(\|\hat{\mathbf{u}}-\hat{\mathbf{v}}\|_2^2=2-2\operatorname{cos}(\mathbf{u},\mathbf{v})\). After this normalization, ranking by highest cosine gives the same result as ranking by shortest Euclidean distance. That equivalence does not generally hold before normalization.
Trying to solve analogies with vectors Lec 6 · pp. 38–39
An analogy \(a:b::c:d\) asks, “\(a\) is to \(b\) as \(c\) is to what?” Imagine moving from \(a\) to \(b\), then making the same move starting at \(c\). If the relationship is captured by a similar direction and distance, the four points roughly form a parallelogram.
The target may not exactly match a stored word vector, so we search for the nearest candidate. This is a nearest-neighbor search. The target and candidate vectors must be nonzero. A common evaluation rule excludes the three words already given in the question.
We want to move from a fruit to the plant it grows on: \(\mathbf{q}=\mathbf{v}_{\text{tree}}-\mathbf{v}_{\text{apple}}+\mathbf{v}_{\text{grape}}\). If the embedding captures this relationship, vine will be near the target. Swapping the subtraction would send us in the opposite direction.
The prose on Lecture 6 p. 38 reverses that subtraction. These notes use the direction given by the slide's parallelogram diagram and general formula.
The GloVe example in the lecture is \(\mathbf{v}_{\text{king}}-\mathbf{v}_{\text{man}}+\mathbf{v}_{\text{woman}}\approx\mathbf{v}_{\text{queen}}\). It shows an approximate pattern in one vector space, rather than an exact algebraic rule that language must obey.
The method has limits. The slides emphasize frequent words, small distances, and certain relationships, such as countries and capitals or grammatical patterns. A convincing two-dimensional plot does not show that every analogy will work. A model may capture useful word similarities and still fail a particular analogy.
What one fixed vector per word cannot yet tell us
- Which meaning is intended: a type vector combines evidence from many appearances. Financial bank and river bank receive the same lookup vector, so that vector alone cannot identify the sentence's intended sense.
- Which relationship holds: similar meanings, related topics, and opposites can all lead to overlapping contexts. Cosine does not label the relationship.
- What a whole sentence means: representing each word does not by itself account for word order, negation, or how the words combine.
- Words with little or no evidence: rare appearances and missing vocabulary entries limit how reliably a word can be represented.
Lecture 6 ends with another direction: contextual embeddings calculate a representation for each word occurrence using its context, rather than always retrieving the same vector. It also names word2vec, GloVe, and SVD as ways to obtain dense vectors, but does not yet derive their full training procedures. The count-based embeddings page covers the construction methods actually developed in Lecture 6; the word2vec page covers the training procedure that Lecture 7 develops.
- Tokens, types, lemmas, and senses distinguish appearances, word forms, dictionary forms, and meanings.
- A word's usual surroundings give us evidence about what it means.
- Similar numerical representations can help a model apply what it learned to new examples.
- Cosine compares the directions of nonzero vectors. A negative value does not define antonymy.
- For \(a:b::c:d\), look near \(\mathbf{v}_b-\mathbf{v}_a+\mathbf{v}_c\). This analogy pattern is approximate and works better for some relationships than others.
These numbers are positions in the PDF file, rather than older slide labels. Lecture 5 covers neighborhood prediction and the transition on pp. 29–32; word concepts and lexical relations on pp. 33–40; semantics and distribution on pp. 41–49; and vectors and generalization on pp. 50–52. Lecture 6 reviews word meaning and distribution on pp. 7–25, covers representations and applications on pp. 26–35, cosine on pp. 36–37, analogies on pp. 38–39, and static versus contextual embeddings on p. 62. The small cosine calculation and implementation notes on this page explain the lecture's formulas.