These notes cover the first seven lectures of CSCI 544. You can read them in the order below or jump to something you want to review. Related ideas stay together on one page, even when the class returns to them in a later lecture.
The labels at the top of each page show which lectures cover that topic. To find notes for a particular class, open the lecture timeline. If you just need a definition or equation, use the glossary and formula sheet. The language button takes you to the same page in Chinese.
What have we learned so far?
The course starts with a simple question: can we predict what someone will say next? It then asks how a model can learn to classify text, how we can represent the meaning of a word with numbers, and how to build a classifier that is more than a single weighted sum.
One idea connects the first two parts: choose model settings that make the observed examples more likely. This is maximum likelihood estimation. In a simple word-counting model, it leads to counting and dividing. In a classifier, it leads to a loss that penalizes low probability for the correct answer.
The same idea returns a third time in word2vec, which trains a logistic regression classifier on word pairs and keeps its weights as word vectors. Both kinds of model also face the same problem: the examples used for training do not cover everything. Smoothing leaves probability for word combinations we have not seen; regularization helps a classifier avoid fitting its training examples too closely.
Choose a place to begin
Start Here
Find out what the course is about and how it is organized.
- 0.1What Can We Do with Language Models?How computers work with language, what language models are useful for, and where they still make mistakes.Lec 1–2
- 0.2Assignments, Grades & Course DetailsFind the grading breakdown, project requirements, quiz rules, Homework 1, deadlines recorded in the slides, and ways to get help.Lec 1–7
Part I · Predicting Words
Build a language model from word counts, test its predictions, and handle combinations missing from the training text.
- 1.1How Language Models Assign ProbabilitiesStart with the chance of the next word, then see how those chances give us the probability of a whole sentence.Lec 2–3
- 1.2Predicting the Next Word with n-gramsCount words in a text collection and use the last few words to predict what comes next.Lec 2–3
- 1.3How Good Are the Predictions? PerplexityUse perplexity to measure how well a model predicts new text, and learn when that score is useful.Lec 2–3
- 1.4Generating Text and Handling Unseen WordsFollow a model as it writes one word at a time, then see why unfamiliar words and word pairs cause trouble.Lec 2–3
- 1.5Leaving Room for Unseen Word PairsSmoothing gives some probability to word combinations missing from training. Compare simple ways to do it.Lec 3–4
Part II · Learning to Classify Text
Follow the whole learning process: prepare data, make a prediction, measure the error, and adjust the model.
- 2.1Turning Text into Inputs a Model Can UsePrepare labeled examples, turn text into numerical features, and see how a classifier learns from them.Lec 3–4
- 2.2Making a Two-Class PredictionLogistic regression adds up evidence, turns the result into a probability, and measures how far its prediction missed the answer.Lec 3–5
- 2.3How a Model Learns from Its MistakesGradient descent tells us which way to adjust the weights. The learning rate controls the size of each step.Lec 4–5
- 2.4Helping a Model Work on New ExamplesA model can fit its training examples too closely. Regularization discourages overly large weights so it can work better on new data.Lec 4–5
- 2.5Choosing among Several Classes with SoftmaxGive each possible class a score, turn the scores into probabilities, and use the correct answer to train the model.Lec 5–6
Part III · Representing Word Meaning
Use the contexts around words to describe their meanings, then compare the resulting vectors.
- 3.1Representing Word Meaning with VectorsWords used in similar settings often have related meanings. Word vectors turn that idea into something we can calculate.Lec 5–7
- 3.2Building Word Vectors from CountsCount where words appear, adjust for common words with TF-IDF or PPMI, and see why shorter vectors can be useful.Lec 6–7
- 3.3Learning Word Vectors by Prediction: word2vecTrain a classifier to guess which words appear near each other, then keep its weights as short, dense word vectors.Lec 7
- 3.4GloVe, Similarity, and What Word Vectors Pick UpCompare vectors by direction, meet GloVe, feed vectors to a classifier, and see how window size and training text shape what the vectors capture.Lec 7
Part IV · Neural Networks
Turn a single classifier into layers of units that can learn boundaries a straight line cannot draw.
What comes next?
The first lecture outlines the rest of the course in roughly historical order:
- Early neural language models, around 2013–2018: learn compact word vectors, then study feed-forward networks, recurrent networks, backpropagation, and encoder–decoder models. Lecture 7 reached the first two of these; backpropagation is scheduled for Sep 17. These show how neural networks can learn from and generate sequences.
- Models from 2018 onward: study attention and Transformers, followed by models that fill in missing words and models that predict the next word.
- Large language models: learn how pretraining, fine-tuning, prompts, instruction tuning, and human feedback shape model behavior. We will also look at generation, incorrect but confident answers, privacy, and other open problems.
The planned homework topics are n-gram models, word vectors with recurrent language models, attention in Transformers, and advanced language-model topics. The course plans to assign two of these, with 3–4 weeks for each; the choices may change.