USC · CSCI 544 · Applied NLP · Fall 2026

Start Here

A guide to the course so far. Find a topic, follow the lessons in order, or look up a formula.

Start with a topic

These notes cover the first seven lectures of CSCI 544. You can read them in the order below or jump to something you want to review. Related ideas stay together on one page, even when the class returns to them in a later lecture.

The labels at the top of each page show which lectures cover that topic. To find notes for a particular class, open the lecture timeline. If you just need a definition or equation, use the glossary and formula sheet. The language button takes you to the same page in Chinese.

What have we learned so far?

The course starts with a simple question: can we predict what someone will say next? It then asks how a model can learn to classify text, how we can represent the meaning of a word with numbers, and how to build a classifier that is more than a single weighted sum.

Lectures 1–3
Predict the next word
Count word sequences, turn the counts into probabilities, and check how well the model predicts new text.
Lectures 3–6
Learn from examples
Give a model text with known labels, measure its mistakes, and adjust its weights to improve its predictions.
Lectures 5–7
Represent word meaning
Look at the words that appear nearby. Use those patterns, by counting or by prediction, to build vectors and compare meanings.
Lecture 7
Add layers
Treat logistic regression as one unit, add a non-linear function, and stack units into a feed-forward network.

One idea connects the first two parts: choose model settings that make the observed examples more likely. This is maximum likelihood estimation. In a simple word-counting model, it leads to counting and dividing. In a classifier, it leads to a loss that penalizes low probability for the correct answer.

The same idea returns a third time in word2vec, which trains a logistic regression classifier on word pairs and keeps its weights as word vectors. Both kinds of model also face the same problem: the examples used for training do not cover everything. Smoothing leaves probability for word combinations we have not seen; regularization helps a classifier avoid fitting its training examples too closely.

Choose a place to begin

Start Here

Find out what the course is about and how it is organized.

Part I · Predicting Words

Build a language model from word counts, test its predictions, and handle combinations missing from the training text.

Part II · Learning to Classify Text

Follow the whole learning process: prepare data, make a prediction, measure the error, and adjust the model.

Part III · Representing Word Meaning

Use the contexts around words to describe their meanings, then compare the resulting vectors.

Part IV · Neural Networks

Turn a single classifier into layers of units that can learn boundaries a straight line cannot draw.

What comes next?

The first lecture outlines the rest of the course in roughly historical order:

  • Early neural language models, around 2013–2018: learn compact word vectors, then study feed-forward networks, recurrent networks, backpropagation, and encoder–decoder models. Lecture 7 reached the first two of these; backpropagation is scheduled for Sep 17. These show how neural networks can learn from and generate sequences.
  • Models from 2018 onward: study attention and Transformers, followed by models that fill in missing words and models that predict the next word.
  • Large language models: learn how pretraining, fine-tuning, prompts, instruction tuning, and human feedback shape model behavior. We will also look at generation, incorrect but confident answers, privacy, and other open problems.

The planned homework topics are n-gram models, word vectors with recurrent language models, attention in Transformers, and advanced language-model topics. The course plans to assign two of these, with 3–4 weeks for each; the choices may change.