Part II · Learning to Classify Text

Turning Text into Inputs a Model Can Use

Prepare labeled examples, turn text into numerical features, and see how a classifier learns from them.

How a model learns from examples Lec 3 · Sep 1 · recapped Lec 4 · Sep 3

Suppose we want a model to tell whether a review is positive or negative. We show it reviews with known answers, let it make predictions, and adjust it when those predictions are poor. This is supervised learning. The same process appears throughout the course, from logistic regression to neural language models, and has five parts:

I · Data
Examples paired with answers
We have \((x^{(i)}, y^{(i)})\), \(i \in \{1 \ldots N\}\). Each input is a vector \([x_1, \ldots, x_d]\), which may use word embeddings.
II · Model
A way to make a prediction
The model computes \(p(y \mid x)\): the probability of an answer given the input. Logistic regression and Naïve Bayes are examples.
III · Loss
A measure of prediction error
A loss function, such as cross-entropy \(L_{CE}\), tells us how much to penalize a prediction.
IV · Optimization
A way to improve the model
An algorithm such as stochastic gradient descent changes the model to reduce the loss.
V · Inference
Use the model on new inputs
After training, we make predictions on test examples and evaluate them.

The first four parts make up training. This page explains how to prepare the data. The following pages cover the model and its loss, how to update the model, and how to avoid fitting the training examples too closely.

What text classification asks us to do Lec 3 · Sep 1

A classifier chooses a category for an input. In language tasks, it might decide whether an email is spam, identify a document's topic or language, or judge the sentiment of a review. Classification is also used in many areas outside NLP.

  • Input: a document \(x\), represented by a list of numerical features. We also choose a fixed set of possible classes, \(C = \{c_1, c_2, \ldots, c_J\}\).
  • Output: one predicted class, written \(\hat y \in C\). The hat marks it as a prediction.
  • Two classes: for binary classification, the training pairs \((x^{(i)}, y^{(i)})\) have labels \(y^{(i)} \in \{0, 1\}\). On a new input \(x_{\text{test}}\), we predict \(\hat y_{\text{test}} \in \{0,1\}\).
  • More than two classes: multiclass, or multinomial, classification chooses from a larger set. For example, labels \(y \in \{0,1,2,3,4\}\) can encode the five possible star ratings.

Why we start with logistic regression

Logistic regression is a useful first model because its steps are easy to inspect and many of the same ideas carry over to neural networks.

  • It is widely used to analyze data in the natural and social sciences.
  • It provides a baseline: a straightforward classifier against which we can compare more complicated models.
  • It can be viewed as a one-layer neural network, making it a starting point for the networks we study later.
  • It learns which label fits an input by modeling \(p(y\mid x)\). This makes it a discriminative classifier; it does not try to explain how the input itself was generated.
  • Other classifiers mentioned in the lecture include Naïve Bayes, k-nearest neighbors, decision trees, and support vector machines (SVMs).

For one example, the model receives features \(x = [x_1, \ldots, x_n]\) and uses a weight for each one, \(w = [w_1, \ldots, w_n]\), to help choose a class. The parameters are sometimes denoted by \(\theta\). Once the model's form is chosen, it has a fixed set of parameters to learn, so we call it a parametric model.

Features describe the text; weights say how much they matter Lec 3 · Sep 1 · recapped Lec 4 · Sep 3

A feature is something we can measure about an input. For a review, a feature might record whether a particular word appears. Its weight tells the model how that evidence should affect the positive-class prediction. Here is the lecture's example:

FeatureWeight
\(x_i\): the review contains awesome+10
\(x_j\): the review contains abysmal−10
\(x_k\): the review contains mediocre−2
\(x_l\): the review contains restaurant≈ 0

In this example, “awesome” provides strong positive evidence, “abysmal” strong negative evidence, and “mediocre” weaker negative evidence. “Restaurant” contributes almost nothing to the sentiment decision. We can design features by hand, a process called feature engineering, or let a model learn useful features automatically.

Preparing the text Lec 3 · Sep 1

Before extracting features, we decide how the text will be cleaned and split into units. These choices still matter today and help us understand which steps modern language models handle automatically.

  • Split the text into tokens. Tokenization identifies the units the model will process. Separating punctuation, for example, changes “Happy New Year!” into “Happy New Year !”. Related cleaning steps may remove extra spaces, external URLs, or non-alphabetical characters when they are unhelpful for the task; these removals are choices, not requirements of tokenization.
  • Optionally standardize the words. We might expand “won't” to “will not”, lowercase “New York” to “new york”, or reduce related word forms. Stemming removes parts of words, as in “tokenization” → “token”; lemmatization aims to recover a dictionary form.
  • Libraries provide many of these operations. The Porter stemmer is one example.

Turning text into numbers Lec 3 · Sep 1

Choose a vocabulary

A vocabulary fixes which words our representation can record and gives each word a position in the vector.

  • Build a dictionary of the words relevant to the task. We may remove very common stop words when they add little useful information; whether a word is useful depends on the task.
  • Assign each word an identifier, for example are → 2.
  • Choose what to do with words outside the vocabulary. We can drop them or place them in an unknown-word slot. The same choice arises for <UNK> in n-gram language models, especially when a new word appears at test time.

Choose what the numbers mean

  • Bag of words records how often each vocabulary word appears. We work through it below.
  • TF-IDF, short for term frequency–inverse document frequency, adjusts those counts so that words common across many documents receive less weight.
  • Word embeddings are learned, dense vectors. Methods such as word2vec and GloVe come later in the course.

Bag of words: count each word Lec 3 · Sep 1 · recapped Lec 4 · Sep 3

Choose a vocabulary of \(k\) words to suit the application. A bag-of-words (BoW) vector has one position per word: \(x = [x_1, \ldots, x_k]\), with \(x_i \in \{0, 1, 2, \ldots\}\). An entry \(x_i = j\) simply means that word \(i\) appeared \(j\) times.

Count the words in a shirt review

Use this vocabulary, in this order: [good, bad, nice, ugly, love, hate, complements, coarse, itchy].

“I love this shirt because it is nice and warm. The fabric is also nice and the color complements my skin tone.”

“Nice” appears twice; “love” and “complements” appear once each. The other vocabulary words do not appear, so the vector is \(x = [0, 0, 2, 0, 1, 0, 1, 0, 0]\).

Why use it?What does it miss?
It is simple to build and can work reasonably well in many settings.Counts lose word order and local context. The representation does not record how “new” combines with “york” or “book”. Most entries are zero, so the vectors are sparse, and common words can dominate the counts. These limits motivate TF-IDF and, when appropriate, stop-word removal.

The idea also extends beyond text. A bag of visual words represents an image by grouping small local image patches into discrete types and counting those types.

What to remember
  • For each model, ask what data it uses, how it predicts, how errors are measured, how it learns, and how it is used on new inputs.
  • Text classification turns a document's features into a class prediction. Logistic regression is a useful first model and comparison point.
  • Prepare the text, choose a vocabulary, and build a numerical representation. Bag of words keeps counts but loses word order.