Start Here

What Can We Do with Language Models?

How computers work with language, what language models are useful for, and where they still make mistakes.

What the course title means Lec 1 · Aug 25

Applied Natural Language Processing means building useful systems that work with human language. Each part of the name tells us something about the course:

  • Natural language means the languages people speak and write, rather than programming languages.
  • Language processing means getting a computer to do the work automatically.
  • Applied means using these ideas to solve problems. The course emphasizes practical methods rather than linguistic theory.
The course's starting point

The opening lecture argues that language models now sit at the center of applied NLP. Some systems also work with images, audio, or other kinds of input, but language remains a central part of them. This is why the course treats NLP as especially relevant.

What does NLP try to do? Lec 1 · Aug 25 · recapped Lec 2 · Aug 27

Natural language processing, or NLP, brings together computer science, artificial intelligence, and linguistics. Machine learning is especially important: it lets a computer learn patterns from examples instead of requiring a person to write every rule.

The aim is to help computers work with language well enough to carry out useful tasks, approaching some aspects of human understanding. The hard part is meaning. Even people can disagree about what a sentence means, so deciding how a computer should represent that meaning is a serious challenge.

Think in terms of an input and an output

We can describe an NLP system as a function that turns an input \(X\) into an output \(Y\). At least one side involves language. Often the input is text, but a system can also turn an image into a written description.

Input \(X\)Output \(Y\)What the system does
Some textThe text that comes nextLanguage modeling
Text in one languageText in another languageTranslation
A piece of textA category labelText classification
A piece of textIts linguistic structureLanguage analysis
An imageA written descriptionImage captioning

A language model starts with autocomplete Lec 1 · Aug 25

Imagine typing “I want to ...”. The next word could be swim, eat, or play. We cannot know which one the writer will choose, but we can make sensible guesses. After “The capital of Nebraska is ...”, Lincoln is a much stronger guess.

This is the basic language-modeling task: use the words already given, called the context, to predict what comes next. Both the input \(X\) and the output \(Y\) are language. Grammar, familiarity, and how often an expression occurs all affect which continuations are likely. The idea that different sentences have different probabilities leads to probabilistic language modeling.

The lecture connects this basic idea to chat systems, more capable autocomplete tools, and assistants that follow instructions. These models assign probabilities to possible responses; when a response is sampled, the result can vary. This is what the terms probabilistic and stochastic describe. A deterministic, rule-based system instead gives the same result when its input and rules are unchanged.

The product examples in the lecture show different parts of NLP at work. Siri follows a conversation and works out which alarm the user means. Google Translate detects languages and rearranges phrases during translation. Google Search can connect “fern” with “indoor plant,” locate a passage that answers a question, and find questions with related meanings.

Why focus on language models? Lec 1 · Aug 25

The course connects artificial intelligence, machine learning, NLP, deep learning, generative AI, and language models. These are related and partly overlapping areas, not a precise chain of nested categories. For example, the course begins with language models that do not use neural networks, then moves toward deep-learning methods.

The lecture presents language models as a foundation for NLP. In the text-based models introduced here, the model learns its internal representations from sequences of text. Those representations can then support a growing range of applications.

  • Language models have been important for decades in speech recognition, spelling correction, and machine translation.
  • Their widespread use makes them a topic that people encounter and discuss well beyond research.
  • Their effects on society are large enough to deserve careful study.
  • The lecture also points to a concentration of resources: as models grow, developing the largest ones becomes feasible for only a few major organizations.

What large language models can do Lec 1 · Aug 25 · repeated Lec 2 · Aug 27

The lectures describe large language models, or LLMs, trained on terabytes of data. This can include selected material collected from the web and text generated by other models. Such training exposes a model to information about both familiar and less familiar people, places, and things.

A sufficiently large model can sometimes answer questions even without a separate training stage devoted to question answering. The lecture uses this as one example of how a broad training task can lead to useful abilities.

The applications listed in class include writing code; summarizing articles, podcasts, and presentations; drafting emails and social posts; suggesting titles; playing games; helping with résumés and cover letters; answering trivia; composing music; finding search-engine optimization (SEO) keywords; writing product descriptions; explaining difficult topics; solving mathematics problems; writing articles, blog posts, and quizzes; and adapting content for a different medium. The slides also cite GPT-4's bar-exam result. Together, these examples illustrate the potential to reduce work that would otherwise require substantial manual effort.

Is the model remembering or applying what it learned?

A model encounters examples of many tasks during training. When it later succeeds, we still need to ask whether it recalled something it had seen or applied a pattern to a new situation. This is the distinction between memorization and generalization. The lecture returns to it in the “Human or AI?” examples.

Useful abilities come with unresolved problems

  • A model can produce a convincing but false statement. This is called a hallucination, and it can spread misinformation.
  • Training and generated output raise questions about privacy and copyright.
  • Models can reproduce biases and create other ethical problems.

The course aims to explain both the basic ideas and the practical methods behind these systems, so that their abilities and limits are easier to understand.

What the course covers Lec 1 · Aug 25

The roadmap in the opening lecture follows the development of language models. The periods below are the course's organizing framework.

PeriodWhat you will study
Before neural language models (through 2013)Start with n-gram models, which use a short word history. Study how context helps, how smoothing handles missing observations, and how to evaluate a model.
Early neural language models (2013–2018)Review machine learning through logistic regression. Then study word embeddings, feed-forward networks, recurrent neural-network (RNN) language models, backpropagation, and encoder–decoder models.
Modern neural language models (2018 onward)Study attention and self-attention, the ideas behind Transformers. Compare masked language models with decoder-only models.
Large language modelsLearn how pretraining and fine-tuning build and adapt models, how models generate text, and how prompting and instruction tuning guide them. Study how human feedback helps align their responses with preferences, along with technical problems such as hallucination and social concerns such as privacy.

By the end of the course, you should understand the foundations of language modeling and build a model through homework and/or the project. You should also be able to connect your model to systems such as ChatGPT and GPT-4, explain the abilities and unresolved problems discussed in class, and recognize new problems worth investigating.

The course does not aim to cover everything in NLP. It will not explore individual tasks such as question answering in detail, classical methods for structured prediction such as logical semantics, lambda calculus, and sequence tagging, or linguistics in depth.

Background you will need: linear algebra, probability and statistics, differential calculus, and gradients. You should know the machine-learning basics: training, validation, and testing; parameters and hyperparameters; sampling; and basic optimization. You also need Python. The course includes a short introduction to PyTorch, but expects you to learn the rest through self-study.