Part IV · Neural Networks

From Logistic Regression to Neural Networks

A neural unit is a weighted sum plus a non-linear function. Stack these units in layers and you get a feed-forward network.

Covered inLec 7Sep 15

Lecture 7 introduces neural networks in its last six slides and stops at the diagram of a multilayer network. The lecture scheduled for Sep 17 continues with backpropagation. This page covers what the slides show and marks the few places where it adds an explanation.

Logistic regression is already a tiny neural network Lec 7 · PDF p. 64

Look again at what a logistic regression classifier does. It takes input features \(x_1,\dots,x_d\), multiplies each by a weight, adds a bias, and passes the total through a sigmoid. The slide draws this as a single neural network unit:

Weighted sum, then a non-linear activation
\[ z=\Big(\sum_{i=1}^{d}w_ix_i\Big)+b=\mathbf{w}\cdot\mathbf{x}+b, \qquad \hat y=\sigma(z) \]

In the diagram, the bias is drawn as an extra input \(x_0=1\) with weight \(b\). Writing it that way lets the sum absorb the bias as one more term. The step that turns \(z\) into an output is called the activation function. Here it is the sigmoid, but any non-linear function can play that role.

The picture resembles a neuron in the brain: several incoming connections with different strengths, one combined signal, and a response that depends on the signal in a non-linear way. That resemblance is where the name comes from. The unit is a piece of arithmetic, not a model of biology.

The activation function is the key ingredient Lec 7 · PDF p. 65

The slide names three common activation functions and shows their curves.

FunctionFormulaOutput rangeValues at \(z=-2,\,0,\,2\)
Sigmoid\(\sigma(z)=\dfrac{1}{1+e^{-z}}\)\((0,1)\)0.119, 0.5, 0.881
tanh\(\tanh(z)=\dfrac{e^{z}-e^{-z}}{e^{z}+e^{-z}}\)\((-1,1)\)−0.964, 0, 0.964
ReLU (rectified linear unit)\(\operatorname{ReLU}(z)=\max(0,z)\)\([0,\infty)\)0, 0, 2

Sigmoid and tanh are both S-shaped. Tanh is a rescaled sigmoid that is centered at zero: \(\tanh(z)=2\sigma(2z)-1\). ReLU is the simplest of the three. It passes positive values through unchanged and replaces negative values with zero. The slide marks ReLU as the most common choice in practice. The three columns on the right are added here to make the shapes concrete.

Why does non-linearity matter? Lec 7 · PDF pp. 66–67

A model that only takes weighted sums can separate two classes with a straight line, or in more dimensions a flat boundary. The slide's first plot shows two clouds of points with a line between them. The second plot shows a harder case: red points in the upper-left and lower-right corners, blue points in the upper-right and lower-left. No single straight line separates red from blue. The data are linearly inseparable.

The next slide applies a tanh transformation to the inputs and plots the result. After the transformation, the red and blue points can be separated by a straight line. A non-linear function reshapes the space so that a linear boundary becomes enough. That is the power of non-linearity.

Added explanation: why the non-linearity cannot be skipped

Suppose a network computed a weighted sum, then another weighted sum of the result, with no activation function in between. Writing the two layers as matrices, \(U(W\mathbf{x})=(UW)\mathbf{x}\), which is a single weighted sum with the combined matrix \(UW\). Stacking linear layers therefore never gets past a linear boundary. The activation function between the layers is what gives the second layer something new to work with.

Stack units into a feed-forward network Lec 7 · PDF p. 68

A feed-forward neural network, also called a multilayer perceptron, arranges units in layers. The slide's diagram has three:

  • An input layer with features \(x_1,\dots,x_{d_0}\), plus the constant \(x_0=1\) for the bias.
  • A hidden layer with units \(h_1,\dots,h_{d_1}\). Each hidden unit is a weighted sum of all the inputs followed by an activation function. The weights form a matrix \(W\) and the biases a vector \(\mathbf{b}\).
  • An output layer with values \(y_1,\dots,y_{d_2}\), computed from the hidden units with another weight matrix \(U\).

Information moves in one direction, from inputs to hidden units to outputs, which is why the network is called feed-forward. Written as equations, one hidden layer looks like this:

Hidden layer: weighted sums of the inputs, then an activation function \(g\) applied to each unit
\[ \mathbf{h}=g(W\mathbf{x}+\mathbf{b}) \]
Output layer: weighted sums of the hidden units, then softmax for a probability over classes
\[ \mathbf{z}=U\mathbf{h}, \qquad \hat{\mathbf{y}}=\operatorname{softmax}(\mathbf{z}) \]

The slide shows the diagram and the matrix labels \(W\), \(\mathbf{b}\), and \(U\); the equations are written out here following the assigned reading, Jurafsky & Martin, Chapter 7. Compare the output step with multinomial logistic regression. It is the same computation, except that the input is the learned hidden vector \(\mathbf{h}\) instead of the raw features. A feed-forward network is multinomial logistic regression stacked on top of a learned, non-linear transformation of its input.

The slide adds that such a network can “technically learn any function.” The claim refers to a result about approximation: with enough hidden units, a network of this shape can approximate a wide class of functions as closely as desired. That result does not say how many units are needed, does not promise that gradient descent will find the right weights, and does not guarantee good behavior on new examples. The qualifications are added here; the slide states only the headline.

The deck ends with the line “let's break it down by revisiting our logistic regression model,” which is where the next lecture picks up.

What to take from this page
  • A neural unit is a weighted sum plus a non-linear activation. Logistic regression is one such unit with a sigmoid.
  • Sigmoid, tanh, and ReLU are the standard activations. ReLU, \(\max(0,z)\), is the most common.
  • Without a non-linearity, stacked layers collapse into one linear model and cannot separate patterns like the four-corner example.
  • A feed-forward network computes \(\mathbf{h}=g(W\mathbf{x}+\mathbf{b})\) and then \(\operatorname{softmax}(U\mathbf{h})\). Its output layer is multinomial logistic regression on learned features.