> For the complete documentation index, see [llms.txt](https://vikram-bajaj.gitbook.io/cs-gy-6923-machine-learning/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://vikram-bajaj.gitbook.io/cs-gy-6923-machine-learning/main-4/types-of-machine-learning/supervised-learning/neural-networks/mlp.md).

# MLP

This is the other name for a neural network.

## Advantages

* Good accuracy on even data that is far from linearly separable
* Can learn complicated functions or concepts

## Disadvantages

* Danger of overfitting
* Slow to train

**Some Notations**

Consider a simple 2 layer neural network (the input layer is not counted as a layer):

1 output unit, **H** hidden nodes (+1 dummy bias unit), **d** input nodes (+1 bias unit)\
**Fully Connected**: every node in a layer is connected to every node in the previous layer.

**(H+1) + H(d+1)** weights to learn.

$$w\_{h\_j}$$is the weight on the edge from input node $$x\_j$$ to hidden node h, $$v\_h$$is the weight on the edhe from hidden node h to the output node.

$$Z\_0, Z\_1,...,Z\_H$$ (with $$Z\_0=1$$) are the activations from the hidden layer, usually sigmoid i.e. $$Z\_h=\frac{1}{1+e^{-w^Tx}}$$

The **Error Function for Regression** is given by:

$$E(W,v)=\frac{1}{2}\sum\_t (r^t-y^t)^2$$ i.e. the *mean squared error*, with W=weights $$w\_{h\_j}$$ and v=weights $$v\_h$$ and $$y=v^TZ$$ i.e. a \_linear activation function \_i.e. $$y = v\_0.1+v\_1Z\_1+...+v\_HZ\_H$$

The **Error Function for Classification** is given by:

$$E(W,v) = -\sum\_t r^tlogy^t + (1-r^t)log(1-y^t)$$ i.e. the *cross entropy loss*

where $$y = \frac{1}{1+e^{-v^TZ}}$$ i.e. the *sigmoid function*

## Batch Gradient Descent

We must find W, v that minimize the error.

```
1. Initialize W and v
2. Repeat unitl convergence:
       compute v^t for each x^t in training
       update each v_h and w_h_j by doing:
       v_h = v_h - eta * dE/dv_h
       w_h_j = w_h_j - eta * dE/dw_h_j
```

We have:

$$\partial E/\partial v\_h = \sum\_t (r^t-y^t)Z\_h^t$$

$$\partial E/\partial w\_{h\_j} = \sum\_t -(r^t-y^t)v\_hz\_h^t(1-z\_h^t)x\_j^t$$

(computed using chain rule i.e. $$\partial E/\partial w\_{h\_j} = \sum\_t \partial E/\partial y \* \partial y/\partial z\_h^t \* \partial z\_h^t/\partial w\_{h\_j}$$)

This technique is called **backpropagation**.
