How neural networks learn
A neural network learns by guessing, measuring how wrong the guess was, and nudging its weights to be a little less wrong next time.
- 1
Layers of neurons
A network is layers of simple units. Every connection has a weight, a number the network learns.
- 2
Forward pass
Input values flow left to right. Each neuron adds up its weighted inputs and applies a simple function.
- 3
Measure the error
The output is compared with the right answer. A loss function turns the gap into a single number.
- 4
Backpropagation
The error flows backward. For every weight, calculus gives how much it contributed to the loss: its gradient.
- 5
Gradient descent
Each weight moves a small step against its gradient. Repeated over many examples, the loss goes down.
Layers of neurons
A network is layers of simple units. Every connection has a weight, a number the network learns.
Forward pass
Input values flow left to right. Each neuron adds up its weighted inputs and applies a simple function.
Measure the error
The output is compared with the right answer. A loss function turns the gap into a single number.
Backpropagation
The error flows backward. For every weight, calculus gives how much it contributed to the loss: its gradient.
Gradient descent
Each weight moves a small step against its gradient. Repeated over many examples, the loss goes down.
In short
- Training is a loop: forward pass, loss, backpropagation and update, repeated over many batches of examples.
- The learning rate sets the step size. Too large and training diverges, too small and it crawls.
- The same idea scales from this toy network to language models with billions of weights.