What Is Gradient Descent in AI? A Plain Explanation
Gradient descent is the step-by-step algorithm that updates an AI model's weights during training. This explainer walks through one real weight update by hand, with actual numbers instead of an analogy.
Gradient descent is the algorithm that lets an AI model learn from its mistakes. Every time a model makes a prediction, gradient descent measures how wrong that prediction was, works out which direction to nudge each internal number, called a weight, to make it less wrong, and moves the weight a small step in that direction. Do this across millions of examples, and a model that started out guessing randomly ends up making accurate predictions. It is the training engine behind almost every AI model in use today, from a simple regression tool to a large language model.
What a loss function actually measures
Gradient descent needs something to minimize, and that is the loss function. A loss function is a formula that turns how wrong a prediction was into a single number. Predict close to the real answer and the loss is small. Predict far off and the loss is large.
A common choice is squared error: take the difference between the prediction and the actual answer, then square it. Squaring makes every error positive and punishes big misses harder than small ones. Google's machine learning course walks through this exact approach on a basic linear model, and the same mechanics scale to far more complex networks.
Once you have a loss number, gradient descent's job is simple to state and slow to do by hand: change the weights so the loss goes down.
How this fits into how AI models learn
A neural network is not one adjustable number, it is millions or billions of them, arranged in layers that each transform the input a little before passing it along. For a closer look at that layered structure, see our explainer on what a neural network actually is.
Training a model runs a loop: send an example through the network to get a prediction (the forward pass), compute the loss, compute the gradient with respect to every weight, then nudge each weight a small step opposite its gradient. Repeat across a training set, sometimes trillions of examples for a large language model, and the weights settle into values that make good predictions. This loop is the mechanical core of how AI models actually learn, whatever the input or output happens to be.
Most models do not stop there. After pretraining with gradient descent, many go through further rounds of fine-tuning steered by human feedback rather than a fixed loss function. That process, reinforcement learning from human feedback, still relies on gradient-based updates, but the signal comes from human preferences instead of a right-or-wrong answer.
The gradient: which way is downhill, and by how much
The word gradient just means slope. For a single weight, the gradient of the loss tells you which direction increasing that weight would push the loss, and how steep that effect is right now. Gradient descent moves opposite that slope, because the goal is to reduce loss.
The walking-downhill metaphor people reach for is not wrong, but it will not show you what actually happens inside the model. Here is what it looks like with real numbers.
A real weight update, step by step
Take the simplest possible model: one input, one weight, no bias term. It predicts a number as prediction = w × x.
Say the true answer for a given input is 10, the input x is 2, the model's current weight w is 1.0, and the learning rate is 0.01.
Given: x = 2, actual y = 10, starting weight w = 1.0, learning rate = 0.01
Step 1, predict: prediction = w * x = 1.0 * 2 = 2.0
Step 2, loss (squared error): (prediction - actual)^2 = (2.0 - 10)^2 = 64
Step 3, gradient of loss with respect to w: 2 * x * (prediction - actual) = 2 * 2 * (2.0 - 10) = -32
Step 4, update the weight: new_w = w - (learning_rate * gradient) = 1.0 - (0.01 * -32) = 1.32
Step 5, recheck: new_prediction = 1.32 * 2 = 2.64, new_loss = (2.64 - 10)^2 = 54.29The loss dropped from 64 to 54.29 in one step, using nothing but arithmetic. Run the same loop again with the new weight and loss keeps falling. Run it thousands of times, across thousands of examples and millions of weights instead of one, and that is what training an AI model actually is: repeated arithmetic, not intuition.
Backpropagation vs gradient descent
These two terms get used interchangeably, and that blurs something worth keeping separate. Backpropagation computes the gradient, the direction and steepness numbers for every weight, by applying the chain rule backward from the output layer to the input layer. Gradient descent takes those gradients and updates the weights.
Put another way, backpropagation answers "what should each weight change by," and gradient descent is the part that changes it. One detailed comparison of the two terms puts it plainly: backpropagation refers only to the method for computing the gradient, while a separate algorithm, such as stochastic gradient descent, is what performs the learning.
Training a neural network runs both every step: forward pass, backpropagation for the gradients, gradient descent (or a variant) to update the weights, then repeat.
Plain gradient descent, as shown above, can update weights after looking at every example before making one move. In practice, most training uses stochastic gradient descent, which updates after each small batch instead of the whole dataset, plus adaptive methods like Adam that adjust the effective learning rate per weight automatically. The underlying idea, measure the loss, compute the gradient, take a small step, stays the same across all of them.
Why this matters if you are building with AI, not researching it
You do not need to compute a gradient by hand to ship a product built on AI. Most tools that let you build an app with AI hide all of this behind an API call. But the mechanics explain a lot of what you will run into anyway: why a model needs a large, varied dataset to generalize instead of memorize, and why a fine-tuning run that will not improve usually points to the learning rate or the data, not bad luck. Knowing what is happening under the hood makes those moments less mysterious.
Frequently asked questions
Is gradient descent the same thing as machine learning?
No. Gradient descent is one optimization algorithm used inside machine learning, the piece that adjusts a model's weights during training. Machine learning is the broader field, and plenty of its methods, like decision trees, do not use gradient descent at all. Most neural networks and large language models do.
What does the learning rate actually control?
The learning rate controls how big a step gradient descent takes on each update. It gets multiplied by the gradient before that amount is subtracted from the weight. A rate of 0.01 versus 0.1 versus 1.0 changes only the size of the step, not the direction it moves.
What happens if the learning rate is too high or too low?
Too high, and updates overshoot the low point of the loss, sometimes making it bounce around instead of shrink. Too low, and training crawls, needing far more steps, or it settles into a shallow dip that is not the best possible answer.
Is backpropagation a type of gradient descent?
No, they are separate steps in the same loop. Backpropagation calculates the gradient for every weight using the chain rule. Gradient descent then uses that gradient to actually update the weights. Neither one produces a trained model without the other.
Does every AI model use gradient descent?
Nearly every neural network does, including the models behind modern chatbots, image generators, and recommendation systems. Some other approaches, like random forests or k-nearest neighbors, learn a different way and skip gradient descent entirely.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


