Dashboard

What Is a Loss Function in AI?

A loss function is the number an AI model minimizes during training. Here's a plain-English explanation with a worked example showing how loss is calculated and why it drops as a model learns.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
8 September 20261 min read

What Is a Loss Function in AI?

A loss function is the formula an AI model uses to measure how wrong one prediction is. It compares the model's output to the correct answer and returns a single number: the loss. A high loss means the guess was far off. A low loss means it was close. Training a model is the repeated process of nudging its internal numbers, called weights, in whatever direction makes that number smaller. Nothing more mysterious happens underneath it: the model guesses, the loss function scores the guess, the weights get a small correction, and the cycle repeats millions of times until the loss stops shrinking much.

A worked example: computing loss by hand

Picture a model built to predict a house's sale price, in thousands of dollars. The house actually sold for 350. Early in training, before the weights have adjusted much, the model predicts 320. One of the simplest and most common loss functions, squared error, is calculated like this:

text
loss = (predicted - actual)^2
loss = (320 - 350)^2
loss = (-30)^2
loss = 900

That 900 is the loss for this one prediction. The unit, thousands of dollars squared, doesn't matter on its own. What matters is the size of that number relative to other predictions the model makes. After more training steps adjust the weights, the same model might predict 340 for the same house:

text
loss = (340 - 350)^2 = 100

The loss dropped from 900 to 100. That drop is what people mean when they say a model "got better" during training. It is literally this number going down, step after step, prediction after prediction. Averaged across many training examples instead of just one, this exact calculation has a name: mean squared error (MSE), and it's the default loss for most models that predict a continuous number.

How does an AI model learn

A neural network turns an input into a prediction by passing numbers through layers of weighted connections, each followed by a nonlinear step called an activation function. None of those weights start out useful. They begin as small random numbers, which is why an untrained model's first guesses look close to noise. See how AI models work for the layer-by-layer mechanics. The loss function is what turns a random guess into a number worth improving. Without it, there would be nothing to correct against.

Gradient descent and the loss function

Knowing the loss for a single prediction only tells you how wrong the model was. Gradient descent is the algorithm that decides how to fix it. It calculates the gradient, essentially the slope, of the loss function with respect to each weight, then nudges every weight a small step in the direction that makes the loss smaller. Do that across millions of examples and thousands of small steps, and the weights settle into values that produce accurate predictions. This is also why training a model is computationally expensive: every step requires recalculating the loss and its gradient before any weight moves at all.

Loss function vs accuracy

Loss and accuracy measure different things, and mixing them up causes confusion. Accuracy is simple: for a classification task, what percentage of predictions were exactly right. It's a blunt number, but it can't be used to actually train a model, because it doesn't change smoothly. A model that predicts "dog" with 51% confidence and one that predicts "dog" with 99% confidence on the same mislabeled image get the identical accuracy score: wrong. Loss doesn't have that problem. A confident wrong answer produces a much higher loss than an unsure wrong answer, which gives gradient descent something real to push against. That's why training logs report loss at every step, while accuracy is usually checked only periodically, more like a scoreboard than a steering signal.

Different tasks use different loss functions

Squared error works well for predicting numbers, house prices, temperatures, sensor readings, but it's the wrong tool for classifying an image or predicting the next word in a sentence. Those tasks typically use cross-entropy loss, which measures how far a predicted probability distribution is from the correct one. A diffusion model, the kind behind most AI image generators, uses a variant: its loss measures how well the model predicts the noise it was trained to remove from an image at each step. The formula changes based on what the model outputs, but the job stays the same: produce a number that goes down as predictions get better.

Why this matters even if you never train a model

Most people interacting with AI never train a model, they just prompt one. But understanding what the model minimized during training explains a lot of prompting behavior. A model was trained to reduce loss on patterns that showed up heavily in its training data, so prompts structured the way that data was structured tend to produce more reliable output. That's part of why prompt engineering, phrasing a request in a way the model was effectively rewarded for handling well, makes such a measurable difference in the quality of what comes back.

Frequently asked questions

What does a low loss value mean?

A low loss means the model's predictions are close to the correct answers, on average, across whatever data it was measured on. It doesn't by itself mean the model generalizes well. A model can drive its training loss close to zero and still perform poorly on new data it wasn't trained on. That gap between training performance and real-world performance is called overfitting.

Is a loss function the same thing as a cost function?

They're closely related and often used interchangeably, but there's a technical distinction. A loss function usually refers to the error on one single prediction. A cost function is the average of that loss across an entire training batch or dataset. Mean squared error, as commonly used, is technically a cost function since it's averaged; the per-example calculation underneath it is the loss.

Can a model's loss actually reach zero?

In practice, no, and if it does, that's usually a warning sign rather than a success. A loss of exactly zero across an entire training set typically means the model has memorized the training examples rather than learned general patterns, a form of overfitting. Well-trained models settle at some small, nonzero loss.

Why does loss sometimes go up during training instead of down?

A short-term rise in loss during training is common and not automatically a problem. It can happen when the learning rate, the size of each weight update, is a bit too high and the weights overshoot. A loss that keeps climbing over many steps, rather than just a step or two, usually points to a learning rate set too high or a problem in how data is being fed to the model.

What loss function should be used for classification versus regression?

Regression tasks, predicting a number, most commonly use mean squared error or a close relative like mean absolute error. Classification tasks, predicting a category, most commonly use cross-entropy loss, sometimes called log loss. Matching the loss function to the type of output the model produces is one of the more consequential setup decisions in training any model.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.

What Is a Loss Function in AI? | swarmz.net