What Is a FLOP in AI Training? The Number in the Law
A FLOP is one arithmetic operation. Total them across a training run and you get a number that now decides which regulatory tier a model falls into.
A FLOP is one floating-point operation: a single multiplication or addition on decimal numbers. Training compute is measured in total FLOPs, the count of all such operations performed across an entire training run. When a model card says a run used 10^25 FLOPs, it means ten septillion arithmetic operations were executed to produce those weights.
The reason this obscure unit now appears in press releases is that it became a legal threshold before it became a marketing one.
What is a FLOP in AI training, and why count them
Every forward and backward pass through a neural network is arithmetic, which is the part of the model pipeline where almost all the money goes. Multiply an input by a weight, add a bias, repeat across every parameter for every token in the batch, then do it backwards to compute gradients. Total those operations across every step of the run and you have the training compute.
A rough and widely used approximation for a dense transformer is:
training FLOPs is about 6 times parameters times training tokens
The factor of 6 comes from two operations per parameter in the forward pass and roughly four in the backward pass. It is an estimate, not an audit, but it is close enough that you can sanity-check a vendor's claim. A 70 billion parameter model trained on 2 trillion tokens lands near 8.4 x 10^23 FLOPs.
Note what this does not measure: wall-clock time, energy, or cost. Those depend on hardware, utilisation, and how many times the run had to be restarted. FLOPs count the arithmetic that the successful run required, which is why the number is comparable across labs and hardware generations in a way that "we used 10,000 chips for three months" is not.
Why 10^25 is the number you keep seeing
The EU AI Act uses training compute as a bright line. General-purpose AI models trained above a compute threshold expressed in FLOPs are presumed to carry systemic risk, which pulls in a heavier set of obligations than the baseline transparency duties that came into force in August 2026.
This is an unusual piece of regulatory design and worth understanding on its own terms. The law does not attempt to measure capability, because capability is contested and hard to define in statute. It measures the input instead, on the assumption that compute correlates with capability well enough to serve as a proxy. The European Commission's own AI Act policy page is the primary source for the current obligations and dates.
Whether a compute proxy is a good proxy is genuinely debatable. Algorithmic efficiency improves, so a model trained today at a given FLOP count is more capable than one trained at the same count two years ago. A threshold fixed in law and a moving efficiency frontier do not stay aligned. That tension is real and unresolved.
What FLOPs tell you as a builder, and what they do not
They tell you:
Roughly which weight class a model is in, which sets expectations for cost per token and latency.
Whether a model falls inside a regulatory tier that will affect the documentation its provider owes you.
How a lab's investment in one model compares to another, on a scale that survives hardware changes.
They do not tell you:
Whether the model is good at your task. Compute buys general capability, not task fit. Your own evaluation on your own inputs remains the only thing that answers that.
How the compute was spent. A well-balanced run following the compute-to-data balance described by scaling laws beats a badly balanced run at the same FLOP count, sometimes substantially.
Anything about inference cost. Training compute is a one-time cost to the lab. What you pay depends on inference, which is a different calculation entirely.
That last point catches people out. A model with a huge training budget can still be cheap to serve if it is small, and a modestly trained model can be expensive to serve if it is large. The two numbers are related but not the same, which is the underlying reason bigger models cost more per token while training cost barely enters your bill at all.
Reading a FLOP figure without being fooled
Three habits help.
First, check whether the figure is stated or estimated. Labs rarely publish exact training compute. Third-party figures are usually derived from the 6-times-parameters-times-tokens approximation using disclosed counts, and if either count is undisclosed, the estimate is a guess wearing a decimal point.
Second, watch the exponent, not the mantissa. The difference between 2 x 10^24 and 8 x 10^24 is real but small in practice. The difference between 10^24 and 10^26 is two orders of magnitude and a different class of system.
Third, separate the number from the claim attached to it. "Trained with more compute than any previous model" is a statement about spending. It is not a statement about whether the result is better, and it is certainly not a statement about whether it is better for you.
Frequently asked questions
Is a FLOP the same as FLOPS?
No, and the plural is an unfortunate accident of notation. FLOPs with a lowercase s is a count of operations, used for total training compute. FLOPS in capitals is floating-point operations per second, a measure of hardware speed. One is a distance, the other is a speed.
How many FLOPs did a specific model use?
Unless the provider published it, nobody outside the lab knows precisely. Estimates circulate and are often reasonable, but treat any specific figure without a primary source as an approximation.
Does more training compute mean a better model?
On average and across large gaps, yes. Within a narrow range, no. Data quality, architecture, and post-training all move capability at a fixed compute budget, which is why models with similar training compute can differ noticeably in practice.
Why does the EU measure compute instead of capability?
Because compute is countable and capability is not, at least not in a way that holds up in law. It is a deliberate proxy, chosen for being administrable rather than for being precise.
Do I need to care about this if I just use an API?
Mostly no. It becomes relevant when a compute threshold changes what documentation and risk disclosures your provider owes you, which is exactly the material you will be asked for during a client security review.
How did this land?
About the author

Senior Editor, AI & Product
Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.


