Dashboard

What Is LoRA in AI? Low-Rank Adaptation Explained

LoRA, short for low-rank adaptation, is a way of customising a large model by training a small set of extra weights instead of changing the model itself. The base model stays frozen. You train two ...

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
8 September 20261 min read

What Is LoRA in AI? Low-Rank Adaptation Explained

LoRA, short for low-rank adaptation, is a way of customising a large model by training a small set of extra weights instead of changing the model itself. The base model stays frozen. You train two thin matrices per layer, keep those, and throw nothing else away. The result is a file measured in megabytes that steers a model measured in gigabytes, and that size difference is the entire reason the technique matters.

It answers a specific question: how do you get a model to behave differently for your use case without paying to retrain it or store your own copy of it.

The arithmetic that makes it small

Inside a transformer, a weight matrix might be 4096 by 4096, which is about 16.8 million numbers. Full fine-tuning updates every one of them, and you end up owning a complete second copy of the model.

LoRA starts from an observation in the original 2021 paper by Hu and colleagues: the update you want to apply during fine-tuning has low intrinsic rank. It can be approximated well by a much simpler matrix. So instead of learning a 4096 by 4096 update, you learn two matrices, one 4096 by r and one r by 4096, and multiply them together to reconstruct something the right shape.

Set r to 8 and the count goes from 16.8 million numbers to about 65,500. That is roughly a 256-fold reduction on one matrix, and it compounds across every layer you adapt. The paper reports reducing trainable parameters by a factor of 10,000 against full fine-tuning of GPT-3 175B, with about a threefold cut in GPU memory.

The rank r is the dial. Low values, 4 to 16, capture style, tone, and format reliably. Higher values, 32 to 128, are what you reach for when you need the model to absorb genuinely new patterns rather than to behave differently with knowledge it already has. Going higher costs training time and adapter size, and past a point it stops helping.

What you get that full fine-tuning does not give you

The size is the headline; the consequences are the point.

  • Adapters are swappable. One base model in memory can serve many customers, each with their own adapter loaded on demand, which is impossible if every customer needs their own full copy of the weights.

  • Adapters are cheap to keep. Storing fifty variants of a behaviour costs less than storing one extra copy of the model.

  • Adapters can be merged. At inference time the two matrices can be folded back into the base weights, so a merged adapter adds no latency at all. You trade the ability to hot-swap for speed.

  • Training is affordable. Memory is dominated by optimiser state, and optimiser state scales with trainable parameters, which is why LoRA fits on hardware where full fine-tuning does not.

QLoRA is the obvious extension: quantise the frozen base model to 4-bit, then train LoRA adapters on top. Since the base is never updated, quantisation error does not accumulate through training the way it would if you were changing those weights. It is how people fine-tune large models on a single consumer GPU. If the quantisation half is unfamiliar, what quantization means in AI covers it separately.

When LoRA is the wrong tool

LoRA changes behaviour. It is poor at installing facts.

You want

Reach for

Consistent output format or house voice

LoRA

A domain style the base model handles awkwardly

LoRA

Answers grounded in your documents

Retrieval, not LoRA

Facts that change weekly

Retrieval, not LoRA

A capability the base model lacks entirely

Full fine-tuning or a different model

The failure people hit most often is trying to teach a model their product documentation with an adapter. It half works, which is worse than not working: the model produces confident text in the right shape with the details wrong. The comparison in RAG, fine-tuning, and long context is the right decision framework before you train anything at all.

How it fits with the rest of the stack

An adapter is a set of weights, so everything true of model weights in AI is true of it, including that it carries a license and that publishing one may carry obligations from the base model it was trained against. If the layers and matrices in the section above went past too quickly, how AI models work covers the structure LoRA is attaching itself to.

It also sits in the same family as the other post-training techniques. Fine-tuning in AI is the general category, LoRA is the parameter-efficient member of it, and the practical toolchain most people use is Hugging Face's PEFT library, which implements LoRA alongside several relatives.

For anyone running models on their own hardware, adapters are the reason a single local base model can cover several jobs, which is a large part of what makes running an ai coding model locally practical rather than theoretical.

FAQ

How big is a LoRA adapter in practice?

Typically single-digit to low hundreds of megabytes, depending on rank and how many layers you adapt, against a base model of several gigabytes or more. The ratio is usually three or four orders of magnitude.

Can you use more than one LoRA at a time?

Yes, adapters can be stacked or weighted, and serving frameworks support loading several against one base model. Results get less predictable as you combine more, because nothing guarantees two independently trained adapters compose cleanly.

Does LoRA make inference slower?

Only if you keep the adapter separate so it can be swapped, which adds a small amount of extra computation. Merge it into the base weights and the merged model runs at exactly base speed.

What rank should I start with?

Start at 8 or 16 for style and formatting work, and only raise it if evaluation shows the adapter is underfitting. Higher rank costs training time and size, and it is rarely the reason a run failed.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.