What is Fine-Tuning?
You have a pre-trained language model with billions of parameters that knows a lot about the world. But you want it to be great at a specific task — writing legal briefs, coding in Rust, or speaking like your brand. Fine-tuning adapts the model by continuing training on your specialized data.
The problem: full fine-tuning means updating ALL parameters.
For a 70B parameter model, that means storing and updating 70 billion weights. You need a full copy of the model in memory, plus optimizer states (2-3x the model size). That's hundreds of gigabytes of VRAM — expensive, slow, and impractical for most teams.
The LoRA Insight
LoRA (Low-Rank Adaptation) is based on a key observation: when you fine-tune a model, the weight updates tend to be low-rank. Instead of updating a huge d×d weight matrix W directly, you decompose the update as ΔW = A × B, where A is d×r and B is r×d, with r much smaller than d.
LoRA Matrix Decomposition
Adjust the rank r to see how LoRA decomposes a large weight update into two small matrices.
Why LoRA is Easy to Train
By only training the small A and B matrices while keeping the base model frozen, LoRA dramatically reduces memory, compute, and storage requirements.
Count parameters and storage
Example configuration: 7 billion frozen parameters and 64 target matrices of 4096 × 4096 entries each. Adapter weights and gradients are FP16; the two Adam states are FP32. These are explicit accounting assumptions, not a hardware recommendation.
64 × r × (4096 + 4096) = 8,388,608
- Raw base weights
- 14.000 GB
- Adapter weights
- 0.017 GB
- Adapter gradients
- 0.017 GB
- Adam states
- 0.067 GB
Decimal GB. Excludes activations, temporary buffers, quantization scales and any FP32 master weights. Batch size, context length, target modules and optimizer change actual training memory. QLoRA quantizes the frozen base; LoRA itself does not require a quantized base.
The original base checkpoint stays recoverable
LoRA freezes the base weights, but the active adapter changes the effective weights and can reduce performance on earlier tasks. Additive updates can have positive and negative components. Removing the adapter restores the base path; it does not prove unchanged capabilities while the adapter is active.
Use Cases
LoRA adapters are used everywhere to specialize foundation models:
Task-Specific Adaptation
Train adapters for coding, medical diagnosis, legal analysis, or customer support. Each domain gets its own small adapter.
Style & Tone Adaptation
Match a specific brand voice, switch between formal and casual, or adapt writing style without retraining the whole model.
Language Adaptation
Improve performance in underrepresented languages by training a LoRA on language-specific data.
Instruction Following
Make a base model follow instructions better by training an adapter on instruction-response pairs.
When NOT to Use LoRA
LoRA is powerful, but it's not the right tool for every job:
Prompt Engineering Would Suffice
If you can get the behavior you want with a good system prompt or few-shot examples, don't train an adapter. It's cheaper, faster, and easier to iterate on.
You Need Broad New Knowledge
LoRA is great for style and behavior, but struggles to inject large amounts of factual knowledge. Use RAG (retrieval) instead for knowledge-heavy tasks.
Your Dataset Is Tiny or Noisy
Small or repetitive datasets can overfit. There is no universal minimum example count: task diversity, label quality, model and regularization matter. Keep a held-out evaluation set.
You Need Real-Time Adaptation
LoRA requires a training step. If your use case needs the model to adapt on-the-fly to new information, use in-context learning or RAG instead.
What the rank constraint permits
A fixed low-rank update limits the directions available to each adapted matrix. This is a parameter-efficiency choice, not a theorem that new knowledge or pretraining is impossible with low-rank methods.
What a rank limit means mathematically
A constructed target matrix diag(4, 2, 1, 0.5). Its best rank-r approximation in Frobenius norm keeps the r largest diagonal entries. We measure a matrix error, not language ability or knowledge retention.
‖ΔW − B A‖F = √Σᵢ₎ᵣ σᵢ² = 1.118
LoRA learns A and B during training; it does not know an optimal target matrix in advance. Higher rank permits more independent update directions, but does not guarantee better evaluation results. Data, target modules and training must be evaluated too.
Update rank is bounded
For ΔW = AB, rank(ΔW) is at most r. The effective matrix W + ΔW can still have full rank. Do not confuse the rank of an update with the rank or knowledge of the entire model.
Target modules and data matter
A rank-8 adapter has fewer independent update directions than a full matrix update. Whether that is enough depends on the task and the matrices adapted; measure validation performance and regressions.
Larger rank changes the cost
Trainable parameters grow as r(d_in + d_out) per target matrix. Higher rank allows more expressive updates but adds storage and optimization cost. There is no universal best rank or guaranteed quality curve.
LoRA Variants & Evolution
The original LoRA paper spawned a family of improvements. Click each card to learn more.
QLoRA
▼Quantized base model + LoRA adapters = fine-tuning on consumer GPUs.
DoRA (Weight-Decomposed LoRA)
▼Separates weight magnitude from direction for better training dynamics.
LoRA+
▼Different learning rates for A and B matrices = faster convergence.
Key Takeaways
- 1LoRA parameterizes an update as two low-rank factors; count savings from the actual dimensions and chosen rank.
- 2The base weights stay frozen, but active adapters can change or degrade behavior. Evaluate them; removing an adapter restores the base path.
- 3Low-rank updates can be effective for adaptation, but suitable rank and target modules depend on the task.
- 4QLoRA keeps a quantized frozen base and trains adapters. Activation and optimizer memory still depend on the training setup.
- 5A fixed adapter rank constrains updates. Do not generalize that constraint into a prohibition on all low-rank pretraining methods.