5 min read
Which GPU to pick for LoRA training: SDXL, FLUX and 7B models
Choosing a card for LoRA training on SDXL, FLUX and 7B LLMs: memory requirements, QLoRA versus LoRA, and how to estimate training time before you start.
LoRA trains a small add-on to a frozen model, so it needs far less memory than full training. Choosing a card comes down to three things: base model size, precision and batch size.
Memory guide
- SDXL LoRA: 16 GB or more.
- FLUX LoRA: 24 GB or more; less is possible with offloading, but slower.
- QLoRA for a 7B model at 4 bits: 16 GB or more.
- LoRA for a 7B model at fp16: 24 GB or more.
- Full training of a 7B model: about 112 GB, so several GPUs.
Estimating time before you start
- Run 20–50 steps on your dataset with your batch size.
- Measure the seconds per step.
- Multiply by the total step count and add 20% for saving and validation.
That tells you the cost before you commit to the full run. For short training, a slightly more expensive but faster card often comes out cheaper overall.