For training, memory matters as much as speed. LoRA and QLoRA let you fine-tune a 7–8B model on one card, while full fine-tuning needs several GPUs. Start with a LoRA on a short dataset and time one step.
In short
Fine-tuning (LoRA, QLoRA) needs memory for the weights, gradients and optimizer: 24 GB or more per GPU. On gpu.nz such servers start from 21 ₽ per hour; save checkpoints so that stopping a server does not lose progress.
QLoRA for a 7B model needs 16 GB or more, fp16 LoRA for 7B needs 24 GB, SDXL LoRA needs 16 GB. Full fine-tuning of a 7B model with Adam needs roughly 112 GB, which means two 80 GB cards.
Video memory
24 GB+
Templates
Kohya SS
How to start
Rent a server sized for your model: 24 GB for a 7B LoRA, 16 GB for QLoRA or an SDXL LoRA.
Pick the Kohya template for images or PyTorch + Jupyter for LLMs, and prepare the dataset.
Run a short test, time one step, multiply by the step count, then save the checkpoint and stop the server.
Kohya SS: Train LoRA and DreamBooth models for SDXL and FLUX.