6 min read
How much VRAM an LLM needs: a table and a rule of thumb
How to estimate memory for model weights at fp16, int8 and int4, how much 8B, 14B, 32B and 70B models need, and why context also uses memory.
The main rule: the model weights have to fit entirely in video memory. If they do not, some layers spill into system RAM and generation becomes many times slower.
The formula
Memory for weights is roughly the parameter count times bytes per parameter: 2 bytes at fp16, 1 byte at int8, 0.5 bytes at int4. Add about 20% for the KV cache, activations and buffers.
memory ≈ parameters (billions) × bytes/param × 1.2
example: 14B at fp16 = 14 × 2 × 1.2 ≈ 34 GB
14B at int4 = 14 × 0.5 × 1.2 ≈ 8.4 GBA table for common sizes
- 8B: fp16 ≈ 19 GB, int8 ≈ 10 GB, int4 ≈ 5 GB. On 24 GB, fp16 or int8 with room to spare.
- 14B: fp16 ≈ 34 GB (does not fit 24 GB), int8 ≈ 17 GB, int4 ≈ 8 GB.
- 32B: int8 ≈ 38 GB, int4 ≈ 19 GB. On 24 GB, int4 only.
- 70B: int4 ≈ 42 GB, int8 ≈ 84 GB. One 80 GB card: int4; int8 needs a second card.
Context uses memory too
Long chats or documents grow the KV cache. For an 8B model at a few thousand tokens of context that is a GB or two; at tens of thousands it is much more. Leave headroom and do not fill the card to the limit.
What to pick
- Chat and code with 7–8B models: a 24 GB card.
- 32–34B models at good quality: 24 GB at int4, or 48 GB and up.
- 70B: 80 GB at int4, or two cards.
These numbers are estimates for short context. The exact value depends on the framework, quantisation format and settings. Check by running the model on a rented server: hourly billing lets you test in a few minutes.