Run LLMs on a rented GPU: Llama, Qwen, DeepSeek, Mistral
Large language models are limited by video memory: the weights have to sit on the GPU, or speed drops by an order of magnitude. Rent a server sized for your model, start it in one command, and pay only for the minutes you use.
In short
Language models run on a graphics card when the weights fit entirely in its memory: most models need 24 GB or more per GPU. On gpu.nz such servers start from 21 ₽ per hour; the Ollama with Open WebUI or vLLM template usually launches in 1–5 minutes.
8–14B parameter models fit in 24 GB: 14B in int8, 8B at full precision. 32–34B models only fit 24 GB at int4. 70B needs an 80 GB card at int4, or two cards.
Video memory
24 GB+
Templates
Ollama + Open WebUI
How to start
Pick a card with 24 GB or more in the catalog and rent a server with the Ollama + Open WebUI template.
Open the chat interface from the server card and pull a model, for example qwen2.5:7b or llama3.1:8b.
For an API, connect over SSH and call http://localhost:11434 (Ollama), or start vLLM with vllm serve.
Ollama + Open WebUI: Your own ChatGPT: Llama, Qwen, DeepSeek, Gemma. Browser chat and an API.
An 8B model runs comfortably on 24 GB at full precision, or on 8–12 GB at int4. Leave extra memory for long context.
Can I run DeepSeek or Qwen 32B?
Yes. On 24 GB a 32–34B model runs at int4. On 48 GB or more you get a higher-quality quantisation.
Do I need my own model API key?
No. Open-weight models are downloaded from their public repositories. You pay only for the GPU server.
What happens to the model if I stop the server?
Files stay on the disk as long as the server exists. The disk is billed while the server is stopped, so remove what you do not need or save results first.