7 min read
Ollama and vLLM on a rented GPU: chat and API in 10 minutes
Step by step: a model in Ollama with chat and API, and vLLM with an OpenAI-compatible server on a rented GPU. Commands, checks and common mistakes.
There are two paths. Ollama is simpler: it pulls quantised models with one command and gives you a chat and a simple API. vLLM is for production load: fast batching and an OpenAI-compatible API. Start with Ollama and move to vLLM when speed becomes the limit.
Ollama: chat and API
Rent a server with 24 GB and the Ollama + Open WebUI template. Open the chat from the server card, then pull a model in the terminal:
ollama pull qwen2.5:7b
ollama run qwen2.5:7b "Write a Python sorting function"The API listens on port 11434 on the server itself. Check from the same server:
curl http://localhost:11434/api/generate -d '{"model":"qwen2.5:7b","prompt":"Hello","stream":false}'To call the API from your own app, forward the port over SSH. This keeps the model off the public internet.
ssh -L 11434:localhost:11434 <user>@<server address> -p <port>vLLM: an OpenAI-compatible server
The vLLM template serves the API on port 8000. Start a model with vllm serve; --max-model-len caps the context and saves memory:
vllm serve Qwen/Qwen2.5-7B-Instruct --port 8000 --max-model-len 8192OpenAI SDK clients work unchanged: set base_url to the server address with port 8000. Only protect the endpoint with a key if you configured one on the server, and do not expose an unauthenticated port.
Common mistakes
- The model does not fit: shorten the context, use a quantised build, or pick a bigger card.
- The API is unreachable from outside: Ollama listens on localhost by default. Use an SSH tunnel rather than opening the port.
- The download is slow: large models are tens of gigabytes. Run pull inside screen or tmux.
Stop the server when you are done. The disk holding the model is billed while the server is stopped, so delete what you do not need.