4 min read
What interruptible GPU servers are and when to use them
An interruptible server costs less than on-demand, but it can be paused if someone outbids you. Which tasks suit it and how to avoid losing work.
An interruptible server runs on a bid: the price is lower than on-demand, but it can be paused if a higher bid comes in for the same card. An on-demand server runs until you stop it.
When it suits
- Batch jobs that can resume from a checkpoint: training, bulk generation, rendering.
- Experiments where losing the last few minutes does not matter.
- Work that saves results to disk regularly.
When on-demand is better
- Interactive work: chatting with a model, a notebook, a client demo.
- An API that has to answer without interruption.
- A long run without checkpoints.
How not to lose work
Save checkpoints and results to disk regularly, and write training scripts so they can resume from the last save. In the catalog, compare the bid with the on-demand price: the gap decides whether the risk is worth it.