GPU Sandboxes
Attach dedicated GPUs to a sandbox for CUDA, model inference, and training workloads.
A GPU sandbox is an ordinary sandbox with one or more GPUs attached. It behaves like any other sandbox — same isolation, same toolbox, same create/exec/delete surface — except that code inside it can see and use real GPUs. This makes GPU sandboxes a good fit for CUDA workloads, model inference and fine-tuning, and any batch job that needs acceleration on demand without managing your own GPU hosts.
You request a GPU by setting the number of GPUs you want at create time. Each GPU is dedicated to the sandbox for its lifetime, and the platform selects the physical hardware and places the sandbox on a capable host for you.
Requesting a GPU
Set resources.gpu to the number of GPUs to attach (the REST field is gpu_host). Omit it, or set it to 0, for an ordinary CPU-only sandbox. Pair it with an image that ships the NVIDIA userspace tools — a CUDA base image such as nvidia/cuda:12.4.1-base-ubuntu22.04 is the simplest starting point.
How GPU allocation works
Requesting a GPU is a count, not a hardware pick. You ask for N GPUs; the platform chooses which physical cards to assign and which host to run on, based on where capacity is available. This keeps your code portable — you never hardcode a device — and lets the platform place the sandbox on suitable hardware.
Each attached GPU is dedicated to your sandbox for its lifetime; GPUs are not shared between sandboxes. Because a GPU sandbox has to land on a GPU-capable host with free cards, creation can take longer than a plain cached-image sandbox, and a request that cannot currently be placed will fail rather than wait indefinitely — retry, or request fewer GPUs.
Note: The image must provide the NVIDIA userspace libraries and tools (a CUDA image does). A GPU attached to a sandbox whose image lacks the CUDA runtime is still present, but
nvidia-smiand CUDA programs will not be available until you install them.
Choosing the GPU count
Request the number of GPUs your workload actually parallelizes across:
- 1 GPU — single-GPU inference, fine-tuning, and most interactive/agent workloads. Start here.
- Multiple GPUs — multi-GPU training or inference that shards a model across cards. Only request more than one if your framework uses them; extra dedicated GPUs bill whether or not your code touches them.
Checkpoint and restore
A GPU sandbox can be checkpointed: the platform freezes the sandbox — including its live CUDA state and VRAM — into a named, durable artifact, and later restores it into a fresh sandbox with the workload resumed exactly where it left off. This is the GPU form of a VM snapshot: capture a warmed-up model or a long-running job once, then bring it back on demand instead of cold-booting and re-loading weights every time.
You do not ask for a "checkpoint" explicitly. Calling create_snapshot on a GPU sandbox automatically captures a checkpoint (the platform selects the checkpoint path for you because the sandbox has GPUs); the kind argument is ignored for GPU sandboxes.
Capturing a checkpoint
The sandbox must be running. The call returns immediately with a snapshot id; capture happens asynchronously and the snapshot moves pending → ready. During capture the sandbox's VRAM is evicted to the checkpoint artifact, so the snapshot embeds the exact device memory of your workload.
Snapshot names must be unique within your project; reusing a name returns 409 Conflict. Poll the snapshot until it is ready before restoring (use ?wait=ready on the snapshot GET to long-poll, as shown in VM Snapshots).
Restoring a checkpoint
Restore by creating a new sandbox whose image references the checkpoint by name — the same image source used for any VM snapshot. The new sandbox comes up with the captured processes and VRAM already in place; you do not set resources.gpu on a restore (the GPU count comes from the checkpoint).
A restore lands on any capable host in the checkpoint's region that has the same GPU model and enough free VRAM — the platform picks the host and the physical card for you. The restored device memory is identical to what was captured, even when the sandbox lands on a different physical GPU than the one it was checkpointed on.
Three things to know about restore placement:
- Same GPU model. A checkpoint restores only onto the GPU model it was captured on; you cannot checkpoint on one class of card and restore onto another.
- Region-pinned. A checkpoint restores in the region it was captured in; the artifact is not replicated across regions.
- VRAM must be available. If no host in the region currently has enough free VRAM of the right model, the restore fails rather than waiting — retry, or free capacity.
Note: GPU checkpointing requires the CUDA checkpoint tooling to be present inside the sandbox image. The platform's CUDA base images include it. If you bring a custom image without it,
create_snapshoton a GPU sandbox will report that the sandbox is not checkpointable.
A warm-model workflow
A common pattern: load a model once, checkpoint the warmed-up sandbox, then restore a ready-to-serve copy per request or per task — skipping the multi-minute weight load each time.
Limits and current scope
A few of the durability features documented elsewhere behave differently for GPU sandboxes:
- Use checkpoints instead of pause/resume. A GPU sandbox does not support warm pause and resume. To preserve GPU state across a teardown, checkpoint it and restore into a new sandbox when you need it again.
- Count only. You request a number of GPUs; you cannot pin a specific model or device from the SDK. If you need a particular class of hardware, contact us about region and hardware availability.
- The reported GPU count on the sandbox object is not populated. Inspect the GPUs from inside the sandbox with
nvidia-smirather than reading a field off the sandbox object.
For everything else — running commands, files, ports, environment and secrets — a GPU sandbox works exactly like any other sandbox.