Preemptible compute is in public preview for Kubernetes and Slurm clusters. There is no minimum-lifetime guarantee for preemptible nodes during preview, so design workloads that can survive losing nodes at any time.
- Standard nodes (the default) are provisioned up front (synchronously) when you create or scale a cluster, and are never preempted.
- Preemptible nodes fill in over time (asynchronously). You set a target, and Together provisions toward it as spare capacity becomes available. The target is not guaranteed, and preemptible nodes can be preempted at any time.
num_gpus), and you cannot convert a node between standard and preemptible in place.
On Kubernetes clusters, preemptible nodes carry the label together.ai/compute-class=preemptible. Node names encode the compute type: gpu-dp is standard, gpu-preemptible-dp is preemptible.
On Slurm clusters, preemptible nodes are identified by the Slurm feature preemptible. Use --constraint=preemptible in your #SBATCH directives to target them.
How preemption works
When Together reclaims a preemptible node, it drains the node for at most five minutes before removing it. Together then provisions replacement preemptible nodes toward your target as capacity becomes available, so you don’t need to re-request capacity. Here’s how preemption works for each cluster type.Kubernetes
- T+0: The node is cordoned, a
TogetherPreemptionNotifiedKubernetes event (type: Warning) is emitted on the node, and pods on the node receive SIGTERM. - Pods that set
terminationGracePeriodSeconds(capped at 300 seconds) get up to the full window to checkpoint and exit. The node is reclaimed as soon as your pods exit, so checkpoint and exit promptly instead of sleeping through the window. - T+5:00: The node is removed, regardless of pod status. Five minutes is a hard maximum, not a guarantee that pods finish.
Slurm
- T+0: Slurm drains the node. No new jobs start on it, and jobs already running keep running.
- Jobs that finish within the window exit normally.
- T+5:00: The node is removed, and any jobs still running on it are killed. Slurm requeues the killed jobs that were submitted with
--requeue.
Request preemptible capacity
Request a preemptible GPU target alongside the standard count, at cluster create or update, from the console, CLI, or API. See the cluster create API reference for the full schema.
Create a cluster with preemptible capacity. Preemptible capacity requires a project, so pass the global
--project flag (or set TOGETHER_PROJECT_ID):
Schedule workloads on preemptible nodes
Target preemptible nodes explicitly for interruptible workers, and keep coordinators and serving replicas on standard nodes.Kubernetes
Slurm
To run a job on preemptible nodes, add--constraint=preemptible to its #SBATCH directives. Add --requeue too, so Slurm resubmits the job if a node is reclaimed while it runs:
Key directives
Detect preemption
Kubernetes
Preemption surfaces through three channels:-
In-pod (recommended): A SIGTERM handler or
preStophook reacts automatically, with no polling. See Handle preemption for examples. -
Kubernetes events: Watch for
TogetherPreemptionNotifiedevents: -
Together API (no kubeconfig needed): Reading a cluster returns a
node_lifecycle_eventsarray (72-hour retention, deduplicated by node and reason). Filter forreason == "TogetherPreemptionNotified". Related reasons includeTogetherNodeAddedandTogetherScaledDown.
Slurm
On Slurm clusters,TogetherPreemptionNotified events appear in the cluster’s node_lifecycle_events array in the Together API and in the console’s Event Timeline, the same as on Kubernetes.
Inside a job, check SLURM_RESTART_COUNT at startup. Slurm sets it when it requeues a job, so a value above 0 means the job is restarting, for example after its node was reclaimed, and should resume from its latest checkpoint.
Handle preemption
Kubernetes
Pod-level: checkpoint on SIGTERM. This pattern is per-workload and requires no extra infrastructure. Claim the full grace window withterminationGracePeriodSeconds and checkpoint when the signal arrives:
The
preStop hook runs first, then SIGTERM, and both count against the same grace period. Use one or the other as the checkpoint trigger, not both doing duplicate work. The preStop variant matters for containers whose main process can’t trap signals, or where PID 1 swallows them.TogetherPreemptionNotified events and reacts, for example by draining a job queue, triggering a coordinated checkpoint, or removing the node from a Ray or torchrun worker pool:
Slurm
A job still running when its node is removed loses everything since its last checkpoint. Write checkpoints to shared storage under/home at a regular interval, and resume from the latest one when Slurm requeues the job. Checkpoint often enough that losing one interval of work is acceptable.
The following job script resumes from the latest checkpoint whenever SLURM_RESTART_COUNT shows that Slurm requeued it. train.py stands in for your training script, which saves a checkpoint to --checkpoint-dir every --checkpoint-interval steps and loads the latest one when passed --resume-from-latest:
When to use preemptible
Use preemptible compute for:- Checkpointed training, fine-tuning, evals, distillation, batch inference, and hyperparameter sweeps. Ray Train and PyTorch Lightning restart from checkpoints out of the box.
- Bursty research: raise the target for a sweep and drop it after. The cluster stays up.
- The only copy of a multi-day run with no checkpoints.
- User-facing serving replicas with no fallback capacity.
- Workloads with strict SLOs.