SLURM Job Script Examples for PyTorch (Single GPU, Arrays, Multi-GPU)

Three working SLURM job scripts for ML research: one GPU, a job array across seeds, and multi-GPU with torchrun, plus the mistakes that waste the most time.

By Dr Raktim Mondol · 29 September 2026 · 2 min read

Most university clusters use SLURM to schedule jobs. You write a shell script that describes what resources you need and what to run, submit it with sbatch, and the scheduler starts it when a machine is free. Here are the three job scripts you will use most for PyTorch research, and how to avoid the mistakes that cost people days.

1. A single-GPU training job

#!/bin/bash
#SBATCH --job-name=train_baseline
#SBATCH --partition=gpu              # your cluster's GPU partition (see: sinfo)
#SBATCH --gres=gpu:1                 # 1 GPU (some clusters use --gpus=1)
#SBATCH --cpus-per-task=8            # CPU cores for data loading
#SBATCH --mem=32G                    # system RAM, not GPU memory
#SBATCH --time=04:00:00              # hh:mm:ss
#SBATCH --output=logs/%x_%j.out      # %x = job name, %j = job ID
#SBATCH --error=logs/%x_%j.err

set -euo pipefail

module purge
module load cuda                     # module names differ per cluster
source ~/miniconda3/etc/profile.d/conda.sh
conda activate myenv

echo "Job $SLURM_JOB_ID running on $(hostname)"
nvidia-smi

python train.py --config configs/baseline.yaml --seed 0 \
    --output_dir "runs/$SLURM_JOB_ID"

Submit it with mkdir -p logs && sbatch single_gpu.sbatch. The mkdir matters: SLURM opens the output file before your script runs, so if the logs/ folder does not exist, the job can fail immediately with no output and no explanation.

2. A job array for several seeds

Reviewers want results averaged over multiple random seeds. Rather than submitting five jobs by hand, use a job array. SLURM runs the same script several times, each with a different SLURM_ARRAY_TASK_ID.

#SBATCH --array=0-4%2                # tasks 0..4, at most 2 running at once
#SBATCH --output=logs/%x_%A_%a.out   # %A = array job ID, %a = task ID

SEED=$SLURM_ARRAY_TASK_ID

python train.py --config configs/baseline.yaml --seed "$SEED" \
    --output_dir "runs/seed_$SEED"
Add these lines to the header and body of the script above. The %2 limits concurrency so you don't hog the queue.

3. Multi-GPU on one node with torchrun

For PyTorch DistributedDataParallel on a single machine, request all the GPUs on one node and let torchrun start one worker per GPU. Your training script must be written for distributed training (for example, it reads LOCAL_RANK and wraps the model in DistributedDataParallel).

#SBATCH --nodes=1
#SBATCH --ntasks-per-node=1          # one launcher; torchrun starts one worker per GPU
#SBATCH --gres=gpu:4
#SBATCH --cpus-per-task=32
#SBATCH --mem=128G

GPUS_PER_NODE=4                      # keep equal to --gres above

torchrun --standalone --nnodes=1 --nproc_per_node="$GPUS_PER_NODE" \
    train_ddp.py --config configs/baseline.yaml --output_dir "runs/$SLURM_JOB_ID"
Multi-node training needs extra rendezvous settings; get single-node working first.

Habits that save the most time

  1. Test small first. Run a couple of epochs on a subset in an interactive session (srun --pty bash with a GPU request) before submitting a 12-hour job.
  2. Estimate `--time` sensibly. Too short and your job is killed; far too long and it waits longer in the queue. Time a short run and request about 30% more.
  3. Checkpoint regularly and support resuming, so a killed job does not lose everything.
  4. Give every run its own output folder, named after the job ID or seed, so runs never overwrite each other.
  5. Watch GPU utilisation. If nvidia-smi shows the GPU sitting near 20 percent, data loading is your bottleneck. Raise --cpus-per-task and your dataloader's num_workers.
  6. Never run heavy work on the login node. It is shared by everyone. Use a job or an interactive session.

Common errors

SymptomLikely cause
Job fails instantly with no logThe logs/ folder did not exist when you submitted.
Invalid account or partitionWrong --partition, or a required --account is missing.
CUDA out of memoryBatch size too large for the GPU. Reduce it, or use gradient accumulation.
“Killed” or an out-of-memory kill in the logRan out of system RAM. Raise --mem or load less data into memory.
Job pending for a long timeNormal on busy clusters. Smaller or shorter requests usually start sooner.

Free download

SLURM Starter Scripts for ML Research

Three job scripts (single GPU, job array, multi-GPU) and a cheat sheet for university HPC clusters.

Want guidance, not just a guide?

HPC & Cloud GPUs for ML Research

SSH, SLURM, and cloud GPUs — run experiments at scale.

3 Hours live online · Second Wednesday of every month · $249 USD

View the course