Kalam
Kalam

ClusterScheduler

Sends your job to whichever cluster will start it soonest.

You might have time on a few clusters.

Picking one to send your job to means knowing things about each: how busy it is right now, how long jobs like yours usually wait, whether it even has the GPUs you need. Wouldn't it be nice not to have to think about any of that?

You don't have to. rclust reads the #SBATCH lines at the top of your job script, like these:

#!/bin/bash#SBATCH --gpus-per-node=2    # two GPUs#SBATCH --mem=64G#SBATCH --time=12:00:00      # for up to 12 hourspython train.py

It skips the clusters that can't run it, asks the rest when a two-GPU, twelve-hour job would start, and submits it to the soonest.

Some clusters require two-factor authentication for each new SSH connection. rclust reuses authenticated connections, so you complete authentication once per cluster and subsequent checks reuse that connection until it expires.

Choosing a cluster

historyWhere jobs like yours have started soonest over the past week, unless a cluster can start it right now. The default, once you've run rclust learn.
earliestThe soonest start by Slurm's own estimate.
balancedFavours clusters where your fairshare is high and the queue is light.

Install

uv tool install git+https://github.com/narunraman/rclustrclust config        # add your clustersrclust discover      # lists each one's GPUs, to paste into the config

Settings live in ~/.config/rclust/config.yaml, or wherever --config or $RCLUST_CONFIG points.

Use

rclust connect --persist 8h                # authenticate once per clusterrclust learn                              # how long jobs have been waitingrclust suggest --gpus 2 --time 12:00:00   # where would it go?rclust submit train.sh                    # send it there