Operator guide for getting NVFlow running on a Slurm cluster: install the client, stage containers and models, configure your cluster, and verify. NVFlow's containers are self-sufficient — all dependencies are pre-installed, so once images and models are staged, the pipeline runs fully offline.
This guide sets up the cluster side. How you run the
nflowclient — local install or the airgappednvflow-clientcontainer, and whether it submits directly or over an SSH tunnel — is summarized in Choose your client setup just below.
Building artifacts? Producing the container images is a maintainer task and lives under
docs/maintainers/, not on this page.
nflow only submits Slurm jobs — the heavy work runs on the cluster. Two
independent choices decide how you run it: how you provision the client, and
how it reaches Slurm.
You run nflow on… |
Provision the client | Reach Slurm | Guide |
|---|---|---|---|
| Cluster login/dev node (internet) | uv sync |
direct | README |
| Laptop / dev box (internet) | uv sync |
SSH tunnel | remote-launch |
| Anywhere, airgapped / no install | nvflow-client container |
direct or SSH tunnel | remote-launch |
Steps 1–6 below are the cluster side (stage images/models, write
my_cluster.yaml, verify) and apply to every row above.
Six steps, top to bottom. Each step below opens with a Goal and ends with a ✅ Done when check so you always know where you are.
| Step | What it does | Who needs it |
|---|---|---|
| 1. Prerequisites | Confirm cluster access + required tools | Everyone |
| 2. Setup Containers | Stage the five .sqsh images on the cluster |
Everyone |
| 3. Download Models | Pre-stage the HF models your workflows use | Everyone |
| 4. GRPO Prerequisites | SEC cache prefetch | GRPO only — else skip |
| 5. Configure Your Cluster | Write cluster_configs/my_cluster.yaml |
Everyone |
| 6. Verify Installation | Sanity-check the whole setup | Everyone |
Shortcut: if a maintainer already staged the
.sqshimages and models for you, you only need Steps 1, 5, and 6.
Step 1 of 6 · Goal: confirm you can reach the cluster and have the tools the setup needs.
Note: This guide assumes you've already installed the client (see the README: install
uv, clone the repo, runuv sync).🔌 Airgapped / no internet on the install host? Skip the local install and drive
nflowfrom the prebuiltnvflow-clientcontainer (CLI + venv baked in, nouv sync, no client internet) — see docs/remote-launch.md. You still stage the worker images and models on the cluster (Steps 2–3 below); only the install differs.
Required:
- Slurm cluster access with SSH keys (or run directly from a login node)
- enroot - on cluster nodes (
enroot version)
Only needed for the parallel container conversion script:
- yq - YAML parser (install guide)
- curl - for downloading configs
Note: If you already have
.sqshcontainer images staged on the cluster, skip to Configure Your Cluster.
Install yq (only needed for the parallel conversion script)
# Check if installed
yq --version
# If not installed:
# macOS
brew install yq
# Linux (auto-detects architecture)
# Supported platforms: linux_amd64, linux_arm64, linux_arm, linux_386, etc.
mkdir -p $HOME/bin
ARCH=$(uname -m); case "$ARCH" in x86_64) ARCH=amd64 ;; aarch64) ARCH=arm64 ;; armv7l) ARCH=arm ;; i686) ARCH=386 ;; esac
wget "https://github.com/mikefarah/yq/releases/latest/download/yq_linux_${ARCH}" -O $HOME/bin/yq
chmod +x $HOME/bin/yq
# Add to PATH (if $HOME/bin not already in PATH)
echo 'export PATH="$HOME/bin:$PATH"' >> $HOME/.bashrc
source $HOME/.bashrccurl & enroot:
curl --version # Usually pre-installed
enroot version # Run on cluster nodeGet your cluster info:
- Slurm account:
sacctmgr show associations user=$USER(look for the Account column) - Available partitions:
sinfo - Slurm version:
scontrol show config | grep SLURM_VERSION(25.x needs the enroot Ray template fix - see Troubleshooting) - Storage paths for data/models/containers
✅ Done when: enroot version works on a cluster node and you know your Slurm account, a partition, and your storage paths.
Step 2 of 6 · Goal: have the five .sqsh container images staged on your cluster, with their paths in hand.
NVFlow runs its cluster jobs inside five .sqsh container images. As an operator you only need the .sqsh files staged on your cluster and their paths recorded in your cluster config.
- Already have
.sqshfiles staged (by a maintainer or a previous setup)? Note their paths and skip to Configure Your Cluster. - Need to build / convert them yourself? See docs/maintainers/containers.md — build host requirements,
docker build, push/save, andenroot importto.sqsh.
Your cluster config references the images by fixed keys — nemo-rl, nemo-skills, vllm, vllm-grpo, and sglang. The build guide's "Update Container Config" step explains how to set them; Configure Your Cluster ties them into your run config.
nemo-gym(needed for GRPO & DG-SDG): the Gym-only stages — GRPOprepare_data/prefetch_cacheand the DG-SDG gym stages — run in a dedicated CPU-onlynemo-gymimage (dockerfiles/Dockerfile.nemo-gym). Stage this sixth image if you run GRPO or DG-SDG; SFT-only and eval-only runs don't need it.
✅ Done when: five .sqsh files exist on the cluster and you have their absolute paths for the config.
Step 3 of 6 · Goal: pre-download the models your chosen workflows need to the cluster's HF models directory.
⚠️ Important: Pre-download models to your cluster storage before running workflows. The runtime setsHF_HUB_OFFLINE=1, so any model not already on disk will fail at job time.Why this matters:
- Avoids wasting expensive GPU time on downloads
- Prevents race conditions when multiple jobs start simultaneously
- Large models (10-100+ GB) can take hours to download
- Network failures during jobs cause workflow failures
Note: hf CLI is included with nemo-skills (via huggingface-hub). Some models are gated and require authentication -- export your HuggingFace token before downloading:
export HF_TOKEN=<YOUR_HF_TOKEN>Download models to your cluster's HuggingFace models directory. The examples below show the models used by the finance recipe workflows -- download only the ones you need:
# GRPO policy model (Qwen3-30B-A3B, MoE — used in grpo/qwen3_30b_a3b.yaml)
uv run hf download Qwen/Qwen3-30B-A3B \
--local-dir /path/to/models/hf_models/Qwen/Qwen3-30B-A3B
# GRPO / eval judge model (GPT-OSS-120B — used for rollout judging and eval)
uv run hf download openai/gpt-oss-120b \
--local-dir /path/to/models/hf_models/openai/gpt-oss-120bFor the quick-start demo (see quick-start.md), download these additional models:
# Demo policy model (Qwen3-4B — used in sft/qwen3_4b.yaml and grpo/qwen3_4b.yaml)
uv run hf download Qwen/Qwen3-4B \
--local-dir /path/to/models/hf_models/Qwen/Qwen3-4B
# Demo SDG generation + eval baseline (GPT-OSS-20B)
uv run hf download openai/gpt-oss-20b \
--local-dir /path/to/models/hf_models/openai/gpt-oss-20b
# Eval baseline (Gemma 3 4B IT)
uv run hf download google/gemma-3-4b-it \
--local-dir /path/to/models/hf_models/google/gemma-3-4b-itIf you enable the finance GroundingVerifier stage, pre-stage its two public models at the pinned revisions used by the workflow:
# Claim-to-evidence routing model
uv run hf download sentence-transformers/all-MiniLM-L6-v2 \
config.json model.safetensors special_tokens_map.json \
tokenizer.json tokenizer_config.json vocab.txt \
--revision 1110a243fdf4706b3f48f1d95db1a4f5529b4d41 \
--local-dir /path/to/models/hf_models/sentence-transformers/all-MiniLM-L6-v2
# Natural-language-inference model
uv run hf download MoritzLaurer/DeBERTa-v3-base-mnli-fever-anli \
added_tokens.json config.json model.safetensors special_tokens_map.json \
spm.model tokenizer.json tokenizer_config.json \
--revision 6f5cf0a2b59cabb106aca4c287eed12e357e90eb \
--local-dir /path/to/models/hf_models/MoritzLaurer/DeBERTa-v3-base-mnli-fever-anliStorage location: Models should go in your mounted HuggingFace models directory (see cluster config mounts section).
Ensure your cluster config has the models directory mounted (cluster config creation is explained in the Configure Your Cluster section below):
mounts:
- /cluster/path/to/models/hf_models:/hf_models # Maps to /hf_models inside containersReference models using the container mount path (/hf_models):
stage_kwargs:
model: /hf_models/Qwen/Qwen3-4B # Path inside container
server_type: sglangWhich models does each workflow need?
| Model | Demo SDG | Demo SFT | Demo GRPO | Demo Eval | Production GRPO | GroundingVerifier |
|---|---|---|---|---|---|---|
Qwen/Qwen3-4B |
✓ | ✓ | ✓ | |||
openai/gpt-oss-20b |
✓ | ✓ | ||||
google/gemma-3-4b-it |
✓ | |||||
openai/gpt-oss-120b |
✓ | ✓ | ||||
Qwen/Qwen3-30B-A3B |
✓ | |||||
sentence-transformers/all-MiniLM-L6-v2 |
✓ | |||||
MoritzLaurer/DeBERTa-v3-base-mnli-fever-anli |
✓ |
Tip: Download commonly used models once and reuse across all workflows.
A handful of stages legitimately need internet on first run to pull benchmark / seed datasets from HuggingFace or SEC EDGAR. Run them on a connected node (or off-cluster) and ship the resulting artifacts to the cluster - they're reused by every subsequent run.
| Stage | Pulls from | Why |
|---|---|---|
workflow-2 download_sec_filings |
SEC EDGAR | Filings aren't on HF |
workflow-3 step-0 create_seed_data |
HF nogabenyoash/SecQue |
Seed dataset |
workflow-1 step-0 prepare_data (eval) |
HF secque, financebench |
Benchmark data |
workflow-5 step-4 prepare_data (GRPO) |
HF | Only if should_download: true |
For these stages, temporarily clear the three HF offline flags (HF_HUB_OFFLINE, HF_DATASETS_OFFLINE, TRANSFORMERS_OFFLINE) in your cluster config. UV_OFFLINE is unrelated — leave it at its default (unset); none of these stages invoke uv.
Note:
huggingface_hubinterpretsTRANSFORMERS_OFFLINE=1asHF_HUB_OFFLINE=1, so all three need to be off (or unset) for HF dataset pulls to succeed.
✅ Done when: the models for your workflow are on disk under your mounted hf_models directory.
Step 4 of 6 · GRPO only — skip this entire step if you're not running GRPO.
There are no sources to clone. The nvflow-nemo-rl trainer image bakes NeMo-RL together with one prebuilt NeMo-Gym venv per component, so GRPO training resolves no packages at job runtime and needs no bind-mount. The Gym-only stages (collect_rollouts, compute_rewards, prefetch_cache, prepare_data) run on the equally self-contained nvflow-nemo-gym image.
How that image is built, and its internals, are covered in docs/development/nemo-rl-gym.md.
If using the finance_sec_search NeMo-Gym environment, you must prefetch the SEC filings cache to a shared mounted path. The default ~/.cache does not work inside Slurm containers.
The GRPO workflow includes a dedicated prefetch_cache stage that runs on a connected node and populates the cache under your workflow-5-grpo/ output directory. See docs/recipes/finance/workflows/06-grpo.md for the full prefetch flow.
✅ Done when: (GRPO users) the SEC filings cache is prefetched if you use finance_sec_search. Everyone else: nothing to do — move on.
Step 5 of 6 · Goal: create and fill in cluster_configs/my_cluster.yaml.
# Copy template
cp cluster_configs/template-slurm.yaml cluster_configs/my_cluster.yamlEdit cluster_configs/my_cluster.yaml and replace all <PLACEHOLDER> values:
- SSH settings - Your cluster login node, username, SSH key path (ONLY for remote job submission from local machine)
- Slurm account/partition - Run
sacctmgr show associations user=$USERandsinfo - Container paths - Copy from
outputs/logs/slurm-containers-<jobid>.outafter running setup_containers.sh - Mount points - Map your cluster paths to container paths (at minimum
<hf_models>:/hf_modelsand a writable data dir<workspace_data>:/workspacefor outputs + caches). Recipe code and assets ship via the nemo-run packaged snapshot (/nemo_run/code), so the repo is not mounted — see Mount Points. - Environment variables - Set
HF_HOMEto a path visible inside the container (see env_vars docs) and any API keys
template-slurm.yaml ships with the offline flags pre-populated - leave them on:
env_vars:
# --- AIR-GAPPED ENFORCEMENT (recommended) ---
- HF_HUB_OFFLINE=1
- HF_DATASETS_OFFLINE=1
- TRANSFORMERS_OFFLINE=1
# UV_OFFLINE: left UNSET (global flag). The images bake every venv they need,
# so no stage resolves packages at runtime either way.
# - UV_OFFLINE=true
# Pre-baked tiktoken / openai_harmony cache (set as ENV in vllm/vllm-grpo
# already; setting here applies them uniformly to nemo-skills and nemo-rl)
- TIKTOKEN_CACHE_DIR=/opt/tiktoken_cache
- TIKTOKEN_RS_CACHE_DIR=/opt/tiktoken_cache
- TIKTOKEN_ENCODINGS_BASE=/opt/tiktoken_cacheFor the one-time connected-node stages listed in Download Models above, comment out the three HF_*_OFFLINE flags just for that submission, then re-enable.
Note: The template includes detailed comments for each section. Your personal config (
my_cluster.yaml) is gitignored to protect secrets.📖 For detailed documentation of all configuration fields, see the Cluster Configuration Guide.
✅ Done when: my_cluster.yaml has no remaining <PLACEHOLDER> values and keeps the air-gap enforcement block.
Step 6 of 6 · Goal: confirm the whole setup before running a real workflow.
# 1. Test NeMo-Skills import
uv run python -c "from nemo_skills.pipeline.cli import generate; print('✅ OK')"
# 2. Check containers exist
ls -lh <PATH_TO_CONTAINERS>/*.sqsh
# 3. Test SSH to cluster (only if submitting from a local machine)
ssh -i <PATH_TO_SSH_KEY> <YOUR_USERNAME>@<YOUR_CLUSTER_LOGIN_NODE> "echo '✅ SSH OK'"
# 4. Test cluster config loads
uv run python -c "from omegaconf import OmegaConf; OmegaConf.load('cluster_configs/my_cluster.yaml'); print('✅ Config OK')"
# 5. List available stages
uv run nflow list-stages✅ Done when: all five checks above pass. You're ready to run a workflow — see Next Steps.
# UV can install it for you
uv python install 3.12
uv syncuv sync --reinstallyq is only needed for the parallel .sqsh conversion script. Install it per the Prerequisites → Install yq block above.
# On cluster
module load enroot # if available
# Or contact your cluster adminBuilding images, docker login nvcr.io, and enroot import quirks (filename colon, # separator for nvcr.io) are covered in docs/maintainers/containers.md.
chmod 600 <PATH_TO_SSH_KEY>
ssh -i <PATH_TO_SSH_KEY> <YOUR_USERNAME>@<YOUR_CLUSTER_LOGIN_NODE>Use absolute paths in cluster config and verify each .sqsh exists (ls -l <PATH_TO_CONTAINERS>/*.sqsh). Re-run container setup if needed.
HF_HOME (and every other path-valued env var) must resolve inside the container -- use a mount destination (e.g. /workspace/cache/huggingface) or a transparently-mounted host path (e.g. /shared/... when - /shared:/shared is in mounts). $HOME and ~/.cache will not resolve.
For symptoms specific to the self-sufficient runtime - GRPO installation_command failing with "No such file or directory", OfflineModeIsEnabled, uv trying to reach PyPI, tiktoken / openai_harmony failing offline - see the Offline Runtime section in the finance troubleshooting guide.
- Verify account:
sacctmgr show associations user=$USER - Check partition:
sinfo -p <YOUR_GPU_PARTITION> - Look at job logs in
ssh_tunnel.job_dir
If training jobs hang after "Starting Ray cluster" (or you see execve(): bad interpreter: No such file or directory), your Slurm likely needs the enroot Ray template: set ray_template: "ray_enroot.sub.j2" in my_cluster.yaml. Full symptoms, cause, and SLURM-version notes are in cluster-configuration.md → Ray Cluster Configuration.
✅ Cluster setup complete!
You can now run workflows on your cluster. Set the config directory:
export NEMO_SKILLS_CONFIG_DIR=/path/to/nvflow/cluster_configsThen head back to the README.md Quick Start section to run your first workflow.
- Build the containers:
docs/maintainers/containers.md - NeMo-RL / NeMo-Gym trainer image & Gym venvs:
docs/development/nemo-rl-gym.md - NVFlow Dockerfiles:
dockerfiles/README.md - NVFlow Self-Sufficient Build / Deploy Guide:
dockerfiles/docker_instructions.md - Cluster Configuration Guide:
docs/cluster-configuration.md - NeMo-Skills: https://github.com/NVIDIA-NeMo/Skills
- NeMo-RL: https://github.com/NVIDIA-NeMo/RL
- NeMo-Gym: https://github.com/NVIDIA-NeMo/Gym
- Official Container Config (NeMo-Skills): https://github.com/NVIDIA-NeMo/Skills/blob/main/cluster_configs/example-slurm.yaml
- Slurm Docs: https://slurm.schedmd.com/
- Enroot: https://github.com/NVIDIA/enroot
Need help? Check the Cluster Configuration Guide or ask your team/cluster admin.