ras/slurm's modify() maps PMIX_ALLOC_EXTEND onto an implementation that
is, by PMIx's own definition, PMIX_ALLOC_NEW.
From pmix_common.h:
#define PMIX_ALLOC_NEW 1 // new allocation is being requested. The resulting allocation will be
// disjoint (i.e., not connected in a job sense) from the requesting allocation
#define PMIX_ALLOC_EXTEND 2 // extend the existing allocation, either in time or as additional resources
What the component does today
prte_ras_slurm_serve_extend_req() → prte_ras_slurm_launch_expander_job()
(src/mca/ras/slurm/ras_slurm_modify_extend.c) runs
salloc --no-shell --exclusive --job-name=prrte --nodes=<N> \
[--account=…] [--partition=…] [--qos=…] [--chdir=…] \
[--mem-per-cpu=… | --mem=…] [--time=…] [--threads-per-core=…]
copying the optional fields from the parent job's scontrol show job --json.
The result is a second, independent Slurm job. Verified on a live Slurm
24.11.6 (contrib/slurmswarm), inside a 2-node allocation, after
elastic extend 1:
JOBID NAME NODES NODELIST STATE
5 prrte 1 node3 RUNNING
4 prte-dvm 2 node[1-2] RUNNING
job 5: Dependency=(null) NumNodes=1
job 4: Dependency=(null) NumNodes=2
Two jobs, no dependency between them, nothing merged. That is precisely
PMIX_ALLOC_NEW's contract — "disjoint, not connected in a job sense" — and
it is wired to PMIX_ALLOC_EXTEND. The internal name launch_expander_job
suggests Slurm's job-expansion feature was the intent, but the argv carries
no --dependency=expand: and no merge step ever runs.
This is not cosmetic. The two directives mean different things to the caller
and to the site:
- Each grow is a separate accounting record, against whatever per-user job
and QoS limits the site enforces.
--time is copied from the parent's time_limit, not its remaining
time, so an expander created an hour into a 2-hour job gets a fresh
2 hours. With no dependency in either direction, it can outlive the
allocation it was grown from, and the parent can hit its wall while the
expander keeps running.
--exclusive is unconditional, however the parent was allocated.
- The DVM ends up spanning N Slurm jobs, which is why the release path has to
work out which job holds a node before it can resize it.
What a genuine PMIX_ALLOC_EXTEND looks like on Slurm
EXTEND is defined as "either in time or as additional resources", so there
are two things to attempt, and they have very different prospects. Measured
on Slurm 24.11.6 with SelectType=select/cons_tres:
Time — works, and is entirely unimplemented today.
$ scontrol update JobId=$P TimeLimit=30
before: TimeLimit=00:10:00
after: TimeLimit=00:30:00
This is the clearest reading of "extend the existing allocation ... in time",
it is a one-command implementation, and ras/slurm does nothing with it.
(Increases are privileged; an unprivileged request is refused by Slurm, which
is a fine answer to hand back.)
Additional resources — the documented recipe, which modern Slurm refuses.
Slurm's way to grow a running allocation is to submit a new job that expands
the original and then merge it in:
salloc --no-shell --dependency=expand:<parent_jobid> --nodes=1 …
scontrol update JobId=<new_jobid> NumNodes=0 # merges into the parent, ends the new job
Slurm then writes slurm_job_<jobid>_resize.{sh,csh} for the environment
update — the very files prte_ras_slurm_cleanup_resize_scripts() already
removes on the release path, so the tree has met this machinery from the
shrink side.
On this cluster both routes are refused outright:
$ salloc --no-shell --dependency=expand:6 --nodes=1
salloc: error: Job submit/allocate failed: Requested operation not supported on this system
$ scontrol update JobId=6 NumNodes=3
Requested operation not supported on this system for job 6
while an in-place shrink — what the release path already does — works fine:
$ scontrol update job $P ReqNodeList=node1
before: NodeList=node[1-2] NumNodes=2
after: NodeList=node1 NumNodes=1
So on a current select/cons_tres Slurm you can extend time and shrink nodes
in place, but you cannot grow nodes in place. A disjoint new allocation is
the only way to get more nodes — which is exactly why the existing expander
code is valuable, and exactly why it belongs under PMIX_ALLOC_NEW.
Suggested shape
- Move the expander-job implementation to serve
PMIX_ALLOC_NEW. It is a
correct NEW and the component does not handle NEW at all today.
- Implement
PMIX_ALLOC_EXTEND as a real extension: scontrol update JobId=<id> TimeLimit=… for the time case, and the
--dependency=expand: + NumNodes=0 merge for the resource case, with a
clean PMIX_ERR_NOT_SUPPORTED where the configuration refuses it — rather
than silently substituting a disjoint allocation.
- A caller that genuinely wants a second allocation asks for
NEW and gets
today's behavior, knowingly.
Why this surfaced
ras/slurm not handling PMIX_ALLOC_NEW is a live bug on its own. modify()
returns PMIX_ERR_NOT_SUPPORTED for it, which prte_ras_base_modify() reads
as "ask the next module" — and ras/hosts (priority 1, last) claims
NEW/EXTEND/RELEASE on the directive alone, regardless of whether it
found any hosts. So under Slurm a PMIX_ALLOC_NEW naming hostnames is served
by ras/hosts, and the named nodes are added to the DVM without the scheduler
ever being asked:
>>> PHASE 1: allocation request returned PMIX_SUCCESS
>>> ALLOC_ID prte-node1-282@1.1 <-- we minted an allocation id
squeue: 3 node[1-3] <-- Slurm allocated nothing
PRRTE then tried to start a prted on the un-allocated node, failed, and took
the DVM down. Dropping ras/hosts from the module set gives the right answer
(PMIX_ERR_NOT_SUPPORTED, DVM survives), which isolates the fall-through as
the cause. That half is being addressed separately, but a ras/slurm that
handled NEW would also close it.
ras/slurm'smodify()mapsPMIX_ALLOC_EXTENDonto an implementation thatis, by PMIx's own definition,
PMIX_ALLOC_NEW.From
pmix_common.h:What the component does today
prte_ras_slurm_serve_extend_req()→prte_ras_slurm_launch_expander_job()(
src/mca/ras/slurm/ras_slurm_modify_extend.c) runscopying the optional fields from the parent job's
scontrol show job --json.The result is a second, independent Slurm job. Verified on a live Slurm
24.11.6 (
contrib/slurmswarm), inside a 2-node allocation, afterelastic extend 1:Two jobs, no dependency between them, nothing merged. That is precisely
PMIX_ALLOC_NEW's contract — "disjoint, not connected in a job sense" — andit is wired to
PMIX_ALLOC_EXTEND. The internal namelaunch_expander_jobsuggests Slurm's job-expansion feature was the intent, but the argv carries
no
--dependency=expand:and no merge step ever runs.This is not cosmetic. The two directives mean different things to the caller
and to the site:
and QoS limits the site enforces.
--timeis copied from the parent'stime_limit, not its remainingtime, so an expander created an hour into a 2-hour job gets a fresh
2 hours. With no dependency in either direction, it can outlive the
allocation it was grown from, and the parent can hit its wall while the
expander keeps running.
--exclusiveis unconditional, however the parent was allocated.work out which job holds a node before it can resize it.
What a genuine
PMIX_ALLOC_EXTENDlooks like on SlurmEXTENDis defined as "either in time or as additional resources", so thereare two things to attempt, and they have very different prospects. Measured
on Slurm 24.11.6 with
SelectType=select/cons_tres:Time — works, and is entirely unimplemented today.
This is the clearest reading of "extend the existing allocation ... in time",
it is a one-command implementation, and
ras/slurmdoes nothing with it.(Increases are privileged; an unprivileged request is refused by Slurm, which
is a fine answer to hand back.)
Additional resources — the documented recipe, which modern Slurm refuses.
Slurm's way to grow a running allocation is to submit a new job that expands
the original and then merge it in:
Slurm then writes
slurm_job_<jobid>_resize.{sh,csh}for the environmentupdate — the very files
prte_ras_slurm_cleanup_resize_scripts()alreadyremoves on the release path, so the tree has met this machinery from the
shrink side.
On this cluster both routes are refused outright:
while an in-place shrink — what the release path already does — works fine:
So on a current
select/cons_tresSlurm you can extend time and shrink nodesin place, but you cannot grow nodes in place. A disjoint new allocation is
the only way to get more nodes — which is exactly why the existing expander
code is valuable, and exactly why it belongs under
PMIX_ALLOC_NEW.Suggested shape
PMIX_ALLOC_NEW. It is acorrect
NEWand the component does not handleNEWat all today.PMIX_ALLOC_EXTENDas a real extension:scontrol update JobId=<id> TimeLimit=…for the time case, and the--dependency=expand:+NumNodes=0merge for the resource case, with aclean
PMIX_ERR_NOT_SUPPORTEDwhere the configuration refuses it — ratherthan silently substituting a disjoint allocation.
NEWand getstoday's behavior, knowingly.
Why this surfaced
ras/slurmnot handlingPMIX_ALLOC_NEWis a live bug on its own.modify()returns
PMIX_ERR_NOT_SUPPORTEDfor it, whichprte_ras_base_modify()readsas "ask the next module" — and
ras/hosts(priority 1, last) claimsNEW/EXTEND/RELEASEon the directive alone, regardless of whether itfound any hosts. So under Slurm a
PMIX_ALLOC_NEWnaming hostnames is servedby
ras/hosts, and the named nodes are added to the DVM without the schedulerever being asked:
PRRTE then tried to start a
prtedon the un-allocated node, failed, and tookthe DVM down. Dropping
ras/hostsfrom the module set gives the right answer(
PMIX_ERR_NOT_SUPPORTED, DVM survives), which isolates the fall-through asthe cause. That half is being addressed separately, but a
ras/slurmthathandled
NEWwould also close it.