Summary
The query relay added for allocation state sends a key the local daemon cannot
answer to the DVM master. That is the right destination for allocation
state, because the master is the only process that holds it. It is the wrong
destination for a whole class of keys whose answer is a property of a
particular node or a particular process — those have to reach the daemon
on that node, or the daemon hosting that proc.
Nothing implements that second relay today. Recording it so it is not
rediscovered the hard way when someone fills in an arm that needs it.
Background
_query() in src/prted/pmix/pmix_server_queries.c is answered by whichever
daemon the client is connected to. Three kinds of key live in it:
- Answerable anywhere — the nidmap and the launch message replicate what
they need to every daemon. Node names, aliases, nodeids, daemon vpids, job
and proc placement. PMIX_QUERY_NAMESPACES, PMIX_JOB_SIZE,
PMIX_QUERY_PROC_TABLE, PMIX_QUERY_RESOLVE_PEERS,
PMIX_QUERY_RESOLVE_NODE, qualified PMIX_SERVER_URI.
- Answerable only by the master — slot counts, node state, the session
table. A daemon holds none of it; the nidmap ships node identity and
never slots, slots_max, slots_inuse or node state, and
prte_sessions on a prted holds the default session and nothing more.
These now relay to the master over PRTE_RML_TAG_QUERY, gated by
prte_get_allocated_nodes() / prte_get_allocation_session() /
prte_get_allocation_sessions().
- Answerable only by one specific daemon — the subject of this issue.
Category 3 already has members that are correct today only because they are
never relayed: PMIX_HWLOC_XML_V1 / _V2 export
prte_hwloc_topology, an unqualified PMIX_SERVER_URI returns this daemon's
own URI, and PMIX_QUERY_LOCAL_PROC_TABLE means the procs this daemon is
hosting. All four mean "about me", and the per-key deferral deliberately
leaves them local so a mixed query does not answer them about the master.
The gap
The problem is the qualified form of a category-3 key: a client asks about
some other node or some other proc. There is no mechanism for that.
The two clearest future cases are the resource-usage arms, which exist and are
empty (pmix_server_queries.c, and docs/todo.rst):
PMIX_QUERY_NODE_RESOURCE_USAGE (pmix.qry.nres) — CPU/memory consumption
of a named node. Only the daemon on that node can sample it.
PMIX_QUERY_PROC_RESOURCE_USAGE (pmix.qry.pres) — the same for a named
process. Only the daemon that forked it can sample it.
Neither can be satisfied by the local daemon, and relaying them to the master
would be just as wrong: the master has no more idea what node4 is consuming
than node2 does. They have to go to the owning daemon.
PMIX_HWLOC_XML_V1 / _V2 are the same shape if a hostname or nodeid
qualifier is ever honored — today they ignore one and export the local
topology, which is a wrong answer to a question that named a different node.
Why this is easy to get wrong
The failure is silent in the same way the slot-count one was. An arm that
reads local state to answer a question about a remote node does not fault —
it returns this node's numbers, formatted correctly, with
PMIX_SUCCESS. The client cannot tell. Whoever fills in a resource-usage arm
will naturally reach for the local sampler, and it will look like it works
whenever the test happens to ask about the node it is running on.
What is needed
A second relay destination, selected the same way the first one is — by what
the answer is a property of, not by a list of keys:
- resolve the qualifier (
PMIX_HOSTNAME / PMIX_NODEID for a node,
PMIX_PROCID for a proc) to the owning daemon. A daemon can already do this
locally: the node pool gives node->daemon for a node, and
daemons->procs[pptr->parent] gives the hosting daemon for a proc, both
populated by the nidmap and the launch message.
- if that daemon is us, answer locally as now;
- otherwise defer the key to that daemon rather than to the master.
The transport can be PRTE_RML_TAG_QUERY unchanged — pmix_server_query_request()
is not master-specific in anything but its receive registration, which is
currently inside the if (PRTE_PROC_IS_MASTER) block in pmix_server_init().
Registering it on every daemon and choosing the destination rank per key
covers both cases with one mechanism. The per-key deferral, the tracker, and
the merge of local and remote results all already work this way.
Worth noting the ordering constraint: a key deferred to a peer daemon and a
key deferred to the master can appear in the same PMIx_Query_info, so the
completion accounting has to wait for more than one reply. Today it waits for
exactly one.
Not urgent
Nothing currently reaches this gap: the category-3 keys that exist are all
unqualified, and the two that would need it are empty arms. This is a note for
whoever implements them.
Summary
The query relay added for allocation state sends a key the local daemon cannot
answer to the DVM master. That is the right destination for allocation
state, because the master is the only process that holds it. It is the wrong
destination for a whole class of keys whose answer is a property of a
particular node or a particular process — those have to reach the daemon
on that node, or the daemon hosting that proc.
Nothing implements that second relay today. Recording it so it is not
rediscovered the hard way when someone fills in an arm that needs it.
Background
_query()insrc/prted/pmix/pmix_server_queries.cis answered by whicheverdaemon the client is connected to. Three kinds of key live in it:
they need to every daemon. Node names, aliases, nodeids, daemon vpids, job
and proc placement.
PMIX_QUERY_NAMESPACES,PMIX_JOB_SIZE,PMIX_QUERY_PROC_TABLE,PMIX_QUERY_RESOLVE_PEERS,PMIX_QUERY_RESOLVE_NODE, qualifiedPMIX_SERVER_URI.table. A daemon holds none of it; the nidmap ships node identity and
never
slots,slots_max,slots_inuseor nodestate, andprte_sessionson a prted holds the default session and nothing more.These now relay to the master over
PRTE_RML_TAG_QUERY, gated byprte_get_allocated_nodes()/prte_get_allocation_session()/prte_get_allocation_sessions().Category 3 already has members that are correct today only because they are
never relayed:
PMIX_HWLOC_XML_V1/_V2exportprte_hwloc_topology, an unqualifiedPMIX_SERVER_URIreturns this daemon'sown URI, and
PMIX_QUERY_LOCAL_PROC_TABLEmeans the procs this daemon ishosting. All four mean "about me", and the per-key deferral deliberately
leaves them local so a mixed query does not answer them about the master.
The gap
The problem is the qualified form of a category-3 key: a client asks about
some other node or some other proc. There is no mechanism for that.
The two clearest future cases are the resource-usage arms, which exist and are
empty (
pmix_server_queries.c, anddocs/todo.rst):PMIX_QUERY_NODE_RESOURCE_USAGE(pmix.qry.nres) — CPU/memory consumptionof a named node. Only the daemon on that node can sample it.
PMIX_QUERY_PROC_RESOURCE_USAGE(pmix.qry.pres) — the same for a namedprocess. Only the daemon that forked it can sample it.
Neither can be satisfied by the local daemon, and relaying them to the master
would be just as wrong: the master has no more idea what node4 is consuming
than node2 does. They have to go to the owning daemon.
PMIX_HWLOC_XML_V1/_V2are the same shape if a hostname or nodeidqualifier is ever honored — today they ignore one and export the local
topology, which is a wrong answer to a question that named a different node.
Why this is easy to get wrong
The failure is silent in the same way the slot-count one was. An arm that
reads local state to answer a question about a remote node does not fault —
it returns this node's numbers, formatted correctly, with
PMIX_SUCCESS. The client cannot tell. Whoever fills in a resource-usage armwill naturally reach for the local sampler, and it will look like it works
whenever the test happens to ask about the node it is running on.
What is needed
A second relay destination, selected the same way the first one is — by what
the answer is a property of, not by a list of keys:
PMIX_HOSTNAME/PMIX_NODEIDfor a node,PMIX_PROCIDfor a proc) to the owning daemon. A daemon can already do thislocally: the node pool gives
node->daemonfor a node, anddaemons->procs[pptr->parent]gives the hosting daemon for a proc, bothpopulated by the nidmap and the launch message.
The transport can be
PRTE_RML_TAG_QUERYunchanged —pmix_server_query_request()is not master-specific in anything but its receive registration, which is
currently inside the
if (PRTE_PROC_IS_MASTER)block inpmix_server_init().Registering it on every daemon and choosing the destination rank per key
covers both cases with one mechanism. The per-key deferral, the tracker, and
the merge of local and remote results all already work this way.
Worth noting the ordering constraint: a key deferred to a peer daemon and a
key deferred to the master can appear in the same
PMIx_Query_info, so thecompletion accounting has to wait for more than one reply. Today it waits for
exactly one.
Not urgent
Nothing currently reaches this gap: the category-3 keys that exist are all
unqualified, and the two that would need it are empty arms. This is a note for
whoever implements them.