Using CUDA driver 580+ causes a performance regression when opal_cuda_support is set to 1, for both Host and Device memory. This has been observed on a Cray GH200 system with OpenMPI 5.0.10 and CUDA driver 580.65.06. Here are OSU BIBW results with cuda support turned off (D2D transfers handled by libfabric and the CXI provider, so benchmark works):
export FI_SHM_USE_XPMEM=1
export FI_PROVIDER=cxi
export OMPI_MCA_opal_common_ofi_provider_include=cxi
export OMPI_MCA_mtl_ofi_av=table
export PRTE_MCA_ras_base_launch_orted_on_hn=1
export OMPI_MCA_pml=cm
export OMPI_MCA_mtl=ofi
mpirun --mca opal_cuda_support 0 $OSU_HOME/mpi/pt2pt/osu_bibw -d cuda -b multiple D D
# OSU MPI-CUDA Bi-Directional Bandwidth Test v7.5.2
# Datatype: MPI_CHAR.
# Size Bandwidth (MB/s)
1 1.81
2 4.41
4 9.03
8 7.27
16 36.42
32 73.13
64 143.76
128 283.76
256 567.56
512 1174.61
1024 2318.62
2048 2060.47
4096 5778.80
8192 15388.42
16384 14109.38
32768 33479.20
65536 36490.61
131072 37478.29
262144 44447.93
524288 32736.82
1048576 42223.74
2097152 44032.52
4194304 46440.23
and with cuda support turned on:
mpirun --mca opal_cuda_support 1 $OSU_HOME/mpi/pt2pt/osu_bibw -d cuda -b multiple D D
# OSU MPI-CUDA Bi-Directional Bandwidth Test v7.5.2
# Datatype: MPI_CHAR.
# Size Bandwidth (MB/s)
1 0.29
2 0.95
4 1.82
8 2.36
16 7.58
32 15.18
64 18.95
128 61.51
256 122.86
512 190.08
1024 489.67
2048 742.58
4096 1971.91
8192 3856.40
16384 7648.14
32768 15236.80
65536 25348.43
131072 16056.73
262144 38744.02
524288 35436.30
1048576 37642.22
2097152 42957.92
4194304 44777.00
Notice large performance penalty especially for small and medium messages.
According to HPE and CSCS, there seems to be a large overhead when checking the pointers with cuDeviceGetAttribute. Similar performance degradation is observed with Cray MPI. In that case the problem can be worked around by dropping cuDeviceGetAttribute in favor of cuDeviceGetAttributes (see an explanation by CSCS: https://docs.cscs.ch/software/communication/cray-mpich/#slow-host-buffer-communication-with-nvidia-driver-version-590-and-later).
Notice that the same workaround does not work in OpenMPI: most of the time cuDeviceGetAttributes is called already on this system. Removing the last remaining call to cuDeviceGetAttribute in function accelerator_cuda_check_mpool does not fix the problem. What does 'fix' the performance is removing support for CUDA_VMM:
--- openmpi-5.0.10.orig/opal/mca/accelerator/cuda/accelerator_cuda.c 2026-02-23 21:21:10.000000000 +0100
+++ openmpi-5.0.10/opal/mca/accelerator/cuda/accelerator_cuda.c 2026-08-20 13:00:00.000000000 +0200
@@ -15,7 +15,7 @@
*/
#include "opal_config.h"
-
+#undef OPAL_CUDA_VMM_SUPPORT
#include <cuda.h>
#include "accelerator_cuda.h"
After doing this performance is back to normal:
mpirun --mca opal_cuda_support 1 $OSU_HOME/mpi/pt2pt/osu_bibw -d cuda -b multiple D D
# OSU MPI-CUDA Bi-Directional Bandwidth Test v7.5.2
# Datatype: MPI_CHAR.
# Size Bandwidth (MB/s)
1 1.69
2 3.53
4 7.15
8 4.81
16 28.70
32 57.51
64 100.62
128 227.92
256 470.38
512 937.63
1024 1859.79
2048 1692.79
4096 7538.69
8192 12201.98
16384 22641.49
32768 31860.60
65536 22120.21
131072 26377.77
262144 44012.68
524288 33077.65
1048576 41367.42
2097152 43505.83
4194304 46407.97
which at least gives an indication of where the problem is coming from. Fiddling around with the code showed that the below calls are responsible for the slowness:
is_vmm = accelerator_cuda_check_vmm(dbuf, &vmm_mem_type, &vmm_dev_id);
is_mpool_ptr = accelerator_cuda_check_mpool(dbuf, &mpool_mem_type, &mpool_dev_id);
Since the problem is rooted inside CUDA and the driver, it's likely it will show on other systems, not only the studied Cray GH200 architecture.
Using CUDA driver 580+ causes a performance regression when
opal_cuda_supportis set to 1, for both Host and Device memory. This has been observed on a Cray GH200 system with OpenMPI 5.0.10 and CUDA driver 580.65.06. Here are OSU BIBW results with cuda support turned off (D2D transfers handled by libfabric and the CXI provider, so benchmark works):and with cuda support turned on:
Notice large performance penalty especially for small and medium messages.
According to HPE and CSCS, there seems to be a large overhead when checking the pointers with
cuDeviceGetAttribute. Similar performance degradation is observed with Cray MPI. In that case the problem can be worked around by droppingcuDeviceGetAttributein favor ofcuDeviceGetAttributes(see an explanation by CSCS: https://docs.cscs.ch/software/communication/cray-mpich/#slow-host-buffer-communication-with-nvidia-driver-version-590-and-later).Notice that the same workaround does not work in OpenMPI: most of the time
cuDeviceGetAttributesis called already on this system. Removing the last remaining call tocuDeviceGetAttributein functionaccelerator_cuda_check_mpooldoes not fix the problem. What does 'fix' the performance is removing support forCUDA_VMM:After doing this performance is back to normal:
which at least gives an indication of where the problem is coming from. Fiddling around with the code showed that the below calls are responsible for the slowness:
Since the problem is rooted inside CUDA and the driver, it's likely it will show on other systems, not only the studied Cray GH200 architecture.