Skip to content

Performance regression with cuda drivers 580+ #14337

Description

@angainor

Using CUDA driver 580+ causes a performance regression when opal_cuda_support is set to 1, for both Host and Device memory. This has been observed on a Cray GH200 system with OpenMPI 5.0.10 and CUDA driver 580.65.06. Here are OSU BIBW results with cuda support turned off (D2D transfers handled by libfabric and the CXI provider, so benchmark works):

export FI_SHM_USE_XPMEM=1                                                                                                  
export FI_PROVIDER=cxi                                                                                                     
export OMPI_MCA_opal_common_ofi_provider_include=cxi                                                                       
export OMPI_MCA_mtl_ofi_av=table                                                                                           
export PRTE_MCA_ras_base_launch_orted_on_hn=1                                                                              
export OMPI_MCA_pml=cm                                                                                                     
export OMPI_MCA_mtl=ofi                                                                                                                                                                              

mpirun --mca opal_cuda_support 0 $OSU_HOME/mpi/pt2pt/osu_bibw -d cuda -b multiple D D

# OSU MPI-CUDA Bi-Directional Bandwidth Test v7.5.2
# Datatype: MPI_CHAR.
# Size      Bandwidth (MB/s)
1                       1.81
2                       4.41
4                       9.03
8                       7.27
16                     36.42
32                     73.13
64                    143.76
128                   283.76
256                   567.56
512                  1174.61
1024                 2318.62
2048                 2060.47
4096                 5778.80
8192                15388.42
16384               14109.38
32768               33479.20
65536               36490.61
131072              37478.29
262144              44447.93
524288              32736.82
1048576             42223.74
2097152             44032.52
4194304             46440.23

and with cuda support turned on:

mpirun --mca opal_cuda_support 1 $OSU_HOME/mpi/pt2pt/osu_bibw -d cuda -b multiple D D

# OSU MPI-CUDA Bi-Directional Bandwidth Test v7.5.2
# Datatype: MPI_CHAR.
# Size      Bandwidth (MB/s)
1                       0.29
2                       0.95
4                       1.82
8                       2.36
16                      7.58
32                     15.18
64                     18.95
128                    61.51
256                   122.86
512                   190.08
1024                  489.67
2048                  742.58
4096                 1971.91
8192                 3856.40
16384                7648.14
32768               15236.80
65536               25348.43
131072              16056.73
262144              38744.02
524288              35436.30
1048576             37642.22
2097152             42957.92
4194304             44777.00

Notice large performance penalty especially for small and medium messages.

According to HPE and CSCS, there seems to be a large overhead when checking the pointers with cuDeviceGetAttribute. Similar performance degradation is observed with Cray MPI. In that case the problem can be worked around by dropping cuDeviceGetAttribute in favor of cuDeviceGetAttributes (see an explanation by CSCS: https://docs.cscs.ch/software/communication/cray-mpich/#slow-host-buffer-communication-with-nvidia-driver-version-590-and-later).

Notice that the same workaround does not work in OpenMPI: most of the time cuDeviceGetAttributes is called already on this system. Removing the last remaining call to cuDeviceGetAttribute in function accelerator_cuda_check_mpool does not fix the problem. What does 'fix' the performance is removing support for CUDA_VMM:

--- openmpi-5.0.10.orig/opal/mca/accelerator/cuda/accelerator_cuda.c	2026-02-23 21:21:10.000000000 +0100
+++ openmpi-5.0.10/opal/mca/accelerator/cuda/accelerator_cuda.c	2026-08-20 13:00:00.000000000 +0200
@@ -15,7 +15,7 @@
  */
 
 #include "opal_config.h"
-
+#undef OPAL_CUDA_VMM_SUPPORT
 #include <cuda.h>
 
 #include "accelerator_cuda.h"

After doing this performance is back to normal:

mpirun --mca opal_cuda_support 1 $OSU_HOME/mpi/pt2pt/osu_bibw -d cuda -b multiple D D

# OSU MPI-CUDA Bi-Directional Bandwidth Test v7.5.2
# Datatype: MPI_CHAR.
# Size      Bandwidth (MB/s)
1                       1.69
2                       3.53
4                       7.15
8                       4.81
16                     28.70
32                     57.51
64                    100.62
128                   227.92
256                   470.38
512                   937.63
1024                 1859.79
2048                 1692.79
4096                 7538.69
8192                12201.98
16384               22641.49
32768               31860.60
65536               22120.21
131072              26377.77
262144              44012.68
524288              33077.65
1048576             41367.42
2097152             43505.83
4194304             46407.97

which at least gives an indication of where the problem is coming from. Fiddling around with the code showed that the below calls are responsible for the slowness:

    is_vmm = accelerator_cuda_check_vmm(dbuf, &vmm_mem_type, &vmm_dev_id);
    is_mpool_ptr = accelerator_cuda_check_mpool(dbuf, &mpool_mem_type, &mpool_dev_id);

Since the problem is rooted inside CUDA and the driver, it's likely it will show on other systems, not only the studied Cray GH200 architecture.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions