To compile and run CUDA code, you need to use an external Kokkos
installation that has been configured with CUDA support.
Note that deal.II requires a compiler that is compatible with the
backends enabled in Kokkos. In particular, you need to explicitly set `CMAKE_CXX_COMPILER` to either Kokkos'
nvcc_wrapper or clang++ when enabling the
Cuda backend.
Several MPI implementations are able to perform MPI operations with data
located in device memory directly without the need to copy to CPU memory
explicitly first. This feature is commonly known as "CUDA-aware MPI".
If deal.II is configured with -DDEAL_II_WITH_MPI=ON then
the configuration system queries the MPI library for this capability and
sets DEAL_II_MPI_WITH_DEVICE_SUPPORT accordingly.
The check is a heuristic based on preprocessor symbols exported by the
mpi-ext.h header, which not every MPI implementation
provides. If the detection fails for your MPI library you can override
the result by configuring with
-DDEAL_II_WITH_MPI=ON
-DDEAL_II_MPI_WITH_DEVICE_SUPPORT=ON
Note, that this bypasses the check. Setting the flag for an MPI library
that is not CUDA-aware will lead to segmentation faults in MPI calls.
Using CUDA in combination with architecture-specific C++ compiler flags
like -march=native is known to be fragile and there might be
compatibility issues with other libraries, e.g. using CUDA 10.1 with
-DDEAL_II_WITH_TBB=ON and
-DDEAL_II_CXX_FLAGS=-march=native results in compile time
errors like:
/usr/lib/gcc/x86_64-linux-gnu/7/include/avx512fintrin.h(11265): error: identifier "__builtin_ia32_scalefsd_round" is undefined
/usr/lib/gcc/x86_64-linux-gnu/7/include/avx512fintrin.h(11274): error: identifier "__builtin_ia32_scalefss_round" is undefined
Since vectorization in VectorizedArray is disabled when compiling with
CUDA support anyway, it is recommended to drop the compile flag in that
case.