Skip to content

fix(cuda): support CUDA 13.x / CCCL 3.x toolchains (removed CUB iterators, MKL/nvcc include, GCC 15 header warning) - #783

Merged
PabloCarmona merged 2 commits into
masterfrom
fix/cuda-13-cccl3-compat
Aug 28, 2026
Merged

PabloCarmona merged 2 commits into
masterfrom
fix/cuda-13-cccl3-compat

Conversation

@PabloCarmona

@PabloCarmona PabloCarmona commented Jul 8, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Building aihwkit with USE_CUDA=ON against a CUDA 13.3 toolchain (CCCL/CUB 3.3.4, GCC 15.2 host compiler, PyTorch 2.12.1+cu130) fails during GPU compilation. This PR fixes three independent incompatibilities uncovered by that toolchain, plus one latent bug in a CUB reduction call that CUB 3.x newly rejects.

None of these changes affect behavior on older CUDA/GCC toolchains — every fix is version-guarded and falls back to a no-op there.

Details

1. CUB "fancy" iterators removed in CCCL 3.0

cub::TransformInputIterator and cub::CountingInputIterator were removed upstream in CCCL 3.0 (shipped with CUDA 13.0+), breaking maximizer.cu, weight_clipper_cuda.cu, and noise_manager.cu with namespace "cub" has no member "TransformInputIterator".

Fix: add thin compatibility aliases in src/rpucuda/cuda/rpu_cub.h, implemented on top of the equivalent thrust::transform_iterator / thrust::counting_iterator (which is what CUB itself now uses internally). Guarded by #if CUB_VERSION >= 300000, so no call sites need to change and older toolkits are unaffected.

2. MKL/OpenBLAS include path not forwarded to nvcc

find_package(MKL) / find_package(OpenBLAS) register their include dir via a global include_directories(SYSTEM ...), but this isn't reliably picked up by the CUDA host compilation on all toolchains (e.g. a conda-provided nvcc) — .cpp files only compiled because conda separately injects its own include dir into CMAKE_CXX_FLAGS, which nvcc never sees. Result: every .cu file that transitively includes math_util.h failed with mkl.h: No such file or directory.

Fix: capture the resolved BLAS include dir in RPU_BLAS_INCLUDE_DIRS (cmake/dependencies.cmake) and explicitly forward it to nvcc via CMAKE_CUDA_FLAGS in CMakeLists.txt.

3. GCC 15 -Wtemplate-body hard error in PyTorch headers

With a GCC 15 host compiler, nvcc's cudafe front-end re-emits ATen/core/List_inl.h in a form that trips GCC 15's new -Wtemplate-body check, which defaults to a hard error (need 'typename' before ... because ... is a dependent scope). The header is valid C++ — plain GCC 15 compiles it without complaint — so this is specifically a cudafe re-emission artifact. Affects any .cu file including torch headers (io_manager.cu, forward_backward_pass.cu, update_management_helper.cu, etc.).

Fix: pass -Xcompiler=-Wno-error=template-body when compiling CUDA sources, guarded to CMAKE_CXX_COMPILER_ID STREQUAL "GNU" AND VERSION_GREATER_EQUAL 15 (the flag doesn't exist before GCC 15).

4. Latent int/float mismatch in a CUB reduction call

Once the above are fixed, noise_manager.cu:242 fails CUB 3.x's stricter template deduction in DeviceSegmentedReduce::Reduce: it passes a bare 0 (int) as the initial value while every sibling call in the same file passes (T)0 / (T)0.0. CUB 3.x deduces the accumulator type from the init value and no longer tolerates the mismatch.

Fix: change 0 to (T)0 to match the pattern used elsewhere in the file.

Verification

Built and ran on CUDA 13.3 / driver 13.3 / RTX 2080 Ti (sm_75) / GCC 15.2 / PyTorch 2.12.1+cu130:

rpu_base cuda.is_compiled(): True
GPU forward output device: cuda:0  shape: (2, 4)
SUCCESS: analog GPU tile forward ran on cuda:0

Full RPU_GPU + rpu_base extension build completes with zero errors, and an AnalogTile.cuda() forward pass runs correctly end-to-end.

Test plan

  • CI: confirm existing CUDA build jobs (older CUDA/CCCL versions) still pass unaffected (all fixes are version-guarded)
  • Manually verify on a CUDA 12.x toolchain that the build/behavior is unchanged
  • Run the existing GPU test suite (RPU_GPU_TEST_SRCS) under CUDA 13.x

…tors, MKL/nvcc include, GCC 15 header warning)

Signed-off-by: Pablo Carmona Gonzalez <pablocarmonagonzalez@gmail.com>
@PabloCarmona
PabloCarmona requested review from anu-pub and maljoras July 8, 2026 13:54
@PabloCarmona PabloCarmona self-assigned this Jul 8, 2026
@PabloCarmona PabloCarmona added bug Something isn't working build Issues related to build system and compilation/installing labels Jul 9, 2026
@PabloCarmona
PabloCarmona merged commit 8165a6a into master Aug 28, 2026
7 checks passed
@PabloCarmona
PabloCarmona deleted the fix/cuda-13-cccl3-compat branch August 28, 2026 16:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working build Issues related to build system and compilation/installing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant