fix(cuda): support CUDA 13.x / CCCL 3.x toolchains (removed CUB iterators, MKL/nvcc include, GCC 15 header warning) - #783
Merged
Conversation
…tors, MKL/nvcc include, GCC 15 header warning) Signed-off-by: Pablo Carmona Gonzalez <pablocarmonagonzalez@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Building aihwkit with
USE_CUDA=ONagainst a CUDA 13.3 toolchain (CCCL/CUB 3.3.4, GCC 15.2 host compiler, PyTorch 2.12.1+cu130) fails during GPU compilation. This PR fixes three independent incompatibilities uncovered by that toolchain, plus one latent bug in a CUB reduction call that CUB 3.x newly rejects.None of these changes affect behavior on older CUDA/GCC toolchains — every fix is version-guarded and falls back to a no-op there.
Details
1. CUB "fancy" iterators removed in CCCL 3.0
cub::TransformInputIteratorandcub::CountingInputIteratorwere removed upstream in CCCL 3.0 (shipped with CUDA 13.0+), breakingmaximizer.cu,weight_clipper_cuda.cu, andnoise_manager.cuwithnamespace "cub" has no member "TransformInputIterator".Fix: add thin compatibility aliases in
src/rpucuda/cuda/rpu_cub.h, implemented on top of the equivalentthrust::transform_iterator/thrust::counting_iterator(which is what CUB itself now uses internally). Guarded by#if CUB_VERSION >= 300000, so no call sites need to change and older toolkits are unaffected.2. MKL/OpenBLAS include path not forwarded to nvcc
find_package(MKL)/find_package(OpenBLAS)register their include dir via a globalinclude_directories(SYSTEM ...), but this isn't reliably picked up by the CUDA host compilation on all toolchains (e.g. a conda-provided nvcc) —.cppfiles only compiled because conda separately injects its ownincludedir intoCMAKE_CXX_FLAGS, which nvcc never sees. Result: every.cufile that transitively includesmath_util.hfailed withmkl.h: No such file or directory.Fix: capture the resolved BLAS include dir in
RPU_BLAS_INCLUDE_DIRS(cmake/dependencies.cmake) and explicitly forward it to nvcc viaCMAKE_CUDA_FLAGSinCMakeLists.txt.3. GCC 15
-Wtemplate-bodyhard error in PyTorch headersWith a GCC 15 host compiler, nvcc's
cudafefront-end re-emitsATen/core/List_inl.hin a form that trips GCC 15's new-Wtemplate-bodycheck, which defaults to a hard error (need 'typename' before ... because ... is a dependent scope). The header is valid C++ — plain GCC 15 compiles it without complaint — so this is specifically a cudafe re-emission artifact. Affects any.cufile including torch headers (io_manager.cu,forward_backward_pass.cu,update_management_helper.cu, etc.).Fix: pass
-Xcompiler=-Wno-error=template-bodywhen compiling CUDA sources, guarded toCMAKE_CXX_COMPILER_ID STREQUAL "GNU" AND VERSION_GREATER_EQUAL 15(the flag doesn't exist before GCC 15).4. Latent int/float mismatch in a CUB reduction call
Once the above are fixed,
noise_manager.cu:242fails CUB 3.x's stricter template deduction inDeviceSegmentedReduce::Reduce: it passes a bare0(int) as the initial value while every sibling call in the same file passes(T)0/(T)0.0. CUB 3.x deduces the accumulator type from the init value and no longer tolerates the mismatch.Fix: change
0to(T)0to match the pattern used elsewhere in the file.Verification
Built and ran on CUDA 13.3 / driver 13.3 / RTX 2080 Ti (sm_75) / GCC 15.2 / PyTorch 2.12.1+cu130:
Full
RPU_GPU+rpu_baseextension build completes with zero errors, and anAnalogTile.cuda()forward pass runs correctly end-to-end.Test plan
RPU_GPU_TEST_SRCS) under CUDA 13.x