From 126c2a296c1389fcb1b3a516573056e9edd51089 Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Wed, 22 May 2024 09:34:39 -0700 Subject: [PATCH 01/34] docs --- README.md | 36 +++++++++++++++++++++++++++++++++++- 1 file changed, 35 insertions(+), 1 deletion(-) diff --git a/README.md b/README.md index ffe91c0..9f986db 100644 --- a/README.md +++ b/README.md @@ -1 +1,35 @@ -# Micromagnetics \ No newline at end of file +# MagneX +MagneX is a massively parallel, 3D micromagnetics solver for modeling magnetic materials. +MagneX solves the Landau-Lifshitz-Gilbert (LLG) equations, including exchange, anisotropy, demagnetization, and Dzyaloshinskii-Moriya interaction (DMI) coupling. +The algorithm is implemented using Exascale Computing Project software framework, AMReX, which provides effective scalability on manycore and GPU-based supercomputing architectures. + +# Installation +## Download AMReX Repository +``` git clone git@github.com:AMReX-Codes/amrex.git ``` +## Download MagneX Repository +``` git@github.com:AMReX-Microelectronics/MagneX.git ``` +## Build +Make sure that the AMReX and MagneX are cloned in the same location in their filesystem. Navogate to the Exec folder of MagneX and execute +```make -j 4``` + +# Running MagneX +Example input scripts are located in `Exec/standard_problem_inputs/` directory. +## Simple Testcase +You can run the following to simulate muMAG Standard Problem 4 dynamics: +## For pure MPI build (but with a single MPI rank) +```./main3d.gnu.MPI.ex standard_problem_inputs/inputs_std4``` +# Visualization and Data Analysis +Refer to the following link for several visualization tools that can be used for AMReX plotfiles. + +[Visualization](https://amrex-codes.github.io/amrex/docs_html/Visualization_Chapter.html) + +### Data Analysis in Python using yt +You can extract the data in numpy array format using yt (you can refer to this for installation and usage of [yt](https://yt-project.org/). After you have installed yt, you can do something as follows, for example, to get variable 'Pz' (z-component of polarization) +``` +import yt +ds = yt.load('./plt00001000/') # for data at time step 1000 +ad0 = ds.covering_grid(level=0, left_edge=ds.domain_left_edge, dims=ds.domain_dimensions) +P_array = ad0['Pz'].to_ndarray() +``` +# Publications +1. Z. Yao, P. Kumar, J. C. LePelch, and A. Nonaka, MagneX: An Exascale-Enabled Micromagnetics Solver for Spintronic Systems, in preparation. From bfe6903e8fb19b5935a8e17cbaa38cec2a3536b1 Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Wed, 22 May 2024 09:42:30 -0700 Subject: [PATCH 02/34] docs update --- README.md | 11 +++++++---- 1 file changed, 7 insertions(+), 4 deletions(-) diff --git a/README.md b/README.md index 9f986db..7c40f02 100644 --- a/README.md +++ b/README.md @@ -4,12 +4,15 @@ MagneX solves the Landau-Lifshitz-Gilbert (LLG) equations, including exchange, a The algorithm is implemented using Exascale Computing Project software framework, AMReX, which provides effective scalability on manycore and GPU-based supercomputing architectures. # Installation -## Download AMReX Repository +## Download AMReX and MagneX Repositories +Make sure that AMReX and MagneX are cloned at the same root location. ``` git clone git@github.com:AMReX-Codes/amrex.git ``` -## Download MagneX Repository -``` git@github.com:AMReX-Microelectronics/MagneX.git ``` +``` git clone git@github.com:AMReX-Microelectronics/MagneX.git ``` +## Dependencies +Beyond a standard Ubuntu22 installation, the Ubuntu packages libfftw3-dev and libfftw3-mpi-dev are required. +Also heFFTe is required. ## Build -Make sure that the AMReX and MagneX are cloned in the same location in their filesystem. Navogate to the Exec folder of MagneX and execute + Navogate to the Exec folder of MagneX and execute ```make -j 4``` # Running MagneX From 84f19dcc4660c5e337d33786e7548c38485a16bc Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Wed, 22 May 2024 09:43:32 -0700 Subject: [PATCH 03/34] docs update --- README.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/README.md b/README.md index 7c40f02..f4f7401 100644 --- a/README.md +++ b/README.md @@ -5,8 +5,8 @@ The algorithm is implemented using Exascale Computing Project software framework # Installation ## Download AMReX and MagneX Repositories -Make sure that AMReX and MagneX are cloned at the same root location. -``` git clone git@github.com:AMReX-Codes/amrex.git ``` +Make sure that AMReX and MagneX are cloned at the same root location. \ +``` git clone git@github.com:AMReX-Codes/amrex.git ``` \ ``` git clone git@github.com:AMReX-Microelectronics/MagneX.git ``` ## Dependencies Beyond a standard Ubuntu22 installation, the Ubuntu packages libfftw3-dev and libfftw3-mpi-dev are required. From d71eb798575603413bf53335b5c4919cfd76ea77 Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Wed, 22 May 2024 09:49:12 -0700 Subject: [PATCH 04/34] docs update --- README.md | 18 +++++++++++++----- 1 file changed, 13 insertions(+), 5 deletions(-) diff --git a/README.md b/README.md index f4f7401..08ac0b7 100644 --- a/README.md +++ b/README.md @@ -4,15 +4,23 @@ MagneX solves the Landau-Lifshitz-Gilbert (LLG) equations, including exchange, a The algorithm is implemented using Exascale Computing Project software framework, AMReX, which provides effective scalability on manycore and GPU-based supercomputing architectures. # Installation +Here are instructions for a basic, pure-MPI (no GPU) installation. More detailed instructions for GPU systems are in the full documentation. ## Download AMReX and MagneX Repositories Make sure that AMReX and MagneX are cloned at the same root location. \ -``` git clone git@github.com:AMReX-Codes/amrex.git ``` \ -``` git clone git@github.com:AMReX-Microelectronics/MagneX.git ``` +``` git clone https://github.com/AMReX-Codes/amrex.git ``` \ +``` git clone https://AMReX-Microelectronics/MagneX.git ``` ## Dependencies -Beyond a standard Ubuntu22 installation, the Ubuntu packages libfftw3-dev and libfftw3-mpi-dev are required. -Also heFFTe is required. +Beyond a standard Ubuntu22 installation, the Ubuntu packages libfftw3-dev, libfftw3-mpi-dev, and cmake are required. +Also heFFTe is required. At the same level that AMReX and MagneX are cloned, run: \ +``` git clone https://github.com/icl-utk-edu/heffte.git ```\ +``` cd heffte ```\ +``` mkdir build ```\ +``` cd build ```\ +``` cmake -DCMAKE_BUILD_TYPE=Release -DCMAKE_CXX_STANDARD=17 -DBUILD_SHARED_LIBS=OFF -DCMAKE_INSTALL_PREFIX=. -DHeffte_ENABLE_FFTW=ON -DHeffte_ENABLE_CUDA=OFF .. ```\ +``` make -j4 ```\ +``` make install``` ## Build - Navogate to the Exec folder of MagneX and execute + Navigate to the Exec folder of MagneX and execute ```make -j 4``` # Running MagneX From 7fe43b2c412a3281fa9492d56959b2863e787510 Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Wed, 22 May 2024 09:56:54 -0700 Subject: [PATCH 05/34] use forward euler to simple build works --- Exec/standard_problem_inputs/inputs_std4 | 14 +++++++------- 1 file changed, 7 insertions(+), 7 deletions(-) diff --git a/Exec/standard_problem_inputs/inputs_std4 b/Exec/standard_problem_inputs/inputs_std4 index 7bc126c..47f4ea5 100644 --- a/Exec/standard_problem_inputs/inputs_std4 +++ b/Exec/standard_problem_inputs/inputs_std4 @@ -6,10 +6,10 @@ max_grid_size_x = 160 max_grid_size_y = 40 max_grid_size_z = 1 dt = 5.e-15 # forward Euler (685s), 1.e-14 unstable -dt = 2.5e-13 # Predictor-corrector (70s), 5.e-13 unstable -dt = 1.e-13 # trapezoidal (63s), 2.5e-13 unstable -dt = 2.5e-13 # SSPRK3 (36s), 5.e-13 unstable -dt = 5.e-13 # RK4 (24s), 1.e-12 unstable +#dt = 2.5e-13 # Predictor-corrector (70s), 5.e-13 unstable +#dt = 1.e-13 # trapezoidal (63s), 2.5e-13 unstable +#dt = 2.5e-13 # SSPRK3 (36s), 5.e-13 unstable +#dt = 5.e-13 # RK4 (24s), 1.e-12 unstable ############## # 1.5625nm case @@ -73,10 +73,10 @@ DMI_coupling = 0 # INTEGRATION -#TimeIntegratorOption = 1 #Forward Euler +TimeIntegratorOption = 1 #Forward Euler #TimeIntegratorOption = 2 #Predictor-corrector #TimeIntegratorOption = 3 #2nd order artemis way -TimeIntegratorOption = 4 #amrex/sundials backend integrators +#TimeIntegratorOption = 4 #amrex/sundials backend integrators # tolerance threshold (L_inf change between iterations) for TimeIntegrationOption 2 and 3 iterative_tolerance = 1.e-9 @@ -98,7 +98,7 @@ integration.type = RungeKutta ### 2 = Trapezoid Method ### 3 = SSPRK3 Method ### 4 = RK4 Method -integration.rk.type = 4 +integration.rk.type = 1 ## If using a user-specified Butcher Tableau, then ## set nodes, weights, and table entries here: From d80cf56ff3025f24c2c46306d01631c07b1bd83c Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Wed, 22 May 2024 10:00:05 -0700 Subject: [PATCH 06/34] Update README.md --- README.md | 31 +++++++++++++++---------------- 1 file changed, 15 insertions(+), 16 deletions(-) diff --git a/README.md b/README.md index 08ac0b7..bef5634 100644 --- a/README.md +++ b/README.md @@ -7,28 +7,27 @@ The algorithm is implemented using Exascale Computing Project software framework Here are instructions for a basic, pure-MPI (no GPU) installation. More detailed instructions for GPU systems are in the full documentation. ## Download AMReX and MagneX Repositories Make sure that AMReX and MagneX are cloned at the same root location. \ -``` git clone https://github.com/AMReX-Codes/amrex.git ``` \ -``` git clone https://AMReX-Microelectronics/MagneX.git ``` +``` >> git clone https://github.com/AMReX-Codes/amrex.git ``` \ +``` >> git clone https://AMReX-Microelectronics/MagneX.git ``` ## Dependencies Beyond a standard Ubuntu22 installation, the Ubuntu packages libfftw3-dev, libfftw3-mpi-dev, and cmake are required. Also heFFTe is required. At the same level that AMReX and MagneX are cloned, run: \ -``` git clone https://github.com/icl-utk-edu/heffte.git ```\ -``` cd heffte ```\ -``` mkdir build ```\ -``` cd build ```\ -``` cmake -DCMAKE_BUILD_TYPE=Release -DCMAKE_CXX_STANDARD=17 -DBUILD_SHARED_LIBS=OFF -DCMAKE_INSTALL_PREFIX=. -DHeffte_ENABLE_FFTW=ON -DHeffte_ENABLE_CUDA=OFF .. ```\ -``` make -j4 ```\ -``` make install``` +``` >> git clone https://github.com/icl-utk-edu/heffte.git ```\ +``` >> cd heffte ```\ +``` >> mkdir build ```\ +``` >> cd build ```\ +``` >> cmake -DCMAKE_BUILD_TYPE=Release -DCMAKE_CXX_STANDARD=17 -DBUILD_SHARED_LIBS=OFF -DCMAKE_INSTALL_PREFIX=. -DHeffte_ENABLE_FFTW=ON -DHeffte_ENABLE_CUDA=OFF .. ```\ +``` >> make -j4 ```\ +``` >> make install``` ## Build - Navigate to the Exec folder of MagneX and execute -```make -j 4``` + Navigate to MagneX/Exec/ and run:\ +```>> make -j4``` # Running MagneX Example input scripts are located in `Exec/standard_problem_inputs/` directory. ## Simple Testcase -You can run the following to simulate muMAG Standard Problem 4 dynamics: -## For pure MPI build (but with a single MPI rank) -```./main3d.gnu.MPI.ex standard_problem_inputs/inputs_std4``` +You can run the following to simulate muMAG Standard Problem 4 dynamics:\ +```>> ./main3d.gnu.MPI.ex standard_problem_inputs/inputs_std4``` # Visualization and Data Analysis Refer to the following link for several visualization tools that can be used for AMReX plotfiles. @@ -38,9 +37,9 @@ Refer to the following link for several visualization tools that can be used for You can extract the data in numpy array format using yt (you can refer to this for installation and usage of [yt](https://yt-project.org/). After you have installed yt, you can do something as follows, for example, to get variable 'Pz' (z-component of polarization) ``` import yt -ds = yt.load('./plt00001000/') # for data at time step 1000 +ds = yt.load('./plt00010000/') # for data at time step 10000 ad0 = ds.covering_grid(level=0, left_edge=ds.domain_left_edge, dims=ds.domain_dimensions) -P_array = ad0['Pz'].to_ndarray() +Px_array = ad0['Mx'].to_ndarray() ``` # Publications 1. Z. Yao, P. Kumar, J. C. LePelch, and A. Nonaka, MagneX: An Exascale-Enabled Micromagnetics Solver for Spintronic Systems, in preparation. From 09256d6b8929583e72ce027b659d854d451a7e7b Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Wed, 22 May 2024 10:04:25 -0700 Subject: [PATCH 07/34] Update README.md --- README.md | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/README.md b/README.md index bef5634..32e2939 100644 --- a/README.md +++ b/README.md @@ -10,8 +10,9 @@ Make sure that AMReX and MagneX are cloned at the same root location. \ ``` >> git clone https://github.com/AMReX-Codes/amrex.git ``` \ ``` >> git clone https://AMReX-Microelectronics/MagneX.git ``` ## Dependencies -Beyond a standard Ubuntu22 installation, the Ubuntu packages libfftw3-dev, libfftw3-mpi-dev, and cmake are required. -Also heFFTe is required. At the same level that AMReX and MagneX are cloned, run: \ +Beyond a standard Ubuntu22 installation, the Ubuntu packages libfftw3-dev, libfftw3-mpi-dev, and cmake are required.\ +SUNDIALS is optional an enabled Runge-Kutta, implicit, and multirate integrators (more detailed instructions in the full documentation).\ +heFFTe is required. At the same level that AMReX and MagneX are cloned, run: \ ``` >> git clone https://github.com/icl-utk-edu/heffte.git ```\ ``` >> cd heffte ```\ ``` >> mkdir build ```\ From fe8ecfdc3faa6b3886b49c0130ca1769c3d86594 Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Wed, 22 May 2024 10:08:43 -0700 Subject: [PATCH 08/34] Update README.md --- README.md | 2 -- 1 file changed, 2 deletions(-) diff --git a/README.md b/README.md index 32e2939..0a89797 100644 --- a/README.md +++ b/README.md @@ -25,8 +25,6 @@ heFFTe is required. At the same level that AMReX and MagneX are cloned, run: \ ```>> make -j4``` # Running MagneX -Example input scripts are located in `Exec/standard_problem_inputs/` directory. -## Simple Testcase You can run the following to simulate muMAG Standard Problem 4 dynamics:\ ```>> ./main3d.gnu.MPI.ex standard_problem_inputs/inputs_std4``` # Visualization and Data Analysis From 1df8bdbab9d6b21a4da39600df354e57eebca386 Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Thu, 23 May 2024 11:06:49 -0700 Subject: [PATCH 09/34] whitespace --- README.md | 2 -- 1 file changed, 2 deletions(-) diff --git a/README.md b/README.md index 220f6e4..ea52767 100644 --- a/README.md +++ b/README.md @@ -2,11 +2,9 @@ MagneX is a massively parallel, 3D micromagnetics solver for modeling magnetic materials. MagneX solves the Landau-Lifshitz-Gilbert (LLG) equations, including exchange, anisotropy, demagnetization, and Dzyaloshinskii-Moriya interaction (DMI) coupling. The algorithm is implemented using Exascale Computing Project software framework, AMReX, which provides effective scalability on manycore and GPU-based supercomputing architectures. - # Documentation and Getting Help More extensive documentation is available [HERE](https://amrex-microelectronics.github.io). Our community is here to help. Please report installation problems or general questions about the code in the github [Issues](https://github.com/AMReX-Microelectronics/MagneX/issues) tab above. - # Installation Here are instructions for a basic, pure-MPI (no GPU) installation. More detailed instructions for GPU systems are in the full documentation. ## Download AMReX and MagneX Repositories From 15a993ce8f8f9dd2b6be8e231f86a4e4b978daed Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Fri, 31 May 2024 11:57:44 -0700 Subject: [PATCH 10/34] enable DMI and anis in scaling test --- Exec/inputs_performance | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/Exec/inputs_performance b/Exec/inputs_performance index 3b291c2..3512d27 100644 --- a/Exec/inputs_performance +++ b/Exec/inputs_performance @@ -31,16 +31,16 @@ Ms_parser(x,y,z) = "(x>0.)*(x<512.e-9)*(y>0.)*(y<512.e-9)*(z>0.)*(z<512. gamma_parser(x,y,z) = "(x>0.)*(x<512.e-9)*(y>0.)*(y<512.e-9)*(z>0.)*(z<512.e-9) * -1.759e11" exchange_parser(x,y,z) = "(x>0.)*(x<512.e-9)*(y>0.)*(y<512.e-9)*(z>0.)*(z<512.e-9) * 1.3e-11" anisotropy_parser(x,y,z) = "(x>0.)*(x<512.e-9)*(y>0.)*(y<512.e-9)*(z>0.)*(z<512.e-9) * -139.26" -DMI_parser(x,y,z) = "0." +DMI_parser(x,y,z) = "(x>0.)*(x<512.e-9)*(y>0.)*(y<512.e-9)*(z>0.)*(z<512.e-9) * -4.5e-3" precession = 1 demag_coupling = 1 FFT_solver = 1 M_normalization = 1 # 0 = unsaturated case; 1 = saturated case exchange_coupling = 1 -anisotropy_coupling = 0 -anisotropy_axis = 0.0 1.0 0.0 -DMI_coupling = 0 +anisotropy_coupling = 1 +anisotropy_axis = 0.0 0.0 1.0 +DMI_coupling = 1 # INTEGRATION From 535bee5a0a2bfe1cceff0a5ca5c6227be2c5a73c Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Wed, 12 Jun 2024 18:10:17 -0700 Subject: [PATCH 11/34] skyrmion diagnostic post processing utility --- Exec/DMI_diagnostic/DMI_diagnostic.cpp | 144 +++++++++++++++++++++++++ Exec/DMI_diagnostic/GNUmakefile | 33 ++++++ Exec/DMI_diagnostic/Make.package | 20 ++++ 3 files changed, 197 insertions(+) create mode 100644 Exec/DMI_diagnostic/DMI_diagnostic.cpp create mode 100644 Exec/DMI_diagnostic/GNUmakefile create mode 100644 Exec/DMI_diagnostic/Make.package diff --git a/Exec/DMI_diagnostic/DMI_diagnostic.cpp b/Exec/DMI_diagnostic/DMI_diagnostic.cpp new file mode 100644 index 0000000..70ffc58 --- /dev/null +++ b/Exec/DMI_diagnostic/DMI_diagnostic.cpp @@ -0,0 +1,144 @@ + +#include +#include + +#include +#include + +using namespace std; +using namespace amrex; + +static +void +PrintUsage (const char* progName) +{ + Print() << std::endl + << "Use this utility to extract a 1D profile along the x-axis of Mz and theta for a DMI skyrmion test." << std::endl << std::endl; + Print() << "Usage:" << '\n'; + Print() << progName << " infile=inputFileName" << '\n' << '\n'; + + exit(1); +} + +int +main (int argc, + char* argv[]) +{ + amrex::Initialize(argc,argv); + + if (argc == 1) { + PrintUsage(argv[0]); + } + + // plotfile name + std::string iFile; + + // read in parameters from inputs file + ParmParse pp; + + // read in plotfile name + pp.query("infile", iFile); + if (iFile.empty()) + amrex::Abort("You must specify `infile'"); + + // for the Header + std::string Header = iFile; + Header += "/Header"; + + // open header + ifstream x; + x.open(Header.c_str(), ios::in); + + // read in first line of header + string str; + x >> str; + + // read in number of components from header + int ncomp; + x >> ncomp; + + // read in variable names from header + for (int n=0; n> str; + } + + // read in dimensionality from header + int dim; + x >> dim; + + if (dim != AMREX_SPACEDIM) { + Print() << "\nError: you are using a " << AMREX_SPACEDIM << "D build to open a " + << dim << "D plotfile\n\n"; + Abort(); + } + + int lev = 0; + + do { + + if (lev > 9) { + Abort("Utility only works for 10 levels of refinement or less"); + } + + // storage for the MultiFab + MultiFab mf; + + std::string iFile_lev = iFile; + + std::string levX = "/Level_"+to_string(lev)+"/Cell"; + std::string levXX = "/Level_0"+to_string(lev)+"/Cell"; + + // now read in the plotfile data + // check to see whether the user pointed to the plotfile base directory + // or the data itself + if (amrex::FileExists(iFile+levX+"_H")) { + iFile_lev += levX; + } else if (amrex::FileExists(iFile+levXX+"_H")) { + iFile_lev += levXX; + } else { + break; // terminate while loop + } + + // read in plotfile to MultiFab + VisMF::Read(mf, iFile_lev); + + if (lev == 0) { + ncomp = mf.nComp(); + Print() << "Number of components in the plotfile = " << ncomp << std::endl; + Print() << "Nodality of plotfile = " << mf.ixType().toIntVect() << std::endl; + } + + // get boxArray to compute number of grid points at the level + BoxArray ba = mf.boxArray(); + Print() << "Number of grid points at level " << lev << " = " << ba.numPts() << std::endl; + + for ( MFIter mfi(mf,false); mfi.isValid(); ++mfi ) { + + const Box& bx = mfi.validbox(); + const auto lo = amrex::lbound(bx); + const auto hi = amrex::ubound(bx); + + const Array4& mfdata = mf.array(mfi); + + Real offset = 0.; + + int k = (hi.z+1)/2; + int j = (hi.y+1)/2; + for (auto i = (hi.x+1)/2; i <= hi.x; ++i) { + std::cout << i << " " << mfdata(i,j,k,3)/1.1e6 << " " << mfdata(i,j,k,16)+offset << "\n"; + // 2pi cyclic fix + if (i5. && mfdata(i+1,j,k,16)<1.) { + offset = 2.*M_PI; + } + } + } + + } // end MFIter + + // proceed to next level of refinement + ++lev; + + } while(true); + +} diff --git a/Exec/DMI_diagnostic/GNUmakefile b/Exec/DMI_diagnostic/GNUmakefile new file mode 100644 index 0000000..7ead0dd --- /dev/null +++ b/Exec/DMI_diagnostic/GNUmakefile @@ -0,0 +1,33 @@ +AMREX_HOME ?= ../../../amrex + +DEBUG = TRUE +DIM = 3 + +COMP = gcc + +PRECISION = DOUBLE + +USE_MPI = FALSE +USE_OMP = FALSE + +################################################### + +EBASE = DMI_diagnostic + +include $(AMREX_HOME)/Tools/GNUMake/Make.defs + +include ./Make.package +include $(AMREX_HOME)/Src/Base/Make.package + +vpath %.c : . $(vpathdir) +vpath %.h : . $(vpathdir) +vpath %.cpp : . $(vpathdir) +vpath %.H : . $(vpathdir) +vpath %.F : . $(vpathdir) +vpath %.f : . $(vpathdir) +vpath %.f90 : . $(vpathdir) + +include $(AMREX_HOME)/Tools/GNUMake/Make.rules + +clean:: + $(SILENT) $(RM) particle_compare.exe diff --git a/Exec/DMI_diagnostic/Make.package b/Exec/DMI_diagnostic/Make.package new file mode 100644 index 0000000..159417d --- /dev/null +++ b/Exec/DMI_diagnostic/Make.package @@ -0,0 +1,20 @@ +CEXE_sources += ${EBASE}.cpp + + +ifneq ($(EBASE), particle_compare) +ifneq ($(EBASE), WritePlotfileToASCII) + + INCLUDE_LOCATIONS += $(AMREX_HOME)/Src/Base + include $(AMREX_HOME)/Src/Base/Make.package + vpathdir += $(AMREX_HOME)/Src/Base + + INCLUDE_LOCATIONS += $(AMREX_HOME)/Src/Extern/amrdata + include $(AMREX_HOME)/Src/Extern/amrdata/Make.package + vpathdir += $(AMREX_HOME)/Src/Extern/amrdata + + ifeq ($(NEEDS_f90_SRC),TRUE) + f90EXE_sources += ${EBASE}_nd.f90 + endif + +endif +endif From a32bd3184978107a6a250cb763df301117277388 Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Thu, 13 Jun 2024 13:47:04 -0700 Subject: [PATCH 12/34] renname file --- Exec/{input_2D_DMI => standard_problem_inputs/inputs_dmi1} | 0 1 file changed, 0 insertions(+), 0 deletions(-) rename Exec/{input_2D_DMI => standard_problem_inputs/inputs_dmi1} (100%) diff --git a/Exec/input_2D_DMI b/Exec/standard_problem_inputs/inputs_dmi1 similarity index 100% rename from Exec/input_2D_DMI rename to Exec/standard_problem_inputs/inputs_dmi1 From c5f33493b3f0efcf8a788b74ec8de8ab8be040be Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Thu, 13 Jun 2024 17:37:05 -0700 Subject: [PATCH 13/34] restart bugfix --- Source/main.cpp | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/Source/main.cpp b/Source/main.cpp index d10a091..519c072 100644 --- a/Source/main.cpp +++ b/Source/main.cpp @@ -90,7 +90,7 @@ void main_main () int start_step = 1; Real time = 0.0; - if (restart > 0) { + if (restart >= 0) { start_step = restart+1; From 23f20f61243bbaef53da10ceeebd37beaf822579 Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Mon, 15 Jul 2024 20:12:55 -0700 Subject: [PATCH 14/34] sundials 7.1.1 --- Exec/README_sundials | 50 +++++++++++++++++++++++++++----------------- 1 file changed, 31 insertions(+), 19 deletions(-) diff --git a/Exec/README_sundials b/Exec/README_sundials index 7d7356d..3fb088e 100644 --- a/Exec/README_sundials +++ b/Exec/README_sundials @@ -3,13 +3,19 @@ https://computing.llnl.gov/projects/sundials/faq#inst Installation -# Download sundials-x.y.z.tar.gz file for SUNDIALS and extract it at the same level as amrex. -# Update 6/11/24 - v7.0.0. tested -https://computing.llnl.gov/projects/sundials/sundials-software - +# You need SUNDIALS v7.1.1 or later. +# Check https://computing.llnl.gov/projects/sundials/sundials-software to see if it's available for download. +# If so, download sundials-x.y.z.tar.gz and extract it at the same level as amrex using >> tar -xzvf sundials-x.y.z.tar.gz # where x.y.z is the version of sundials +>> mv sundials-x.y.z sundials-src + +# If v7.1.1. is not available on the website, clone the git repo directly and use the latest version +# At the same level that amrex is cloned, do: + +>> git clone https://github.com/LLNL/sundials.git +>> mv sundials sundials-src -# at the same level that amrex is cloned, do: +# Next >> mkdir sundials >> cd sundials @@ -22,7 +28,7 @@ HOST BUILD >> mkdir builddir >> cd builddir ->> cmake -DCMAKE_INSTALL_PREFIX=/pathto/sundials/instdir -DEXAMPLES_INSTALL_PATH=/pathto/sundials/instdir/examples -DENABLE_MPI=ON ../../sundials-x.y.z +>> cmake -DCMAKE_INSTALL_PREFIX=/pathto/sundials/instdir -DEXAMPLES_INSTALL_PATH=/pathto/sundials/instdir/examples -DENABLE_MPI=ON ../../sundials-src ###################### NVIDIA/CUDA BUILD @@ -34,7 +40,7 @@ NVIDIA/CUDA BUILD >> mkdir builddir_cuda >> cd builddir_cuda ->> cmake -DCMAKE_INSTALL_PREFIX=/pathto/sundials/instdir_cuda -DEXAMPLES_INSTALL_PATH=/pathto/sundials/instdir_cuda/examples -DENABLE_CUDA=ON -DENABLE_MPI=ON ../../sundials-x.y.z +>> cmake -DCMAKE_INSTALL_PREFIX=/pathto/sundials/instdir_cuda -DEXAMPLES_INSTALL_PATH=/pathto/sundials/instdir_cuda/examples -DENABLE_CUDA=ON -DENABLE_MPI=ON ../../sundials-src ###################### @@ -59,16 +65,14 @@ NVIDIA/CUDA BUILD TimeIntegratorOption = 4 #amrex/sundials backend integrators -## *** Selecting the integrator backend *** -## integration.type can take on the following string or int values: -## (without the quotation marks) -## "ForwardEuler" or "0" = Native Forward Euler Integrator -## "RungeKutta" or "1" = Native Explicit Runge Kutta -## "SUNDIALS" or "2" = SUNDIALS ARKODE Integrator -## for example: -integration.type = +# INTEGRATION +## integration.type can take on the following values: +## 0 or "ForwardEuler" => Native AMReX Forward Euler integrator +## 1 or "RungeKutta" => Native AMReX Explicit Runge Kutta controlled by integration.rk.type +## 2 or "SUNDIALS" => SUNDIALS backend controlled by integration.sundials.strategy +integration.type = SUNDIALS -## *** Parameters Needed For Native Explicit Runge-Kutta *** +## Native AMReX Explicit Runge-Kutta parameters # ## integration.rk.type can take the following values: ### 0 = User-specified Butcher Tableau @@ -76,7 +80,15 @@ integration.type = ### 2 = Trapezoid Method ### 3 = SSPRK3 Method ### 4 = RK4 Method -integration.rk.type = +integration.rk.type = 1 + +# Set the SUNDIALS method type: +# ERK = Explicit Runge-Kutta method +# DIRK = Diagonally Implicit Runge-Kutta method +# +# Optionally select a specific SUNDIALS method by name, see the SUNDIALS +# documentation for the supported method names -integration.sundials.strategy = ERK -#integration.sundials.erk.method = SSPRK3 +# Use forward Euler (fixed step sizes only) +integration.sundials.type = ERK +integration.sundials.method = ARKODE_FORWARD_EULER_1_1 From 437c39b0e8d508f88e1c25268008b6267b035c54 Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Fri, 2 Aug 2024 16:25:22 -0700 Subject: [PATCH 15/34] inputs file for paper demag plot --- Exec/inputs_demagviz | 115 +++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 115 insertions(+) create mode 100644 Exec/inputs_demagviz diff --git a/Exec/inputs_demagviz b/Exec/inputs_demagviz new file mode 100644 index 0000000..28b6b22 --- /dev/null +++ b/Exec/inputs_demagviz @@ -0,0 +1,115 @@ +n_cell = 64 64 64 +max_grid_size_x = 64 +max_grid_size_y = 64 +max_grid_size_z = 64 + +dt = 1.e-15 +nsteps = 1 +plot_int = 1 +chk_int = -1 +restart = -1 + +prob_lo = 0. 0. 0. +prob_hi = 256.e-9 256.e-9 256.e-9 + +mu0 = 1.25663e-6 + +Mx_parser(x,y,z) = "8.e5 * (x>64.e-9)*(x<192.e-9)*(y>64.e-9)*(y<192.e-9)*(z>64.e-9)*(z<192.e-9)" +My_parser(x,y,z) = "0." +Mz_parser(x,y,z) = "0." + +# Field 1: mu_0 Hx=-24.6 mT, mu_0 Hy= 4.3 mT, mu_0 Hz= 0.0 mT +# which is a field approximately 25 mT, directed 170 degrees counterclockwise from the positive x axis +timedependent_Hbias = 0 +Hx_bias_parser(x,y,z,t) = "1.e5" +Hy_bias_parser(x,y,z,t) = "1.e5" +Hz_bias_parser(x,y,z,t) = "1.e5" + +timedependent_alpha = 0 +alpha_parser(x,y,z,t) = "(x>64.e-9)*(x<192.e-9)*(y>64.e-9)*(y<192.e-9)*(z>64.e-9)*(z<192.e-9) * 0.5" +Ms_parser(x,y,z) = "(x>64.e-9)*(x<192.e-9)*(y>64.e-9)*(y<192.e-9)*(z>64.e-9)*(z<192.e-9) * 8.e5" +gamma_parser(x,y,z) = "(x>64.e-9)*(x<192.e-9)*(y>64.e-9)*(y<192.e-9)*(z>64.e-9)*(z<192.e-9) * -1.759e11" +exchange_parser(x,y,z) = "(x>64.e-9)*(x<192.e-9)*(y>64.e-9)*(y<192.e-9)*(z>64.e-9)*(z<192.e-9) * 1.3e-11" +anisotropy_parser(x,y,z) = "(x>64.e-9)*(x<192.e-9)*(y>64.e-9)*(y<192.e-9)*(z>64.e-9)*(z<192.e-9) * -139.26" +DMI_parser(x,y,z) = "(x>64.e-9)*(x<192.e-9)*(y>64.e-9)*(y<192.e-9)*(z>64.e-9)*(z<192.e-9) * -4.5e-3" + +precession = 1 +demag_coupling = 1 +FFT_solver = 1 +M_normalization = 1 # 0 = unsaturated case; 1 = saturated case +exchange_coupling = 1 +anisotropy_coupling = 1 +anisotropy_axis = 0.0 0.0 1.0 +DMI_coupling = 1 + +# INTEGRATION + +TimeIntegratorOption = 1 #Forward Euler +#TimeIntegratorOption = 2 #Predictor-corrector +#TimeIntegratorOption = 3 #2nd order artemis way +#TimeIntegratorOption = 4 #amrex/sundials backend integrators + +# tolerance threshold (L_inf change between iterations) for TimeIntegrationOption 2 and 3 +iterative_tolerance = 1.e-9 + +## amrex/sundials backend integrators +## *** Selecting the integrator backend *** +## integration.type can take on the following string or int values: +## (without the quotation marks) +## "ForwardEuler" or "0" = Native Forward Euler Integrator +## "RungeKutta" or "1" = Native Explicit Runge Kutta controlled by integration.rk.type +## "SUNDIALS" or "2" = SUNDIALS Integrators controlled by integration.sundials.strategy +integration.type = RungeKutta + +## *** Parameters Needed For Native Explicit Runge-Kutta *** +# +## integration.rk.type can take the following values: +### 0 = User-specified Butcher Tableau +### 1 = Forward Euler +### 2 = Trapezoid Method +### 3 = SSPRK3 Method +### 4 = RK4 Method +integration.rk.type = 3 + +## If using a user-specified Butcher Tableau, then +## set nodes, weights, and table entries here: +# +## The Butcher Tableau is read as a flattened, +## lower triangular matrix (but including the diagonal) +## in row major format. +integration.rk.weights = 1 +integration.rk.nodes = 0 +integration.rk.tableau = 0.0 + +## *** Parameters Needed For SUNDIALS Integrators *** +## integration.sundials.strategy specifies which ARKODE strategy to use. +## The available options are (without the quotations): +## "ERK" = Explicit Runge Kutta +## "MRI" = Multirate Integrator +## "MRITEST" = Tests the Multirate Integrator by setting a zero-valued fast RHS function +## for example: +integration.sundials.strategy = ERK + +## *** Parameters Specific to SUNDIALS ERK Strategy *** +## (Requires integration.type=SUNDIALS and integration.sundials.strategy=ERK) +## integration.sundials.erk.method specifies which explicit Runge Kutta method +## for SUNDIALS to use. The following options are supported: +## "SSPRK3" = 3rd order strong stability preserving RK (default) +## "Trapezoid" = 2nd order trapezoidal rule +## "ForwardEuler" = 1st order forward euler +## for example: +integration.sundials.erk.method = SSPRK3 + +## *** Parameters Specific to SUNDIALS MRI Strategy *** +## (Requires integration.type=SUNDIALS and integration.sundials.strategy=MRI) +## integration.sundials.mri.implicit_inner specifies whether or not to use an implicit inner solve +## integration.sundials.mri.outer_method specifies which outer (slow) method to use +## integration.sundials.mri.inner_method specifies which inner (fast) method to use +## The following options are supported for both the inner and outer methods: +## "KnothWolke3" = 3rd order Knoth-Wolke method (default for outer method) +## "Trapezoid" = 2nd order trapezoidal rule +## "ForwardEuler" = 1st order forward euler (default for inner method) +## for example: +integration.sundials.mri.implicit_inner = false +integration.sundials.mri.outer_method = KnothWolke3 +integration.sundials.mri.inner_method = Trapezoid From 2fe1a87e413328bb9c1f66ec1080a40ffe566c2a Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Mon, 19 Aug 2024 16:44:21 -0700 Subject: [PATCH 16/34] license and copyright --- Legal.txt | 17 +++++++++++++++++ license.txt | 43 +++++++++++++++++++++++++++++++++++++++++++ 2 files changed, 60 insertions(+) create mode 100644 Legal.txt create mode 100644 license.txt diff --git a/Legal.txt b/Legal.txt new file mode 100644 index 0000000..3fe4c70 --- /dev/null +++ b/Legal.txt @@ -0,0 +1,17 @@ +*** Copyright Notice *** + +Exascale-Enabled Models and Algorithms for Microelectronics Applications +(MicroEleX) Copyright (c) 2024, The Regents of the University of California, +through Lawrence Berkeley National Laboratory (subject to receipt of any +required approvals from the U.S. Dept. of Energy). All rights reserved. + +If you have questions about your rights to use or distribute this software, +please contact Berkeley Lab's Intellectual Property Office at +IPO@lbl.gov. + +NOTICE. This Software was developed under funding from the U.S. Department +of Energy and the U.S. Government consequently retains certain rights. As +such, the U.S. Government has been granted for itself and others acting on +its behalf a paid-up, nonexclusive, irrevocable, worldwide license in the +Software to reproduce, distribute copies to the public, prepare derivative +works, and perform publicly and display publicly, and to permit others to do so. diff --git a/license.txt b/license.txt new file mode 100644 index 0000000..e3eed5c --- /dev/null +++ b/license.txt @@ -0,0 +1,43 @@ +*** License Agreement *** + +Exascale-Enabled Models and Algorithms for Microelectronics Applications +(MicroEleX) Copyright (c) 2024, The Regents of the University of California, +through Lawrence Berkeley National Laboratory (subject to receipt of any +required approvals from the U.S. Dept. of Energy). All rights reserved. + +Redistribution and use in source and binary forms, with or without +modification, are permitted provided that the following conditions are met: + +(1) Redistributions of source code must retain the above copyright notice, +this list of conditions and the following disclaimer. + +(2) Redistributions in binary form must reproduce the above copyright +notice, this list of conditions and the following disclaimer in the +documentation and/or other materials provided with the distribution. + +(3) Neither the name of the University of California, Lawrence Berkeley +National Laboratory, U.S. Dept. of Energy nor the names of its contributors +may be used to endorse or promote products derived from this software +without specific prior written permission. + +THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS AND CONTRIBUTORS "AS IS" AND +ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED +WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE +ARE DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT OWNER OR CONTRIBUTORS BE LIABLE +FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL +DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR +SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER +CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY, +OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE +OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE. + +You are under no obligation whatsoever to provide any bug fixes, patches, +or upgrades to the features, functionality or performance of the source +code ("Enhancements") to anyone; however, if you choose to make your +Enhancements available either publicly, or directly to Lawrence Berkeley +National Laboratory, without imposing a separate written license agreement +for such Enhancements, then you hereby grant the following license: a +non-exclusive, royalty-free perpetual license to install, use, modify, +prepare derivative works, incorporate into other computer software, +distribute, and sublicense such enhancements or derivative works thereof, +in binary and source code form. From 24a4e0988b0caa7fbdcbf6ab3cb608b3076d636e Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Mon, 30 Sep 2024 08:24:33 -0700 Subject: [PATCH 17/34] patches to get regression tests running again some benchmarks will be changing --- Exec/regression_inputs/inputs_PSSW | 2 +- Exec/regression_inputs/inputs_compare_subdomain | 2 +- Source/main.cpp | 4 ++-- 3 files changed, 4 insertions(+), 4 deletions(-) diff --git a/Exec/regression_inputs/inputs_PSSW b/Exec/regression_inputs/inputs_PSSW index 8141fff..0f1bcd7 100644 --- a/Exec/regression_inputs/inputs_PSSW +++ b/Exec/regression_inputs/inputs_PSSW @@ -32,7 +32,7 @@ anisotropy_parser(x,y,z) = "(x>-2587.e-9)*(x<2587.e-9)*(y>-2587.e-9)*(y<2587.e-9 DMI_parser(x,y,z) = "0.0" precession = 1 -demag_coupling = 1 +demag_coupling = 0 FFT_solver = 1 M_normalization = 1 # 0 = unsaturated case; 1 = saturated case exchange_coupling = 1 diff --git a/Exec/regression_inputs/inputs_compare_subdomain b/Exec/regression_inputs/inputs_compare_subdomain index 0e826d0..527b929 100644 --- a/Exec/regression_inputs/inputs_compare_subdomain +++ b/Exec/regression_inputs/inputs_compare_subdomain @@ -32,7 +32,7 @@ anisotropy_parser(x,y,z) = "(x>-1.5625e-8)*(x<1.5625e-8)*(y>-15.625e-9)*(y<15.62 DMI_parser(x,y,z) = "0.0" precession = 1 -demag_coupling = 1 +demag_coupling = 0 FFT_solver = 1 M_normalization = 1 # 0 = unsaturated case; 1 = saturated case exchange_coupling = 1 diff --git a/Source/main.cpp b/Source/main.cpp index 0e9162b..c9c429a 100644 --- a/Source/main.cpp +++ b/Source/main.cpp @@ -211,11 +211,11 @@ void main_main () if (demag_coupling == 1) { const Real* dx = geom.CellSize(); - if (dx[0] != dx[1]) { + if (!almostEqual(dx[0], dx[1], 10)) { Abort("Demag requires dx=dy"); } #if (AMREX_SPACEDIM==3) - if (dx[0] != dx[2]) { + if (!almostEqual(dx[0], dx[2], 10)) { Abort("Demag requires dx=dy=dz"); } #endif From afe92a244f3c91d64da29bf16f34b5f3c8ba7dd0 Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Wed, 2 Oct 2024 11:59:32 -0700 Subject: [PATCH 18/34] fix sundials settings --- Exec/inputs_performance | 79 ++++++++++++++++++----------------------- 1 file changed, 35 insertions(+), 44 deletions(-) diff --git a/Exec/inputs_performance b/Exec/inputs_performance index 3512d27..9381c61 100644 --- a/Exec/inputs_performance +++ b/Exec/inputs_performance @@ -43,7 +43,6 @@ anisotropy_axis = 0.0 0.0 1.0 DMI_coupling = 1 # INTEGRATION - TimeIntegratorOption = 1 #Forward Euler #TimeIntegratorOption = 2 #Predictor-corrector #TimeIntegratorOption = 3 #2nd order artemis way @@ -52,24 +51,19 @@ TimeIntegratorOption = 1 #Forward Euler # tolerance threshold (L_inf change between iterations) for TimeIntegrationOption 2 and 3 iterative_tolerance = 1.e-9 -## amrex/sundials backend integrators -## *** Selecting the integrator backend *** -## integration.type can take on the following string or int values: -## (without the quotation marks) -## "ForwardEuler" or "0" = Native Forward Euler Integrator -## "RungeKutta" or "1" = Native Explicit Runge Kutta controlled by integration.rk.type -## "SUNDIALS" or "2" = SUNDIALS Integrators controlled by integration.sundials.strategy +# for TimeIntegratorOption = 4, integration.type can take on the following values: +## 0 or "ForwardEuler" => Native AMReX Forward Euler integrator +## 1 or "RungeKutta" => Native AMReX Explicit Runge Kutta controlled by integration.rk.type +## 2 or "SUNDIALS" => SUNDIALS backend controlled by integration.sundials.type integration.type = RungeKutta -## *** Parameters Needed For Native Explicit Runge-Kutta *** -# -## integration.rk.type can take the following values: +## for integration.type = RungeKutta, integration.rk.type can take the following values: ### 0 = User-specified Butcher Tableau ### 1 = Forward Euler ### 2 = Trapezoid Method ### 3 = SSPRK3 Method ### 4 = RK4 Method -integration.rk.type = 3 +integration.rk.type = 4 ## If using a user-specified Butcher Tableau, then ## set nodes, weights, and table entries here: @@ -81,35 +75,32 @@ integration.rk.weights = 1 integration.rk.nodes = 0 integration.rk.tableau = 0.0 -## *** Parameters Needed For SUNDIALS Integrators *** -## integration.sundials.strategy specifies which ARKODE strategy to use. -## The available options are (without the quotations): -## "ERK" = Explicit Runge Kutta -## "MRI" = Multirate Integrator -## "MRITEST" = Tests the Multirate Integrator by setting a zero-valued fast RHS function -## for example: -integration.sundials.strategy = ERK - -## *** Parameters Specific to SUNDIALS ERK Strategy *** -## (Requires integration.type=SUNDIALS and integration.sundials.strategy=ERK) -## integration.sundials.erk.method specifies which explicit Runge Kutta method -## for SUNDIALS to use. The following options are supported: -## "SSPRK3" = 3rd order strong stability preserving RK (default) -## "Trapezoid" = 2nd order trapezoidal rule -## "ForwardEuler" = 1st order forward euler -## for example: -integration.sundials.erk.method = SSPRK3 - -## *** Parameters Specific to SUNDIALS MRI Strategy *** -## (Requires integration.type=SUNDIALS and integration.sundials.strategy=MRI) -## integration.sundials.mri.implicit_inner specifies whether or not to use an implicit inner solve -## integration.sundials.mri.outer_method specifies which outer (slow) method to use -## integration.sundials.mri.inner_method specifies which inner (fast) method to use -## The following options are supported for both the inner and outer methods: -## "KnothWolke3" = 3rd order Knoth-Wolke method (default for outer method) -## "Trapezoid" = 2nd order trapezoidal rule -## "ForwardEuler" = 1st order forward euler (default for inner method) -## for example: -integration.sundials.mri.implicit_inner = false -integration.sundials.mri.outer_method = KnothWolke3 -integration.sundials.mri.inner_method = Trapezoid +## for integration.type = SUNDIALS, set the SUNDIALS method type: +# ERK = Explicit Runge-Kutta method +# DIRK = Diagonally Implicit Runge-Kutta method +# IMEX-RK = Implicit-Explicit Additive Runge-Kutta method +# EX-MRI = Explicit Multirate Infatesimal method +# IM-MRI = Implicit Multirate Infatesimal method +# IMEX-MRI = Implicit-Explicit Multirate Infatesimal method +integration.sundials.type = EX-MRI + +# *** Select a specific SUNDIALS ERK or DIRK method *** +# https://sundials.readthedocs.io/en/latest/arkode/Butcher_link.html#explicit-butcher-tables +# https://sundials.readthedocs.io/en/latest/arkode/Butcher_link.html#implicit-butcher-tables +#integration.sundials.method = ARKODE_FORWARD_EULER_1_1 +#integration.sundials.method = ARKODE_BACKWARD_EULER_1_1 + +# *** Select a specific SUNDIALS IMEX-RK method *** +# https://sundials.readthedocs.io/en/latest/arkode/Butcher_link.html#additive-butcher-tables +#integration.sundials.method_i = ARKODE_ARK2_DIRK_3_1_2 +#integration.sundials.method_e = ARKODE_ARK2_ERK_3_1_2 + +# *** Select a specific SUNDIALS EX-MRI, IM-MRI, or IMEX-MRI method *** +# https://sundials.readthedocs.io/en/latest/arkode/Usage/MRIStep/MRIStepCoupling.html#mri-coupling-tables +# *** Select a specific SUNDIALS MRI slow method coupling table *** +integration.sundials.method = ARKODE_MIS_KW3 +# *** Select a specific SUNDIALS ERK or DIRK fast method *** +integration.sundials.fast_type = ERK # ERK or DIRK +integration.sundials.fast_method = ARKODE_KNOTH_WOLKE_3_3 + +fast_exchange = 1 From 52cf85cc2874758e19f61b0aab604bf5bedf087f Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Fri, 1 Nov 2024 10:09:28 -0700 Subject: [PATCH 19/34] cleanup per Weiqun --- Exec/GNUmakefile | 12 +----------- Source/Demagnetization.cpp | 2 +- 2 files changed, 2 insertions(+), 12 deletions(-) diff --git a/Exec/GNUmakefile b/Exec/GNUmakefile index 8cbbc1e..6c9ceb2 100644 --- a/Exec/GNUmakefile +++ b/Exec/GNUmakefile @@ -8,6 +8,7 @@ USE_CUDA = FALSE USE_HIP = FALSE COMP = gnu DIM = 3 +USE_FFT = TRUE USE_SUNDIALS = FALSE @@ -47,17 +48,6 @@ ifeq ($(USE_SUNDIALS),TRUE) endif -ifeq ($(USE_CUDA),TRUE) - libraries += -lcufft -else ifeq ($(USE_HIP),TRUE) - # Use rocFFT. ROC_PATH is defined in amrex - INCLUDE_LOCATIONS += $(ROC_PATH)/rocfft/include - LIBRARY_LOCATIONS += $(ROC_PATH)/rocfft/lib - LIBRARIES += -L$(ROC_PATH)/rocfft/lib -lrocfft -else - libraries += -lfftw3_mpi -lfftw3f -lfftw3 -endif - include $(AMREX_HOME)/Src/Base/Make.package include $(AMREX_HOME)/Src/FFT/Make.package include $(AMREX_HOME)/Src/Boundary/Make.package diff --git a/Source/Demagnetization.cpp b/Source/Demagnetization.cpp index c495a34..d24a6ff 100644 --- a/Source/Demagnetization.cpp +++ b/Source/Demagnetization.cpp @@ -60,7 +60,7 @@ void Demagnetization::define() Geometry cgeom_large(cdomain_large, real_box_large, CoordSys::cartesian, is_periodic); auto cba_large = amrex::decompose(cdomain_large, ParallelContext::NProcsSub(), {AMREX_D_DECL(true,true,false)}); - DistributionMapping cdm_large = amrex::FFT::detail::make_iota_distromap(cba_large.size()); + DistributionMapping cdm_large(cba_large); Kxx_fft.define(cba_large, cdm_large, 1, 0); Kxy_fft.define(cba_large, cdm_large, 1, 0); From 97d152a61fe686babbe6f0c4bf45bf4324d0332d Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Mon, 10 Mar 2025 21:25:49 -0700 Subject: [PATCH 20/34] logic for different physics wasn't right for single-rate case --- Source/main.cpp | 48 ++++++++++++++++++++++++++++-------------------- 1 file changed, 28 insertions(+), 20 deletions(-) diff --git a/Source/main.cpp b/Source/main.cpp index 7ad879b..4893c33 100644 --- a/Source/main.cpp +++ b/Source/main.cpp @@ -558,39 +558,47 @@ void main_main () } // exchange - if ( (using_MRI==1 && fast_exchange) || (using_IMEX==1 && implicit_exchange) ) { - for (int idim=0; idim Date: Tue, 29 Jul 2025 13:56:49 -0700 Subject: [PATCH 21/34] add SUNDIALS_HOME --- Exec/GNUmakefile | 31 +------------------------------ 1 file changed, 1 insertion(+), 30 deletions(-) diff --git a/Exec/GNUmakefile b/Exec/GNUmakefile index 6c9ceb2..0379b6f 100644 --- a/Exec/GNUmakefile +++ b/Exec/GNUmakefile @@ -11,6 +11,7 @@ DIM = 3 USE_FFT = TRUE USE_SUNDIALS = FALSE +SUNDIALS_HOME ?= ../../sundials/instdir include $(AMREX_HOME)/Tools/GNUMake/Make.defs @@ -18,36 +19,6 @@ include ../Source/Make.package VPATH_LOCATIONS += ../Source INCLUDE_LOCATIONS += ../Source -ifeq ($(USE_SUNDIALS),TRUE) - ifeq ($(USE_CUDA),TRUE) - SUNDIALS_ROOT ?= $(TOP)../../sundials/instdir_cuda - else - SUNDIALS_ROOT ?= $(TOP)../../sundials/instdir - endif - ifeq ($(NERSC_HOST),perlmutter) - SUNDIALS_LIB_DIR ?= $(SUNDIALS_ROOT)/lib64 - else - SUNDIALS_LIB_DIR ?= $(SUNDIALS_ROOT)/lib - endif - - USE_CVODE_LIBS ?= TRUE - USE_ARKODE_LIBS ?= TRUE - - DEFINES += -DAMREX_USE_SUNDIALS - INCLUDE_LOCATIONS += $(SUNDIALS_ROOT)/include - LIBRARY_LOCATIONS += $(SUNDIALS_LIB_DIR) - - LIBRARIES += -L$(SUNDIALS_LIB_DIR) -lsundials_cvode - LIBRARIES += -L$(SUNDIALS_LIB_DIR) -lsundials_arkode - LIBRARIES += -L$(SUNDIALS_LIB_DIR) -lsundials_nvecmanyvector - LIBRARIES += -L$(SUNDIALS_LIB_DIR) -lsundials_core - - ifeq ($(USE_CUDA),TRUE) - LIBRARIES += -L$(SUNDIALS_LIB_DIR) -lsundials_nveccuda - endif - -endif - include $(AMREX_HOME)/Src/Base/Make.package include $(AMREX_HOME)/Src/FFT/Make.package include $(AMREX_HOME)/Src/Boundary/Make.package From b95ceba2359aa2d9f47bb741cf3b7b698738964b Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Wed, 30 Jul 2025 09:56:08 -0700 Subject: [PATCH 22/34] sundials update --- Exec/README_sundials | 96 ++------------------------------------------ 1 file changed, 3 insertions(+), 93 deletions(-) diff --git a/Exec/README_sundials b/Exec/README_sundials index 3fb088e..3b6830a 100644 --- a/Exec/README_sundials +++ b/Exec/README_sundials @@ -1,94 +1,4 @@ -SUNDIALS installation guide: -https://computing.llnl.gov/projects/sundials/faq#inst +Refer to https://amrex-codes.github.io/amrex/docs_html/TimeIntegration_Chapter.html +for SNUDIALS installation, build, and usage instructions. -Installation - -# You need SUNDIALS v7.1.1 or later. -# Check https://computing.llnl.gov/projects/sundials/sundials-software to see if it's available for download. -# If so, download sundials-x.y.z.tar.gz and extract it at the same level as amrex using ->> tar -xzvf sundials-x.y.z.tar.gz # where x.y.z is the version of sundials ->> mv sundials-x.y.z sundials-src - -# If v7.1.1. is not available on the website, clone the git repo directly and use the latest version -# At the same level that amrex is cloned, do: - ->> git clone https://github.com/LLNL/sundials.git ->> mv sundials sundials-src - -# Next - ->> mkdir sundials ->> cd sundials - -###################### -HOST BUILD -###################### - ->> mkdir instdir ->> mkdir builddir ->> cd builddir - ->> cmake -DCMAKE_INSTALL_PREFIX=/pathto/sundials/instdir -DEXAMPLES_INSTALL_PATH=/pathto/sundials/instdir/examples -DENABLE_MPI=ON ../../sundials-src - -###################### -NVIDIA/CUDA BUILD -###################### - -# Navigate back to the 'sundials' directory and do: - ->> mkdir instdir_cuda ->> mkdir builddir_cuda ->> cd builddir_cuda - ->> cmake -DCMAKE_INSTALL_PREFIX=/pathto/sundials/instdir_cuda -DEXAMPLES_INSTALL_PATH=/pathto/sundials/instdir_cuda/examples -DENABLE_CUDA=ON -DENABLE_MPI=ON ../../sundials-src - -###################### - ->> make -j4 ->> make install - -# in your .bashrc or preferred configuration file, add the following (and then "source ~/.bashrc") - -# If you have a CPU build: - ->> export LD_LIBRARY_PATH=$LD_LIBRARY_PATH:/pathto/sundials/instdir/lib/ - -# If you have a NVIDIA/CUDA build: - ->> export LD_LIBRARY_PATH=$LD_LIBRARY_PATH:/pathto/sundials/instdir_cuda/lib/ - -# now you are ready to compile MagneX with: - ->> make -j4 USE_SUNDIALS=TRUE # optional to have 'USE_CUDA=TRUE' as well - -# in your inputs file, you will need to have: - -TimeIntegratorOption = 4 #amrex/sundials backend integrators - -# INTEGRATION -## integration.type can take on the following values: -## 0 or "ForwardEuler" => Native AMReX Forward Euler integrator -## 1 or "RungeKutta" => Native AMReX Explicit Runge Kutta controlled by integration.rk.type -## 2 or "SUNDIALS" => SUNDIALS backend controlled by integration.sundials.strategy -integration.type = SUNDIALS - -## Native AMReX Explicit Runge-Kutta parameters -# -## integration.rk.type can take the following values: -### 0 = User-specified Butcher Tableau -### 1 = Forward Euler -### 2 = Trapezoid Method -### 3 = SSPRK3 Method -### 4 = RK4 Method -integration.rk.type = 1 - -# Set the SUNDIALS method type: -# ERK = Explicit Runge-Kutta method -# DIRK = Diagonally Implicit Runge-Kutta method -# -# Optionally select a specific SUNDIALS method by name, see the SUNDIALS -# documentation for the supported method names - -# Use forward Euler (fixed step sizes only) -integration.sundials.type = ERK -integration.sundials.method = ARKODE_FORWARD_EULER_1_1 +Make sure that SUNDIALS_HOME in the GNUmakefile points to the installation directory. From a689f78c3ee836dfc38f1ee85b31a076019de56b Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Fri, 2 Jan 2026 08:30:31 -0800 Subject: [PATCH 23/34] comments in inputs file --- Exec/standard_problem_inputs/inputs_std4_IMEX | 120 ++++++++++++++++++ Exec/standard_problem_inputs/inputs_std4_MRI | 2 +- Exec/standard_problem_inputs/inputs_std4_RK4 | 4 +- 3 files changed, 123 insertions(+), 3 deletions(-) create mode 100644 Exec/standard_problem_inputs/inputs_std4_IMEX diff --git a/Exec/standard_problem_inputs/inputs_std4_IMEX b/Exec/standard_problem_inputs/inputs_std4_IMEX new file mode 100644 index 0000000..c0cbe9f --- /dev/null +++ b/Exec/standard_problem_inputs/inputs_std4_IMEX @@ -0,0 +1,120 @@ +amrex.use_gpu_aware_mpi=1 + +############## +# 0.78125nm case +############## +n_cell = 640 160 4 +max_grid_size_x = 640 +max_grid_size_y = 160 +max_grid_size_z = 4 +dt = 5.e-15 # IMEX, 1.e-14 unstable + +stop_time = 4.e-9 + +plot_int = 25000 +chk_int = -1 +restart = -1 +diag_type = 4 + +prob_lo = 0. 0. 0. +prob_hi = 500.e-9 125.e-9 3.125e-9 + +mu0 = 1.25663e-6 + +Mx_parser(x,y,z) = "8.e5 * (x>0.)*(x<500.e-9)*(y>0.)*(y<125.e-9)*(z>0.)*(z<3.125e-9)" +My_parser(x,y,z) = "0." +Mz_parser(x,y,z) = "0." + +# Field 1: mu_0 Hx=-24.6 mT, mu_0 Hy= 4.3 mT, mu_0 Hz= 0.0 mT +# which is a field approximately 25 mT, directed 170 degrees counterclockwise from the positive x axis +timedependent_Hbias = 1 +Hx_bias_parser(x,y,z,t) = "(t<=2.e-11)*1.e5 + (t>2.e-11)*(t<=3.e-11)*(3.e-11-t)*1.e16 - (t>1.e-9)*19576" +Hy_bias_parser(x,y,z,t) = "(t<=2.e-11)*1.e5 + (t>2.e-11)*(t<=3.e-11)*(3.e-11-t)*1.e16 + (t>1.e-9)*3422" +Hz_bias_parser(x,y,z,t) = "(t<=2.e-11)*1.e5 + (t>2.e-11)*(t<=3.e-11)*(3.e-11-t)*1.e16" + +# Field 2: mu_0 Hx=-35.5 mT, mu_0 Hy=-6.3 mT, mu_0 Hz= 0.0 mT +# which is a field approximately 36 mT, directed 190 degrees counterclockwise from the positive x axis +#timedependent_Hbias = 1 +#Hx_bias_parser(x,y,z,t) = "(t<=2.e-11)*1.e5 + (t>2.e-11)*(t<=3.e-11)*(3.e-11-t)*1.e16 - (t>1.e-9)*28259" +#Hy_bias_parser(x,y,z,t) = "(t<=2.e-11)*1.e5 + (t>2.e-11)*(t<=3.e-11)*(3.e-11-t)*1.e16 - (t>1.e-9)*5013" +#Hz_bias_parser(x,y,z,t) = "(t<=2.e-11)*1.e5 + (t>2.e-11)*(t<=3.e-11)*(3.e-11-t)*1.e16" + +timedependent_alpha = 1 +alpha_parser(x,y,z,t) = " (x>0.)*(x<500.e-9)*(y>0.)*(y<125.e-9)*(z>0.)*(z<3.125e-9) * (0.5*(t<=1.e-9) + 0.02*(t>1.e-9))" +Ms_parser(x,y,z) = " (x>0.)*(x<500.e-9)*(y>0.)*(y<125.e-9)*(z>0.)*(z<3.125e-9) * 8.e5" +gamma_parser(x,y,z) = " (x>0.)*(x<500.e-9)*(y>0.)*(y<125.e-9)*(z>0.)*(z<3.125e-9) * -1.759e11" +exchange_parser(x,y,z) = " (x>0.)*(x<500.e-9)*(y>0.)*(y<125.e-9)*(z>0.)*(z<3.125e-9)* 1.3e-11" +anisotropy_parser(x,y,z) = "0." +DMI_parser(x,y,z) = "0." + +precession = 1 +demag_coupling = 1 +FFT_solver = 1 +M_normalization = 1 # 0 = unsaturated case; 1 = saturated case +exchange_coupling = 1 +anisotropy_coupling = 0 +anisotropy_axis = 0.0 1.0 0.0 +DMI_coupling = 0 + +# INTEGRATION +#TimeIntegratorOption = 1 #Forward Euler +#TimeIntegratorOption = 2 #Predictor-corrector +#TimeIntegratorOption = 3 #2nd order artemis way +TimeIntegratorOption = 4 #amrex/sundials backend integrators + +# tolerance threshold (L_inf change between iterations) for TimeIntegrationOption 2 and 3 +iterative_tolerance = 1.e-9 + +# for TimeIntegratorOption = 4, integration.type can take on the following values: +## 0 or "ForwardEuler" => Native AMReX Forward Euler integrator +## 1 or "RungeKutta" => Native AMReX Explicit Runge Kutta controlled by integration.rk.type +## 2 or "SUNDIALS" => SUNDIALS backend controlled by integration.sundials.type +integration.type = SUNDIALS + +## for integration.type = RungeKutta, integration.rk.type can take the following values: +### 0 = User-specified Butcher Tableau +### 1 = Forward Euler +### 2 = Trapezoid Method +### 3 = SSPRK3 Method +### 4 = RK4 Method +integration.rk.type = 4 + +## If using a user-specified Butcher Tableau, then +## set nodes, weights, and table entries here: +# +## The Butcher Tableau is read as a flattened, +## lower triangular matrix (but including the diagonal) +## in row major format. +integration.rk.weights = 1 +integration.rk.nodes = 0 +integration.rk.tableau = 0.0 + +## for integration.type = SUNDIALS, set the SUNDIALS method type: +# ERK = Explicit Runge-Kutta method +# DIRK = Diagonally Implicit Runge-Kutta method +# IMEX-RK = Implicit-Explicit Additive Runge-Kutta method +# EX-MRI = Explicit Multirate Infatesimal method +# IM-MRI = Implicit Multirate Infatesimal method +# IMEX-MRI = Implicit-Explicit Multirate Infatesimal method +integration.sundials.type = IMEX-RK + +# *** Select a specific SUNDIALS ERK or DIRK method *** +# https://sundials.readthedocs.io/en/latest/arkode/Butcher_link.html#explicit-butcher-tables +# https://sundials.readthedocs.io/en/latest/arkode/Butcher_link.html#implicit-butcher-tables +#integration.sundials.method = ARKODE_FORWARD_EULER_1_1 +#integration.sundials.method = ARKODE_BACKWARD_EULER_1_1 + +# *** Select a specific SUNDIALS IMEX-RK method *** +# https://sundials.readthedocs.io/en/latest/arkode/Butcher_link.html#additive-butcher-tables +integration.sundials.method_i = ARKODE_ARK2_DIRK_3_1_2 +integration.sundials.method_e = ARKODE_ARK2_ERK_3_1_2 + +# *** Select a specific SUNDIALS EX-MRI, IM-MRI, or IMEX-MRI method *** +# https://sundials.readthedocs.io/en/latest/arkode/Usage/MRIStep/MRIStepCoupling.html#mri-coupling-tables +# *** Select a specific SUNDIALS MRI slow method coupling table *** +#integration.sundials.method = ARKODE_MIS_KW3 +# *** Select a specific SUNDIALS ERK or DIRK fast method *** +#integration.sundials.fast_type = ERK # ERK or DIRK +#integration.sundials.fast_method = ARKODE_KNOTH_WOLKE_3_3 + +implicit_exchange = 1 diff --git a/Exec/standard_problem_inputs/inputs_std4_MRI b/Exec/standard_problem_inputs/inputs_std4_MRI index 9b15ce3..b9770da 100644 --- a/Exec/standard_problem_inputs/inputs_std4_MRI +++ b/Exec/standard_problem_inputs/inputs_std4_MRI @@ -7,7 +7,7 @@ n_cell = 640 160 4 max_grid_size_x = 640 max_grid_size_y = 160 max_grid_size_z = 4 -dt = 1.25e-13 # SUNDIALS MRI +dt = 1.25e-13 # SUNDIALS MRI, 2.5e-13 unstable stop_time = 4.e-9 diff --git a/Exec/standard_problem_inputs/inputs_std4_RK4 b/Exec/standard_problem_inputs/inputs_std4_RK4 index becd5c2..13b08db 100644 --- a/Exec/standard_problem_inputs/inputs_std4_RK4 +++ b/Exec/standard_problem_inputs/inputs_std4_RK4 @@ -7,11 +7,11 @@ n_cell = 640 160 4 max_grid_size_x = 640 max_grid_size_y = 160 max_grid_size_z = 4 -dt = 2.5e-14 # RK4 (438s) 5.e-14 unstable +dt = 2.5e-14 # RK4, 5.e-14 unstable stop_time = 4.e-9 -plot_int = 10000 +plot_int = 5000 chk_int = -1 restart = -1 diag_type = 4 From e2cba04ae0365f066608fe8a1528cdfa24dc1eca Mon Sep 17 00:00:00 2001 From: tyh0123 Date: Mon, 5 Jan 2026 22:47:01 -0800 Subject: [PATCH 24/34] initial commit --- Exec/GNUmakefile | 28 +++++++++- Source/Demag_ml.cpp | 112 ++++++++++++++++++++++++++++++++++++++ Source/MagneX.H | 9 +++ Source/MagneX.cpp | 7 +++ Source/MagneX_namespace.H | 2 + Source/Make.package | 1 + Source/main.cpp | 68 +++++++++++++++++++++-- 7 files changed, 221 insertions(+), 6 deletions(-) create mode 100644 Source/Demag_ml.cpp diff --git a/Exec/GNUmakefile b/Exec/GNUmakefile index 6c9ceb2..6b84926 100644 --- a/Exec/GNUmakefile +++ b/Exec/GNUmakefile @@ -1,5 +1,5 @@ # AMREX_HOME defines the directory in which we will find all the AMReX code. -AMREX_HOME ?= ../../amrex +AMREX_HOME = ../../amrex DEBUG = FALSE USE_MPI = TRUE @@ -12,6 +12,32 @@ USE_FFT = TRUE USE_SUNDIALS = FALSE + +# Pytorch directories +ifeq ($(USE_CUDA),TRUE) + PYTORCH_ROOT := ../../libtorch_cuda +else + PYTORCH_ROOT := ../../libtorch_cpu +endif +TORCH_LIBPATH = $(PYTORCH_ROOT)/lib + +ifeq ($(USE_CUDA),TRUE) + TORCH_LIBS = -ltorch -ltorch_cpu -lc10 -lc10_cuda -lcuda +else + TORCH_LIBS = -ltorch -ltorch_cpu -lc10 +endif + +INCLUDE_LOCATIONS += $(PYTORCH_ROOT)/include \ + $(PYTORCH_ROOT)/include/torch/csrc/api/include +LIBRARY_LOCATIONS += $(TORCH_LIBPATH) + +DEFINES += -D_GLIBCXX_USE_CXX11_ABI=1 +ifeq ($(USE_CUDA),TRUE) + LDFLAGS += -Xlinker "--no-as-needed,-rpath $(TORCH_LIBPATH) $(TORCH_LIBS)" +else + LDFLAGS += -Wl,--no-as-needed,-rpath=$(TORCH_LIBPATH) $(TORCH_LIBS) +endif + include $(AMREX_HOME)/Tools/GNUMake/Make.defs include ../Source/Make.package diff --git a/Source/Demag_ml.cpp b/Source/Demag_ml.cpp new file mode 100644 index 0000000..3502a07 --- /dev/null +++ b/Source/Demag_ml.cpp @@ -0,0 +1,112 @@ +#include "MagneX.H" +#include + +using namespace amrex; + +void CalculateH_demag_ML(const Array& Mfield, + torch::jit::script::Module& x_norm_module, + torch::jit::script::Module& ml_module, + torch::jit::script::Module& y_norm_module, + Array& H_demagfield) +{ + BL_PROFILE_VAR("CalculateH_demag_ML()", CalculateH_demag_ML); + + for (MFIter mfi(Mfield[0], TilingIfNotGPU()); mfi.isValid(); ++mfi) { + + const Box& bx = mfi.validbox(); + + const auto& Mx = Mfield[0].const_array(mfi); + const auto& My = Mfield[1].const_array(mfi); + const auto& Mz = Mfield[2].const_array(mfi); + + auto Hx_demag = H_demagfield[0].array(mfi); + auto Hy_demag = H_demagfield[1].array(mfi); + auto Hz_demag = H_demagfield[2].array(mfi); + + const IntVect bx_lo = bx.smallEnd(); + const IntVect nbox = bx.size(); + +#if AMREX_SPACEDIM == 2 + const int ncell = nbox[0] * nbox[1]; +#else + const int ncell = nbox[0] * nbox[1] * nbox[2]; +#endif + + // Host-visible (Managed) buffers filled on GPU + amrex::Gpu::ManagedVector aux_Mx(ncell); + amrex::Gpu::ManagedVector aux_My(ncell); + amrex::Gpu::ManagedVector aux_Mz(ncell); + + Real* AMREX_RESTRICT auxPtr_Mx = aux_Mx.dataPtr(); + Real* AMREX_RESTRICT auxPtr_My = aux_My.dataPtr(); + Real* AMREX_RESTRICT auxPtr_Mz = aux_Mz.dataPtr(); + + // Fill aux buffers from MultiFab on GPU + amrex::ParallelFor(bx, [=] AMREX_GPU_DEVICE(int i, int j, int k) noexcept { + const int ii = i - bx_lo[0]; + const int jj = j - bx_lo[1]; + +#if AMREX_SPACEDIM == 2 + const int index = jj + ii * nbox[1]; +#else + const int kk = k - bx_lo[2]; + const int index = kk + jj * nbox[2] + ii * nbox[2] * nbox[1]; +#endif + + auxPtr_Mx[index] = Mx(i, j, k); + auxPtr_My[index] = My(i, j, k); + auxPtr_Mz[index] = Mz(i, j, k); + }); + + // Make sure aux buffers are ready for from_blob on host side + amrex::Gpu::streamSynchronize(); + + // Wrap buffers as CPU tensors (no copy) + at::Tensor inputs_torch_Mx = torch::from_blob(auxPtr_Mx, {ncell, 1}, torch::kFloat64); + at::Tensor inputs_torch_My = torch::from_blob(auxPtr_My, {ncell, 1}, torch::kFloat64); + at::Tensor inputs_torch_Mz = torch::from_blob(auxPtr_Mz, {ncell, 1}, torch::kFloat64); + + // Reshape (assumes ncell == 128*128*4) + at::Tensor reshaped_Mx = inputs_torch_Mx.reshape({128, 128, 4}); + at::Tensor reshaped_My = inputs_torch_My.reshape({128, 128, 4}); + at::Tensor reshaped_Mz = inputs_torch_Mz.reshape({128, 128, 4}); + + // Stack into [3, 128, 128, 4] then add batch dim -> [1, 3, 128, 128, 4] + at::Tensor final_tensor_M = torch::stack({reshaped_Mx, reshaped_My, reshaped_Mz}, 0); + final_tensor_M = final_tensor_M.to(torch::kCUDA).to(torch::kFloat32); + final_tensor_M = final_tensor_M.unsqueeze(0); + + // Normalize -> model -> denormalize + at::Tensor norm_torch = x_norm_module.get_method("encode")({final_tensor_M}).toTensor(); + at::Tensor outputs_torch = ml_module.forward({norm_torch}).toTensor(); + at::Tensor denorm_torch = y_norm_module.get_method("decode")({outputs_torch}).toTensor(); + + // Convert to float64 for accessor usage (same as your original) + denorm_torch = denorm_torch.to(torch::kFloat64); + + // Extract H components: denorm_torch shape assumed [1, 3, 128, 128, 4] + at::Tensor denorm_torch_Hx = denorm_torch.select(0, 0).select(0, 0).flatten(); + at::Tensor denorm_torch_Hy = denorm_torch.select(0, 0).select(0, 1).flatten(); + at::Tensor denorm_torch_Hz = denorm_torch.select(0, 0).select(0, 2).flatten(); + +#ifdef AMREX_USE_CUDA + auto denorm_torch_Hx_acc = denorm_torch_Hx.packed_accessor64(); + auto denorm_torch_Hy_acc = denorm_torch_Hy.packed_accessor64(); + auto denorm_torch_Hz_acc = denorm_torch_Hz.packed_accessor64(); +#endif + + // Copy tensor data back into demag field + amrex::ParallelFor(bx, [=] AMREX_GPU_DEVICE(int i, int j, int k) noexcept { + const int ii = i - bx_lo[0]; + const int jj = j - bx_lo[1]; + const int kk = k - bx_lo[2]; + const int index = kk + jj * nbox[2] + ii * nbox[2] * nbox[1]; + + Hx_demag(i, j, k) = denorm_torch_Hx_acc[index]; + Hy_demag(i, j, k) = denorm_torch_Hy_acc[index]; + Hz_demag(i, j, k) = denorm_torch_Hz_acc[index]; + }); + + amrex::Gpu::streamSynchronize(); + } +} \ No newline at end of file diff --git a/Source/MagneX.H b/Source/MagneX.H index b741dd2..6d7d4ba 100644 --- a/Source/MagneX.H +++ b/Source/MagneX.H @@ -1,5 +1,6 @@ #ifdef AMREX_USE_CUDA #include +#include #else #include #ifdef AMREX_USE_MPI @@ -184,3 +185,11 @@ void WritePlotfile(MultiFab& Ms, const Geometry& geom, const Real& time, const int& plt_step); + + +void CalculateH_demag_ML(const Array< MultiFab, AMREX_SPACEDIM> & Mfield, + torch::jit::script::Module& x_norm_module, + torch::jit::script::Module& ml_module, + torch::jit::script::Module& y_norm_module, + Array< MultiFab, AMREX_SPACEDIM> & H_demagfield); + \ No newline at end of file diff --git a/Source/MagneX.cpp b/Source/MagneX.cpp index ba35f1c..d7e0ec2 100644 --- a/Source/MagneX.cpp +++ b/Source/MagneX.cpp @@ -132,6 +132,10 @@ AMREX_GPU_MANAGED int MagneX::demag_coupling; // 0 = FFTW (single-MPI), 1 = heFFTe (distributed) AMREX_GPU_MANAGED int MagneX::FFT_solver; +// ML flag +int MagneX::ml_enable; + + void InitializeMagneXNamespace() { BL_PROFILE_VAR("InitializeMagneXNamespace()",InitializeMagneXNameSpace); @@ -228,6 +232,9 @@ void InitializeMagneXNamespace() { restart = -1; pp.query("restart",restart); + ml_enable = 0; + pp.query("ml_enable",ml_enable); + diag_type = -1; pp.query("diag_type",diag_type); diff --git a/Source/MagneX_namespace.H b/Source/MagneX_namespace.H index 0660dce..18bfe82 100644 --- a/Source/MagneX_namespace.H +++ b/Source/MagneX_namespace.H @@ -52,6 +52,8 @@ namespace MagneX { extern int diag_type; + extern int ml_enable; + extern int timedependent_Hbias; extern int timedependent_alpha; diff --git a/Source/Make.package b/Source/Make.package index f99179f..4277239 100644 --- a/Source/Make.package +++ b/Source/Make.package @@ -15,3 +15,4 @@ CEXE_headers += CartesianAlgorithm_K.H CEXE_headers += Demagnetization.H CEXE_headers += MagneX.H CEXE_headers += MagneX_namespace.H +CEXE_sources += Demag_ml.cpp \ No newline at end of file diff --git a/Source/main.cpp b/Source/main.cpp index a2263c1..f54d182 100644 --- a/Source/main.cpp +++ b/Source/main.cpp @@ -1,15 +1,17 @@ #include "MagneX.H" #include "Demagnetization.H" - +#include #include #include - +#include #ifdef AMREX_USE_SUNDIALS #include #endif #include +#include // for at::cuda::setDevice +#include using namespace amrex; using namespace MagneX; @@ -57,6 +59,10 @@ void main_main () Array LLG_RHS; Array LLG_RHS_pre; Array LLG_RHS_avg; + torch::jit::script::Module ml_module; + torch::jit::script::Module x_norm_module; + torch::jit::script::Module y_norm_module; + // Declare variables for hysteresis Real normalized_Mx; @@ -100,6 +106,42 @@ void main_main () } + // ********************************** + // // LOAD PYTORCH MODEL + + BL_PROFILE_VAR("LoadPytorch",LoadPytorch); + + // Load pytorch module via torch script + + + std::string ml_model_name; + std::string x_normalizer_name; + std::string y_normalizer_name; + + ParmParse pp_ml; + pp_ml.query("ml_model_name", ml_model_name); + pp_ml.query("x_normalizer_name", x_normalizer_name); + pp_ml.query("y_normalizer_name", y_normalizer_name); + + amrex::Print()<<"\n"< Date: Mon, 23 Feb 2026 22:53:03 -0800 Subject: [PATCH 25/34] updated ML part of the MagneX by adding: 1. USE_ML flag for during compile. 2. Added 'ml_enable' flag in the input file to control inference with ML or not --- Exec/GNUmakefile | 79 +++++++--- Source/Demag_ml.cpp | 367 ++++++++++++++++++++++++++++++++++---------- Source/MagneX.H | 45 +++++- Source/main.cpp | 133 ++++++++++++---- 4 files changed, 492 insertions(+), 132 deletions(-) diff --git a/Exec/GNUmakefile b/Exec/GNUmakefile index 6b84926..8436003 100644 --- a/Exec/GNUmakefile +++ b/Exec/GNUmakefile @@ -9,35 +9,72 @@ USE_HIP = FALSE COMP = gnu DIM = 3 USE_FFT = TRUE - +USE_ML = FALSE USE_SUNDIALS = FALSE +TINY_PROFILE = FALSE +PROFILE = FALSE +ifeq ($(USE_ML),TRUE) + # Define a macro for the C++ preprocessor + DEFINES += -DML_ENABLE -D_GLIBCXX_USE_CXX11_ABI=1 -# Pytorch directories -ifeq ($(USE_CUDA),TRUE) - PYTORCH_ROOT := ../../libtorch_cuda -else - PYTORCH_ROOT := ../../libtorch_cpu -endif -TORCH_LIBPATH = $(PYTORCH_ROOT)/lib + # Pytorch root directory selection + ifeq ($(USE_CUDA),TRUE) + PYTORCH_ROOT := ../../libtorch_cuda + else + PYTORCH_ROOT := ../../libtorch_cpu + endif + + TORCH_LIBPATH = $(PYTORCH_ROOT)/lib -ifeq ($(USE_CUDA),TRUE) - TORCH_LIBS = -ltorch -ltorch_cpu -lc10 -lc10_cuda -lcuda -else - TORCH_LIBS = -ltorch -ltorch_cpu -lc10 -endif + # Library definitions + ifeq ($(USE_CUDA),TRUE) + # Note: Modern LibTorch often requires both torch_cuda and torch_cpu + TORCH_LIBS = -ltorch -ltorch_cuda -ltorch_cpu -lc10 -lc10_cuda -lcuda + else + TORCH_LIBS = -ltorch -ltorch_cpu -lc10 + endif -INCLUDE_LOCATIONS += $(PYTORCH_ROOT)/include \ - $(PYTORCH_ROOT)/include/torch/csrc/api/include -LIBRARY_LOCATIONS += $(TORCH_LIBPATH) + # Header search paths + INCLUDE_LOCATIONS += $(PYTORCH_ROOT)/include \ + $(PYTORCH_ROOT)/include/torch/csrc/api/include + + # Library search paths + LIBRARY_LOCATIONS += $(TORCH_LIBPATH) -DEFINES += -D_GLIBCXX_USE_CXX11_ABI=1 -ifeq ($(USE_CUDA),TRUE) - LDFLAGS += -Xlinker "--no-as-needed,-rpath $(TORCH_LIBPATH) $(TORCH_LIBS)" -else - LDFLAGS += -Wl,--no-as-needed,-rpath=$(TORCH_LIBPATH) $(TORCH_LIBS) + # Linker flags (rpath ensures the .so files are found at runtime) + ifeq ($(USE_CUDA),TRUE) + LDFLAGS += -Xlinker "--no-as-needed,-rpath,$(TORCH_LIBPATH)" $(TORCH_LIBS) + else + LDFLAGS += -Wl,--no-as-needed,-rpath=$(TORCH_LIBPATH) $(TORCH_LIBS) + endif endif +# # Pytorch directories +# ifeq ($(USE_CUDA),TRUE) +# PYTORCH_ROOT := ../../libtorch_cuda +# else +# PYTORCH_ROOT := ../../libtorch_cpu +# endif +# TORCH_LIBPATH = $(PYTORCH_ROOT)/lib + +# ifeq ($(USE_CUDA),TRUE) +# TORCH_LIBS = -ltorch -ltorch_cpu -lc10 -lc10_cuda -lcuda +# else +# TORCH_LIBS = -ltorch -ltorch_cpu -lc10 +# endif + +# INCLUDE_LOCATIONS += $(PYTORCH_ROOT)/include \ +# $(PYTORCH_ROOT)/include/torch/csrc/api/include +# LIBRARY_LOCATIONS += $(TORCH_LIBPATH) + +# DEFINES += -D_GLIBCXX_USE_CXX11_ABI=1 +# ifeq ($(USE_CUDA),TRUE) +# LDFLAGS += -Xlinker "--no-as-needed,-rpath $(TORCH_LIBPATH) $(TORCH_LIBS)" +# else +# LDFLAGS += -Wl,--no-as-needed,-rpath=$(TORCH_LIBPATH) $(TORCH_LIBS) +# endif + include $(AMREX_HOME)/Tools/GNUMake/Make.defs include ../Source/Make.package diff --git a/Source/Demag_ml.cpp b/Source/Demag_ml.cpp index 3502a07..e2cb215 100644 --- a/Source/Demag_ml.cpp +++ b/Source/Demag_ml.cpp @@ -1,112 +1,321 @@ +// MagneX_ML_Infer_Dynamic.cpp #include "MagneX.H" #include - +#include +// #include +#include +#include +#include using namespace amrex; -void CalculateH_demag_ML(const Array& Mfield, - torch::jit::script::Module& x_norm_module, - torch::jit::script::Module& ml_module, - torch::jit::script::Module& y_norm_module, - Array& H_demagfield) +// ------------------------------------------------------------ +// Helper: Move all parameters + buffers of a TorchScript module +// to the same device (e.g., cuda:3) to avoid device mismatch. +// ------------------------------------------------------------ +void MoveModuleToDevice(torch::jit::script::Module& m, + const torch::Device& device) { - BL_PROFILE_VAR("CalculateH_demag_ML()", CalculateH_demag_ML); + m.to(device); +// for (auto& p : m.named_parameters(true)) { +// p.value().set_data(p.value().to(device)); +// } +// for (auto& b : m.named_buffers(true)) { +// b.value().set_data(b.value().to(device)); +// } +} - for (MFIter mfi(Mfield[0], TilingIfNotGPU()); mfi.isValid(); ++mfi) { +// ------------------------------------------------------------ +// Helper: read expected_spatial = [nx, ny, nz] from normalizer. +// normalizer must have exported method get_expected_spatial(). +// ------------------------------------------------------------ +amrex::IntVect GetExpectedSpatial(torch::jit::script::Module& x_norm_module) +{ + at::Tensor t = x_norm_module.get_method("get_expected_spatial")({}).toTensor(); + t = t.to(torch::kCPU).to(torch::kLong).contiguous(); - const Box& bx = mfi.validbox(); + AMREX_ALWAYS_ASSERT_WITH_MESSAGE(t.numel() == 3, + "get_expected_spatial() must return a tensor of 3 elements: [nx, ny, nz]"); - const auto& Mx = Mfield[0].const_array(mfi); - const auto& My = Mfield[1].const_array(mfi); - const auto& Mz = Mfield[2].const_array(mfi); + const int nx = static_cast(t[0].item()); + const int ny = static_cast(t[1].item()); + const int nz = static_cast(t[2].item()); - auto Hx_demag = H_demagfield[0].array(mfi); - auto Hy_demag = H_demagfield[1].array(mfi); - auto Hz_demag = H_demagfield[2].array(mfi); + AMREX_ALWAYS_ASSERT_WITH_MESSAGE(nx > 0 && ny > 0 && nz > 0, + "expected_spatial invalid (must be >0)"); - const IntVect bx_lo = bx.smallEnd(); - const IntVect nbox = bx.size(); + return amrex::IntVect(nx, ny, nz); +} -#if AMREX_SPACEDIM == 2 - const int ncell = nbox[0] * nbox[1]; -#else - const int ncell = nbox[0] * nbox[1] * nbox[2]; +// ------------------------------------------------------------ +// Bundle: holds model + normalizers + expected shape meta. +// ------------------------------------------------------------ +struct MLBundle +{ + torch::jit::script::Module x_norm; + torch::jit::script::Module y_norm; + torch::jit::script::Module model; + + amrex::IntVect expected_spatial{0}; // (nx,ny,nz) + int device_id = 0; + + bool initialized = false; +}; + +// ------------------------------------------------------------ +// Load bundle from .pt files; move everything to cuda:device_id; +// read expected_spatial from x_norm. +// ------------------------------------------------------------ +static inline MLBundle LoadMLBundle(const std::string& x_norm_pt, + const std::string& y_norm_pt, + const std::string& model_pt, + int device_id) +{ + MLBundle b; + b.device_id = device_id; + + // Load modules (default loads to CPU) + b.x_norm = torch::jit::load(x_norm_pt); + b.y_norm = torch::jit::load(y_norm_pt); + b.model = torch::jit::load(model_pt); + + // Move to target GPU + torch::Device dev(torch::kCUDA, device_id); + MoveModuleToDevice(b.x_norm, dev); + MoveModuleToDevice(b.y_norm, dev); + MoveModuleToDevice(b.model, dev); + + // Read expected spatial shape from normalizer + b.expected_spatial = GetExpectedSpatial(b.x_norm); + + b.initialized = true; + + amrex::Print() << "[MLBundle] Loaded modules on cuda:" << device_id + << " expected_spatial = (" + << b.expected_spatial[0] << ", " + << b.expected_spatial[1] << ", " + << b.expected_spatial[2] << ")\n"; + + return b; +} + +// ------------------------------------------------------------ +// Pack Mfield MultiFab (Mx,My,Mz) into Torch tensor [1,3,nx,ny,nz] +// Dynamically sized using bx.size() and expected_spatial. +// ------------------------------------------------------------ +at::Tensor PackMfieldToTensorDynamic( + const Array& Mfield, + const MFIter& mfi, + const Box& bx, + const amrex::IntVect& expected_spatial, + int device_id) +{ +#if AMREX_SPACEDIM != 3 + AMREX_ALWAYS_ASSERT_WITH_MESSAGE(false, "PackMfieldToTensorDynamic expects 3D"); #endif - // Host-visible (Managed) buffers filled on GPU - amrex::Gpu::ManagedVector aux_Mx(ncell); - amrex::Gpu::ManagedVector aux_My(ncell); - amrex::Gpu::ManagedVector aux_Mz(ncell); + const auto& Mx = Mfield[0].const_array(mfi); + const auto& My = Mfield[1].const_array(mfi); + const auto& Mz = Mfield[2].const_array(mfi); + + const IntVect bx_lo = bx.smallEnd(); + const IntVect nbox = bx.size(); // (nx,ny,nz) + + const int nx = nbox[0]; + const int ny = nbox[1]; + const int nz = nbox[2]; + const int ncell = nx * ny * nz; + + AMREX_ALWAYS_ASSERT_WITH_MESSAGE( + nx == expected_spatial[0] && ny == expected_spatial[1] && nz == expected_spatial[2], + "PackMfieldToTensorDynamic: bx size != expected_spatial from normalizer" + ); + + amrex::Gpu::ManagedVector aux_Mx(ncell), aux_My(ncell), aux_Mz(ncell); + Real* AMREX_RESTRICT auxPtr_Mx = aux_Mx.dataPtr(); + Real* AMREX_RESTRICT auxPtr_My = aux_My.dataPtr(); + Real* AMREX_RESTRICT auxPtr_Mz = aux_Mz.dataPtr(); + + amrex::ParallelFor(bx, [=] AMREX_GPU_DEVICE(int i, int j, int k) noexcept { + const int ii = i - bx_lo[0]; + const int jj = j - bx_lo[1]; + const int kk = k - bx_lo[2]; + + // flatten index consistent with (nx,ny,nz) reshape below + const int index = kk + jj * nz + ii * nz * ny; + + auxPtr_Mx[index] = Mx(i, j, k); + auxPtr_My[index] = My(i, j, k); + auxPtr_Mz[index] = Mz(i, j, k); + }); + + amrex::Gpu::streamSynchronize(); + + // Wrap managed memory; then we will immediately copy to CUDA tensor + at::Tensor tMx = torch::from_blob(auxPtr_Mx, {ncell}, torch::kFloat64); + at::Tensor tMy = torch::from_blob(auxPtr_My, {ncell}, torch::kFloat64); + at::Tensor tMz = torch::from_blob(auxPtr_Mz, {ncell}, torch::kFloat64); + + at::Tensor rMx = tMx.reshape({nx, ny, nz}); + at::Tensor rMy = tMy.reshape({nx, ny, nz}); + at::Tensor rMz = tMz.reshape({nx, ny, nz}); + + at::Tensor M = torch::stack({rMx, rMy, rMz}, 0); // [3,nx,ny,nz] - Real* AMREX_RESTRICT auxPtr_Mx = aux_Mx.dataPtr(); - Real* AMREX_RESTRICT auxPtr_My = aux_My.dataPtr(); - Real* AMREX_RESTRICT auxPtr_Mz = aux_Mz.dataPtr(); + torch::Device dev(torch::kCUDA, device_id); + M = M.to(dev).to(torch::kFloat32).unsqueeze(0); // [1,3,nx,ny,nz] - // Fill aux buffers from MultiFab on GPU - amrex::ParallelFor(bx, [=] AMREX_GPU_DEVICE(int i, int j, int k) noexcept { - const int ii = i - bx_lo[0]; - const int jj = j - bx_lo[1]; + return M; +} -#if AMREX_SPACEDIM == 2 - const int index = jj + ii * nbox[1]; +// ------------------------------------------------------------ +// Normalize, Forward, Denormalize +// ------------------------------------------------------------ +at::Tensor NormalizeInput(const at::Tensor& M_cuda_f32, + torch::jit::script::Module& x_norm_module) +{ + BL_PROFILE("NormalizeInput"); + return x_norm_module.get_method("encode")({M_cuda_f32}).toTensor(); +} + +// at::Tensor MLForwardOnly(const at::Tensor& norm_tensor, +// torch::jit::script::Module& ml_module) +// { +// BL_PROFILE("MLForwardOnly"); +// auto out = ml_module.forward({norm_tensor}).toTensor(); + +// #ifdef AMREX_USE_CUDA +// // 把 forward 的 GPU 时间“结算”在这里(用于 profiling 验证) +// amrex::Gpu::streamSynchronize(); +// #endif + +// return out; +// } +at::Tensor MLForwardOnly(const at::Tensor& norm_tensor, + torch::jit::script::Module& ml_module) +{ + BL_PROFILE("MLForwardOnly"); + +#ifdef AMREX_USE_CUDA + at::cuda::CUDAEvent start(/*enable_timing=*/true); + at::cuda::CUDAEvent stop (/*enable_timing=*/true); + + auto stream = at::cuda::getDefaultCUDAStream(); + + stream.synchronize(); + + start.record(stream); + + auto out = ml_module.forward({norm_tensor}).toTensor(); + + stop.record(stream); + stop.synchronize(); + + float ms = start.elapsed_time(stop); + amrex::Print() << "Forward-only time (ms) = " << ms << "\n"; + + return out; #else - const int kk = k - bx_lo[2]; - const int index = kk + jj * nbox[2] + ii * nbox[2] * nbox[1]; + return ml_module.forward({norm_tensor}).toTensor(); #endif +} - auxPtr_Mx[index] = Mx(i, j, k); - auxPtr_My[index] = My(i, j, k); - auxPtr_Mz[index] = Mz(i, j, k); - }); - // Make sure aux buffers are ready for from_blob on host side - amrex::Gpu::streamSynchronize(); +at::Tensor DenormalizeOutput(const at::Tensor& y_tensor, + torch::jit::script::Module& y_norm_module) +{ + BL_PROFILE("DenormalizeOutput"); + // Keep device, convert to float64 for AMReX Real=double path + return y_norm_module.get_method("decode")({y_tensor}).toTensor().to(torch::kFloat64); +} - // Wrap buffers as CPU tensors (no copy) - at::Tensor inputs_torch_Mx = torch::from_blob(auxPtr_Mx, {ncell, 1}, torch::kFloat64); - at::Tensor inputs_torch_My = torch::from_blob(auxPtr_My, {ncell, 1}, torch::kFloat64); - at::Tensor inputs_torch_Mz = torch::from_blob(auxPtr_Mz, {ncell, 1}, torch::kFloat64); +// ------------------------------------------------------------ +// Unpack Torch tensor [1,3,nx,ny,nz] into H_demagfield MultiFabs. +// ------------------------------------------------------------ +void UnpackTensorToHfieldDynamic( + const at::Tensor& denorm_torch_f64, // [1,3,nx,ny,nz] + Array& H_demagfield, + const MFIter& mfi, + const Box& bx, + const amrex::IntVect& expected_spatial) +{ +#if AMREX_SPACEDIM != 3 + AMREX_ALWAYS_ASSERT_WITH_MESSAGE(false, "UnpackTensorToHfieldDynamic expects 3D"); +#endif - // Reshape (assumes ncell == 128*128*4) - at::Tensor reshaped_Mx = inputs_torch_Mx.reshape({128, 128, 4}); - at::Tensor reshaped_My = inputs_torch_My.reshape({128, 128, 4}); - at::Tensor reshaped_Mz = inputs_torch_Mz.reshape({128, 128, 4}); + auto Hx_demag = H_demagfield[0].array(mfi); + auto Hy_demag = H_demagfield[1].array(mfi); + auto Hz_demag = H_demagfield[2].array(mfi); - // Stack into [3, 128, 128, 4] then add batch dim -> [1, 3, 128, 128, 4] - at::Tensor final_tensor_M = torch::stack({reshaped_Mx, reshaped_My, reshaped_Mz}, 0); - final_tensor_M = final_tensor_M.to(torch::kCUDA).to(torch::kFloat32); - final_tensor_M = final_tensor_M.unsqueeze(0); + const IntVect bx_lo = bx.smallEnd(); + const IntVect nbox = bx.size(); - // Normalize -> model -> denormalize - at::Tensor norm_torch = x_norm_module.get_method("encode")({final_tensor_M}).toTensor(); - at::Tensor outputs_torch = ml_module.forward({norm_torch}).toTensor(); - at::Tensor denorm_torch = y_norm_module.get_method("decode")({outputs_torch}).toTensor(); + const int nx = nbox[0]; + const int ny = nbox[1]; + const int nz = nbox[2]; - // Convert to float64 for accessor usage (same as your original) - denorm_torch = denorm_torch.to(torch::kFloat64); + AMREX_ALWAYS_ASSERT_WITH_MESSAGE( + nx == expected_spatial[0] && ny == expected_spatial[1] && nz == expected_spatial[2], + "UnpackTensorToHfieldDynamic: bx size != expected_spatial from normalizer" + ); - // Extract H components: denorm_torch shape assumed [1, 3, 128, 128, 4] - at::Tensor denorm_torch_Hx = denorm_torch.select(0, 0).select(0, 0).flatten(); - at::Tensor denorm_torch_Hy = denorm_torch.select(0, 0).select(0, 1).flatten(); - at::Tensor denorm_torch_Hz = denorm_torch.select(0, 0).select(0, 2).flatten(); + AMREX_ALWAYS_ASSERT_WITH_MESSAGE(denorm_torch_f64.dim() == 5, "denorm must be [1,3,nx,ny,nz]"); + AMREX_ALWAYS_ASSERT_WITH_MESSAGE(denorm_torch_f64.size(0) == 1, "batch must be 1"); + AMREX_ALWAYS_ASSERT_WITH_MESSAGE(denorm_torch_f64.size(1) == 3, "channel must be 3"); + AMREX_ALWAYS_ASSERT_WITH_MESSAGE(denorm_torch_f64.size(2) == nx && + denorm_torch_f64.size(3) == ny && + denorm_torch_f64.size(4) == nz, + "denorm tensor spatial mismatch"); + + // Flatten each component to 1D for packed_accessor + at::Tensor Hx = denorm_torch_f64.select(0, 0).select(0, 0).contiguous().view({-1}); + at::Tensor Hy = denorm_torch_f64.select(0, 0).select(0, 1).contiguous().view({-1}); + at::Tensor Hz = denorm_torch_f64.select(0, 0).select(0, 2).contiguous().view({-1}); #ifdef AMREX_USE_CUDA - auto denorm_torch_Hx_acc = denorm_torch_Hx.packed_accessor64(); - auto denorm_torch_Hy_acc = denorm_torch_Hy.packed_accessor64(); - auto denorm_torch_Hz_acc = denorm_torch_Hz.packed_accessor64(); + auto Hx_acc = Hx.packed_accessor64(); + auto Hy_acc = Hy.packed_accessor64(); + auto Hz_acc = Hz.packed_accessor64(); #endif - // Copy tensor data back into demag field - amrex::ParallelFor(bx, [=] AMREX_GPU_DEVICE(int i, int j, int k) noexcept { - const int ii = i - bx_lo[0]; - const int jj = j - bx_lo[1]; - const int kk = k - bx_lo[2]; - const int index = kk + jj * nbox[2] + ii * nbox[2] * nbox[1]; + amrex::ParallelFor(bx, [=] AMREX_GPU_DEVICE(int i, int j, int k) noexcept { + const int ii = i - bx_lo[0]; + const int jj = j - bx_lo[1]; + const int kk = k - bx_lo[2]; + + const int index = kk + jj * nz + ii * nz * ny; + + Hx_demag(i, j, k) = Hx_acc[index]; + Hy_demag(i, j, k) = Hy_acc[index]; + Hz_demag(i, j, k) = Hz_acc[index]; + }); + + amrex::Gpu::streamSynchronize(); +} + +// ------------------------------------------------------------ +// One-call wrapper: pack -> encode -> forward -> decode -> unpack +// ------------------------------------------------------------ +void RunMLDemagOnBox( + MLBundle& b, + const Array& Mfield, + Array& H_demagfield, + const MFIter& mfi, + const Box& bx) +{ + AMREX_ALWAYS_ASSERT_WITH_MESSAGE(b.initialized, "MLBundle not initialized"); + + // 1) Pack raw M to [1,3,nx,ny,nz] on cuda:device_id + at::Tensor M = PackMfieldToTensorDynamic(Mfield, mfi, bx, b.expected_spatial, b.device_id); + + // 2) Normalize (encode) on same GPU + at::Tensor norm = NormalizeInput(M, b.x_norm); + + // 3) Forward + at::Tensor pred = MLForwardOnly(norm, b.model); - Hx_demag(i, j, k) = denorm_torch_Hx_acc[index]; - Hy_demag(i, j, k) = denorm_torch_Hy_acc[index]; - Hz_demag(i, j, k) = denorm_torch_Hz_acc[index]; - }); + // 4) Denormalize + at::Tensor denorm = DenormalizeOutput(pred, b.y_norm); - amrex::Gpu::streamSynchronize(); - } + // 5) Unpack back to MultiFab + UnpackTensorToHfieldDynamic(denorm, H_demagfield, mfi, bx, b.expected_spatial); } \ No newline at end of file diff --git a/Source/MagneX.H b/Source/MagneX.H index 6d7d4ba..f9b817e 100644 --- a/Source/MagneX.H +++ b/Source/MagneX.H @@ -187,9 +187,42 @@ void WritePlotfile(MultiFab& Ms, const int& plt_step); -void CalculateH_demag_ML(const Array< MultiFab, AMREX_SPACEDIM> & Mfield, - torch::jit::script::Module& x_norm_module, - torch::jit::script::Module& ml_module, - torch::jit::script::Module& y_norm_module, - Array< MultiFab, AMREX_SPACEDIM> & H_demagfield); - \ No newline at end of file +// void CalculateH_demag_ML(const Array< MultiFab, AMREX_SPACEDIM> & Mfield, +// torch::jit::script::Module& x_norm_module, +// torch::jit::script::Module& ml_module, +// torch::jit::script::Module& y_norm_module, +// Array< MultiFab, AMREX_SPACEDIM> & H_demagfield); + + +at::Tensor PackMfieldToTensorDynamic( + const amrex::Array& Mfield, + const amrex::MFIter& mfi, + const amrex::Box& bx, + const amrex::IntVect& expected_spatial, + int device_id); + +at::Tensor NormalizeInput( + const at::Tensor& M_cuda_f32, + torch::jit::script::Module& x_norm_module); + +at::Tensor MLForwardOnly( + const at::Tensor& norm_tensor, + torch::jit::script::Module& ml_module); + +at::Tensor DenormalizeOutput( + const at::Tensor& y_tensor, + torch::jit::script::Module& y_norm_module); + +void UnpackTensorToHfieldDynamic( + const at::Tensor& denorm_torch_f64, + amrex::Array& H_demagfield, + const amrex::MFIter& mfi, + const amrex::Box& bx, + const amrex::IntVect& expected_spatial); + +// Move all parameters + buffers of a TorchScript module to a device +void MoveModuleToDevice(torch::jit::script::Module& m, + const torch::Device& device); + +// Read expected_spatial = [nx, ny, nz] from normalizer module +amrex::IntVect GetExpectedSpatial(torch::jit::script::Module& x_norm_module); \ No newline at end of file diff --git a/Source/main.cpp b/Source/main.cpp index f54d182..53064f7 100644 --- a/Source/main.cpp +++ b/Source/main.cpp @@ -4,6 +4,7 @@ #include #include #include +#include #ifdef AMREX_USE_SUNDIALS #include #endif @@ -89,6 +90,10 @@ void main_main () // Count how many times we have incremented Hbias int increment_count = 0; + // ML related variables (Declared here so they are visible to the whole function) + amrex::IntVect expected_spatial(0,0,0); + int device_id = -1; + BoxArray ba; DistributionMapping dm; @@ -108,40 +113,62 @@ void main_main () // ********************************** // // LOAD PYTORCH MODEL + if (ml_enable == 1) { + BL_PROFILE_VAR("LoadPytorch",LoadPytorch); + + // Load pytorch module via torch script + + + std::string ml_model_name; + std::string x_normalizer_name; + std::string y_normalizer_name; + + ParmParse pp_ml; + pp_ml.query("ml_model_name", ml_model_name); + pp_ml.query("x_normalizer_name", x_normalizer_name); + pp_ml.query("y_normalizer_name", y_normalizer_name); - BL_PROFILE_VAR("LoadPytorch",LoadPytorch); + amrex::Print()<<"\n"< tensor + at::Tensor M_cuda_f32 = PackMfieldToTensorDynamic( + Mfield_old, mfi, bx, expected_spatial, device_id + ); + + // + at::Tensor norm = NormalizeInput(M_cuda_f32, x_norm_module); + at::Tensor pred = MLForwardOnly(norm, ml_module); + at::Tensor denorm_f64 = DenormalizeOutput(pred, y_norm_module); + + // unpack: tensor -> MultiFab + UnpackTensorToHfieldDynamic(denorm_f64, H_demagfield, mfi, bx, expected_spatial); + } } else { + // demag_solver.CalculateH_demag(Mfield_old, H_demagfield); + amrex::Gpu::streamSynchronize(); + double start_time = amrex::second(); + demag_solver.CalculateH_demag(Mfield_old, H_demagfield); + + amrex::Gpu::streamSynchronize(); + double end_time = amrex::second(); + + amrex::Print() << "Demag Solver Time: " << (end_time - start_time) * 1000.0 << " ms" << std::endl; } } @@ -545,9 +599,36 @@ void main_main () if (demag_coupling == 1) { // demag_solver.CalculateH_demag(Mfield, H_demagfield); if (ml_enable == 1) { - CalculateH_demag_ML(Mfield, x_norm_module, ml_module, y_norm_module, H_demagfield); + // CalculateH_demag_ML(Mfield, x_norm_module, ml_module, y_norm_module, H_demagfield); + + for (amrex::MFIter mfi(Mfield_old[0], amrex::TilingIfNotGPU()); + mfi.isValid(); ++mfi) + { + const amrex::Box& bx = mfi.validbox(); + + // pack: MultiFab -> tensor + at::Tensor M_cuda_f32 = PackMfieldToTensorDynamic( + Mfield_old, mfi, bx, expected_spatial, device_id + ); + + at::Tensor norm = NormalizeInput(M_cuda_f32, x_norm_module); + at::Tensor pred = MLForwardOnly(norm, ml_module); + at::Tensor denorm_f64 = DenormalizeOutput(pred, y_norm_module); + + // unpack: tensor -> MultiFab + UnpackTensorToHfieldDynamic(denorm_f64, H_demagfield, mfi, bx, expected_spatial); + } } else { + // demag_solver.CalculateH_demag(Mfield, H_demagfield); + amrex::Gpu::streamSynchronize(); + double start_time = amrex::second(); + demag_solver.CalculateH_demag(Mfield, H_demagfield); + + amrex::Gpu::streamSynchronize(); + double end_time = amrex::second(); + + amrex::Print() << "Demag Solver Time: " << (end_time - start_time) * 1000.0 << " ms" << std::endl; } } From f65c0fcdcca2b22ba06a77a83fcf809fac55a62a Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Tue, 24 Feb 2026 05:57:21 -0800 Subject: [PATCH 26/34] clean up makefile --- Exec/GNUmakefile | 34 +++++----------------------------- 1 file changed, 5 insertions(+), 29 deletions(-) diff --git a/Exec/GNUmakefile b/Exec/GNUmakefile index e9ff6b5..2379345 100644 --- a/Exec/GNUmakefile +++ b/Exec/GNUmakefile @@ -10,10 +10,12 @@ COMP = gnu DIM = 3 USE_FFT = TRUE USE_ML = FALSE -USE_SUNDIALS = FALSE TINY_PROFILE = FALSE PROFILE = FALSE +USE_SUNDIALS = FALSE +SUNDIALS_HOME ?= ../../sundials/instdir + ifeq ($(USE_ML),TRUE) # Define a macro for the C++ preprocessor DEFINES += -DML_ENABLE -D_GLIBCXX_USE_CXX11_ABI=1 @@ -24,7 +26,7 @@ ifeq ($(USE_ML),TRUE) else PYTORCH_ROOT := ../../libtorch_cpu endif - + TORCH_LIBPATH = $(PYTORCH_ROOT)/lib # Library definitions @@ -38,7 +40,7 @@ ifeq ($(USE_ML),TRUE) # Header search paths INCLUDE_LOCATIONS += $(PYTORCH_ROOT)/include \ $(PYTORCH_ROOT)/include/torch/csrc/api/include - + # Library search paths LIBRARY_LOCATIONS += $(TORCH_LIBPATH) @@ -50,32 +52,6 @@ ifeq ($(USE_ML),TRUE) endif endif -# # Pytorch directories -# ifeq ($(USE_CUDA),TRUE) -# PYTORCH_ROOT := ../../libtorch_cuda -# else -# PYTORCH_ROOT := ../../libtorch_cpu -# endif -# TORCH_LIBPATH = $(PYTORCH_ROOT)/lib - -# ifeq ($(USE_CUDA),TRUE) -# TORCH_LIBS = -ltorch -ltorch_cpu -lc10 -lc10_cuda -lcuda -# else -# TORCH_LIBS = -ltorch -ltorch_cpu -lc10 -# endif - -# INCLUDE_LOCATIONS += $(PYTORCH_ROOT)/include \ -# $(PYTORCH_ROOT)/include/torch/csrc/api/include -# LIBRARY_LOCATIONS += $(TORCH_LIBPATH) - -# DEFINES += -D_GLIBCXX_USE_CXX11_ABI=1 -# ifeq ($(USE_CUDA),TRUE) -# LDFLAGS += -Xlinker "--no-as-needed,-rpath $(TORCH_LIBPATH) $(TORCH_LIBS)" -# else -# LDFLAGS += -Wl,--no-as-needed,-rpath=$(TORCH_LIBPATH) $(TORCH_LIBS) -# endif -SUNDIALS_HOME ?= ../../sundials/instdir - include $(AMREX_HOME)/Tools/GNUMake/Make.defs include ../Source/Make.package From 7519d1ff6b8e3115beecb3f45087d645bdfe1569 Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Tue, 24 Feb 2026 05:57:55 -0800 Subject: [PATCH 27/34] fix trailing whitespace --- Source/MagneX.H | 2 +- Source/main.cpp | 8 ++++---- 2 files changed, 5 insertions(+), 5 deletions(-) diff --git a/Source/MagneX.H b/Source/MagneX.H index f9b817e..ea95cde 100644 --- a/Source/MagneX.H +++ b/Source/MagneX.H @@ -192,7 +192,7 @@ void WritePlotfile(MultiFab& Ms, // torch::jit::script::Module& ml_module, // torch::jit::script::Module& y_norm_module, // Array< MultiFab, AMREX_SPACEDIM> & H_demagfield); - + at::Tensor PackMfieldToTensorDynamic( const amrex::Array& Mfield, diff --git a/Source/main.cpp b/Source/main.cpp index 53d9be8..c5e50de 100644 --- a/Source/main.cpp +++ b/Source/main.cpp @@ -470,12 +470,12 @@ void main_main () } } else { // demag_solver.CalculateH_demag(Mfield_old, H_demagfield); - amrex::Gpu::streamSynchronize(); + amrex::Gpu::streamSynchronize(); double start_time = amrex::second(); demag_solver.CalculateH_demag(Mfield_old, H_demagfield); - amrex::Gpu::streamSynchronize(); + amrex::Gpu::streamSynchronize(); double end_time = amrex::second(); amrex::Print() << "Demag Solver Time: " << (end_time - start_time) * 1000.0 << " ms" << std::endl; @@ -620,12 +620,12 @@ void main_main () } } else { // demag_solver.CalculateH_demag(Mfield, H_demagfield); - amrex::Gpu::streamSynchronize(); + amrex::Gpu::streamSynchronize(); double start_time = amrex::second(); demag_solver.CalculateH_demag(Mfield, H_demagfield); - amrex::Gpu::streamSynchronize(); + amrex::Gpu::streamSynchronize(); double end_time = amrex::second(); amrex::Print() << "Demag Solver Time: " << (end_time - start_time) * 1000.0 << " ms" << std::endl; From 417749a8d7dc17361b69ff7e94373c99e6b15eb9 Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Tue, 24 Feb 2026 06:02:25 -0800 Subject: [PATCH 28/34] no longer need fftw includes, these are contained in USE_FFT=TRUE --- Source/MagneX.H | 9 --------- 1 file changed, 9 deletions(-) diff --git a/Source/MagneX.H b/Source/MagneX.H index b741dd2..c09c3f6 100644 --- a/Source/MagneX.H +++ b/Source/MagneX.H @@ -1,12 +1,3 @@ -#ifdef AMREX_USE_CUDA -#include -#else -#include -#ifdef AMREX_USE_MPI -#include -#endif -#endif - #include #include "MagneX_namespace.H" From 3a87d3d2f557b13d75a9eff21bfeb804f5363a8a Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Tue, 24 Feb 2026 06:25:22 -0800 Subject: [PATCH 29/34] non ml compile --- Exec/GNUmakefile | 8 +++++--- Source/Demag_ml.cpp | 6 +++++- Source/MagneX.H | 10 +++------- Source/MagneX.cpp | 5 +++++ Source/Make.package | 2 +- Source/main.cpp | 20 +++++++++++++------- 6 files changed, 32 insertions(+), 19 deletions(-) diff --git a/Exec/GNUmakefile b/Exec/GNUmakefile index 2379345..498c4a6 100644 --- a/Exec/GNUmakefile +++ b/Exec/GNUmakefile @@ -1,5 +1,6 @@ # AMREX_HOME defines the directory in which we will find all the AMReX code. -AMREX_HOME = ../../amrex +AMREX_HOME ?= ../../amrex +SUNDIALS_HOME ?= ../../sundials/instdir DEBUG = FALSE USE_MPI = TRUE @@ -12,11 +13,12 @@ USE_FFT = TRUE USE_ML = FALSE TINY_PROFILE = FALSE PROFILE = FALSE - USE_SUNDIALS = FALSE -SUNDIALS_HOME ?= ../../sundials/instdir ifeq ($(USE_ML),TRUE) + + CPPFLAGS += -DAMREX_USE_ML + # Define a macro for the C++ preprocessor DEFINES += -DML_ENABLE -D_GLIBCXX_USE_CXX11_ABI=1 diff --git a/Source/Demag_ml.cpp b/Source/Demag_ml.cpp index e2cb215..c6f5cc5 100644 --- a/Source/Demag_ml.cpp +++ b/Source/Demag_ml.cpp @@ -1,3 +1,5 @@ +#ifdef AMREX_USE_ML + // MagneX_ML_Infer_Dynamic.cpp #include "MagneX.H" #include @@ -318,4 +320,6 @@ void RunMLDemagOnBox( // 5) Unpack back to MultiFab UnpackTensorToHfieldDynamic(denorm, H_demagfield, mfi, bx, b.expected_spatial); -} \ No newline at end of file +} + +#endif diff --git a/Source/MagneX.H b/Source/MagneX.H index 43b732f..4d42701 100644 --- a/Source/MagneX.H +++ b/Source/MagneX.H @@ -176,13 +176,7 @@ void WritePlotfile(MultiFab& Ms, const Real& time, const int& plt_step); - -// void CalculateH_demag_ML(const Array< MultiFab, AMREX_SPACEDIM> & Mfield, -// torch::jit::script::Module& x_norm_module, -// torch::jit::script::Module& ml_module, -// torch::jit::script::Module& y_norm_module, -// Array< MultiFab, AMREX_SPACEDIM> & H_demagfield); - +#ifdef AMREX_USE_ML at::Tensor PackMfieldToTensorDynamic( const amrex::Array& Mfield, @@ -216,3 +210,5 @@ void MoveModuleToDevice(torch::jit::script::Module& m, // Read expected_spatial = [nx, ny, nz] from normalizer module amrex::IntVect GetExpectedSpatial(torch::jit::script::Module& x_norm_module); + +#endif diff --git a/Source/MagneX.cpp b/Source/MagneX.cpp index d7e0ec2..515e376 100644 --- a/Source/MagneX.cpp +++ b/Source/MagneX.cpp @@ -234,6 +234,11 @@ void InitializeMagneXNamespace() { ml_enable = 0; pp.query("ml_enable",ml_enable); +#ifndef AMREX_USE_ML + if (ml_enable == 1) { + amrex::Abort("ml_enable=1 requires USE_ML=TRUE"); + } +#endif diag_type = -1; pp.query("diag_type",diag_type); diff --git a/Source/Make.package b/Source/Make.package index 4277239..24f7df0 100644 --- a/Source/Make.package +++ b/Source/Make.package @@ -1,6 +1,7 @@ CEXE_sources += Checkpoint.cpp CEXE_sources += ComputeLLGRHS.cpp CEXE_sources += Demagnetization.cpp +CEXE_sources += Demag_ml.cpp CEXE_sources += Diagnostics.cpp CEXE_sources += EffectiveAnisotropyField.cpp CEXE_sources += EffectiveDMIField.cpp @@ -15,4 +16,3 @@ CEXE_headers += CartesianAlgorithm_K.H CEXE_headers += Demagnetization.H CEXE_headers += MagneX.H CEXE_headers += MagneX_namespace.H -CEXE_sources += Demag_ml.cpp \ No newline at end of file diff --git a/Source/main.cpp b/Source/main.cpp index c5e50de..53aa946 100644 --- a/Source/main.cpp +++ b/Source/main.cpp @@ -1,6 +1,5 @@ #include "MagneX.H" #include "Demagnetization.H" -#include #include #include #include @@ -11,8 +10,12 @@ #include +#ifdef AMREX_USE_ML +#include #include // for at::cuda::setDevice #include +#endif + using namespace amrex; using namespace MagneX; @@ -60,10 +63,11 @@ void main_main () Array LLG_RHS; Array LLG_RHS_pre; Array LLG_RHS_avg; +#ifdef AMREX_USE_ML torch::jit::script::Module ml_module; torch::jit::script::Module x_norm_module; torch::jit::script::Module y_norm_module; - +#endif // Declare variables for hysteresis Real normalized_Mx; @@ -114,6 +118,7 @@ void main_main () // ********************************** // // LOAD PYTORCH MODEL if (ml_enable == 1) { +#ifdef AMREX_USE_ML BL_PROFILE_VAR("LoadPytorch",LoadPytorch); // Load pytorch module via torch script @@ -164,6 +169,7 @@ void main_main () << expected_spatial[2] << "\n"; Print() << "Model loaded.\n"; +#endif } else { Print() << "ML disabled. Skipping model load.\n"; @@ -449,7 +455,7 @@ void main_main () if (demag_coupling == 1) { // demag_solver.CalculateH_demag(Mfield_old, H_demagfield); if (ml_enable == 1) { - // CalculateH_demag_ML(Mfield_old, x_norm_module, ml_module, y_norm_module, H_demagfield); +#ifdef AMREX_USE_ML for (amrex::MFIter mfi(Mfield_old[0], amrex::TilingIfNotGPU()); mfi.isValid(); ++mfi) { @@ -460,7 +466,6 @@ void main_main () Mfield_old, mfi, bx, expected_spatial, device_id ); - // at::Tensor norm = NormalizeInput(M_cuda_f32, x_norm_module); at::Tensor pred = MLForwardOnly(norm, ml_module); at::Tensor denorm_f64 = DenormalizeOutput(pred, y_norm_module); @@ -468,6 +473,7 @@ void main_main () // unpack: tensor -> MultiFab UnpackTensorToHfieldDynamic(denorm_f64, H_demagfield, mfi, bx, expected_spatial); } +#endif } else { // demag_solver.CalculateH_demag(Mfield_old, H_demagfield); amrex::Gpu::streamSynchronize(); @@ -599,8 +605,7 @@ void main_main () if (demag_coupling == 1) { // demag_solver.CalculateH_demag(Mfield, H_demagfield); if (ml_enable == 1) { - // CalculateH_demag_ML(Mfield, x_norm_module, ml_module, y_norm_module, H_demagfield); - +#ifdef AMREX_USE_ML for (amrex::MFIter mfi(Mfield_old[0], amrex::TilingIfNotGPU()); mfi.isValid(); ++mfi) { @@ -618,6 +623,7 @@ void main_main () // unpack: tensor -> MultiFab UnpackTensorToHfieldDynamic(denorm_f64, H_demagfield, mfi, bx, expected_spatial); } +#endif } else { // demag_solver.CalculateH_demag(Mfield, H_demagfield); amrex::Gpu::streamSynchronize(); @@ -810,7 +816,7 @@ void main_main () if (fast_demag==1) { // demag_solver.CalculateH_demag(ar_state, H_demagfield); if (ml_enable == 1) { - CalculateH_demag_ML(ar_state, x_norm_module, ml_module, y_norm_module, H_demagfield); + amrex::Abort("add ML demag to fast dynamics"); } else { demag_solver.CalculateH_demag(ar_state, H_demagfield); } From 7836fba9d23d270fbd0fa21759a9bd04b9d4c529 Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Tue, 24 Feb 2026 06:43:46 -0800 Subject: [PATCH 30/34] put in ml_demag protections when not implemented --- Source/main.cpp | 33 ++++++++++----------------------- 1 file changed, 10 insertions(+), 23 deletions(-) diff --git a/Source/main.cpp b/Source/main.cpp index 53aa946..d39ec68 100644 --- a/Source/main.cpp +++ b/Source/main.cpp @@ -453,7 +453,6 @@ void main_main () // Evolve H_demag if (demag_coupling == 1) { - // demag_solver.CalculateH_demag(Mfield_old, H_demagfield); if (ml_enable == 1) { #ifdef AMREX_USE_ML for (amrex::MFIter mfi(Mfield_old[0], amrex::TilingIfNotGPU()); @@ -475,16 +474,7 @@ void main_main () } #endif } else { - // demag_solver.CalculateH_demag(Mfield_old, H_demagfield); - amrex::Gpu::streamSynchronize(); - double start_time = amrex::second(); - demag_solver.CalculateH_demag(Mfield_old, H_demagfield); - - amrex::Gpu::streamSynchronize(); - double end_time = amrex::second(); - - amrex::Print() << "Demag Solver Time: " << (end_time - start_time) * 1000.0 << " ms" << std::endl; } } @@ -603,7 +593,6 @@ void main_main () // Poisson solve and H_demag computation with Mfield if (demag_coupling == 1) { - // demag_solver.CalculateH_demag(Mfield, H_demagfield); if (ml_enable == 1) { #ifdef AMREX_USE_ML for (amrex::MFIter mfi(Mfield_old[0], amrex::TilingIfNotGPU()); @@ -625,16 +614,7 @@ void main_main () } #endif } else { - // demag_solver.CalculateH_demag(Mfield, H_demagfield); - amrex::Gpu::streamSynchronize(); - double start_time = amrex::second(); - demag_solver.CalculateH_demag(Mfield, H_demagfield); - - amrex::Gpu::streamSynchronize(); - double end_time = amrex::second(); - - amrex::Print() << "Demag Solver Time: " << (end_time - start_time) * 1000.0 << " ms" << std::endl; } } @@ -739,7 +719,11 @@ void main_main () H_demagfield[idim].setVal(0.); } } else { - demag_solver.CalculateH_demag(ar_state, H_demagfield); + if (ml_enable == 1) { + amrex::Abort("add ML demag to SUNDIALS rhs"); + } else { + demag_solver.CalculateH_demag(ar_state, H_demagfield); + } } } @@ -814,7 +798,6 @@ void main_main () // H_demag if (demag_coupling == 1) { if (fast_demag==1) { - // demag_solver.CalculateH_demag(ar_state, H_demagfield); if (ml_enable == 1) { amrex::Abort("add ML demag to fast dynamics"); } else { @@ -899,7 +882,11 @@ void main_main () // H_demag if (demag_coupling == 1) { if (implicit_demag==1) { - demag_solver.CalculateH_demag(ar_state, H_demagfield); + if (ml_enable == 1) { + amrex::Abort("ML demag for implicit not supported"); + } else { + demag_solver.CalculateH_demag(ar_state, H_demagfield); + } } else { for (int idim=0; idim Date: Tue, 24 Feb 2026 06:47:17 -0800 Subject: [PATCH 31/34] minor --- Exec/GNUmakefile | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/Exec/GNUmakefile b/Exec/GNUmakefile index 498c4a6..e8cfea4 100644 --- a/Exec/GNUmakefile +++ b/Exec/GNUmakefile @@ -1,6 +1,5 @@ # AMREX_HOME defines the directory in which we will find all the AMReX code. AMREX_HOME ?= ../../amrex -SUNDIALS_HOME ?= ../../sundials/instdir DEBUG = FALSE USE_MPI = TRUE @@ -14,6 +13,7 @@ USE_ML = FALSE TINY_PROFILE = FALSE PROFILE = FALSE USE_SUNDIALS = FALSE +SUNDIALS_HOME ?= ../../sundials/instdir ifeq ($(USE_ML),TRUE) From 7a9594f51d61adf4bf8b5e23ed406e59e508905f Mon Sep 17 00:00:00 2001 From: tyh0123 Date: Tue, 24 Feb 2026 10:44:47 -0800 Subject: [PATCH 32/34] fixed MagneX.H header for compiling issue --- Source/MagneX.H | 17 +++++++++++++---- 1 file changed, 13 insertions(+), 4 deletions(-) diff --git a/Source/MagneX.H b/Source/MagneX.H index 4d42701..bfc3135 100644 --- a/Source/MagneX.H +++ b/Source/MagneX.H @@ -1,3 +1,13 @@ +#ifdef AMREX_USE_CUDA +#include +#include +#else +#include +#ifdef AMREX_USE_MPI +#include +#endif +#endif + #include #include "MagneX_namespace.H" @@ -176,7 +186,8 @@ void WritePlotfile(MultiFab& Ms, const Real& time, const int& plt_step); -#ifdef AMREX_USE_ML + + at::Tensor PackMfieldToTensorDynamic( const amrex::Array& Mfield, @@ -209,6 +220,4 @@ void MoveModuleToDevice(torch::jit::script::Module& m, const torch::Device& device); // Read expected_spatial = [nx, ny, nz] from normalizer module -amrex::IntVect GetExpectedSpatial(torch::jit::script::Module& x_norm_module); - -#endif +amrex::IntVect GetExpectedSpatial(torch::jit::script::Module& x_norm_module); \ No newline at end of file From 661895b60715f9fab2c74afb1cd64e22281f7037 Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Tue, 24 Feb 2026 11:14:27 -0800 Subject: [PATCH 33/34] cleanup include statements --- Source/MagneX.H | 15 +++++---------- 1 file changed, 5 insertions(+), 10 deletions(-) diff --git a/Source/MagneX.H b/Source/MagneX.H index bfc3135..7c441b2 100644 --- a/Source/MagneX.H +++ b/Source/MagneX.H @@ -1,11 +1,5 @@ -#ifdef AMREX_USE_CUDA -#include +#ifdef AMREX_USE_ML #include -#else -#include -#ifdef AMREX_USE_MPI -#include -#endif #endif #include @@ -186,8 +180,7 @@ void WritePlotfile(MultiFab& Ms, const Real& time, const int& plt_step); - - +#ifdef AMREX_USE_ML at::Tensor PackMfieldToTensorDynamic( const amrex::Array& Mfield, @@ -220,4 +213,6 @@ void MoveModuleToDevice(torch::jit::script::Module& m, const torch::Device& device); // Read expected_spatial = [nx, ny, nz] from normalizer module -amrex::IntVect GetExpectedSpatial(torch::jit::script::Module& x_norm_module); \ No newline at end of file +amrex::IntVect GetExpectedSpatial(torch::jit::script::Module& x_norm_module); + +#endif From 07c8343b65006dc365369fc8bf296af54ea28063 Mon Sep 17 00:00:00 2001 From: Andy Nonaka Date: Tue, 24 Feb 2026 11:44:32 -0800 Subject: [PATCH 34/34] pytorch install instructions --- Exec/README | 72 ----------------------------------------- Exec/README_md.pytorch | 36 +++++++++++++++++++++ Exec/README_md.sundials | 6 ++++ Exec/README_sundials | 4 --- 4 files changed, 42 insertions(+), 76 deletions(-) delete mode 100644 Exec/README create mode 100644 Exec/README_md.pytorch create mode 100644 Exec/README_md.sundials delete mode 100644 Exec/README_sundials diff --git a/Exec/README b/Exec/README deleted file mode 100644 index ce5b9b2..0000000 --- a/Exec/README +++ /dev/null @@ -1,72 +0,0 @@ -------------------------------------------- - -inputs_compare_std4 - -Standard problem 4. -Entire domain is magnetic. -Matches comapre_std.m -==================== Initial Setup ==================== - demag_coupling = 1 - M_normalization = 1 - exchange_coupling = 1 - DMI_coupling = 0 - anisotropy_coupling = 0 - TimeIntegratorOption = 1 - -------------------------------------------- - -inputs_compare_subdomain - -Test comparison problem of demag solver; to compare initial H_demag to compare_subdomain.m -A cubic block of material within the domain is magnetic. -==================== Initial Setup ==================== - demag_coupling = 1 - M_normalization = 1 - exchange_coupling = 1 - DMI_coupling = 0 - anisotropy_coupling = 0 - TimeIntegratorOption = 1 - -------------------------------------------- - -inputs_exchange - -A block of magnetic material in the center of the domain initialized -so that My = Ms. Only exchange physics is enabled, so the system does -not change due to the interface boundary conditions. -==================== Initial Setup ==================== - demag_coupling = 0 - M_normalization = 1 - exchange_coupling = 1 - DMI_coupling = 0 - anisotropy_coupling = 0 - TimeIntegratorOption = 2 - -------------------------------------------- - -inputs_PSSW - -A block of magnetic material in the center of the domain initialized so that My = Ms. -==================== Initial Setup ==================== - demag_coupling = 1 - M_normalization = 1 - exchange_coupling = 1 - DMI_coupling = 0 - anisotropy_coupling = 0 - TimeIntegratorOption = 1 - -------------------------------------------- - -inputs_restart - -This is the same as inputs_PSSW but with lower resolution and restart -hooks more easily enabled to test this capability with these physics. -==================== Initial Setup ==================== - demag_coupling = 1 - M_normalization = 1 - exchange_coupling = 1 - DMI_coupling = 0 - anisotropy_coupling = 0 - TimeIntegratorOption = 1 - -------------------------------------------- diff --git a/Exec/README_md.pytorch b/Exec/README_md.pytorch new file mode 100644 index 0000000..e84f4c2 --- /dev/null +++ b/Exec/README_md.pytorch @@ -0,0 +1,36 @@ +# PyTorch (libtorch) Download and Setup + +This guide downloads the libtorch CUDA 11.8 C++ distribution and unzips it in the same directory that contains `MagneX`, then renames the extracted folder to `libtorch_cuda`. + +## Steps + +1. Change to the parent directory of `MagneX`: + +```bash +cd +``` + +2. Download the libtorch archive with `wget`: + +```bash +wget https://download.pytorch.org/libtorch/cu118/libtorch-cxx11-abi-shared-with-deps-2.7.1%2Bcu118.zip +``` + +3. Unzip the archive in the same directory: + +```bash +unzip libtorch-cxx11-abi-shared-with-deps-2.7.1+cu118.zip +``` + +4. Rename the extracted folder to `libtorch_cuda`: + +```bash +mv libtorch libtorch_cuda +``` + +## Result + +After the steps above, you should have: + +- `/MagneX` +- `/pytorch_cuda` diff --git a/Exec/README_md.sundials b/Exec/README_md.sundials new file mode 100644 index 0000000..04f7cdd --- /dev/null +++ b/Exec/README_md.sundials @@ -0,0 +1,6 @@ +# SUNDIALS Setup + +Refer to https://amrex-codes.github.io/amrex/docs_html/TimeIntegration_Chapter.html +for SUNDIALS installation, build, and usage instructions. + +Make sure that `SUNDIALS_HOME` in the `GNUmakefile` points to the installation directory. diff --git a/Exec/README_sundials b/Exec/README_sundials deleted file mode 100644 index 3b6830a..0000000 --- a/Exec/README_sundials +++ /dev/null @@ -1,4 +0,0 @@ -Refer to https://amrex-codes.github.io/amrex/docs_html/TimeIntegration_Chapter.html -for SNUDIALS installation, build, and usage instructions. - -Make sure that SUNDIALS_HOME in the GNUmakefile points to the installation directory.