Repository navigation
[gpu-1] End-to-end GPU test of VlasovAmpereOneSpecies on 1 and 2 ranks - #734
Draft
max-models wants to merge 14 commits into
Draft
max-models wants to merge 14 commits into
max-models wants to merge 14 commits into
Conversation
Moves the feectools submodule to struphy-hub/feectools#96 (branch device-kronecker-solve): KroneckerLinearSolver solves device data with dense inverses of the 1D matrices (one GEMM per direction) instead of copying it to the host for LAPACK/SuperLU, so MassMatrixPreconditioner solves on the CuPy backend make no host/device transfers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Point the feectools submodule at struphy-hub/feectools#97 (branch stencil-6d-views): the 3D stencil dot/transpose kernels take the matrix data as 6D views, and StencilMatrix.dot/vdot/transpose no longer raise for matrices with fewer than 2p+1 diagonals. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Move the feectools submodule to struphy-hub/feectools#98: on the CuPy backend inner products return 0-d device arrays and the Krylov solvers copy one scalar per iteration to the host (the convergence test). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The mass-matrix preconditioners could not be created on the CuPy backend: the assembled 1d mass matrices come back from StencilMatrix.toarray() as host (NumPy) arrays, and is_circulant/FFTSolver asserted xp.ndarray (cupy.ndarray on CuPy). The stencil-matrix fill also mixed xp.nonzero with host arrays. The 1d matrices and solvers are now host setup data on every backend (what KroneckerLinearSolver expects, and what feectools#96 builds its device inverses from); only the process-local 1d stencil matrices of the KroneckerStencilMatrix are copied to the device once. The duplicated 1d setup of both preconditioners moves into _solver_and_local_matrix_1d. FFTSolver accepts host or device matrices and solves device right-hand sides on the host. The NumPy path computes the same as before. Adds a fake-CuPy test (subprocess, CUDA launches emulated on the CPU) comparing both preconditioners for M0/M1/M2 (periodic and clamped) with NumPy, and the same check on a GPU. Part of #689 and #650. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Run a small LinearVlasovAmpereOneSpecies case (periodic Cuboid, 8 elements, 320 Sobol markers, 3 steps) on NumPy and CuPy, compare energies, E-field coefficients and markers, and count the host/device transfers of the time loop per phase (none in model.integrate, scalar-sized downloads in the diagnostics, one download per saved dataset in the output). GPU-only part is marked requires_cupy; the NumPy reference runs everywhere and under MPI. Sobol marker loading now generates on the host and copies the result to the active backend: sobol_seq is a scalar loop and passed lists to cupy.transpose. Part of #689. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…6d-views Stack #703 -> #706: feectools submodule points to the head of struphy-hub/feectools#97, which now contains #96. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- New test_gpu_VlasovAmpereOneSpecies.py: a small Landau-damping run (full-f weights, e1_v1 binning plot) on NumPy and on CuPy; compares energies, E and phi coefficients and markers, and checks the transfer budget of the time loop. Passes on the fake CuPy with emulated CUDA launches on 1 and 2 ranks. - Shared helpers in models/tests/gpu_e2e.py (transfer budget, instrumentation, gathering the markers, device context); test_gpu_LinearVlasovAmpereOneSpecies.py uses them. On the fake CuPy every CUDA launch is emulated and CudaKernel.compile is skipped, so both tests now also run without a GPU (CUNUMPY_FAKE_CUPY=1, or a serial child process). - Transfer budget of model.integrate: only scalar downloads (the residual norm of each Krylov iteration, which feectools copies for its convergence test), no arrays or host kernels. - Particles: probe whether MPI is CUDA-aware when the marker array is allocated (setup), not on the first marker exchange inside the first time step. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models
marked this pull request as draft
October 8, 2026 13:34
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Solves the following issue(s):
Part of #689, tracked in #650. This PR covers the items of #689 that are left for
VlasovAmpereOneSpecies: running on 1 and 2 ranks, no host/device transfers in the time loop, and the diagnostics it needs. It does not close #689, because these runs use the CPU emulation; the first run on a real GPU is #687.Stack: this is gpu-1 of a new stack. Merge it after #722 (
cuda-namespaces), #716, #717 and #719. It merges those branches in, so until they are merged, this diff also shows their commits. gpu-2 (the benchmark script) is based on this branch.Core changes:
models/tests/test_gpu_VlasovAmpereOneSpecies.py. A small Landau-damping run on NumPy and on CuPy: 8 cells, 320 Sobol markers, 3 steps, full-f weights, and ane1_v1binning plot so that the diagnostics run on the device too. It compares the energies, thee_fieldandphicoefficients, and the markers by ID (tolerances explained incompare_runs), and checks the transfer budget of the time loop.models/tests/gpu_e2e.py. The transfer budget, the instrumentation, gathering markers across ranks and the device context, moved out of End-to-end GPU test for LinearVlasovAmpereOneSpecies #719's test.test_gpu_LinearVlasovAmpereOneSpecies.pynow uses them.CUNUMPY_FAKE_CUPY=1), every CUDA launch is emulated on the CPU andCudaKernel.compileis skipped.Simulation.compile_cuda_kernelsstill checks that every kernel of the time loop has a CUDA version. Without a GPU and without the fake CuPy,test_fake_cupy_matches_numpyruns the comparison in a serial child process.model.integrate. Only scalar downloads are allowed: the residual norm of each Krylov iteration, which feectools copies for its convergence test (feectools#98 keeps everything else of the solve on the device). No arrays, no host-to-device copies, no host kernels.Particles. Checking whether MPI is CUDA-aware now happens when the marker array is allocated (in the setup), not on the first marker exchange. Before, that check was the only array transfer inside the time loop on 2 ranks.Results (fake CuPy, every CUDA launch emulated on the CPU; macOS; not run on a GPU):
test_gpu_VlasovAmpereOneSpecies.py(NumPy reference + CuPy vs NumPy + transfer budget)test_gpu_LinearVlasovAmpereOneSpecies.py::test_cupy_matches_numpy_without_transfersTransfers in the time loop of the CuPy run:
integrate: one 8-byte residual norm per pcg iteration.Diagnostics: no new CUDA kernels are needed for this model. The scalars use
xpreductions and feectools' device inner products, and binning usesxp.histogramdd.Found along the way (not changed here):
stab_eps=1e-14puts (Monte-Carlo net charge) / 1e-14 into the constant mode ofphi. That makesphiabout 1e14, so E = -grad(phi) loses about 5% to cancellation on either backend. The verification tests use this default. The new test usesstab_eps=1e-6, as End-to-end GPU test for LinearVlasovAmpereOneSpecies #719 does.control_variate=True,kinetic_energyat t=0 is computed with the full-f weights and afterwards with the delta-f weights, sototal_energyjumps after the first step. With full-f weights the total is conserved from step to step to 1e-9, but the t=0 value still differs by about 2e-4 (not investigated).Documentation changes:
CUDA_STRATEGY.md: the end-to-end entry now covers both models, the CPU emulation and the transfer budget.🤖 Generated with Claude Code