Skip to content

[gpu-1] End-to-end GPU test of VlasovAmpereOneSpecies on 1 and 2 ranks - #734

Draft
max-models wants to merge 14 commits into
cuda-namespacesfrom
gpu-1-vlasov-ampere-e2e
Draft

max-models wants to merge 14 commits into
cuda-namespacesfrom
gpu-1-vlasov-ampere-e2e

Conversation

@max-models

Copy link
Copy Markdown
Member

Solves the following issue(s):

Part of #689, tracked in #650. This PR covers the items of #689 that are left for VlasovAmpereOneSpecies: running on 1 and 2 ranks, no host/device transfers in the time loop, and the diagnostics it needs. It does not close #689, because these runs use the CPU emulation; the first run on a real GPU is #687.

Stack: this is gpu-1 of a new stack. Merge it after #722 (cuda-namespaces), #716, #717 and #719. It merges those branches in, so until they are merged, this diff also shows their commits. gpu-2 (the benchmark script) is based on this branch.

Core changes:

  • New test models/tests/test_gpu_VlasovAmpereOneSpecies.py. A small Landau-damping run on NumPy and on CuPy: 8 cells, 320 Sobol markers, 3 steps, full-f weights, and an e1_v1 binning plot so that the diagnostics run on the device too. It compares the energies, the e_field and phi coefficients, and the markers by ID (tolerances explained in compare_runs), and checks the transfer budget of the time loop.
  • Shared helpers in models/tests/gpu_e2e.py. The transfer budget, the instrumentation, gathering markers across ranks and the device context, moved out of End-to-end GPU test for LinearVlasovAmpereOneSpecies #719's test. test_gpu_LinearVlasovAmpereOneSpecies.py now uses them.
  • The GPU tests also run without a GPU. On cunumpy's fake CuPy (CUNUMPY_FAKE_CUPY=1), every CUDA launch is emulated on the CPU and CudaKernel.compile is skipped. Simulation.compile_cuda_kernels still checks that every kernel of the time loop has a CUDA version. Without a GPU and without the fake CuPy, test_fake_cupy_matches_numpy runs the comparison in a serial child process.
  • Transfer budget of model.integrate. Only scalar downloads are allowed: the residual norm of each Krylov iteration, which feectools copies for its convergence test (feectools#98 keeps everything else of the solve on the device). No arrays, no host-to-device copies, no host kernels.
  • Particles. Checking whether MPI is CUDA-aware now happens when the marker array is allocated (in the setup), not on the first marker exchange. Before, that check was the only array transfer inside the time loop on 2 ranks.

Results (fake CuPy, every CUDA launch emulated on the CPU; macOS; not run on a GPU):

test 1 rank 2 ranks
test_gpu_VlasovAmpereOneSpecies.py (NumPy reference + CuPy vs NumPy + transfer budget) passed passed
test_gpu_LinearVlasovAmpereOneSpecies.py::test_cupy_matches_numpy_without_transfers passed not run

Transfers in the time loop of the CuPy run:

  • integrate: one 8-byte residual norm per pcg iteration.
  • Diagnostics: one 24-byte copy of the three scalars per update.
  • Output: one copy per saved dataset.

Diagnostics: no new CUDA kernels are needed for this model. The scalars use xp reductions and feectools' device inner products, and binning uses xp.histogramdd.

Found along the way (not changed here):

  • The initial Poisson solve with the default stab_eps=1e-14 puts (Monte-Carlo net charge) / 1e-14 into the constant mode of phi. That makes phi about 1e14, so E = -grad(phi) loses about 5% to cancellation on either backend. The verification tests use this default. The new test uses stab_eps=1e-6, as End-to-end GPU test for LinearVlasovAmpereOneSpecies #719 does.
  • With control_variate=True, kinetic_energy at t=0 is computed with the full-f weights and afterwards with the delta-f weights, so total_energy jumps after the first step. With full-f weights the total is conserved from step to step to 1e-9, but the t=0 value still differs by about 2e-4 (not investigated).

Documentation changes:

CUDA_STRATEGY.md: the end-to-end entry now covers both models, the CPU emulation and the transfer budget.

🤖 Generated with Claude Code

max-models and others added 14 commits October 7, 2026 18:43
Moves the feectools submodule to struphy-hub/feectools#96 (branch
device-kronecker-solve): KroneckerLinearSolver solves device data with
dense inverses of the 1D matrices (one GEMM per direction) instead of
copying it to the host for LAPACK/SuperLU, so MassMatrixPreconditioner
solves on the CuPy backend make no host/device transfers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Point the feectools submodule at struphy-hub/feectools#97 (branch
stencil-6d-views): the 3D stencil dot/transpose kernels take the matrix
data as 6D views, and StencilMatrix.dot/vdot/transpose no longer raise for
matrices with fewer than 2p+1 diagonals.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Move the feectools submodule to struphy-hub/feectools#98: on the CuPy
backend inner products return 0-d device arrays and the Krylov solvers
copy one scalar per iteration to the host (the convergence test).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The mass-matrix preconditioners could not be created on the CuPy backend:
the assembled 1d mass matrices come back from StencilMatrix.toarray() as
host (NumPy) arrays, and is_circulant/FFTSolver asserted xp.ndarray
(cupy.ndarray on CuPy). The stencil-matrix fill also mixed xp.nonzero
with host arrays.

The 1d matrices and solvers are now host setup data on every backend
(what KroneckerLinearSolver expects, and what feectools#96 builds its
device inverses from); only the process-local 1d stencil matrices of the
KroneckerStencilMatrix are copied to the device once. The duplicated 1d
setup of both preconditioners moves into _solver_and_local_matrix_1d.
FFTSolver accepts host or device matrices and solves device right-hand
sides on the host. The NumPy path computes the same as before.

Adds a fake-CuPy test (subprocess, CUDA launches emulated on the CPU)
comparing both preconditioners for M0/M1/M2 (periodic and clamped) with
NumPy, and the same check on a GPU.

Part of #689 and #650.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Run a small LinearVlasovAmpereOneSpecies case (periodic Cuboid, 8 elements,
320 Sobol markers, 3 steps) on NumPy and CuPy, compare energies, E-field
coefficients and markers, and count the host/device transfers of the time
loop per phase (none in model.integrate, scalar-sized downloads in the
diagnostics, one download per saved dataset in the output). GPU-only part
is marked requires_cupy; the NumPy reference runs everywhere and under MPI.

Sobol marker loading now generates on the host and copies the result to the
active backend: sobol_seq is a scalar loop and passed lists to cupy.transpose.

Part of #689.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…6d-views

Stack #703 -> #706: feectools submodule points to the head of
struphy-hub/feectools#97, which now contains #96.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Stack #703 -> #706 -> #716: feectools submodule points to the head of
struphy-hub/feectools#98, which now contains #96 and #97.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- New test_gpu_VlasovAmpereOneSpecies.py: a small Landau-damping run (full-f weights, e1_v1
  binning plot) on NumPy and on CuPy; compares energies, E and phi coefficients and markers, and
  checks the transfer budget of the time loop. Passes on the fake CuPy with emulated CUDA
  launches on 1 and 2 ranks.
- Shared helpers in models/tests/gpu_e2e.py (transfer budget, instrumentation, gathering the
  markers, device context); test_gpu_LinearVlasovAmpereOneSpecies.py uses them. On the fake CuPy
  every CUDA launch is emulated and CudaKernel.compile is skipped, so both tests now also run
  without a GPU (CUNUMPY_FAKE_CUPY=1, or a serial child process).
- Transfer budget of model.integrate: only scalar downloads (the residual norm of each Krylov
  iteration, which feectools copies for its convergence test), no arrays or host kernels.
- Particles: probe whether MPI is CUDA-aware when the marker array is allocated (setup), not on
  the first marker exchange inside the first time step.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@max-models
max-models marked this pull request as draft October 8, 2026 13:34

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant