Skip to content

End-to-end GPU test for LinearVlasovAmpereOneSpecies - #719

Draft
max-models wants to merge 1 commit into
develfrom
linear-vlasov-ampere-gpu-e2e
Draft

max-models wants to merge 1 commit into
develfrom
linear-vlasov-ampere-gpu-e2e

Conversation

@max-models

@max-models max-models commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Solves the following issue(s):

Part of #689 (run a complete model end to end on the GPU), tracked in #650. This PR adds the end-to-end test for LinearVlasovAmpereOneSpecies. It runs a small case on NumPy and on CuPy, compares the two runs, and checks that the time loop stays on the device. The GPU test is expected to fail until the dependencies below are merged and released. That is why the PR says "Part of #689" and does not close it: the issue also asks for 2-rank GPU runs, CUDA diagnostics kernels, a decision on when to compile the kernels and a CPU/GPU timing.

Depends on (for the GPU test to pass):

Core changes:

  • New test src/struphy/models/tests/test_gpu_LinearVlasovAmpereOneSpecies.py:
    • Setup. The energy-conservation setup of verification/test_verif_LinearVlasovAmpereOneSpecies.py, made smaller: periodic Cuboid(r1=12.56), 8×1×1 elements, degree (3, 1, 1), 320 markers, dt=0.1, 3 steps, no B0 and no E0. A time step is PushEta followed by EfieldWeightsCoupling.
      • Markers are loaded with sobol_standard, so both backends start from the same markers. pseudo_random draws from xp.random, and NumPy and CuPy give different streams.
      • Both solvers use tol=1e-12, so the backends differ only by round-off, not by where pcg stops.
      • No sorting boxes. The model does not need them, and box sorting has no device version yet.
    • test_numpy_reference (runs everywhere). The NumPy run must make physical sense: the field energy changes, en_tot is conserved to 1e-6, and all markers are there. Running the transfer instrumentation on NumPy must record zero events.
    • test_cupy_matches_numpy_without_transfers (requires_cupy, skipped without a GPU). It runs NumPy first, then CuPy, and compares at the end:
      • Velocities and positions: rtol=atol=1e-13. Velocities do not change, and positions get an explicit push, so both backends do the same floating-point operations (up to FMA).
      • Weights and E-field coefficients: atol = 1e-8 · max|ref|. Energy time series: rtol=1e-8. These values come from the Crank–Nicolson Schur solve (pcg to 1e-12, condition number of order 10–100 on this grid) and from the accumulation, which uses atomic adds and so sums in a different order. The expected difference is about 1e-10, so the tolerance leaves a factor of 100. It is still far below the change per step (about 1e-2 in en_E).
      • Matching: markers are gathered from all ranks and sorted by marker ID. Field coefficients are compared per rank, since both runs have the same decomposition.
    • Transfer budget. Transfers are counted with cunumpy.profiling.count_transfers, from the end of the setup (_initialize_hdf5_datasets) to the end of run(), by wrapping the model methods and DataContainer.save_data:
      • model.integrate: no to_host, to_device, kernel_conversion or fallback (device-only copies are allowed).
      • Diagnostics (update_scalar_quantities, update_markers_to_be_saved, update_distr_functions): no host-to-device copies and no host kernels. At most one 8-byte download per scalar per call.
      • Output (save_data): no host-to-device copies and no host kernels. At most one download per saved dataset, and no more than those datasets' bytes in total.
      • Not seen: implicit syncs such as float(device_scalar) in the pcg convergence check. Only transfers made through cunumpy are counted.
    • MPI. The tests are MPI-compatible (mpirun -n 2 python -m pytest --with-mpi ...). Only the NumPy part runs here.
  • pic/sobol_seq.py and Sobol loading in Particles.draw_markers. sobol_seq now always runs on the host (NumPy), and draw_markers copies the result to the active backend with xp.to_cunumpy (setup only, both Sobol branches). Before, sobol_seq used cunumpy and passed Python lists to cupy.transpose, which CuPy rejects. It is a loop over scalars, so it does not belong on the device anyway.

Model-specific changes:

None.

How far the CuPy run gets without a GPU. I ran the CuPy part locally on cunumpy's fake CuPy (CUNUMPY_FAKE_CUPY=1, clean MPI environment, ad-hoc script, not committed):

  1. On this branch alone, the first failure is in Simulation.allocate, while the Derham projectors are assembled (CommutingProjector.__init__ → KroneckerStencilMatrix.transpose → feectools' device transpose kernel): NotImplementedError: fake CuPy cannot compile CUDA kernels. Every CUDA launch fails there.

  2. With every CUDA launch emulated on the CPU (emulated_launches() from CUDA version of the linear_vlasov_ampere accumulation #705's pic/tests/cuda_emulation.py), the Derham, the domain and the markers are set up on the device:

    • Without the Sobol change of this PR, marker loading fails: cupy.transpose rejects lists.
    • With SortingParameters(boxes_per_dim=...) the run also fails in SortingBoxes (round() of a CuPy scalar, then the pyccel initialize_neighbours called with a device array). This test does not use sorting boxes for that reason; that bug is not fixed here.
    • The next failure is EfieldWeightsCoupling.allocate → MassMatrixPreconditioner(M1) → is_circulant: assert isinstance(mat, xp.ndarray). This is the mass-preconditioner-cupy dependency. With precond=None the run still stops at the same assert, because the initial Poisson solve builds an L2Projector, which always creates a MassMatrixPreconditioner.

    So the time loop was not reached locally, and the transfer budget has not been tested against a real CuPy run yet. Once the dependencies are in, expect to tune the budget after the first H100 run, for example if the scalars are reduced on the host.

CI: devel has no GPU job that runs struphy's pytest suite. The old GitLab test_cupy job only runs utils/cupy_vs_numpy.py and does not install struphy, and #710 replaces that file. The test is therefore not wired into CI here. Once #710 is merged, its struphy_gpu_tests job should add:

python -m pytest -v -ra src/struphy/models/tests/test_gpu_LinearVlasovAmpereOneSpecies.py

and its mpi_gpu_tests job:

mpirun -n 2 python -m pytest -v -ra --with-mpi src/struphy/models/tests/test_gpu_LinearVlasovAmpereOneSpecies.py

Documentation changes:

CUDA_STRATEGY.md: one checklist entry for the end-to-end test and its dependencies.

Testing (macOS, no GPU; pyccel kernels compiled with GNU/Fortran; GPU tests not run):

before (devel) after
test_draw_parallel.py test_verif_LinearVlasovAmpereOneSpecies.py + new test 11 passed 12 passed, 1 skipped
same, mpirun -n 2 --with-mpi 11 passed 12 passed, 1 skipped

The skip is the GPU test (CuPy/GPU not available). test_draw_parallel.py covers the Sobol loading paths.

🤖 Generated with Claude Code

Run a small LinearVlasovAmpereOneSpecies case (periodic Cuboid, 8 elements,
320 Sobol markers, 3 steps) on NumPy and CuPy, compare energies, E-field
coefficients and markers, and count the host/device transfers of the time
loop per phase (none in model.integrate, scalar-sized downloads in the
diagnostics, one download per saved dataset in the output). GPU-only part
is marked requires_cupy; the NumPy reference runs everywhere and under MPI.

Sobol marker loading now generates on the host and copies the result to the
active backend: sobol_seq is a scalar loop and passed lists to cupy.transpose.

Part of #689.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant