Repository navigation
End-to-end GPU test for LinearVlasovAmpereOneSpecies - #719
Draft
max-models wants to merge 1 commit into
Draft
max-models wants to merge 1 commit into
max-models wants to merge 1 commit into
Conversation
Run a small LinearVlasovAmpereOneSpecies case (periodic Cuboid, 8 elements, 320 Sobol markers, 3 steps) on NumPy and CuPy, compare energies, E-field coefficients and markers, and count the host/device transfers of the time loop per phase (none in model.integrate, scalar-sized downloads in the diagnostics, one download per saved dataset in the output). GPU-only part is marked requires_cupy; the NumPy reference runs everywhere and under MPI. Sobol marker loading now generates on the host and copies the result to the active backend: sobol_seq is a scalar loop and passed lists to cupy.transpose. Part of #689. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Solves the following issue(s):
Part of #689 (run a complete model end to end on the GPU), tracked in #650. This PR adds the end-to-end test for
LinearVlasovAmpereOneSpecies. It runs a small case on NumPy and on CuPy, compares the two runs, and checks that the time loop stays on the device. The GPU test is expected to fail until the dependencies below are merged and released. That is why the PR says "Part of #689" and does not close it: the issue also asks for 2-rank GPU runs, CUDA diagnostics kernels, a decision on when to compile the kernels and a CPU/GPU timing.Depends on (for the GPU test to pass):
linear_vlasov_ampereaccumulationMassMatrixPreconditioneron CuPy: MassMatrixPreconditioner on the CuPy backend #717Core changes:
src/struphy/models/tests/test_gpu_LinearVlasovAmpereOneSpecies.py:verification/test_verif_LinearVlasovAmpereOneSpecies.py, made smaller: periodicCuboid(r1=12.56), 8×1×1 elements, degree (3, 1, 1), 320 markers,dt=0.1, 3 steps, no B0 and no E0. A time step isPushEtafollowed byEfieldWeightsCoupling.sobol_standard, so both backends start from the same markers.pseudo_randomdraws fromxp.random, and NumPy and CuPy give different streams.tol=1e-12, so the backends differ only by round-off, not by where pcg stops.test_numpy_reference(runs everywhere). The NumPy run must make physical sense: the field energy changes,en_totis conserved to 1e-6, and all markers are there. Running the transfer instrumentation on NumPy must record zero events.test_cupy_matches_numpy_without_transfers(requires_cupy, skipped without a GPU). It runs NumPy first, then CuPy, and compares at the end:rtol=atol=1e-13. Velocities do not change, and positions get an explicit push, so both backends do the same floating-point operations (up to FMA).atol = 1e-8 · max|ref|. Energy time series:rtol=1e-8. These values come from the Crank–Nicolson Schur solve (pcg to 1e-12, condition number of order 10–100 on this grid) and from the accumulation, which uses atomic adds and so sums in a different order. The expected difference is about 1e-10, so the tolerance leaves a factor of 100. It is still far below the change per step (about 1e-2 inen_E).cunumpy.profiling.count_transfers, from the end of the setup (_initialize_hdf5_datasets) to the end ofrun(), by wrapping the model methods andDataContainer.save_data:model.integrate: noto_host,to_device,kernel_conversionorfallback(device-only copies are allowed).update_scalar_quantities,update_markers_to_be_saved,update_distr_functions): no host-to-device copies and no host kernels. At most one 8-byte download per scalar per call.save_data): no host-to-device copies and no host kernels. At most one download per saved dataset, and no more than those datasets' bytes in total.float(device_scalar)in the pcg convergence check. Only transfers made through cunumpy are counted.mpirun -n 2 python -m pytest --with-mpi ...). Only the NumPy part runs here.pic/sobol_seq.pyand Sobol loading inParticles.draw_markers.sobol_seqnow always runs on the host (NumPy), anddraw_markerscopies the result to the active backend withxp.to_cunumpy(setup only, both Sobol branches). Before,sobol_sequsedcunumpyand passed Python lists tocupy.transpose, which CuPy rejects. It is a loop over scalars, so it does not belong on the device anyway.Model-specific changes:
None.
How far the CuPy run gets without a GPU. I ran the CuPy part locally on cunumpy's fake CuPy (
CUNUMPY_FAKE_CUPY=1, clean MPI environment, ad-hoc script, not committed):On this branch alone, the first failure is in
Simulation.allocate, while the Derham projectors are assembled (CommutingProjector.__init__→KroneckerStencilMatrix.transpose→ feectools' devicetransposekernel):NotImplementedError: fake CuPy cannot compile CUDA kernels. Every CUDA launch fails there.With every CUDA launch emulated on the CPU (
emulated_launches()from CUDA version of the linear_vlasov_ampere accumulation #705'spic/tests/cuda_emulation.py), the Derham, the domain and the markers are set up on the device:cupy.transposerejects lists.SortingParameters(boxes_per_dim=...)the run also fails inSortingBoxes(round()of a CuPy scalar, then the pyccelinitialize_neighbourscalled with a device array). This test does not use sorting boxes for that reason; that bug is not fixed here.EfieldWeightsCoupling.allocate→MassMatrixPreconditioner(M1)→is_circulant:assert isinstance(mat, xp.ndarray). This is themass-preconditioner-cupydependency. Withprecond=Nonethe run still stops at the same assert, because the initial Poisson solve builds anL2Projector, which always creates aMassMatrixPreconditioner.So the time loop was not reached locally, and the transfer budget has not been tested against a real CuPy run yet. Once the dependencies are in, expect to tune the budget after the first H100 run, for example if the scalars are reduced on the host.
CI: devel has no GPU job that runs struphy's pytest suite. The old GitLab
test_cupyjob only runsutils/cupy_vs_numpy.pyand does not install struphy, and #710 replaces that file. The test is therefore not wired into CI here. Once #710 is merged, itsstruphy_gpu_testsjob should add:and its
mpi_gpu_testsjob:Documentation changes:
CUDA_STRATEGY.md: one checklist entry for the end-to-end test and its dependencies.Testing (macOS, no GPU; pyccel kernels compiled with GNU/Fortran; GPU tests not run):
devel)test_draw_parallel.py test_verif_LinearVlasovAmpereOneSpecies.py+ new testmpirun -n 2 --with-mpiThe skip is the GPU test (
CuPy/GPU not available).test_draw_parallel.pycovers the Sobol loading paths.🤖 Generated with Claude Code