Repository navigation
Field evaluation and diagnostics on the device - #724
Draft
max-models wants to merge 6 commits into
Draft
max-models wants to merge 6 commits into
max-models wants to merge 6 commits into
Conversation
On the CuPy backend the per-step diagnostics compute on the device and copy only reduced results to the host, explicitly and once per output step: - Scalars: values of a Scalars container are views into one device buffer, copied with one xp.to_numpy per update (Scalars.to_host/host_value); no float() per FunctionScalar; output and printing read the host copy. - DataContainer copies device datasets with xp.to_numpy (counted). - time_state is NumPy (host bookkeeping of the time loop). - Particles.binning works on CuPy (count_nonzero of a list failed) and gathers the valid markers once; saved markers gather without counting the mask. - SplineFunction: host-cached process bounds and masked writes in point flagging. - eval_spline_mpi_tensor_product_fixed moved into its own kernel folder with a CUDA version and parity cases; prepare_eval_tp_fixed computes on the host and copies once. Part of #697, #650. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Stack the PR on #721. Conflicts in CUDA_STRATEGY.md: checklist entries in PR order, kept the notes of every PR. Semantic fix for #721 below: SplineFunction.eval_tp_fixed_loc takes kind, pn and starts from the SplineArguments of the component (self._args_spline) instead of the removed tuple self._args_eval. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models
added a commit
that referenced
this pull request
Oct 7, 2026
…lders Stack the PR on #724. Conflicts in CUDA_STRATEGY.md: checklist entries in PR order, kept the notes of every PR; the kernel-folder row counts the 16 FEEC assembly folders, 4 spline evaluation folders and the 18 kernels with a CUDA version. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d starts from SplineArguments spline_evaluation_arguments returns (_data, args_spline) since #721. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models
added a commit
that referenced
this pull request
Oct 7, 2026
…lders Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…mpi_tensor_product_fixed) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models
added a commit
that referenced
this pull request
Oct 7, 2026
…lders Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models
added a commit
that referenced
this pull request
Oct 7, 2026
Stack the PR on #725. Conflicts: feec/mass.py, boundary_mass.py, basis_projection_ops.py (took #725's kernel folders; their outputs moved into the folders), test_cuda_emulation.py (#723's argument_arrays check on top of the tests added below), CUDA_STRATEGY.md (kept both notes). Semantic fixes for the PRs below: - OUTPUTS for the 16 FEEC assembly folders of #725 (same arguments as the former PyccelKernel(outputs=...) wraps, plus kernel_*d_vec and surface_kernel_3d_vec) and for eval_spline_mpi_tensor_product_fixed (#724). - OUTPUTS of eval_spline_mpi_markers/matrix/sparse_meshgrid point at values in the SplineArguments signatures of #721 (3, 5, 5). - argument_arrays (cuda_parity_cases.py) knows CudaSplineArguments (#721). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models
added a commit
that referenced
this pull request
Oct 7, 2026
Pulls in the two follow-up fixes of #724 (already resolved identically in the previous merge). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models
added a commit
that referenced
this pull request
Oct 8, 2026
Stack the PR on #723. Conflicts: eval_spline_mpi_{markers,matrix,sparse_meshgrid}_cuda.cu (SplineArgs of #721 with the qualified evaluation_kernels_3d::eval_spline_mpi), linalg_kernels.cuh, filler_kernels.cuh, particle_to_mat_kernels.cuh (helpers of #705/#713 moved into their module namespaces). Semantic fixes: the struphy_cuda::<pyccel module> rule applied to the CUDA code of the PRs below: linear_vlasov_ampere (#705), vlasov_maxwell (#713), bstar_parallel_3form (#718) and eval_spline_mpi_tensor_product_fixed (#724) use 'using namespace struphy_cuda;' and qualified helper calls; m_v_fill_b_v1_symm calls filler_kernels:: and evaluation_kernels_3d::; the fill_v1_symm test wrapper calls particle_to_mat_kernels::m_v_fill_b_v1_symm. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 8, 2026
Draft
max-models
added this pull request to stack #728
October 8, 2026 05:45
max-models
removed this pull request from stack #728
October 8, 2026 06:00
max-models
added this pull request to stack #729
October 8, 2026 06:00
max-models
marked this pull request as draft
October 8, 2026 09:13
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack: part 11 of 14, based on #721 (merge that first), next: #725. Full order: #709 → #708 → #711 → #712 → #705 → #713 → #714 → #718 → #720 → #721 → #724 → #725 → #723 → #722.
Stack fixes (for #721 below):
SplineFunction.eval_tp_fixed_locnow takeskind,pnandstartsfromself._args_spline[i], because the tuple_args_evalno longer exists. The parity case ofeval_spline_mpi_tensor_product_fixedunpacksspline_evaluation_argumentsthe same way.test_catalog_signaturesnow counts 4 spline evaluation kernels. Follow-up, not done here: switcheval_spline_mpi_tensor_product_fixedto theSplineArgumentsclass of #721.Solves the following issue(s):
Part of #697 (field evaluation and diagnostics on the device), tracked in #650. On the CuPy backend, the diagnostics of each output step now run on the device. Only the reduced results reach the host: the scalars in one copy per update, and every other saved dataset in one counted copy per output step. The marker diagnostics kernels in
pic/diagnostics/kernels(used by the 5D and hybrid models) are not ported yet. They are listed as follow-up work, so this PR does not close #697.Core changes:
I found the problems below without a GPU, on cunumpy's fake CuPy (
CUNUMPY_FAKE_CUPY=1, clean environment) withcunumpy.profiling.count_transfers. CUDA launches were emulated on the CPU withemulated_launches()from #705's worktree, in ad-hoc probes that are not committed. In those probes, implicit syncs (float(),bool(),__index__,.get()of a device array) were counted by patching the fakendarray.Scalars.update(FunctionScalarPICetc.)float()per function scalar (implicit sync)xp.to_numpyof all values (8 bytes per scalar)DataContainer.save_data.get()per device dataset, not seen bycount_transfers; scalars and time copied one by onexp.to_numpyper device dataset; scalars and time are already on the host, so they need no copyParticles.binningxp.count_nonzero(list)xp.histogramddon the device, no transfersupdate_markers_to_be_savedxp.count_nonzeroused as a slice bound (__index__sync)SplineFunction.__call__(meshgrid/point)bool()syncs per call (process bounds compared as device scalars)xp.copyto(..., where=...)SplineFunction.eval_tp_fixed_locxp.array(degree)uploaded per callDerham.prepare_eval_tp_fixed(setup)models/scalars.py):Scalarskeeps one device buffer. Each contained scalar'svalueis a view into it, soupdate()ends withto_host(), a singlexp.to_numpy.host_value(name)returns a NumPy view of the host copy.FunctionScalarFEEC/PIC/SPHstore the function's result through_scalar_valueinstead offloat(), so a device reduction stays on the device.StruphyModel.print_scalar_quantitiesand the HDF5 setup read the host copy.io/output_handling.py):DataContainer._as_numpy_arrayis nowxp.to_numpy. This is the one place where output copies device data.simulation/sim.py):time_stateis NumPy on both backends. The loop reads it withfloat()every step, which was a device sync on CuPy.TimeDependentSourcestill gets a scalar it can pass toxp.sin/xp.cos.pic/base.py): the valid markers are gathered once, and the Jacobian is evaluated once instead of twice. On NumPy the results are the same values as before.eval_spline_mpi_tensor_product_fixed: moved frombsplines/evaluation_kernels_3d.pyintobsplines/kernels/eval_spline_mpi_tensor_product_fixed/(folder rule). It has a new_cuda.cu(one thread per grid point, reusingeval_spline_mpi_kernel) and parity cases inpic/tests/cuda_parity_cases.py(additive; CUDA version of the linear_vlasov_ampere accumulation #705/CUDA version of the vlasov_maxwell accumulation #713 also add entries there).eval_tp_fixed_loctakeskind/pn/startsfromSplineFunction._args_evalinstead of building them for every call.markers[valid_mks], also insideKineticEnergyPIC), one per gather.BilinearEnergyFEEC/VolumeFormEnergyFEECstill copy one scalar per inner-product block inside feectools until feectools: inner products stay on the device #716 (Inner products stay on the device feectools#98) is in. With feectools: inner products stay on the device #716 these become device scalars, and this PR's container already handles them.Model-specific changes:
None.
LinearVlasovAmpereOneSpecies._compute_en_wnow stays on the device becauseFunctionScalarPICno longer callsfloat().Documentation changes:
CUDA_STRATEGY.md: one checklist entry and a short section "Diagnostics on the device (#697)" with the follow-ups: CUDA versions of the marker diagnostics kernels inpic/diagnostics/kernels/, and SPH kernel-density plots (step 6).Tests:
New
src/struphy/models/tests/test_diagnostics_on_device.pyhas five checks (scalars, binning, saved markers, output, point flagging) with mock particles. They need no CUDA kernel. Each check runs:test_on_numpy);test_on_fake_cupy), compared with NumPy, counting transfers, with implicit syncs turned into errors;test_on_cupy, skipped here).test_eval_tp_fixed_loc_on_cupy(GPU, skipped here) compares the CUDAeval_tp_fixed_locwith NumPy.New parity cases
eval_spline_mpi_tensor_product_fixed(2 cases), run by CPU emulation intest_cuda_emulation.py.Local tests run (macOS, no GPU, pyccel kernels compiled with GNU/Fortran; GPU tests not run):
models/tests/test_diagnostics_on_device.py,io/tests/test_output_handling.py,pic/tests/test_binning.py,models/tests/test_scalars.py: 24 passed, 6 skipped (GPU), 2 failed. The two failures aretest_en_fB_ignores_ghost_markersandtest_guiding_center_en_fB_does_not_divide_by_Npintest_scalars.py. Their mocks have nosave_magnetic_background_energy, whichGuidingCenter._compute_en_fBcalls, and the failures come from code this PR does not touch, so I think they also fail ondevel. I did not confirm this with a separatedevelrun.pic/tests/test_cuda_emulation.py -k tensor_productand the parity-case/signature checks oftest_cuda_parity.py: passed (GPU parity skipped).models/tests/verification/test_verif_LinearVlasovAmpereOneSpecies.py: 2 passed, serial and withmpirun -n 2 ... --with-mpi.test_variational_dissipation.py(the models that useeval_tp_fixed_loc). CI covers these.🤖 Generated with Claude Code