Skip to content

Field evaluation and diagnostics on the device - #724

Draft
max-models wants to merge 6 commits into
spline-function-args-classfrom
field-diagnostics-on-device
Draft

max-models wants to merge 6 commits into
spline-function-args-classfrom
field-diagnostics-on-device

Conversation

@max-models

@max-models max-models commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Stack: part 11 of 14, based on #721 (merge that first), next: #725. Full order: #709 → #708 → #711 → #712 → #705 → #713 → #714 → #718 → #720 → #721 → #724 → #725 → #723 → #722.

Stack fixes (for #721 below): SplineFunction.eval_tp_fixed_loc now takes kind, pn and starts from self._args_spline[i], because the tuple _args_eval no longer exists. The parity case of eval_spline_mpi_tensor_product_fixed unpacks spline_evaluation_arguments the same way. test_catalog_signatures now counts 4 spline evaluation kernels. Follow-up, not done here: switch eval_spline_mpi_tensor_product_fixed to the SplineArguments class of #721.


Solves the following issue(s):

Part of #697 (field evaluation and diagnostics on the device), tracked in #650. On the CuPy backend, the diagnostics of each output step now run on the device. Only the reduced results reach the host: the scalars in one copy per update, and every other saved dataset in one counted copy per output step. The marker diagnostics kernels in pic/diagnostics/kernels (used by the 5D and hybrid models) are not ported yet. They are listed as follow-up work, so this PR does not close #697.

Core changes:

I found the problems below without a GPU, on cunumpy's fake CuPy (CUNUMPY_FAKE_CUPY=1, clean environment) with cunumpy.profiling.count_transfers. CUDA launches were emulated on the CPU with emulated_launches() from #705's worktree, in ad-hoc probes that are not committed. In those probes, implicit syncs (float(), bool(), __index__, .get() of a device array) were counted by patching the fake ndarray.

Path Before (fake CuPy) After
Scalars.update (FunctionScalarPIC etc.) one float() per function scalar (implicit sync) no syncs; one xp.to_numpy of all values (8 bytes per scalar)
DataContainer.save_data one raw .get() per device dataset, not seen by count_transfers; scalars and time copied one by one one counted xp.to_numpy per device dataset; scalars and time are already on the host, so they need no copy
Particles.binning fails: xp.count_nonzero(list) runs xp.histogramdd on the device, no transfers
update_markers_to_be_saved xp.count_nonzero used as a slice bound (__index__ sync) the row count is taken from the gather; no extra sync
SplineFunction.__call__ (meshgrid/point) six bool() syncs per call (process bounds compared as device scalars) bounds cached on the host once; flagging uses xp.copyto(..., where=...)
SplineFunction.eval_tp_fixed_loc fails: pyccel kernel called with device arrays; xp.array(degree) uploaded per call CUDA kernel, no transfers
Derham.prepare_eval_tp_fixed (setup) fails on CuPy (pyccel helpers given device buffers) computed on the host, copied to the device once
  • Scalars (models/scalars.py): Scalars keeps one device buffer. Each contained scalar's value is a view into it, so update() ends with to_host(), a single xp.to_numpy. host_value(name) returns a NumPy view of the host copy. FunctionScalarFEEC/PIC/SPH store the function's result through _scalar_value instead of float(), so a device reduction stays on the device. StruphyModel.print_scalar_quantities and the HDF5 setup read the host copy.
  • Output (io/output_handling.py): DataContainer._as_numpy_array is now xp.to_numpy. This is the one place where output copies device data.
  • Time state (simulation/sim.py): time_state is NumPy on both backends. The loop reads it with float() every step, which was a device sync on CuPy. TimeDependentSource still gets a scalar it can pass to xp.sin/xp.cos.
  • Binning (pic/base.py): the valid markers are gathered once, and the Jacobian is evaluated once instead of twice. On NumPy the results are the same values as before.
  • Kernel eval_spline_mpi_tensor_product_fixed: moved from bsplines/evaluation_kernels_3d.py into bsplines/kernels/eval_spline_mpi_tensor_product_fixed/ (folder rule). It has a new _cuda.cu (one thread per grid point, reusing eval_spline_mpi_kernel) and parity cases in pic/tests/cuda_parity_cases.py (additive; CUDA version of the linear_vlasov_ampere accumulation #705/CUDA version of the vlasov_maxwell accumulation #713 also add entries there). eval_tp_fixed_loc takes kind/pn/starts from SplineFunction._args_eval instead of building them for every call.
  • Not changed, still implicit syncs: boolean-mask gathers (markers[valid_mks], also inside KineticEnergyPIC), one per gather. BilinearEnergyFEEC/VolumeFormEnergyFEEC still copy one scalar per inner-product block inside feectools until feectools: inner products stay on the device #716 (Inner products stay on the device feectools#98) is in. With feectools: inner products stay on the device #716 these become device scalars, and this PR's container already handles them.

Model-specific changes:

None. LinearVlasovAmpereOneSpecies._compute_en_w now stays on the device because FunctionScalarPIC no longer calls float().

Documentation changes:

CUDA_STRATEGY.md: one checklist entry and a short section "Diagnostics on the device (#697)" with the follow-ups: CUDA versions of the marker diagnostics kernels in pic/diagnostics/kernels/, and SPH kernel-density plots (step 6).

Tests:

  • New src/struphy/models/tests/test_diagnostics_on_device.py has five checks (scalars, binning, saved markers, output, point flagging) with mock particles. They need no CUDA kernel. Each check runs:

    • on NumPy (test_on_numpy);
    • on fake CuPy in a subprocess (test_on_fake_cupy), compared with NumPy, counting transfers, with implicit syncs turned into errors;
    • on a real GPU (test_on_cupy, skipped here).

    test_eval_tp_fixed_loc_on_cupy (GPU, skipped here) compares the CUDA eval_tp_fixed_loc with NumPy.

  • New parity cases eval_spline_mpi_tensor_product_fixed (2 cases), run by CPU emulation in test_cuda_emulation.py.

Local tests run (macOS, no GPU, pyccel kernels compiled with GNU/Fortran; GPU tests not run):

  • models/tests/test_diagnostics_on_device.py, io/tests/test_output_handling.py, pic/tests/test_binning.py, models/tests/test_scalars.py: 24 passed, 6 skipped (GPU), 2 failed. The two failures are test_en_fB_ignores_ghost_markers and test_guiding_center_en_fB_does_not_divide_by_Np in test_scalars.py. Their mocks have no save_magnetic_background_energy, which GuidingCenter._compute_en_fB calls, and the failures come from code this PR does not touch, so I think they also fail on devel. I did not confirm this with a separate devel run.
  • pic/tests/test_cuda_emulation.py -k tensor_product and the parity-case/signature checks of test_cuda_parity.py: passed (GPU parity skipped).
  • models/tests/verification/test_verif_LinearVlasovAmpereOneSpecies.py: 2 passed, serial and with mpirun -n 2 ... --with-mpi.
  • Not run locally: the full suite, the other model tests, and test_variational_dissipation.py (the models that use eval_tp_fixed_loc). CI covers these.

🤖 Generated with Claude Code

On the CuPy backend the per-step diagnostics compute on the device and copy only
reduced results to the host, explicitly and once per output step:

- Scalars: values of a Scalars container are views into one device buffer,
  copied with one xp.to_numpy per update (Scalars.to_host/host_value); no
  float() per FunctionScalar; output and printing read the host copy.
- DataContainer copies device datasets with xp.to_numpy (counted).
- time_state is NumPy (host bookkeeping of the time loop).
- Particles.binning works on CuPy (count_nonzero of a list failed) and gathers
  the valid markers once; saved markers gather without counting the mask.
- SplineFunction: host-cached process bounds and masked writes in point flagging.
- eval_spline_mpi_tensor_product_fixed moved into its own kernel folder with a
  CUDA version and parity cases; prepare_eval_tp_fixed computes on the host and
  copies once.

Part of #697, #650.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Stack the PR on #721.

Conflicts in CUDA_STRATEGY.md: checklist entries in PR order, kept the notes of every PR.
Semantic fix for #721 below: SplineFunction.eval_tp_fixed_loc takes kind, pn and starts from the
SplineArguments of the component (self._args_spline) instead of the removed tuple self._args_eval.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@max-models
max-models changed the base branch from devel to spline-function-args-class October 7, 2026 22:25
max-models added a commit that referenced this pull request Oct 7, 2026
…lders

Stack the PR on #724.

Conflicts in CUDA_STRATEGY.md: checklist entries in PR order, kept the notes of every PR; the kernel-folder row
counts the 16 FEEC assembly folders, 4 spline evaluation folders and the 18 kernels with a CUDA version.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d starts from SplineArguments

spline_evaluation_arguments returns (_data, args_spline) since #721.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models added a commit that referenced this pull request Oct 7, 2026
…lders

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…mpi_tensor_product_fixed)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models added a commit that referenced this pull request Oct 7, 2026
…lders

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models added a commit that referenced this pull request Oct 7, 2026
Stack the PR on #725.

Conflicts: feec/mass.py, boundary_mass.py, basis_projection_ops.py (took #725's kernel folders; their outputs moved
into the folders), test_cuda_emulation.py (#723's argument_arrays check on top of the tests added below),
CUDA_STRATEGY.md (kept both notes).
Semantic fixes for the PRs below:
- OUTPUTS for the 16 FEEC assembly folders of #725 (same arguments as the former PyccelKernel(outputs=...)
  wraps, plus kernel_*d_vec and surface_kernel_3d_vec) and for eval_spline_mpi_tensor_product_fixed (#724).
- OUTPUTS of eval_spline_mpi_markers/matrix/sparse_meshgrid point at values in the SplineArguments
  signatures of #721 (3, 5, 5).
- argument_arrays (cuda_parity_cases.py) knows CudaSplineArguments (#721).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models added a commit that referenced this pull request Oct 7, 2026
Pulls in the two follow-up fixes of #724 (already resolved identically in the previous merge).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models added a commit that referenced this pull request Oct 8, 2026
Stack the PR on #723.

Conflicts: eval_spline_mpi_{markers,matrix,sparse_meshgrid}_cuda.cu (SplineArgs of #721 with the qualified
evaluation_kernels_3d::eval_spline_mpi), linalg_kernels.cuh, filler_kernels.cuh, particle_to_mat_kernels.cuh
(helpers of #705/#713 moved into their module namespaces).
Semantic fixes: the struphy_cuda::<pyccel module> rule applied to the CUDA code of the PRs below:
linear_vlasov_ampere (#705), vlasov_maxwell (#713), bstar_parallel_3form (#718) and
eval_spline_mpi_tensor_product_fixed (#724) use 'using namespace struphy_cuda;' and qualified helper calls;
m_v_fill_b_v1_symm calls filler_kernels:: and evaluation_kernels_3d::; the fill_v1_symm test wrapper
calls particle_to_mat_kernels::m_v_fill_b_v1_symm.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@max-models
max-models added this pull request to stack #728 October 8, 2026 05:45
@max-models
max-models removed this pull request from stack #728 October 8, 2026 06:00
@max-models
max-models added this pull request to stack #729 October 8, 2026 06:00
@max-models
max-models marked this pull request as draft October 8, 2026 09:13

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Field evaluation on the device (SplineFunction and diagnostics)

1 participant