Repository navigation
Marker exchange between ranks on the device (CUDA-aware MPI) - #712
Draft
max-models wants to merge 2 commits into
Draft
max-models wants to merge 2 commits into
max-models wants to merge 2 commits into
Conversation
max-models
added a commit
that referenced
this pull request
Oct 7, 2026
Stack the PR on #712. Conflict in CUDA_STRATEGY.md: kept both sections (marker exchange notes, then linear_vlasov_ampere notes). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models
added a commit
that referenced
this pull request
Oct 7, 2026
Stack the PR on #714. Conflicts in CUDA_STRATEGY.md: kept the linear_vlasov_ampere checklist entry and #690's decision in the Next entry (without the steps done below), the marker exchange note of #712, and the polar-splines open question of #714. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 8, 2026
Draft
max-models
added this pull request to stack #728
October 8, 2026 05:45
max-models
removed this pull request from stack #728
October 8, 2026 06:00
max-models
added this pull request to stack #729
October 8, 2026 06:00
On the CuPy backend, Particles._sendrecv_markers hands the marker rows to Isend/Irecv as device buffers through cunumpy.mpi.mpi_buffer (directly after synchronize_for_mpi with CUDA-aware MPI, staged through host memory otherwise). Rows from each rank arrive in a contiguous slice of one device receive buffer, and one device scatter fills the holes. The per-rank counts are small NumPy arrays on every backend (they are host integers already), and a uniform alpha is a Python scalar, so the exchange makes no host/device copy and no implicit host read of a device array. The NumPy path is unchanged. Adds pic/tests/test_sorting_device.py (fake CuPy under mpirun, and a GPU variant) and a CI step running it on the fake CuPy. Solves #698, part of #650. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models
force-pushed
the
cuda-aware-marker-exchange
branch
from
October 8, 2026 07:22
8d3ade3 to
f2e8676
Compare
max-models
marked this pull request as draft
October 8, 2026 09:15
7 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack: part 4 of 14, based on #711 (merge that first), next: #705. Full order: #709 → #708 → #711 → #712 → #705 → #713 → #714 → #718 → #720 → #721 → #724 → #725 → #723 → #722.
Solves the following issue(s):
Solves #698, part of #650. Markers move between ranks without going through the host on the CuPy backend:
Particles.mpi_sort_markersexchanges the marker rows as device buffers with CUDA-aware MPI.Core changes:
Particles._sendrecv_markers_deviceinpic/base.py, called by_sendrecv_markerson CuPy):_sendrecv_get_destinations) and one device receive buffer go toIsend/Irecvthroughcunumpy.mpi.mpi_buffer. With CUDA-aware MPI this passes the device arrays directly, aftercunumpy.mpi.synchronize_for_mpi, as feectools' data exchangers have done since CUDA 2 - MPI with device buffers feectools#86. Without CUDA-aware MPI the same buffers are staged through (pinned) host memory, with a one-time warning, instead of handing device pointers to a host-only MPI.iarrive in a contiguous slice of the receive buffer. The holes are filled in the same order as on NumPy, but by one device scatter instead of one per source rank.Alltoall, so they skip the same messages.xp.mpi.mpi_is_cuda_aware()/set_mpi_cuda_aware(), otherwise a collective probe ofmpi_commat the first exchange. This is safe becausempi_sort_markersis collective. The probe copies 32 bytes to the host once.send_info/recv_infoare small NumPy arrays on every backend. The send counts are sizes ofnonzeroresults, which are host integers already (CuPy synchronizes to size them). So theAlltoallof the counts passes host buffers, and the receive buffer, hole slices and load-imbalance check need no device reads. Before this PR, CuPy used device count arrays: one host write per rank,Alltoallon device buffers, andlist(recv_info),recv_info[i]andfirst_hole[i]read back element by element.alpha:_sendrecv_determine_mtbschecks a tuple or listalphaon the host. Before, it asserted on device comparisons, which made two implicit reads. A uniformalpha(the usual case: 1.0 or 0.5 everywhere) becomes a Python scalar, so the 3-vector is no longer copied to the device. The values are the same on NumPy.xpalready gave on NumPy, and the scalaralpha, which gives the same values.pic/tests/test_sorting_device.py:test_mpi_sort_markers_fake_cupy[bc, cuda_aware]runs undermpirun -n N env CUNUMPY_FAKE_CUPY=1 pytest --with-mpi ...and is skipped unless the fake CuPy is active. It sorts the same markers on NumPy and on the fake CuPy across all ranks, two shifts in a row, so most markers change rank, with periodic and remove boundaries. It compares the whole marker arrays andvalid_mksexactly, then sorts again withdo_test=True.count_transfers()is 0 duringmpi_sort_markersand that no fake CuPy array is read by the host (int/bool/float/__index__/iteration/.get/.item/.tolistare patched to count). Before this PR there were 4 such reads per call.set_mpi_cuda_aware(False)), it checks for exactly oneto_hostper non-empty send buffer, carryingn_send * n_cols * 8bytes, and oneto_devicewhen something is received.__cuda_array_interface__.test_mpi_sort_markers_gpu[bc](requires_cupy,mpirun -n 2with one GPU per rank): the same check with the probed MPI, without the host-read patching.test-PR-unit-mpi.ymlruns the fake CuPy test with the samen-procsmatrix (1–4).Model-specific changes:
None. Not covered by this PR:
_communicate_boxes→_sendrecv_markers_boxes) still passes device arrays to MPI withoutsynchronize_for_mpi, and it reads_recv_info_boxelement by element. It can use the same pattern in a separate PR.apply_kinetic_bc, whichmpi_sort_markerscalls first, was not changed. With the boundaries tested here, it makes no host/device copy and no host read.Documentation changes:
CUDA_STRATEGY.md: the "MPI + GPUs" open question now refers to the new section "Marker exchange implementation notes (Multi-rank marker exchange with CUDA-aware MPI #698)".mpi_sort_markers,_sendrecv_all_to_alland the new_sendrecv_markers_device/_mpi_cuda_aware.Testing (macOS, Open MPI 5.0.9 (not CUDA-aware), pyccel kernels compiled with GNU/Fortran; GPU tests not run):
devel)mpirun -n 2 pytest --with-mpitest_sorting.py test_pushers.py test_draw_parallel.py test_neighbor_ranks.py test_accum_vec_H1.pympirun -n 3mpirun -n 2 pytest --with-mpi test_sph.py test_kernel_setup.py test_tesselation.pympirun -n 2 env CUNUMPY_FAKE_CUPY=1 pytest --with-mpi test_sorting_device.py-n 1,-n 3,-n 4On
devel, the fake CuPy test fails because of the implicit host reads (__bool__×3,__iter__) with CUDA-aware MPI, and because the device buffers are never staged when MPI is not CUDA-aware.🤖 Generated with Claude Code