Skip to content

Marker exchange between ranks on the device (CUDA-aware MPI) - #712

Draft
max-models wants to merge 2 commits into
compile-cuda-kernels-setupfrom
cuda-aware-marker-exchange
Draft

max-models wants to merge 2 commits into
compile-cuda-kernels-setupfrom
cuda-aware-marker-exchange

Conversation

@max-models

@max-models max-models commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Stack: part 4 of 14, based on #711 (merge that first), next: #705. Full order: #709 → #708 → #711 → #712 → #705 → #713 → #714 → #718 → #720 → #721 → #724 → #725 → #723 → #722.


Solves the following issue(s):

Solves #698, part of #650. Markers move between ranks without going through the host on the CuPy backend: Particles.mpi_sort_markers exchanges the marker rows as device buffers with CUDA-aware MPI.

Core changes:

  • Device exchange of marker rows (Particles._sendrecv_markers_device in pic/base.py, called by _sendrecv_markers on CuPy):
    • The rows to send (already gathered on the device per destination by _sendrecv_get_destinations) and one device receive buffer go to Isend/Irecv through cunumpy.mpi.mpi_buffer. With CUDA-aware MPI this passes the device arrays directly, after cunumpy.mpi.synchronize_for_mpi, as feectools' data exchangers have done since CUDA 2 - MPI with device buffers feectools#86. Without CUDA-aware MPI the same buffers are staged through (pinned) host memory, with a one-time warning, instead of handing device pointers to a host-only MPI.
    • Rows from rank i arrive in a contiguous slice of the receive buffer. The holes are filled in the same order as on NumPy, but by one device scatter instead of one per source rank.
    • Messages with a count of zero are skipped. Both sides know every count after the Alltoall, so they skip the same messages.
    • Whether MPI is CUDA-aware: the answer already recorded by xp.mpi.mpi_is_cuda_aware()/set_mpi_cuda_aware(), otherwise a collective probe of mpi_comm at the first exchange. This is safe because mpi_sort_markers is collective. The probe copies 32 bytes to the host once.
  • Counts on the host without copies: send_info/recv_info are small NumPy arrays on every backend. The send counts are sizes of nonzero results, which are host integers already (CuPy synchronizes to size them). So the Alltoall of the counts passes host buffers, and the receive buffer, hole slices and load-imbalance check need no device reads. Before this PR, CuPy used device count arrays: one host write per rank, Alltoall on device buffers, and list(recv_info), recv_info[i] and first_hole[i] read back element by element.
  • alpha: _sendrecv_determine_mtbs checks a tuple or list alpha on the host. Before, it asserted on device comparisons, which made two implicit reads. A uniform alpha (the usual case: 1.0 or 0.5 everywhere) becomes a Python scalar, so the 3-vector is no longer copied to the device. The values are the same on NumPy.
  • NumPy path unchanged: the same messages in the same order, and the same hole filling. The only differences are the NumPy count arrays, which xp already gave on NumPy, and the scalar alpha, which gives the same values.
  • Tests: new pic/tests/test_sorting_device.py:
    • test_mpi_sort_markers_fake_cupy[bc, cuda_aware] runs under mpirun -n N env CUNUMPY_FAKE_CUPY=1 pytest --with-mpi ... and is skipped unless the fake CuPy is active. It sorts the same markers on NumPy and on the fake CuPy across all ranks, two shifts in a row, so most markers change rank, with periodic and remove boundaries. It compares the whole marker arrays and valid_mks exactly, then sorts again with do_test=True.
      • With CUDA-aware MPI (the probe, or forced), it checks that count_transfers() is 0 during mpi_sort_markers and that no fake CuPy array is read by the host (int/bool/float/__index__/iteration/.get/.item/.tolist are patched to count). Before this PR there were 4 such reads per call.
      • With staging (set_mpi_cuda_aware(False)), it checks for exactly one to_host per non-empty send buffer, carrying n_send * n_cols * 8 bytes, and one to_device when something is received.
      • Fake CuPy arrays are host memory, so the non-CUDA-aware Open MPI here takes them through __cuda_array_interface__.
    • test_mpi_sort_markers_gpu[bc] (requires_cupy, mpirun -n 2 with one GPU per rank): the same check with the probed MPI, without the host-read patching.
    • CI: a step in test-PR-unit-mpi.yml runs the fake CuPy test with the same n-procs matrix (1–4).

Model-specific changes:

None. Not covered by this PR:

  • The SPH ghost-box exchange (_communicate_boxes → _sendrecv_markers_boxes) still passes device arrays to MPI without synchronize_for_mpi, and it reads _recv_info_box element by element. It can use the same pattern in a separate PR.
  • apply_kinetic_bc, which mpi_sort_markers calls first, was not changed. With the boundaries tested here, it makes no host/device copy and no host read.
  • A two-rank GPU run (the last item of Multi-rank marker exchange with CUDA-aware MPI #698) has not been done: no GPU here.

Documentation changes:

  • CUDA_STRATEGY.md: the "MPI + GPUs" open question now refers to the new section "Marker exchange implementation notes (Multi-rank marker exchange with CUDA-aware MPI #698)".
  • Docstrings of mpi_sort_markers, _sendrecv_all_to_all and the new _sendrecv_markers_device/_mpi_cuda_aware.

Testing (macOS, Open MPI 5.0.9 (not CUDA-aware), pyccel kernels compiled with GNU/Fortran; GPU tests not run):

before (devel) after
mpirun -n 2 pytest --with-mpi test_sorting.py test_pushers.py test_draw_parallel.py test_neighbor_ranks.py test_accum_vec_H1.py 205 passed 205 passed
same, mpirun -n 3 205 passed 205 passed
mpirun -n 2 pytest --with-mpi test_sph.py test_kernel_setup.py test_tesselation.py – 743 passed
mpirun -n 2 env CUNUMPY_FAKE_CUPY=1 pytest --with-mpi test_sorting_device.py 6 failed, 2 skipped 6 passed, 2 skipped (GPU)
same with -n 1, -n 3, -n 4 – 6 passed, 2 skipped

On devel, the fake CuPy test fails because of the implicit host reads (__bool__ ×3, __iter__) with CUDA-aware MPI, and because the device buffers are never staged when MPI is not CUDA-aware.

🤖 Generated with Claude Code

@max-models
max-models changed the base branch from devel to compile-cuda-kernels-setup October 7, 2026 21:58
max-models added a commit that referenced this pull request Oct 7, 2026
Stack the PR on #712.

Conflict in CUDA_STRATEGY.md: kept both sections (marker exchange notes, then linear_vlasov_ampere notes).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models added a commit that referenced this pull request Oct 7, 2026
Stack the PR on #713.

Conflict in CUDA_STRATEGY.md (porting-order gates): kept #714's removal of the MHD equilibria gate and
#705's note on 6D views; marker sorting is on the device since #712.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models added a commit that referenced this pull request Oct 7, 2026
Stack the PR on #714.

Conflicts in CUDA_STRATEGY.md: kept the linear_vlasov_ampere checklist entry and #690's decision in the
Next entry (without the steps done below), the marker exchange note of #712, and the polar-splines open question of #714.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@max-models
max-models added this pull request to stack #728 October 8, 2026 05:45
@max-models
max-models removed this pull request from stack #728 October 8, 2026 06:00
@max-models
max-models added this pull request to stack #729 October 8, 2026 06:00
On the CuPy backend, Particles._sendrecv_markers hands the marker rows to
Isend/Irecv as device buffers through cunumpy.mpi.mpi_buffer (directly after
synchronize_for_mpi with CUDA-aware MPI, staged through host memory
otherwise). Rows from each rank arrive in a contiguous slice of one device
receive buffer, and one device scatter fills the holes. The per-rank counts
are small NumPy arrays on every backend (they are host integers already), and
a uniform alpha is a Python scalar, so the exchange makes no host/device copy
and no implicit host read of a device array. The NumPy path is unchanged.

Adds pic/tests/test_sorting_device.py (fake CuPy under mpirun, and a GPU
variant) and a CI step running it on the fake CuPy.

Solves #698, part of #650.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant