Skip to content

Cuda 2 mpi sync - #86

Draft
max-models wants to merge 3 commits into
cuda-developmentfrom
cuda-2-mpi-sync
Draft

max-models wants to merge 3 commits into
cuda-developmentfrom
cuda-2-mpi-sync

Conversation

@max-models

@max-models max-models commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Summary

MPI is allowed on the CuPy backend and gives correct results with device buffers. This requires a CUDA-aware MPI library.

Changes

  • feectools.ddm.mpi no longer turns MPI off when ARRAY_BACKEND=cupy. The segfaults that check was guarding against come from MPI libraries that are not CUDA-aware.
  • cunumpy.synchronize_for_mpi is called before every MPI call that uses device buffers. CuPy launches kernels asynchronously and MPI doesn't know about CUDA streams, so a buffer that is still being written would be sent with wrong data and no error. The call is added in:
    • the blocking, non-blocking and interface data exchangers
    • the Allreduce in StencilVectorSpace.inner
    • the Alltoallv calls of the parallel Kronecker solver
  • Dependency: cunumpy>=0.3.0, which provides synchronize_for_mpi.
  • CuPy fixes that only the MPI tests reach:
    • xp.dot/xp.vdot on .flat iterators in StencilInterfaceMatrix._dot and in the pure-Python inner product
    • test_cart_1d assigning Python lists to CuPy arrays
  • New linalg/tests/test_mpi_device.py: compares distributed results against global references, and checks that the exchangers synchronize before calling MPI.

Testing

With 2 MPI ranks and a CUDA-aware Open MPI, all MPI tests in ddm and linalg pass on both backends. Serial tests are unchanged.

Stack

This is step 2 of 4 for CUDA support. The PRs are stacked, and each one targets devel-tiny, so the diff here also contains the earlier steps. Review only commit ccb39f8 in this PR, and merge them in order:

  1. Cuda 1 xp arrays #85 — Run feectools with the CuPy backend
  2. Cuda 2 mpi sync #86 — MPI with device buffers ← this PR
  3. Cuda 3 device binding #87 — Bind each MPI rank to its own GPU
  4. Cuda 4 device kernels #88 — Stencil operations on the device

🤖 Generated with Claude Code

max-models and others added 2 commits September 30, 2026 23:53
Make feectools work when cunumpy's backend is CuPy, without any device
kernels yet: kernels that stay on the host still copy their arrays.

- Wrap all Pyccel kernels (stencil, B-splines, field evaluation, DOF
  kernels) in cunumpy.PyccelKernel, so they accept CuPy arrays.
- Keep host-only metadata on NumPy: MPI/index bookkeeping in ddm (cart,
  partition, petsc) and fem.partitioning, Kronecker solver sizes, and
  index arithmetic with Python ints (compute_diag_len, math.prod).
- Stage data for host-only libraries by array, not by global backend:
  LAPACK/SuperLU direct solvers, SciPy FFT, SciPy sparse products.
- Fix calls that ran host kernels on device arrays: the second
  stencil2coo call in StencilMatrix.tosparse and the conjugate transpose.
- Vectorize the construction of the 1D collocation matrices in the
  global projectors (element-wise indexing was a device round trip per
  entry: 334 s of a 348 s Derham setup on the GPU).
- GMRES: take real scalars from CuPy views before modifying them.
- Fix StencilMatrix._update_ghost_regions_serial: the ghost region is
  pads * shifts wide (wrong whenever shifts > 1, on both backends).
- Tests: work with CuPy arrays; skip PETSc tests without petsc4py.

Serial tests pass on both backends (core, ddm, fem, linalg).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Allow MPI on the CuPy backend and make it correct with device buffers
(requires a CUDA-aware MPI library).

- feectools.ddm.mpi no longer disables MPI when ARRAY_BACKEND=cupy; the
  segfaults it guarded against come from MPI libraries that are not
  CUDA-aware.
- Call cunumpy.synchronize_for_mpi before every MPI call on device
  buffers: CuPy kernels run asynchronously and MPI does not know about
  CUDA streams, so a buffer still being written would be sent silently
  wrong. Covers the blocking, non-blocking and interface data exchangers,
  the Allreduce in StencilVectorSpace.inner and the Alltoallv calls of the
  parallel Kronecker solver. Requires cunumpy >= 0.3.0.
- Fix CuPy incompatibilities reached only by the MPI tests: xp.dot/vdot
  on .flat iterators in StencilInterfaceMatrix._dot and the pure-Python
  inner product, and test_cart_1d assigning Python lists to CuPy arrays.
- Add test_mpi_device.py: distributed results against global references,
  and that the exchangers synchronize before MPI.

With 2 MPI ranks and a CUDA-aware Open MPI, all MPI tests in ddm and
linalg pass on both backends; serial tests are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@max-models
max-models changed the base branch from devel-tiny to cuda-development October 2, 2026 13:04

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant