Repository navigation
Polar splines on the CuPy backend - #720
Draft
max-models wants to merge 5 commits into
Draft
max-models wants to merge 5 commits into
max-models wants to merge 5 commits into
Conversation
Build the polar blocks (PolarExtractionBlocksC1) on the host from a host copy of the control points on every backend, and apply device copies of them (cupyx.scipy.sparse via xp.scipy.sparse, cached per block list) in PolarExtractionOperator/PolarLinearOperator when the CuPy backend is active. Synchronize device buffers before the ring Allreduces. Remove the NotImplementedError for polar_splines=True on CuPy. The restart left inverse of SplineFunction is built on the host. New test_polar_cupy.py compares NumPy and CuPy (fake CuPy without a GPU) for IGAPolarCylinder and IGAPolarTorus. Solves #695, part of #650. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Stack the PR on #718. Conflicts in CUDA_STRATEGY.md: porting-order gates and open questions now say that polar splines (#695) and MHD equilibria (#696) both run on CuPy; kept the marker exchange, linear_vlasov_ampere, vlasov_maxwell and polar splines notes, in PR order. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 8, 2026
Draft
max-models
added this pull request to stack #728
October 8, 2026 05:45
max-models
removed this pull request from stack #728
October 8, 2026 06:00
max-models
added this pull request to stack #729
October 8, 2026 06:00
max-models
marked this pull request as draft
October 8, 2026 09:13
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack: part 9 of 14, based on #718 (merge that first), next: #721. Full order: #709 → #708 → #711 → #712 → #705 → #713 → #714 → #718 → #720 → #721 → #724 → #725 → #723 → #722.
Solves the following issue(s):
Solves #695, part of #650.
Derham(polar_splines=True)raisedNotImplementedErroron CuPy (since PR 19, #683). The polar extraction operators applied SciPy sparse matrices to stencil data that lives on the device. The first failure wascsr_matrixof a device array inPolarExtractionBlocksC1.Core changes:
Blocks on the host (
polar/extraction_operators.py):PolarExtractionBlocksC1builds every block on every backend with NumPy/SciPy, from a host copy of the control points (xp.to_numpy(domain.cx)). This covers the basis and DOF extraction blocks and the polargrad/curl/divblocks. The blocks are setup data, so.T,.toarray()and the restart left inverse keep working on them. The legacyPolarSplines_C0_2D/C1_2Dclasses are unchanged.Products on the device (
polar/linear_operators.py):PolarExtractionOperatorandPolarLinearOperatorkeep theirblocks_*as host SciPy matrices.dotuses device copies:DeviceSparseMatrix, acupyx.scipy.sparse.csr_matrixobtained through cunumpy'sxp.scipy.sparse.3·n2 × 3·n2in eta1–eta2,n3 × n3in eta3). Blocks with no non-zeros are not copied; their products are zeros.kron_matvec_2dpasses transposed views._backend_blocksreturns the SciPy blocks themselves.set_device_sparse_module(module)swaps the sparse backend. This is only for tests without a GPU.MPI: the ring reductions in
dot_inner_tp_rings,dot_parts_of_polarandPolarVector.toarraycallcunumpy.mpi.synchronize_for_mpibeforeAllreduce. On CuPy these are device buffers (CUDA-aware MPI).Derham: theNotImplementedErrorguard for polar splines on CuPy is removed.SplineFunction._restart_extraction_opbuilds its pseudo-inverse with NumPy (host blocks).Tests: new
polar/tests/test_polar_cupy.py. It builds the same polar Derham on NumPy and CuPy forIGAPolarCylinderandIGAPolarTorusand compares:E,P(DOF extraction) and their transposes for all five spacesgrad/curl/divand their transposesPolarVectorarithmetic (+ - * neg += -= *= dot copy toarray)P0..P3SplineFunctioncoefficients (extract_coeffs, restart inverse)M0..MvIt also checks that every result lives on the expected backend.
test_polar_fake_cupy, subprocess withCUNUMPY_FAKE_CUPY=1):cupyx.scipy.sparseis replaced by a dense device stand-in, and kernels run in their host version, because the fake cannot launch CUDA kernels. The mass matrices are left out here: their weights need struphy's CUDA geometry kernels, which take CUDA argument objects. They could be added onceemulated_launches()from CUDA version of the linear_vlasov_ampere accumulation #705 is ondevel. A mutation check confirmed that the test catches host blocks applied to device data.test_polar_on_cupy, skipped here): the full comparison, mass matrices included, with realcupyx.scipy.sparse.test_device_blocks_numpy_backendchecks that the NumPy path uses the SciPy blocks unchanged.test_polar_splines_not_supported_on_cupyis removed.Not covered:
SplineFunction.__call__(field evaluation is not ported to CuPy for any space yet), and the polar mass-matrix preconditioners on CuPy.Model-specific changes:
None.
Documentation changes:
CUDA_STRATEGY.md:Local testing (macOS, no GPU, pyccel kernels compiled with GNU/Fortran; GPU tests not run; GitHub CI covers the rest):
polar/tests/test_polar_cupy.py,feec/tests/test_derham_gpu.py,test_l2_projectors.py::test_l2_projectors_polar: 6 passed, 5 skipped (GPU)polar/tests/test_polar.py(NumPy): 13 passed.test_mass_matrices.py::test_mass_polar(first two cases): 2 passed.mpirun -n 2ontest_polar.py::test_polar_adjoints_small_nel2,test_extraction_ops_and_derivativesandtest_restart_polar: 9 passed.test_mass_matrices.py::test_mass_preconditioner_polarlocally: one case takes about 22 min on this machine (it passed on unchangeddevel). I did not runtest_basis_ops_polarafter the change.🤖 Generated with Claude Code