Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
67 commits
Select commit Hold shift + click to select a range
7f4d062
Added cupy-cuda12x to new [gpu] optional dependency
max-models Sep 30, 2026
7f85df7
Added CudaKernel, Kernel and KernelCatalog classes
max-models Sep 30, 2026
865019d
Added Cuda versions of the argument classes
max-models Sep 30, 2026
b2ca928
Added tests for the cuda kernel class
max-models Sep 30, 2026
0fed27a
Added demo files [skip ci]
max-models Sep 30, 2026
566eb5a
Merge remote-tracking branch 'origin/devel' into cuda-kernel-proof-of…
max-models Sep 30, 2026
606f926
Make SPH linear kernel gradients vanish at r=0 (#475)
max-models Sep 30, 2026
de328c9
Merge remote-tracking branch 'origin/devel' into cuda-kernel-proof-of…
max-models Sep 30, 2026
3f86e93
Removed transform()
max-models Sep 30, 2026
cea5a96
Cleaned up the changes in the core of the code since this branch shou…
max-models Sep 30, 2026
21865b6
Added CUDA_STRATEGY.md [skip ci]
max-models Sep 30, 2026
de851a7
Combined demo files and test file
max-models Sep 30, 2026
8a30ba2
Pass CUDA kernel arguments as they are and n_threads explicitly
max-models Sep 30, 2026
7d00df6
Cast Python scalars and flatten argument classes in CudaKernel again
max-models Sep 30, 2026
ae44cf0
Remove the scalar casts from CudaKernel, keep flattening argument cla…
max-models Sep 30, 2026
49343cb
Load CUDA kernels from <name>_cuda.cu files
max-models Sep 30, 2026
650d020
Add KernelCatalog for packages with one folder per kernel
max-models Sep 30, 2026
d7977f4
Let Pusher accept a Kernel and choose the kernel once at setup
max-models Sep 30, 2026
db9dced
Upgrade cunumpy and use cunumpy.get_backend()
max-models Sep 30, 2026
764ba54
Merge branch 'cuda-pr-1-proof-of-concept' into cuda-pr-2-kernel-files
max-models Sep 30, 2026
63f5eda
Merge branch 'cuda-pr-2-kernel-files' into cuda-pr-3-kernel-catalog
max-models Sep 30, 2026
9a698a3
Merge branch 'cuda-pr-3-kernel-catalog' into cuda-pr-4-pusher
max-models Sep 30, 2026
5d2f672
Merge remote-tracking branch 'origin/devel' into cuda-pr-1-proof-of-c…
max-models Sep 30, 2026
bcc0fa3
Merge branch 'cuda-pr-1-proof-of-concept' into cuda-pr-2-kernel-files
max-models Sep 30, 2026
fbb3340
Merge branch 'cuda-pr-2-kernel-files' into cuda-pr-3-kernel-catalog
max-models Sep 30, 2026
65ce485
Add Domain.cuda_args_domain and fix Domain deepcopy on the CuPy backend
max-models Sep 30, 2026
fb72c7e
Fix Domain.cuda_args_domain tests and backend dependence
max-models Sep 30, 2026
1bd1d80
Merge branch 'devel' into cuda-pr-1-proof-of-concept
max-models Sep 30, 2026
4c29ab6
Merge branch 'devel' into cuda-pr-2-kernel-files
max-models Sep 30, 2026
a09822a
Merge branch 'devel' into cuda-pr-3-kernel-catalog
max-models Sep 30, 2026
cd3ef6b
Merge branch 'devel' into cuda-pr-4-pusher
max-models Sep 30, 2026
aa9d4b8
Merge branch 'devel' into cuda-pr-5-domain
max-models Sep 30, 2026
39a4b9e
Merge branch 'devel' into cuda-pr-1-proof-of-concept
max-models Oct 1, 2026
6925961
Merge branch 'devel' into cuda-pr-1-proof-of-concept
max-models Oct 1, 2026
791960f
Add Argument baseclass
max-models Oct 1, 2026
13cf447
Merge branch 'cuda-pr-1-proof-of-concept' into cuda-pr-2-kernel-files
max-models Oct 1, 2026
060db87
Merge branch 'cuda-pr-2-kernel-files' into cuda-pr-3-kernel-catalog
max-models Oct 1, 2026
7004324
Merge branch 'cuda-pr-3-kernel-catalog' into cuda-pr-4-pusher
max-models Oct 1, 2026
7db6731
Merge branch 'devel' into cuda-pr-3-kernel-catalog
max-models Oct 1, 2026
f85b824
Merge branch 'devel' into cuda-pr-2-kernel-files
max-models Oct 1, 2026
5cf9f49
Merge branch 'devel' into cuda-pr-2-kernel-files
max-models Oct 1, 2026
cad5b83
Merge branch 'devel' into cuda-pr-3-kernel-catalog
max-models Oct 1, 2026
cec2e72
Merge branch 'cuda-pr-2-kernel-files' into cuda-pr-3-kernel-catalog
max-models Oct 1, 2026
9e4a8f7
Merge branch 'cuda-pr-3-kernel-catalog' into cuda-pr-4-pusher
max-models Oct 1, 2026
8e1306d
Merge branch 'devel' into cuda-pr-3-kernel-catalog
max-models Oct 1, 2026
2ced9f7
Merge branch 'cuda-pr-3-kernel-catalog' into cuda-pr-4-pusher
max-models Oct 1, 2026
79d2d6d
Merge branch 'cuda-pr-4-pusher' into cuda-pr-5-domain
max-models Oct 1, 2026
206ace5
Only set _args_domain once, no logic in the property
max-models Oct 1, 2026
9529d59
Renamed the helper methods
max-models Oct 1, 2026
470b43c
Particles on GPU
max-models Oct 1, 2026
6e0015b
Merge branch 'devel' into cuda-pr-6-particles-on-gpu
max-models Oct 1, 2026
cb3e4c6
Merge branch 'devel' into cuda-pr-5-domain
max-models Oct 1, 2026
bb4ccbd
Merge branch 'cuda-pr-5-domain' into cuda-pr-6-particles-on-gpu
max-models Oct 1, 2026
b2da690
Derham on the GPU: create Derham on CuPy, add Derham.cuda_args_derham
max-models Oct 1, 2026
763e3a1
Merge branch 'devel' into cuda-pr-5-domain
max-models Oct 1, 2026
333f810
Merge branch 'cuda-pr-5-domain' into cuda-pr-6-particles-on-gpu
max-models Oct 1, 2026
b785d68
CUDA_STRATEGY: PR 7 depends on feectools#85
max-models Oct 1, 2026
0c4cddc
Set self._args_markers in the __init__
max-models Oct 1, 2026
4d219fb
Merge branch 'devel' into cuda-pr-6-particles-on-gpu
max-models Oct 1, 2026
6d9c3a4
Merge branch 'cuda-pr-6-particles-on-gpu' into cuda-pr-7-derham
max-models Oct 1, 2026
1c72356
Merge branch 'devel' into cuda-pr-6-particles-on-gpu
max-models Oct 1, 2026
9097969
Merge branch 'cuda-pr-6-particles-on-gpu' into cuda-pr-7-derham
max-models Oct 1, 2026
20dfa66
Merge branch 'devel' into cuda-pr-6-particles-on-gpu
max-models Oct 1, 2026
ec35f3a
Merge branch 'cuda-pr-6-particles-on-gpu' into cuda-pr-7-derham
max-models Oct 1, 2026
079084f
Merge branch 'cuda-pr-7-derham' of github.com:struphy-hub/struphy int…
max-models Oct 1, 2026
0517cae
Merge branch 'devel' into cuda-pr-7-derham
max-models Oct 2, 2026
ac7eb6b
Remove cuda_args_derham
max-models Oct 2, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 17 additions & 5 deletions CUDA_STRATEGY.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ The work is split into small PRs that can be reviewed and merged one at a time.
- [x] **PR 4: `Pusher` accepts `Kernel`** — the kernel for the active backend is chosen once, when the pusher is created; a plain `PyccelKernel` is wrapped, so the propagators do not change (no behaviour change on CPU).
- [x] **PR 5: `Domain` on the GPU** — domain arguments are selected and stored at domain construction; CUDA arguments reference device arrays, and deepcopy/unpickling rebuilds the arguments from the copied or restored arrays.
- [x] **PR 6: `Particles` on the GPU** — `Particles` can be created on the CuPy backend, and `Particles.args_markers` is selected as the CUDA or Pyccel argument bundle at construction.
- [ ] **PR 7: `Derham` on the GPU** — `Derham` can be created on the CuPy backend, plus `Derham.cuda_args_derham`.
- [x] **PR 7: `Derham` on the GPU** — `Derham` can be created on the CuPy backend, plus `Derham.cuda_args_derham`.
- [ ] **PR 8: Shared CUDA headers for the argument classes** — one `.cuh` per argument class instead of long flat kernel signatures.
- [ ] **PR 9: One folder per kernel, starting with `pic/pushing`** — pure refactor, no behaviour change.
- [ ] **PR 10: Device versions of helper kernels** — B-spline evaluation, mapping evaluation (per domain), small linear algebra, as `__device__` functions in `.cuh` headers.
Expand Down Expand Up @@ -41,12 +41,12 @@ CUDA kernels can be added one by one. If the code runs on the GPU and needs a ke
- **No silent CPU fallback on the GPU.** A kernel without a CUDA version raises an error on the GPU backend. Falling back would mean copying data to the host and back at every call.
- **Small steps.** Every PR keeps the CPU code path working and tested.

## Current state (PR 6)
## Current state (PR 7)

| File | Content |
|---|---|
| `src/struphy/utils/kernel_backends.py` | `is_cuda_backend()`, `CudaKernel` (wraps a `cupy.RawKernel`, compiled lazily; expands `Argument.get_cuda_args()` and takes `n_threads`), `Kernel` and `KernelCatalog` for backend selection and discovery |
| `src/struphy/utils/cuda_arguments.py` | `Argument` contract plus `CudaMarkerArguments` and `CudaDomainArguments`; CUDA arrays are stored individually and returned in signature order by `get_cuda_args()` |
| `src/struphy/utils/cuda_arguments.py` | `Argument` contract plus `CudaMarkerArguments`, `CudaDerhamArguments` and `CudaDomainArguments`; CUDA arrays are stored individually and returned in signature order by `get_cuda_args()` |
| `src/struphy/geometry/base.py` | `Domain.args_domain` is selected once at construction; CUDA domains use device arrays, while direct Pyccel geometry calls retain a host argument bundle |
| `src/struphy/pic/base.py` | `Particles` arrays and `args_markers` use the backend selected at construction; direct Pyccel methods retain a private host bundle |
| `src/struphy/pic/tests/test_kernel_backends.py` | the demo kernel pair `push_eta_linear` (pyccel function compiled with `epyccel` at test time, CUDA source string) and tests on both backends |
Expand Down Expand Up @@ -126,7 +126,7 @@ kernel = catalog["push_eta_stage"] # Kernel: pyccel or CUDA depending on the ba
- The CUDA argument objects hold references. If an owner reallocates an array (today the markers are allocated once), it must rebuild its CUDA arguments at the same place, exactly like for the pyccel arguments.
- First these classes must be creatable on the CuPy backend at all:
- `Particles`: wrap Python lists in `xp.array` before reductions, and use host buffers for scalar MPI gathers (`pic/base.py`).
- `Derham`: NumPy arrays from feectools reach `cupy.ascontiguousarray`.
- `Derham`: feectools and struphy moved host data (knots, grids, collocation matrices) to CuPy before calling pyccel kernels (see PR 7 below).
- `Domain`: deepcopy and unpickling on CuPy failed (see PR 5 below).

### PR 5: `Domain` on the GPU (complete)
Expand All @@ -142,7 +142,19 @@ kernel = catalog["push_eta_stage"] # Kernel: pyccel or CUDA depending on the ba
- Particle arrays, validity masks, and boundary-condition codes are allocated through `cunumpy`, so they live on CuPy when the CuPy backend is active.
- `args_markers` is built as `CudaMarkerArguments` from device arrays on CuPy, or as `MarkerArguments` on NumPy. A private host bundle remains for direct Pyccel calls.
- Domain decomposition now wraps the Python `nprocs` list with `xp.array` before calling `xp.prod`. Scalar MPI gathers use small NumPy buffers and copy the results back to the active array backend, avoiding unsupported CuPy buffers in MPI calls.
- Full GPU particle pushing still depends on CUDA versions of the required kernels and on PR 7's `Derham` support.
- Full GPU particle pushing still depends on CUDA versions of the required kernels.

### PR 7: `Derham` on the GPU (complete)

- feectools and the struphy code that builds `Derham` followed `xp` everywhere. So on CuPy, the knots, quadrature grids and decomposition metadata became device arrays and then reached pyccel kernels, SciPy or MPI, which only take host arrays.
- feectools: struphy-hub/feectools#85, the first of the feectools CUDA PRs, makes feectools run on the CuPy backend. It must be merged, and the submodule or the feectools version bumped, before `Derham` can be created on CuPy.
- struphy: data that describes the spline spaces is host data on every backend. Only the coefficients (`StencilVector` data) and stencil matrices live on the device. The projection and quadrature grids of `Derham` (`get_pts_and_wts`, ...) and `spline_types_pyccel` are NumPy. `domain_array`, `index_array(_N/_D)` and `neighbours` are gathered with NumPy MPI buffers and then converted with `xp.asarray`, so they are device arrays on CuPy like `Particles.domain_array`.
- `Derham.args_derham` is built from the host knots, degrees and starts on both backends. `Derham.cuda_args_derham` lazily builds `CudaDerhamArguments` with one device copy of these small arrays. On NumPy it raises (host arrays are never copied to the device). The pyccel scratch arrays (`bn1`, ..., `bd3`) are not part of it; they become per-thread local arrays in CUDA (PR 10).
- Not supported on CuPy yet:
- Local projectors (`DerhamOptions.local_projectors=True`) raise `NotImplementedError` when the `Derham` is created. `CommutingProjectorLocal` builds its data with `xp` and calls pyccel kernels on it, like `Derham` did.
- Polar splines need a spline mapping, which cannot be created on CuPy yet (see PR 5).
- Field evaluation (`SplineFunction.__call__`, ...) still calls pyccel kernels with the coefficients, which are device arrays on CuPy. It needs CUDA evaluation kernels (PR 10+).
- Tests: `feec/tests/test_derham_gpu.py`. Without a GPU, a strict host stand-in for CuPy (rejects host/device mixing, cannot run kernels; not part of the repository) was used, on top of struphy-hub/feectools#85. With it, a `Derham` created on the "CuPy" backend matches the NumPy one on 1, 2 and 4 MPI processes.

### PR 8: Argument structs in shared headers

Expand Down
Loading
Loading