Skip to content

Compile the CUDA kernels in the setup stage (Simulation.compile_cuda_kernels) - #711

Open
max-models wants to merge 4 commits into
cuda-fused-pusher-preamblefrom
compile-cuda-kernels-setup
Open

max-models wants to merge 4 commits into
cuda-fused-pusher-preamblefrom
compile-cuda-kernels-setup

Conversation

@max-models

@max-models max-models commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Stack: part 3 of 14, based on #708 (merge that first), next: #712. Full order: #709 → #708 → #711 → #712 → #705 → #713 → #714 → #718 → #720 → #721 → #724 → #725 → #723 → #722.


Solves the following issue(s):

Solves #704, tracked in #650. Until now a CUDA kernel compiled on its first call, which put the compile time into the first time step and its profiling regions. Now Simulation.run() calls the new Simulation.compile_cuda_kernels() after allocate(), so the kernels compile before the time stepping.

Core changes:

  • Simulation.compile_cuda_kernels() (simulation/sim.py):
    • It runs after allocate() in run(), inside a new profiling region, setup: compile cuda kernels.
    • On the CuPy backend it calls Kernel.compile() once for each kernel of Simulation.kernels() and returns the compiled kernels.
    • On the NumPy backend it does nothing and returns (). NumPy runs are unaffected.
  • Fail fast: on CuPy, if any kernel the time stepping calls has no CUDA version, the method raises NotImplementedError before compiling anything.
    • The error lists every such kernel and the path where its .cu file is expected. For example, VlasovAmpereOneSpecies names vlasov_maxwell.
    • Without this check, the run fails at the first call of the first missing kernel, in the time loop.
    • Kernels with missing_cuda="fallback" are not counted as missing.
  • Simulation.kernels(): collects the kernels from the objects that keep and call them, through a new kernels() method on each owner. Each kernel is listed once (utils/kernel_compilation.collect_kernels). The owners:
    • Domain: the four geometry kernels.
    • Derham: the three spline evaluation kernels of SplineFunction, and feectools' stencil_dot_3d, stencil_transpose_3d, stencil_inner_3d and stencil_axpy_3d (called by the FEEC operators and solvers).
    • Particles: reflect, if an axis has reflecting boundary conditions.
    • Pusher: the pusher kernel and the kernels of its init/eval KernelSetups. KernelSetup, Accumulator and AccumulatorVector list their own kernel.
    • Propagator.kernels(): by default it collects from the propagator's attributes. That covers Kernels kept as attributes and objects with a kernels() method, also inside lists, tuples and dicts. It does not go into other objects. A propagator that calls a kernel it does not keep can override the method; none needs to today.
  • Not listed: kernels that only diagnostics or the initial solves call. They run during setup ("setup: initial diagnostics", "model post allocate") anyway.

Model-specific changes:

None. I checked the kernel lists against what the time loop actually calls. For each of 19 models (all particle models, plus Maxwell and LinearMHD), I ran the default parameter file for one step on NumPy while recording every Kernel call during model.integrate. Every kernel called in the time loop is in Simulation.kernels() for all 19 models. (This was an ad-hoc script; the committed test does the same check for VlasovAmpereOneSpecies.)

Documentation changes:

CUDA_STRATEGY.md has a short new section, "Compiling the CUDA kernels at setup (struphy#704)". The PR 16 notes now point to it from the sentence "when to compile them is decided later".

Tests (new simulation/tests/test_compile_cuda_kernels.py):

  • collect_kernels walks owners and containers, lists each kernel once, and missing_cuda finds the kernels without a CUDA version.
  • For a small VlasovAmpereOneSpecies, Simulation.kernels() is exactly the domain + Derham kernels plus push_eta_stage, vlasov_maxwell and push_v_with_efield, and the only kernel missing CUDA is vlasov_maxwell. Reflecting particles add reflect.
  • Every Kernel called in the time loop of a one-step run is listed.
  • On NumPy, compile_cuda_kernels() is a no-op (Kernel.compile patched to fail).
  • With the backend switched to CuPy by monkeypatching: every kernel is compiled, and a missing CUDA version raises before anything is compiled.
  • The same CuPy path runs in a subprocess on cunumpy's fake CuPy (CUNUMPY_FAKE_CUPY=1, clean MPI environment), with CudaKernel.compile recorded.
  • GPU test (skipped here): Vlasov compiles all its kernels at setup, and VlasovAmpereOneSpecies raises naming vlasov_maxwell.

Testing (macOS, no GPU; pyccel kernels compiled with GNU/Fortran; GPU tests not run):

The suites were simulation/tests, pic/tests/{test_kernel_setup,test_reflect,test_pushers,test_kernel_backends,test_cuda_emulation}.py, feec/tests/test_derham_gpu.py and geometry/tests/test_domain.py.

before (devel) after
those suites 336 passed, 55 skipped 343 passed, 56 skipped (7 new passed, the new GPU test skipped)

The "before" numbers are the same suites without the new test file; no existing test file was changed. Not verified without a GPU: the real NVRTC compile through Kernel.compile() and the GPU test. On the CPU, only the call path up to CudaKernel.compile is tested.

🤖 Generated with Claude Code

max-models and others added 2 commits October 7, 2026 20:49
…kernels)

Simulation.run() compiles the CUDA kernels of the time stepping after allocate(), in the
profiling region 'setup: compile cuda kernels', instead of at their first call. The kernels
are listed by the objects that keep them (kernels() on Domain, Derham, Particles, Pusher,
KernelSetup, Accumulator, AccumulatorVector and Propagator). On CuPy, kernels without a CUDA
version raise NotImplementedError at setup, all listed; on NumPy nothing happens.

Solves #704.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Stack the PR on #708.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@max-models
max-models changed the base branch from devel to cuda-fused-pusher-preamble October 7, 2026 21:56
max-models added a commit that referenced this pull request Oct 7, 2026
Stack the PR on #711.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models added a commit that referenced this pull request Oct 7, 2026
Stack the PR on #705.

Conflicts: CUDA_STRATEGY.md (kept both PRs' notes; vlasov_maxwell row stays done), cuda_parity_cases.py and
test_cuda_parity.py (kept the linear_vlasov_ampere and vlasov_maxwell cases and imports).
Semantic fix for #711 below: vlasov_maxwell has a CUDA version now, so test_compile_cuda_kernels expects
VlasovAmpereOneSpecies to have no missing CUDA kernels and checks fail-fast with push_bxu_Hdiv instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@max-models
max-models added this pull request to stack #728 October 8, 2026 05:45
@max-models
max-models removed this pull request from stack #728 October 8, 2026 06:00
@max-models
max-models added this pull request to stack #729 October 8, 2026 06:00
@max-models
max-models marked this pull request as draft October 8, 2026 09:15
@max-models
max-models marked this pull request as ready for review October 8, 2026 11:42

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants