Repository navigation
Compile the CUDA kernels in the setup stage (Simulation.compile_cuda_kernels) - #711
Open
max-models wants to merge 4 commits into
Open
max-models wants to merge 4 commits into
max-models wants to merge 4 commits into
Conversation
…kernels) Simulation.run() compiles the CUDA kernels of the time stepping after allocate(), in the profiling region 'setup: compile cuda kernels', instead of at their first call. The kernels are listed by the objects that keep them (kernels() on Domain, Derham, Particles, Pusher, KernelSetup, Accumulator, AccumulatorVector and Propagator). On CuPy, kernels without a CUDA version raise NotImplementedError at setup, all listed; on NumPy nothing happens. Solves #704. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Stack the PR on #708. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models
added a commit
that referenced
this pull request
Oct 7, 2026
Stack the PR on #711. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models
added a commit
that referenced
this pull request
Oct 7, 2026
Stack the PR on #705. Conflicts: CUDA_STRATEGY.md (kept both PRs' notes; vlasov_maxwell row stays done), cuda_parity_cases.py and test_cuda_parity.py (kept the linear_vlasov_ampere and vlasov_maxwell cases and imports). Semantic fix for #711 below: vlasov_maxwell has a CUDA version now, so test_compile_cuda_kernels expects VlasovAmpereOneSpecies to have no missing CUDA kernels and checks fail-fast with push_bxu_Hdiv instead. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 8, 2026
Draft
max-models
added this pull request to stack #728
October 8, 2026 05:45
max-models
removed this pull request from stack #728
October 8, 2026 06:00
max-models
added this pull request to stack #729
October 8, 2026 06:00
max-models
marked this pull request as draft
October 8, 2026 09:15
max-models
marked this pull request as ready for review
October 8, 2026 11:42
spossann
approved these changes
Oct 8, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack: part 3 of 14, based on #708 (merge that first), next: #712. Full order: #709 → #708 → #711 → #712 → #705 → #713 → #714 → #718 → #720 → #721 → #724 → #725 → #723 → #722.
Solves the following issue(s):
Solves #704, tracked in #650. Until now a CUDA kernel compiled on its first call, which put the compile time into the first time step and its profiling regions. Now
Simulation.run()calls the newSimulation.compile_cuda_kernels()afterallocate(), so the kernels compile before the time stepping.Core changes:
Simulation.compile_cuda_kernels()(simulation/sim.py):allocate()inrun(), inside a new profiling region,setup: compile cuda kernels.Kernel.compile()once for each kernel ofSimulation.kernels()and returns the compiled kernels.(). NumPy runs are unaffected.NotImplementedErrorbefore compiling anything..cufile is expected. For example,VlasovAmpereOneSpeciesnamesvlasov_maxwell.missing_cuda="fallback"are not counted as missing.Simulation.kernels(): collects the kernels from the objects that keep and call them, through a newkernels()method on each owner. Each kernel is listed once (utils/kernel_compilation.collect_kernels). The owners:Domain: the four geometry kernels.Derham: the three spline evaluation kernels ofSplineFunction, and feectools'stencil_dot_3d,stencil_transpose_3d,stencil_inner_3dandstencil_axpy_3d(called by the FEEC operators and solvers).Particles:reflect, if an axis has reflecting boundary conditions.Pusher: the pusher kernel and the kernels of its init/evalKernelSetups.KernelSetup,AccumulatorandAccumulatorVectorlist their own kernel.Propagator.kernels(): by default it collects from the propagator's attributes. That coversKernels kept as attributes and objects with akernels()method, also inside lists, tuples and dicts. It does not go into other objects. A propagator that calls a kernel it does not keep can override the method; none needs to today.Model-specific changes:
None. I checked the kernel lists against what the time loop actually calls. For each of 19 models (all particle models, plus Maxwell and LinearMHD), I ran the default parameter file for one step on NumPy while recording every
Kernelcall duringmodel.integrate. Every kernel called in the time loop is inSimulation.kernels()for all 19 models. (This was an ad-hoc script; the committed test does the same check forVlasovAmpereOneSpecies.)Documentation changes:
CUDA_STRATEGY.mdhas a short new section, "Compiling the CUDA kernels at setup (struphy#704)". The PR 16 notes now point to it from the sentence "when to compile them is decided later".Tests (new
simulation/tests/test_compile_cuda_kernels.py):collect_kernelswalks owners and containers, lists each kernel once, andmissing_cudafinds the kernels without a CUDA version.VlasovAmpereOneSpecies,Simulation.kernels()is exactly the domain + Derham kernels pluspush_eta_stage,vlasov_maxwellandpush_v_with_efield, and the only kernel missing CUDA isvlasov_maxwell. Reflecting particles addreflect.Kernelcalled in the time loop of a one-step run is listed.compile_cuda_kernels()is a no-op (Kernel.compilepatched to fail).CUNUMPY_FAKE_CUPY=1, clean MPI environment), withCudaKernel.compilerecorded.Vlasovcompiles all its kernels at setup, andVlasovAmpereOneSpeciesraises namingvlasov_maxwell.Testing (macOS, no GPU; pyccel kernels compiled with GNU/Fortran; GPU tests not run):
The suites were
simulation/tests,pic/tests/{test_kernel_setup,test_reflect,test_pushers,test_kernel_backends,test_cuda_emulation}.py,feec/tests/test_derham_gpu.pyandgeometry/tests/test_domain.py.devel)The "before" numbers are the same suites without the new test file; no existing test file was changed. Not verified without a GPU: the real NVRTC compile through
Kernel.compile()and the GPU test. On the CPU, only the call path up toCudaKernel.compileis tested.🤖 Generated with Claude Code