Repository navigation
Declare kernel outputs for PyccelKernels - #723
Draft
max-models wants to merge 6 commits into
Draft
max-models wants to merge 6 commits into
max-models wants to merge 6 commits into
Conversation
Each kernel folder's __init__.py declares the arguments its kernel writes to
(OUTPUTS, passed to cunumpy as host_options={"outputs": OUTPUTS}); the
PyccelKernel wraps in feec/mass.py, boundary_mass.py and
basis_projection_ops.py pass outputs= too. The parity tests compare the
declared outputs; new tests check the declarations on the NumPy backend
and in the CPU emulation of the CUDA kernels.
Solves #342.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Stack the PR on #725. Conflicts: feec/mass.py, boundary_mass.py, basis_projection_ops.py (took #725's kernel folders; their outputs moved into the folders), test_cuda_emulation.py (#723's argument_arrays check on top of the tests added below), CUDA_STRATEGY.md (kept both notes). Semantic fixes for the PRs below: - OUTPUTS for the 16 FEEC assembly folders of #725 (same arguments as the former PyccelKernel(outputs=...) wraps, plus kernel_*d_vec and surface_kernel_3d_vec) and for eval_spline_mpi_tensor_product_fixed (#724). - OUTPUTS of eval_spline_mpi_markers/matrix/sparse_meshgrid point at values in the SplineArguments signatures of #721 (3, 5, 5). - argument_arrays (cuda_parity_cases.py) knows CudaSplineArguments (#721). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Pulls in the two follow-up fixes of #724 (already resolved identically in the previous merge). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models
added a commit
that referenced
this pull request
Oct 8, 2026
Stack the PR on #723. Conflicts: eval_spline_mpi_{markers,matrix,sparse_meshgrid}_cuda.cu (SplineArgs of #721 with the qualified evaluation_kernels_3d::eval_spline_mpi), linalg_kernels.cuh, filler_kernels.cuh, particle_to_mat_kernels.cuh (helpers of #705/#713 moved into their module namespaces). Semantic fixes: the struphy_cuda::<pyccel module> rule applied to the CUDA code of the PRs below: linear_vlasov_ampere (#705), vlasov_maxwell (#713), bstar_parallel_3form (#718) and eval_spline_mpi_tensor_product_fixed (#724) use 'using namespace struphy_cuda;' and qualified helper calls; m_v_fill_b_v1_symm calls filler_kernels:: and evaluation_kernels_3d::; the fill_v1_symm test wrapper calls particle_to_mat_kernels::m_v_fill_b_v1_symm. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 8, 2026
max-models
added this pull request to stack #728
October 8, 2026 05:45
max-models
removed this pull request from stack #728
October 8, 2026 06:00
max-models
added a commit
that referenced
this pull request
Oct 8, 2026
max-models
marked this pull request as draft
October 8, 2026 09:13
max-models
added a commit
that referenced
this pull request
Oct 8, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack: part 13 of 14, based on #725 (merge that first), next: #722. Full order: #709 → #708 → #711 → #712 → #705 → #713 → #714 → #718 → #720 → #721 → #724 → #725 → #723 → #722.
Stack fixes: added
OUTPUTSto the 16 FEEC assembly folders of #725, using the same arguments as the formerPyccelKernel(..., outputs=...)wraps inmass.py/boundary_mass.py/basis_projection_ops.py, pluskernel_*d_vecandsurface_kernel_3d_vec. Also addedOUTPUTSforeval_spline_mpi_tensor_product_fixed(#724). The indices ofeval_spline_mpi_{markers,matrix,sparse_meshgrid}were moved to theSplineArgumentssignatures of #721 (3, 5, 5).argument_arraysnow knowsCudaSplineArguments. Thetest_cuda_emulation.pymerge keeps the tests added below (test_emulated_v1_symm_filler). Local re-run on the stacked branch:test_cuda_parity.py,test_kernel_backends.py,test_evaluation_cuda.py(59 passed, 336 skipped),test_cuda_emulation.py(74 passed).Solves the following issue(s):
Solves #342, part of the CUDA work tracked in #650. Struphy's
PyccelKernels now declare which arguments they write to (outputs=), which cunumpy 0.6.1 already supports (PyccelKernel(..., outputs=...),Kernel.from_folder(..., host_options={"outputs": ...}), andassert_kernels_agreeuseskernel.host_kernel.outputsby default). Before this PR, no struphy kernel declared outputs. Without a declaration, the CuPy fallback copies every converted array back to the device after each call, and the parity tests compare every array.Core changes:
Kernel folders (all 90): each
<package>/kernels/<name>/__init__.pydeclaresThe declarations come from a static analysis of the pyccel sources, which follows writes through aliases (
markers = args_markers.markers) and through called kernel helpers. I reviewed the results by hand. The breakdown:args_markers(only itsmarkersfield).mat*/vec*arrays.markersorres.out/values/mat_f/f_eval*argument.The scratch arrays
bn1 … bd3ofDerhamArgumentsare written by some kernels but are not outputs. They are not in the CUDA struct either.Other
PyccelKernel(...)wraps passoutputs=too:WeightedMassOperatorassembly (kernel_{1,2,3}d_mat,-1),eval_quad(kernel_*d_eval,-1), the matrix-freedot(kernel_3d_matrixfree,-2, which writesout._data; the last argument isv._data) anddiagonal(kernel_3d_diag,-1), infeec/mass.pyBoundaryMassOperator(surface_kernel_3d_mat,-1), infeec/boundary_mass.pyBasisProjectionOperator(assemble_dofs_for_weighted_basisfuns_*d,0), infeec/basis_projection_ops.pyNegative indices keep one declaration valid for every dimension. The test-only wraps in
test_mat_vec_filler.pyandtest_kernel_setup.pyare unchanged.Parity tests use the declarations:
test_cuda_parity.py:KernelCatalog.from_packagebuilds its ownKernels and does not see the folders'OUTPUTS. The test therefore takes the kernels from the folder modules (DECLARED), and the GPUtest_parity(assert_kernels_agree) now compares only the declared outputs. A side effect, found by reading the code but not run on a GPU: before,assert_kernels_agreealso collected the hostDerhamArgumentsscratch arrays (argument i.bn1, …). The CUDA struct has no such fields, so the host and CUDA key sets could not match for kernels takingargs_derham.test_kernels_declare_outputs(no GPU needed): every kernel folder declares non-empty, in-rangeOUTPUTS.test_declared_outputs(NumPy backend, every parity case): the pyccel kernel changes at least one declared output and no other argument array. Argument objects are compared through their CUDA struct fields, so scratch is ignored. The helperargument_arrayslives incuda_parity_cases.py.test_cuda_emulation.test_emulated_paritynow compares only the outputs with the pyccel kernel, and also asserts that the emulated CUDA kernel leaves every non-output array unchanged. I checked both new checks by mutation: declaringpush_eta_stagewithOUTPUTS = (4,)makes both tests fail.CUDA_STRATEGY.md: short note onOUTPUTS.Model-specific changes:
None. Behaviour on the NumPy backend is unchanged, because
PyccelKernelusesoutputsonly when it converts CuPy arrays.Documentation changes:
A short note in
CUDA_STRATEGY.md.Local tests run (macOS, no GPU; GitHub CI covers the rest):
pic/tests/test_cuda_parity.py: 18 passed, 213 skipped (the GPU tests)pic/tests/test_cuda_emulation.py -k emulated_parity: all 15 kernels passed. This ran as two invocations: the first hit my 5-minute limit after 13 kernels, and the remaining two (kernel_pullpush_pic,kernel_pullpush) passed separately in 80 s.pic/tests/test_kernel_backends.pyandtest_kernel_setup.py: 89 passed, 22 skippedfeec/tests/test_boundary_integrals.pyandtest_mass_reassembly.py: 25 passedGPU tests were not run. On the GPU,
test_paritynow compares only the declared outputs. Theoutputs=of the FEEC mass/projection wraps only take effect on the CuPy fallback path. No local test exercises them, neither on a GPU nor with the fake CuPy.🤖 Generated with Claude Code