Skip to content

Declare kernel outputs for PyccelKernels - #723

Draft
max-models wants to merge 6 commits into
feec-assembly-kernel-foldersfrom
pyccel-kernel-outputs
Draft

max-models wants to merge 6 commits into
feec-assembly-kernel-foldersfrom
pyccel-kernel-outputs

Conversation

@max-models

@max-models max-models commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Stack: part 13 of 14, based on #725 (merge that first), next: #722. Full order: #709 → #708 → #711 → #712 → #705 → #713 → #714 → #718 → #720 → #721 → #724 → #725 → #723 → #722.

Stack fixes: added OUTPUTS to the 16 FEEC assembly folders of #725, using the same arguments as the former PyccelKernel(..., outputs=...) wraps in mass.py/boundary_mass.py/basis_projection_ops.py, plus kernel_*d_vec and surface_kernel_3d_vec. Also added OUTPUTS for eval_spline_mpi_tensor_product_fixed (#724). The indices of eval_spline_mpi_{markers,matrix,sparse_meshgrid} were moved to the SplineArguments signatures of #721 (3, 5, 5). argument_arrays now knows CudaSplineArguments. The test_cuda_emulation.py merge keeps the tests added below (test_emulated_v1_symm_filler). Local re-run on the stacked branch: test_cuda_parity.py, test_kernel_backends.py, test_evaluation_cuda.py (59 passed, 336 skipped), test_cuda_emulation.py (74 passed).


Solves the following issue(s):

Solves #342, part of the CUDA work tracked in #650. Struphy's PyccelKernels now declare which arguments they write to (outputs=), which cunumpy 0.6.1 already supports (PyccelKernel(..., outputs=...), Kernel.from_folder(..., host_options={"outputs": ...}), and assert_kernels_agree uses kernel.host_kernel.outputs by default). Before this PR, no struphy kernel declared outputs. Without a declaration, the CuPy fallback copies every converted array back to the device after each call, and the parity tests compare every array.

Core changes:

  • Kernel folders (all 90): each <package>/kernels/<name>/__init__.py declares

    # the arguments the kernel writes to, by position: args_markers
    OUTPUTS = (2,)
    
    push_eta_stage = Kernel.from_folder(__name__, structs=CUDA_STRUCTS, host_options={"outputs": OUTPUTS})

    The declarations come from a static analysis of the pyccel sources, which follows writes through aliases (markers = args_markers.markers) and through called kernel helpers. I reviewed the results by hand. The breakdown:

    • 41 pushers write only args_markers (only its markers field).
    • The accumulation kernels write their mat*/vec* arrays.
    • The diagnostics write markers or res.
    • Evaluation, geometry, SPH and local-projector kernels write their out/values/mat_f/f_eval* argument.

    The scratch arrays bn1 … bd3 of DerhamArguments are written by some kernels but are not outputs. They are not in the CUDA struct either.

  • Other PyccelKernel(...) wraps pass outputs= too:

    • WeightedMassOperator assembly (kernel_{1,2,3}d_mat, -1), eval_quad (kernel_*d_eval, -1), the matrix-free dot (kernel_3d_matrixfree, -2, which writes out._data; the last argument is v._data) and diagonal (kernel_3d_diag, -1), in feec/mass.py
    • BoundaryMassOperator (surface_kernel_3d_mat, -1), in feec/boundary_mass.py
    • BasisProjectionOperator (assemble_dofs_for_weighted_basisfuns_*d, 0), in feec/basis_projection_ops.py

    Negative indices keep one declaration valid for every dimension. The test-only wraps in test_mat_vec_filler.py and test_kernel_setup.py are unchanged.

  • Parity tests use the declarations:

    • test_cuda_parity.py: KernelCatalog.from_package builds its own Kernels and does not see the folders' OUTPUTS. The test therefore takes the kernels from the folder modules (DECLARED), and the GPU test_parity (assert_kernels_agree) now compares only the declared outputs. A side effect, found by reading the code but not run on a GPU: before, assert_kernels_agree also collected the host DerhamArguments scratch arrays (argument i.bn1, …). The CUDA struct has no such fields, so the host and CUDA key sets could not match for kernels taking args_derham.
    • New test_kernels_declare_outputs (no GPU needed): every kernel folder declares non-empty, in-range OUTPUTS.
    • New test_declared_outputs (NumPy backend, every parity case): the pyccel kernel changes at least one declared output and no other argument array. Argument objects are compared through their CUDA struct fields, so scratch is ignored. The helper argument_arrays lives in cuda_parity_cases.py.
    • test_cuda_emulation.test_emulated_parity now compares only the outputs with the pyccel kernel, and also asserts that the emulated CUDA kernel leaves every non-output array unchanged. I checked both new checks by mutation: declaring push_eta_stage with OUTPUTS = (4,) makes both tests fail.
  • CUDA_STRATEGY.md: short note on OUTPUTS.

Model-specific changes:

None. Behaviour on the NumPy backend is unchanged, because PyccelKernel uses outputs only when it converts CuPy arrays.

Documentation changes:

A short note in CUDA_STRATEGY.md.

Local tests run (macOS, no GPU; GitHub CI covers the rest):

  • pic/tests/test_cuda_parity.py: 18 passed, 213 skipped (the GPU tests)
  • pic/tests/test_cuda_emulation.py -k emulated_parity: all 15 kernels passed. This ran as two invocations: the first hit my 5-minute limit after 13 kernels, and the remaining two (kernel_pullpush_pic, kernel_pullpush) passed separately in 80 s.
  • pic/tests/test_kernel_backends.py and test_kernel_setup.py: 89 passed, 22 skipped
  • feec/tests/test_boundary_integrals.py and test_mass_reassembly.py: 25 passed

GPU tests were not run. On the GPU, test_parity now compares only the declared outputs. The outputs= of the FEEC mass/projection wraps only take effect on the CuPy fallback path. No local test exercises them, neither on a GPU nor with the fake CuPy.

🤖 Generated with Claude Code

max-models and others added 3 commits October 7, 2026 22:55
Each kernel folder's __init__.py declares the arguments its kernel writes to
(OUTPUTS, passed to cunumpy as host_options={"outputs": OUTPUTS}); the
PyccelKernel wraps in feec/mass.py, boundary_mass.py and
basis_projection_ops.py pass outputs= too. The parity tests compare the
declared outputs; new tests check the declarations on the NumPy backend
and in the CPU emulation of the CUDA kernels.

Solves #342.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Stack the PR on #725.

Conflicts: feec/mass.py, boundary_mass.py, basis_projection_ops.py (took #725's kernel folders; their outputs moved
into the folders), test_cuda_emulation.py (#723's argument_arrays check on top of the tests added below),
CUDA_STRATEGY.md (kept both notes).
Semantic fixes for the PRs below:
- OUTPUTS for the 16 FEEC assembly folders of #725 (same arguments as the former PyccelKernel(outputs=...)
  wraps, plus kernel_*d_vec and surface_kernel_3d_vec) and for eval_spline_mpi_tensor_product_fixed (#724).
- OUTPUTS of eval_spline_mpi_markers/matrix/sparse_meshgrid point at values in the SplineArguments
  signatures of #721 (3, 5, 5).
- argument_arrays (cuda_parity_cases.py) knows CudaSplineArguments (#721).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Pulls in the two follow-up fixes of #724 (already resolved identically in the previous merge).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@max-models
max-models changed the base branch from devel to feec-assembly-kernel-folders October 7, 2026 22:58
max-models added a commit that referenced this pull request Oct 8, 2026
Stack the PR on #723.

Conflicts: eval_spline_mpi_{markers,matrix,sparse_meshgrid}_cuda.cu (SplineArgs of #721 with the qualified
evaluation_kernels_3d::eval_spline_mpi), linalg_kernels.cuh, filler_kernels.cuh, particle_to_mat_kernels.cuh
(helpers of #705/#713 moved into their module namespaces).
Semantic fixes: the struphy_cuda::<pyccel module> rule applied to the CUDA code of the PRs below:
linear_vlasov_ampere (#705), vlasov_maxwell (#713), bstar_parallel_3form (#718) and
eval_spline_mpi_tensor_product_fixed (#724) use 'using namespace struphy_cuda;' and qualified helper calls;
m_v_fill_b_v1_symm calls filler_kernels:: and evaluation_kernels_3d::; the fill_v1_symm test wrapper
calls particle_to_mat_kernels::m_v_fill_b_v1_symm.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@max-models
max-models added this pull request to stack #728 October 8, 2026 05:45
@max-models
max-models removed this pull request from stack #728 October 8, 2026 06:00
@max-models
max-models marked this pull request as draft October 8, 2026 09:13

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant