Skip to content

Move FEEC assembly kernels into kernel folders - #725

Draft
max-models wants to merge 6 commits into
field-diagnostics-on-devicefrom
feec-assembly-kernel-folders
Draft

max-models wants to merge 6 commits into
field-diagnostics-on-devicefrom
feec-assembly-kernel-folders

Conversation

@max-models

@max-models max-models commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Stack: part 12 of 14, based on #724 (merge that first), next: #723. Full order: #709 → #708 → #711 → #712 → #705 → #713 → #714 → #718 → #720 → #721 → #724 → #725 → #723 → #722.


Solves the following issue(s):

Solves #411, tracked in #650. The FEEC assembly kernels are the last entry kernels that were still called as PyccelKernel(getattr(module, name)) from shared modules. They now follow the folder rule of CUDA_STRATEGY.md, like the pusher, accumulation and geometry kernels. This is a pure restructuring: no numerical change.

Core changes:

  • One folder per kernel: the 16 entry kernels of feec/mass_kernels.py and feec/basis_projection_kernels.py each move into feec/kernels/<name>/, which holds <name>_kernels.py and an __init__.py declaring <name> = Kernel.from_folder(__name__, structs=CUDA_STRUCTS, missing_cuda="fallback"). There is no .cu file, because none of these kernels has a CUDA port yet. Both old modules are removed.
    • From mass_kernels.py: kernel_{1,2,3}d_mat, kernel_{1,2,3}d_vec, kernel_{1,2,3}d_eval, kernel_3d_matrixfree, kernel_3d_diag, surface_kernel_3d_mat, surface_kernel_3d_vec
    • From basis_projection_kernels.py: assemble_dofs_for_weighted_basisfuns_{1,2,3}d
  • The function bodies are unchanged. I checked this with an AST comparison of all 16 functions against devel. Each file imports only what it uses (numpy as np and/or shape).
  • Call sites in feec/mass.py, feec/boundary_mass.py and feec/basis_projection_ops.py import the kernels. For kernels that depend on the dimension, they pick the right one from a small module-level dict (_MAT_KERNELS[ldim], _EVAL_KERNELS, _MATRIXFREE_KERNELS, _DIAG_KERNELS, _ASSEMBLY_KERNELS) instead of getattr on a module. The direct pyccel calls of mass_kernels.kernel_3d_vec in L2Projector also go through the Kernel now.
  • missing_cuda="fallback": on the CuPy backend these kernels keep the behaviour of the former PyccelKernel calls: they run on host copies of the arrays, with a warning the first time. With the default ("raise"), FEEC setups that currently assemble mass matrices on CuPy would start to fail. When a kernel is ported, its port removes the fallback for that kernel.
  • collect_kernel_files() picks up the new files with no change, since their names contain kernels. It now lists 142 files, 17 of them under feec/kernels.
  • Tests: test_catalog_signatures expects 17 kernels in feec.kernels (was 1). feec.kernels was already in test_cuda_parity.PACKAGES, so the new folders get the signature check and the folder-declaration check.

Model-specific changes:

None.

Documentation changes:

  • feec/kernels/__init__.py: the docstring lists the assembly kernels and the fallback.
  • CUDA_STRATEGY.md:
    • a checklist entry for this PR
    • the kernel count in "Current state"
    • a short section "FEEC assembly kernels (Put each kernel into its own folder and file #411)" saying that these kernels are now ready for CUDA ports (add <name>_cuda.cu and parity cases, then drop the fallback)

Testing (macOS, no GPU; pyccel kernels compiled in the worktree with GNU/Fortran; GPU tests not run; following the agreed scope, I ran only a few directly relevant tests, and CI covers the rest):

  • feec/tests/test_mass_reassembly.py: 2 passed
  • feec/tests/test_basis_ops.py::test_some_basis_ops and ::test_transposed_update_weights_drops_zero_blocks: passed
  • feec/tests/test_boundary_integrals.py -k "unit_cube_constant or hcurl_per_face": 12 passed
  • feec/tests/test_utilities_kernels.py, pic/tests/test_kernel_backends.py::test_catalog_signatures, pic/tests/test_cuda_parity.py::test_folders_declare_their_kernels: passed
  • An ad-hoc script (not committed) on Colella with 6×5×4 elements and degree (2,3,2):
    • Matrix-free M0–M3 dot agree with the assembled operators to about 1e-15 relative (kernel_3d_matrixfree against kernel_3d_mat).
    • Matrix-free and assembled diagonal() agree exactly (kernel_3d_diag).
    • L2Projector.get_dofs (kernel_3d_vec) and eval_quad (kernel_3d_eval) run.

Not run locally: the full test suites, test_mass_matrices.py (32³ and 48³ grids, slow) and MPI runs. No MPI code changed.

🤖 Generated with Claude Code

max-models and others added 2 commits October 7, 2026 22:59
The 16 entry kernels of feec/mass_kernels.py and feec/basis_projection_kernels.py
(used by WeightedMassOperator, StencilMatrixFreeMassOperator, L2Projector,
BoundaryMassOperator and BasisProjectionOperator) move into one folder each under
feec/kernels/<name>/ with <name>_kernels.py and an __init__.py declaring the
cunumpy Kernel. Call sites import the kernels instead of wrapping
PyccelKernel(getattr(module, name)). The function bodies are unchanged.

The kernels have no CUDA version yet and are declared with
missing_cuda="fallback", which keeps the former PyccelKernel behaviour on the
CuPy backend.

Solves #411.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lders

Stack the PR on #724.

Conflicts in CUDA_STRATEGY.md: checklist entries in PR order, kept the notes of every PR; the kernel-folder row
counts the 16 FEEC assembly folders, 4 spline evaluation folders and the 18 kernels with a CUDA version.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@max-models
max-models changed the base branch from devel to field-diagnostics-on-device October 7, 2026 22:27
max-models and others added 2 commits October 8, 2026 00:43
…lders

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lders

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models added a commit that referenced this pull request Oct 7, 2026
Stack the PR on #725.

Conflicts: feec/mass.py, boundary_mass.py, basis_projection_ops.py (took #725's kernel folders; their outputs moved
into the folders), test_cuda_emulation.py (#723's argument_arrays check on top of the tests added below),
CUDA_STRATEGY.md (kept both notes).
Semantic fixes for the PRs below:
- OUTPUTS for the 16 FEEC assembly folders of #725 (same arguments as the former PyccelKernel(outputs=...)
  wraps, plus kernel_*d_vec and surface_kernel_3d_vec) and for eval_spline_mpi_tensor_product_fixed (#724).
- OUTPUTS of eval_spline_mpi_markers/matrix/sparse_meshgrid point at values in the SplineArguments
  signatures of #721 (3, 5, 5).
- argument_arrays (cuda_parity_cases.py) knows CudaSplineArguments (#721).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models added a commit that referenced this pull request Oct 7, 2026
Pulls in the two follow-up fixes of #724 (already resolved identically in the previous merge).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@max-models
max-models added this pull request to stack #728 October 8, 2026 05:45
@max-models
max-models removed this pull request from stack #728 October 8, 2026 06:00
@max-models
max-models added this pull request to stack #729 October 8, 2026 06:00

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant