Repository navigation
Argument class for SplineFunction kernel arguments - #721
Draft
max-models wants to merge 4 commits into
Draft
max-models wants to merge 4 commits into
max-models wants to merge 4 commits into
Conversation
SplineFunction builds one SplineArguments (CudaSplineArguments on CuPy) per component, holding kind, pn, tn1, tn2, tn3 and starts. The eval_spline_mpi_* entry kernels (pyccel and CUDA) take it as args_spline instead of six loose arguments. New struct SplineArgs in kernel_arguments/spline_args.cuh, generated by write_spline_header(). Solves #678. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models
added a commit
that referenced
this pull request
Oct 7, 2026
Stack the PR on #721. Conflicts in CUDA_STRATEGY.md: checklist entries in PR order, kept the notes of every PR. Semantic fix for #721 below: SplineFunction.eval_tp_fixed_loc takes kind, pn and starts from the SplineArguments of the component (self._args_spline) instead of the removed tuple self._args_eval. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models
added a commit
that referenced
this pull request
Oct 7, 2026
…d starts from SplineArguments spline_evaluation_arguments returns (_data, args_spline) since #721. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models
added a commit
that referenced
this pull request
Oct 7, 2026
Stack the PR on #725. Conflicts: feec/mass.py, boundary_mass.py, basis_projection_ops.py (took #725's kernel folders; their outputs moved into the folders), test_cuda_emulation.py (#723's argument_arrays check on top of the tests added below), CUDA_STRATEGY.md (kept both notes). Semantic fixes for the PRs below: - OUTPUTS for the 16 FEEC assembly folders of #725 (same arguments as the former PyccelKernel(outputs=...) wraps, plus kernel_*d_vec and surface_kernel_3d_vec) and for eval_spline_mpi_tensor_product_fixed (#724). - OUTPUTS of eval_spline_mpi_markers/matrix/sparse_meshgrid point at values in the SplineArguments signatures of #721 (3, 5, 5). - argument_arrays (cuda_parity_cases.py) knows CudaSplineArguments (#721). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models
added a commit
that referenced
this pull request
Oct 8, 2026
Stack the PR on #723. Conflicts: eval_spline_mpi_{markers,matrix,sparse_meshgrid}_cuda.cu (SplineArgs of #721 with the qualified evaluation_kernels_3d::eval_spline_mpi), linalg_kernels.cuh, filler_kernels.cuh, particle_to_mat_kernels.cuh (helpers of #705/#713 moved into their module namespaces). Semantic fixes: the struphy_cuda::<pyccel module> rule applied to the CUDA code of the PRs below: linear_vlasov_ampere (#705), vlasov_maxwell (#713), bstar_parallel_3form (#718) and eval_spline_mpi_tensor_product_fixed (#724) use 'using namespace struphy_cuda;' and qualified helper calls; m_v_fill_b_v1_symm calls filler_kernels:: and evaluation_kernels_3d::; the fill_v1_symm test wrapper calls particle_to_mat_kernels::m_v_fill_b_v1_symm. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 8, 2026
Draft
max-models
added this pull request to stack #728
October 8, 2026 05:45
max-models
removed this pull request from stack #728
October 8, 2026 06:00
max-models
added this pull request to stack #729
October 8, 2026 06:00
max-models
marked this pull request as draft
October 8, 2026 09:13
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack: part 10 of 14, based on #720 (merge that first), next: #724. Full order: #709 → #708 → #711 → #712 → #705 → #713 → #714 → #718 → #720 → #721 → #724 → #725 → #723 → #722.
Stack merge: in
SplineFunction.__init__(feec/psydac_derham.py), theSplineArguments/CudaSplineArgumentsare now built after #709's degree check and its host-to-device copy of the knots.Solves the following issue(s):
Solves #678 (the remaining point: #667 (comment)). Part of #650.
Core changes:
SplineFunctionno longer builds(kind, pn, tn1, tn2, tn3, starts)tuples for the spline evaluation kernels. It now uses an argument class pair, likeDerhamArguments/CudaDerhamArguments:kernel_arguments/spline_args_kernels.SplineArguments(kind, pn, tn1, tn2, tn3, starts)and its CUDA twinkernel_arguments/spline_args_cuda.CudaSplineArguments(aCudaStructArgumentswith structSplineArgs). The headerkernel_arguments/spline_args.cuhis generated by the newwrite_spline_header().SPLINE_STRUCTSis added toCUDA_STRUCTS.SplineFunction.__init__creates one argument object per component:SplineArgumentson NumPy,CudaSplineArgumentson CuPy. They are stored in_args_spline, which replaces_args_eval. The CUDA degree check (degrees 1 to 8) moves fromSplineFunctionintoCudaSplineArguments, the same as inCudaDerhamArguments.eval_spline_mpi_markers,eval_spline_mpi_matrixandeval_spline_mpi_sparse_meshgrid(pyccel and CUDA, still 1:1) takeargs_splineinstead of the six loose arguments. The shared helpereval_spline_mpi(pyccel and__device__) keeps its flat signature, because other kernels call it too.CudaSplineArgumentsaccepts only C-contiguous knot vectors.SplineFunctionalready passes contiguous ones, sotest_evaluation_paritynow uses contiguous knots instead of strided views.No CUDA kernel other than these three takes the spline arguments, and pushers and accumulation already use
DerhamArguments. One possible follow-up:eval_spline_mpi_tensor_product_fixedis a pyccel-only kernel inevaluation_kernels_3d.pythatSplineFunctionstill calls with loosekind, pn, starts(no knots).Tests:
test_kernel_backends.py: the new pair is added toARGUMENT_PAIRS, so the mirror check (and the struct layout check on a GPU) covers it. New testtest_generated_spline_header. The owner test checks thatSplineFunctioncreatesSplineArgumentson NumPy.kernel_test_args.spline_evaluation_argumentsreturns(_data, args_spline), so the parity and emulation cases now pass the argument object.test_cuda_emulation.CUDA_CLASSESincludesCudaSplineArguments.Local tests (macOS, no GPU), after recompiling the changed kernels:
bsplines/tests,feec/tests/test_eval_field.py,feec/tests/test_derham_gpu.py,pic/tests/test_cuda_emulation.py,pic/tests/test_kernel_backends.py,pic/tests/test_cuda_parity.py,pic/tests/test_pushers.py,pic/tests/test_accum_vec_H1.py: 258 passed, 315 skipped. On devel, the same files gave 256 passed, 314 skipped; the differences are the 2 new tests and 1 new GPU-only skip.mpirun -n 2onbsplines/tests/test_eval_spline_mpi.py,feec/tests/test_eval_field.py,pic/tests/test_pushers.py,pic/tests/test_accum_vec_H1.py: 139 passed (139 on devel).test_emulated_parity[eval_spline_mpi_*]).CudaSplineArgumentsand packed the struct, and the degree check and the rejection of host arrays both work. The fake CuPy cannot construct a fullDerham, soSplineFunctionon CuPy was not run end-to-end.GPU tests were not run (no GPU available). This affects the struct layout, the parity tests and
test_spline_function_backends[cupy].Model-specific changes:
None.
Documentation changes:
CUDA_STRATEGY.md: the new pair is added to the argument class table and the current-state table, plus a short note in the PR 16 implementation notes.🤖 Generated with Claude Code