Repository navigation
Hand-written CUDA vs code generation: port bstar_parallel_3form both ways (decision for #690) - #718
Draft
max-models wants to merge 4 commits into
Draft
max-models wants to merge 4 commits into
max-models wants to merge 4 commits into
Conversation
…ways (decision for #690) - bstar_parallel_3form_cuda.cu: hand-written CUDA version, 1:1 with pyccel, on the existing device helpers; new device helper eval_0form_spline_mpi in evaluation_kernels_3d.cuh. - 14 parity cases (every mapping, alpha = 1/0/mixed, wrap-around in mod, hole) in cuda_parity_cases.py; passes the CPU emulation. - pic/tests/codegen_spike: prototype generator translating the pyccel source (kernel and all helpers) into numba.cuda; agrees with pyccel in the numba.cuda simulator and under njit. Test-only, skipped without numba. - CUDA_STRATEGY.md: decision section (hand-written stays), comparison and what would change it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Stack the PR on #714. Conflicts in CUDA_STRATEGY.md: kept the linear_vlasov_ampere checklist entry and #690's decision in the Next entry (without the steps done below), the marker exchange note of #712, and the polar-splines open question of #714. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models
added a commit
that referenced
this pull request
Oct 7, 2026
Stack the PR on #718. Conflicts in CUDA_STRATEGY.md: porting-order gates and open questions now say that polar splines (#695) and MHD equilibria (#696) both run on CuPy; kept the marker exchange, linear_vlasov_ampere, vlasov_maxwell and polar splines notes, in PR order. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models
added a commit
that referenced
this pull request
Oct 8, 2026
Stack the PR on #723. Conflicts: eval_spline_mpi_{markers,matrix,sparse_meshgrid}_cuda.cu (SplineArgs of #721 with the qualified evaluation_kernels_3d::eval_spline_mpi), linalg_kernels.cuh, filler_kernels.cuh, particle_to_mat_kernels.cuh (helpers of #705/#713 moved into their module namespaces). Semantic fixes: the struphy_cuda::<pyccel module> rule applied to the CUDA code of the PRs below: linear_vlasov_ampere (#705), vlasov_maxwell (#713), bstar_parallel_3form (#718) and eval_spline_mpi_tensor_product_fixed (#724) use 'using namespace struphy_cuda;' and qualified helper calls; m_v_fill_b_v1_symm calls filler_kernels:: and evaluation_kernels_3d::; the fill_v1_symm test wrapper calls particle_to_mat_kernels::m_v_fill_b_v1_symm. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 8, 2026
max-models
added this pull request to stack #728
October 8, 2026 05:45
max-models
removed this pull request from stack #728
October 8, 2026 06:00
max-models
added this pull request to stack #729
October 8, 2026 06:00
max-models
marked this pull request as draft
October 8, 2026 09:13
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack: part 8 of 14, based on #714 (merge that first), next: #720. Full order: #709 → #708 → #711 → #712 → #705 → #713 → #714 → #718 → #720 → #721 → #724 → #725 → #723 → #722.
Solves the following issue(s):
Solves #690 (the decision part), tracked in #650. Decides whether the guiding-center kernels (step 3 of the porting order in
CUDA_STRATEGY.md) are ported by hand or generated from the pyccel source. To decide, one guiding-center kernel,bstar_parallel_3form, was ported both ways.Decision: hand-written CUDA stays. The comparison table, the reasoning and what would change the decision are in the new
CUDA_STRATEGY.mdsection Hand-written CUDA vs. code generation (decision for #690). Short version:numpy.absin the SPH helpers.CudaKernel..cu. numba's simulator only interprets the Python.cupyx.jitwould need the same rewriting of the pyccel source. Its reference lists no per-thread arrays and no struct arguments. Not run here (no CuPy).pyccel-cudafork (last commit January 2025) still has open issues for thread indexing and passing arrays..cufiles in place, behind the same tests..cutext (not numba), reusing the prototype's AST transformations.Core changes:
pic/pushing/kernels/bstar_parallel_3form/bstar_parallel_3form_cuda.cu, 1:1 with pyccel: same names and arguments, one thread per marker row, holes skipped as in pyccel (markers[ip, 0] == -1).numpy.mod(eta, 1.0)becomeseta - floor(eta)(FortranMODULO, which pyccel generates).df(every mapping),det,get_spans/SplineScratchandeval_spline_mpi_kernel.eval_0form_spline_mpiinbsplines/evaluation_kernels_3d.cuh.KernelSetupin the guiding-center propagators) are unchanged.pic/tests/cuda_parity_cases.py.eta + shiftwrap around inmod, row 0 has a negativeeta_nand row 2 is a hole.args_markers.n_markers.modfrom the CUDA kernel makes the emulation test fail.pic/tests/codegen_spike/(numba_codegen.py,compare.py,test_numba_codegen.py).NUMBA_ENABLE_CUDASIM=1) and compiled withnumba.njit.kernelsin the module names, sostruphy compileignores it.Model-specific changes:
None.
bstar_parallel_3formis one of the init kernels ofPushGuidingCenterBxEstarandPushGuidingCenterParallel. Neither propagator runs on the GPU yet: the rest of their chain is not ported (driftkinetic_hamiltonian,unit_b_1form,bstar_2form, the discrete-gradient pushers,gc_density_0form).Documentation changes:
CUDA_STRATEGY.md:Testing (macOS, no GPU; pyccel kernels compiled with GNU/Fortran; GPU tests not run, so the new kernel has not run on a GPU):
devel)test_cuda_emulation.py,test_cuda_parity.py,test_kernel_backends.py,test_device_helpers.pympirun -n 1 pytest test_kernel_setup.py test_pushers.pycodegen_spike/test_numba_codegen.pyThe "before" pass/skip split is derived from the collected tests: this PR adds exactly one emulation test and 14 GPU parity tests.
🤖 Generated with Claude Code