Skip to content

CUDA version of the linear_vlasov_ampere accumulation - #705

Draft
max-models wants to merge 9 commits into
cuda-aware-marker-exchangefrom
cuda-linear-vlasov-ampere
Draft

max-models wants to merge 9 commits into
cuda-aware-marker-exchangefrom
cuda-linear-vlasov-ampere

Conversation

@max-models

@max-models max-models commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Stack: part 5 of 14, based on #712 (merge that first), next: #713. Full order: #709 → #708 → #711 → #712 → #705 → #713 → #714 → #718 → #720 → #721 → #724 → #725 → #723 → #722.


Solves the following issue(s):

Part of #688 (blocked matrix accumulations), tracked in #650. First matrix accumulation on the GPU: linear_vlasov_ampere, the accumulation of EfieldWeightsCoupling in LinearVlasovAmpereOneSpecies (step 2 of the porting order in CUDA_STRATEGY.md). It uses the Array6D<double> views of cunumpy 0.6.1.

Core changes:

  • CUDA kernel pic/accumulation/kernels/linear_vlasov_ampere/linear_vlasov_ampere_cuda.cu: same arguments in the same order as the pyccel kernel. It launches one thread per marker row and skips holes and boundary particles. Per marker it calls df, matrix_inv, matrix_vector, outer and then m_v_fill_b_v1_symm. The six blocks mat11 … mat33 are Array6D<double> (strided), the vectors are Array3D<double>, and f0_values is const double*, as in push_weights_with_efield_lin_va.

  • Device helpers, 1:1 with pyccel (same names, arguments and loop order):

    • fill_mat and fill_mat_vec in filler_kernels.cuh
    • m_v_fill_b_v1_symm in particle_to_mat_kernels.cuh
    • outer in linalg_kernels.cuh

    Every addition into matrix or vector data is a cunumpy_atomic_add, as in charge_density_0form. The spline values live in the thread-local SplineScratch.

  • Accumulator matrix path on CuPy: no code change was needed, and the new test checks it. On CuPy, StencilMatrix._data/StencilVector._data are device arrays. The accumulator zeroes them in place and passes the same arrays to the kernel. The ghost-region exchange, the transposed blocks of the symmetric matrix (feectools' device transpose), the scaled operator built by EfieldWeightsCoupling for SchurSolver, its dot and the Schur solve all run on them.

  • Tests:

    • PARITY_CASES["linear_vlasov_ampere"] covers 129 markers (with a hole and a boundary particle) in Cuboid, Colella, HollowTorus and a 3d spline mapping. The data are shaped like the real accumulation matrices (v1_symm_accumulation_data() in kernel_test_args.py). Tolerances are rtol=1e-12, atol=1e-10: atomics change the summation order and entries reach 1e4. The CPU emulation catches a transposed DF.
    • m_v_fill_b_v1_symm is compared with pyccel through a wrapper kernel: test_v1_symm_filler on a GPU, test_emulated_v1_symm_filler by CPU emulation.
    • New test_accum_matrix_cupy.py accumulates as EfieldWeightsCoupling does (Colella, clamped in eta1) and compares with NumPy:
      • It checks the nine kernel arrays, the dense blocks, the accumulated vector, BC.dot(x) and three CG iterations of the Schur solve.
      • It also checks that the kernel writes into the blocks' own _data (no copies), and that nothing is copied to the host from the accumulation up to BC.dot(x).
      • On a GPU it runs with assert_no_transfers. Without a GPU it runs in a subprocess on cunumpy's fake CuPy (CUNUMPY_FAKE_CUPY=1, clean MPI environment). There, every CudaKernel launch is emulated on the CPU in place on the fake device arrays: the new emulated_launches() in pic/tests/cuda_emulation.py covers struphy's and feectools' kernels, struct arguments included. This test also fails on a transposed DF.

Model-specific changes:

None. LinearVlasovAmpereOneSpecies cannot run end to end on the GPU yet. Still needed:

  • CG inner products: feectools' StencilVector.inner returns its result on the host (xp.to_numpy), so each CG inner product copies a scalar. The Schur solve is outside the transfer guard because of this.
  • Default preconditioner: MassMatrixPreconditioner (the default of EfieldWeightsCoupling) cannot be created on CuPy, because is_circulant asserts a host array. The device Kronecker solve is a separate PR.
  • Parallel feectools PRs: the stencil 6D views and the device Kronecker solve.
  • H100 run: the first run on an H100 (First GPU run of the CUDA tests (H100) #687), then the model end to end (Run VlasovAmpereOneSpecies end to end on the GPU #689).

Documentation changes:

CUDA_STRATEGY.md:

  • linear_vlasov_ampere is marked as ported in the checklist and the porting order.
  • 6D views are no longer a gate.
  • New section: implementation notes for linear_vlasov_ampere.

Fix added after review (commit 1979ea9): the test hole was not a hole for accumulation kernels.
kernel_test_args.marker_arguments() marked the hole (row 0) only in column 8 (first_init_idx), which the pushers test.
Accumulation kernels, including linear_vlasov_ampere, test column 0, so this PR's parity case had no real hole.
Row 0 is now -1 in every column, as struphy marks holes. After the fix: test_cuda_emulation.py + test_accum_matrix_cupy.py 72 passed, 1 skipped; test_cuda_parity.py -k linear_vlasov_ampere passed on NumPy.

Testing (macOS, no GPU; pyccel kernels compiled with GNU/Fortran; GPU tests not run):

before (devel) after
test_kernel_backends.py test_cuda_parity.py test_device_helpers.py test_cuda_emulation.py 83 passed, 317 skipped 85 passed, 322 skipped
test_accum_matrix_cupy.py (new) – 1 passed (fake CuPy), 1 skipped (GPU)
mpirun -n 2: test_accum_vec_H1.py test_pushers.py test_verif_LinearVlasovAmpereOneSpecies.py 110 passed 110 passed

The new skips are GPU tests: 4 parity cases, test_v1_symm_filler and test_accumulator_matrix_on_cupy.

🤖 Generated with Claude Code

@max-models
max-models changed the base branch from devel to cuda-aware-marker-exchange October 7, 2026 21:59
max-models added a commit that referenced this pull request Oct 7, 2026
Stack the PR on #705.

Conflicts: CUDA_STRATEGY.md (kept both PRs' notes; vlasov_maxwell row stays done), cuda_parity_cases.py and
test_cuda_parity.py (kept the linear_vlasov_ampere and vlasov_maxwell cases and imports).
Semantic fix for #711 below: vlasov_maxwell has a CUDA version now, so test_compile_cuda_kernels expects
VlasovAmpereOneSpecies to have no missing CUDA kernels and checks fail-fast with push_bxu_Hdiv instead.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models added a commit that referenced this pull request Oct 7, 2026
Stack the PR on #713.

Conflict in CUDA_STRATEGY.md (porting-order gates): kept #714's removal of the MHD equilibria gate and
#705's note on 6D views; marker sorting is on the device since #712.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models added a commit that referenced this pull request Oct 8, 2026
Stack the PR on #723.

Conflicts: eval_spline_mpi_{markers,matrix,sparse_meshgrid}_cuda.cu (SplineArgs of #721 with the qualified
evaluation_kernels_3d::eval_spline_mpi), linalg_kernels.cuh, filler_kernels.cuh, particle_to_mat_kernels.cuh
(helpers of #705/#713 moved into their module namespaces).
Semantic fixes: the struphy_cuda::<pyccel module> rule applied to the CUDA code of the PRs below:
linear_vlasov_ampere (#705), vlasov_maxwell (#713), bstar_parallel_3form (#718) and
eval_spline_mpi_tensor_product_fixed (#724) use 'using namespace struphy_cuda;' and qualified helper calls;
m_v_fill_b_v1_symm calls filler_kernels:: and evaluation_kernels_3d::; the fill_v1_symm test wrapper
calls particle_to_mat_kernels::m_v_fill_b_v1_symm.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@max-models
max-models added this pull request to stack #728 October 8, 2026 05:45
@max-models
max-models removed this pull request from stack #728 October 8, 2026 06:00
@max-models
max-models added this pull request to stack #729 October 8, 2026 06:00
max-models and others added 2 commits October 8, 2026 09:25
- linear_vlasov_ampere_cuda.cu: same arguments as pyccel, one thread per marker
  row, matrix blocks as Array6D<double> views (cunumpy 0.6.1)
- device fill_mat, fill_mat_vec (filler_kernels.cuh), m_v_fill_b_v1_symm
  (particle_to_mat_kernels.cuh) and outer (linalg_kernels.cuh), with atomic adds
- parity cases (Cuboid, Colella, HollowTorus, 3d spline) and a device-helper
  test of m_v_fill_b_v1_symm, on a GPU and by CPU emulation
- test_accum_matrix_cupy.py: the matrix path of Accumulator on CuPy vs NumPy
  (device data, ghost regions, transposed blocks, Schur operator and solve),
  without a GPU on the fake CuPy with emulated kernel launches
  (emulated_launches() in pic/tests/cuda_emulation.py)
- CUDA_STRATEGY.md: porting order and implementation notes

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The hole in marker_arguments() was only marked in column 8 (first_init_idx),
which the pushers test. Accumulation kernels test column 0, so the
linear_vlasov_ampere parity case had no real hole.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
## Summary

`struphy lint` only enforced `ruff format`, ruff's import sorting (`ruff
check --select I`) and a check for `# $` OpenMP pragmas. `struphy
format` only ran those two ruff commands in a loop. This PR drops both
commands and uses ruff directly.

## Changes

- **Removed `struphy format` and `struphy lint`.** This deletes
`src/struphy/console/format.py` (~1550 lines) and its test.
- **Removed the OpenMP `# $` check.** pyccel now accepts a space after
`#` in OpenMP pragmas.
- **`struphy build-init-files` moved** to
`src/struphy/console/build_init_files.py`. It still regenerates
`models/`, `propagators/` and `geometry/domains/` `__init__.py`, then
sorts imports and formats them with ruff. Its `--linters`,
`--iterations` and `-y` options are removed. The package folders are now
found relative to the installed package rather than the current working
directory.
- **Import sorting is done by ruff** (`extend-select = ["I"]` in
`pyproject.toml`), replacing isort.
- **Dev extra:** removed `autopep8`, `isort`, `flake8`, `pylint`,
`ssort`, `add-trailing-comma` and `tabulate`, and added `ruff==0.15.0`.
Also removed the `[tool.autopep8]`, `[tool.isort]` and `[tool.pylint]`
config sections and the `[flake8]` section in `setup.cfg`.
- **CI:**
- GitHub: removed the `struphy_lint_all` and `isort` jobs. The `ruff`
job now runs `ruff format --check` and `ruff check`; until now `ruff
format --check` was commented out.
  - GitLab: the four `struphy lint` jobs are replaced by one `ruff` job.
  - Neither job needs a struphy install anymore.
- **One-time formatting:** `struphy lint` only covered `src/`, so I ran
`ruff check --fix` and `ruff format` on `tutorials/`, `doc/`, `utils/`
and `profiling/`. That's 25 files, all automatic fixes.
- **Docs:** the developer's guide has a new section on formatting and
import sorting, plus a section on `struphy build-init-files`.

## How to format now

```
ruff check --fix   # lint + sort imports
ruff format        # format
```

## Testing

- `ruff format --check` and `ruff check` pass on the whole repo.
- `struphy build-init-files` reproduces the committed `__init__.py`
files exactly.
- The console tests `test_state.py` and `test_params_profile.py` pass.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant