Skip to content

Fuse the pusher preamble into one CUDA kernel - #708

Open
max-models wants to merge 3 commits into
cuda-bugfixesfrom
cuda-fused-pusher-preamble
Open

max-models wants to merge 3 commits into
cuda-bugfixesfrom
cuda-fused-pusher-preamble

Conversation

@max-models

@max-models max-models commented Oct 7, 2026 •

Copy link
Copy Markdown
Member

Stack: part 2 of 14, based on #709 (merge that first), next: #711. Full order: #709 → #708 → #711 → #712 → #705 → #713 → #714 → #718 → #720 → #721 → #724 → #725 → #723 → #722.


Fuse the pusher preamble into one CUDA kernel

On the CuPy backend, the three column-slice assignments at the start of Pusher._push

markers[:, first_pusher_idx:first_shift_idx] = markers[:, : 3 + vdim]
markers[:, first_shift_idx:residual_idx] = 0.0
markers[:, residual_idx:-2] = 0.0

each sweep the whole row-major marker buffer, because the slices are strided. With 4M markers (an 8M-row buffer of 1.66 GB), this took 4.6 ms per pusher call, more than the push kernels themselves.

Changes

  • New src/struphy/pic/pushing/prepare_push_cuda.cu: one thread per entry of the contiguous column range [first_pusher_idx, n_cols - 2), so all three assignments are done in a single coalesced pass.
  • Pusher._push calls it on the CuPy backend; the kernel is compiled once, on first use. The NumPy path is unchanged.

This roughly speeds up the preamble by 2x.

The next big improvement for the markers would be #707

Stack the PR on #709.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@max-models
max-models marked this pull request as ready for review October 8, 2026 11:42

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant