Repository navigation
feectools: inner products stay on the device - #716
Merged
Merged
Conversation
Moves the feectools submodule to struphy-hub/feectools#96 (branch device-kronecker-solve): KroneckerLinearSolver solves device data with dense inverses of the 1D matrices (one GEMM per direction) instead of copying it to the host for LAPACK/SuperLU, so MassMatrixPreconditioner solves on the CuPy backend make no host/device transfers. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Point the feectools submodule at struphy-hub/feectools#97 (branch stencil-6d-views): the 3D stencil dot/transpose kernels take the matrix data as 6D views, and StencilMatrix.dot/vdot/transpose no longer raise for matrices with fewer than 2p+1 diagonals. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Move the feectools submodule to struphy-hub/feectools#98: on the CuPy backend inner products return 0-d device arrays and the Krylov solvers copy one scalar per iteration to the host (the convergence test). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 7, 2026
…6d-views Stack #703 -> #706: feectools submodule points to the head of struphy-hub/feectools#97, which now contains #96. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 7, 2026
Merged
max-models
marked this pull request as draft
October 8, 2026 09:13
max-models
marked this pull request as ready for review
October 8, 2026 09:14
spossann
previously approved these changes
Oct 8, 2026
spossann
approved these changes
Oct 9, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack: part 3 of 3 — based on #706 (
feectools-stencil-6d-views, merge that first, after #703), next: none (top of the stack).Corresponding update in feectools: struphy-hub/feectools#98. Mirrors the feectools stack struphy-hub/feectools#96 → struphy-hub/feectools#97 → struphy-hub/feectools#98.
feectools-stencil-6d-viewsis merged into this branch (no rebase), and thefeectoolssubmodule now points at the stacked head of struphy-hub/feectools#98,6d88806, which contains feectools#96 and #97. Against its base (#706) this PR only moves the submoduled16a6ad→6d88806. After each feectools PR of the stack is merged intodevel-tiny, thefeectoolssubmodule of the struphy stack must be moved to the corresponding mergeddevel-tinycommit (thepr-feectools-submodulecheck fails until then).Solves the following issue(s):
Part of #689 (
VlasovAmpereOneSpeciesend to end on the GPU, no host/device transfers in the time loop), tracked in #650. This PR moves thefeectoolssubmodule to struphy-hub/feectools#98 ("Inner products stay on the device"). On the CuPy backend, every CG iteration of a struphy solve copied two scalars from the device to the host; after this change it copies one (the convergence test).The
pr-feectools-submodulecheck fails until struphy-hub/feectools#98 is merged intodevel-tiny. After that, the submodule should point at the merge commit.Core changes:
feectoolssubmodule:d16a6ad(base feectools: stencil 3D device kernels with 6D views #706) →6d88806(branchinner-on-deviceof Inner products stay on the device feectools#98, stacked on feectools#96 and Remove deprecated console commands #97; originallya15a8e8). No struphy code changes.What changes in feectools, and only on the CuPy backend:
StencilVector.inner/BlockVector.inner/dot_innerreturn a 0-d device array, in serial and with MPI.axpyaccepts such a scalar without copying it to the host.On NumPy, results and types are unchanged.
Copies to the host per iteration, counted under cunumpy's fake CuPy:
struphy call sites of
.innerthat run on CuPy now get a 0-d device array. I checked them:models/scalars.pywrite it intoxpbuffers (local_value[0] = ...), which works on the device;linear_mhd.py,shear_alfven.py, ...) return it, and the scalar machinery handles it like the existingdot_innerresults;PolarVector.dotadds a NumPy scalar to it, which works, but polar splines are not on CuPy yet (Polar splines on the CuPy backend #695);.inner. It works with device scalars: the arithmetic stays on the device, and the comparisons with 0 copy to the host implicitly.This removes the first blocker listed in CUDA version of the linear_vlasov_ampere accumulation #705 ("CG inner products ... The Schur solve is outside the transfer guard because of this"). The remaining copy per CG iteration still counts as a transfer for
assert_no_transfers. The transfer guard could then cover the Schur solve with an allowance of one 8-byte copy per iteration.Model-specific changes:
None.
Documentation changes:
None in struphy. The feectools PR documents the change in feectools'
CUDA_STRATEGY.md(section "Inner products on the device").Testing (macOS, no GPU; struphy kernels compiled with GNU/Fortran; GPU tests not run):
feectools (see Inner products stay on the device feectools#98):
ddm/tests/test_cart_*d.py;mpirun -n 2linalg MPI suite: 738 → 739 passed;struphy, solver tests on the NumPy backend, with feectools
devel-tiny(before) and this branch (after):propagators/tests/test_poisson.py -m "not mpi" -k "not multigrid": 40 passed before, 40 passed after;linear_algebra/tests/test_saddlepoint_massmatrices.py: the first case (SaddlePointSolverUzawaNumpy) passed before and after.I stopped the remaining saddle-point and multigrid tests: with the feectools kernels uncompiled in the checkouts, each test took more than 10 minutes. On NumPy, feectools' change is limited to identity helpers and an equivalent loop structure in BiCGStab, and the feectools solver tests check the same iteration counts and results.
🤖 Generated with Claude Code