Repository navigation
CUDA 2 - MPI with device buffers - #86
Merged
Merged
Conversation
Make feectools work when cunumpy's backend is CuPy, without any device kernels yet: kernels that stay on the host still copy their arrays. - Wrap all Pyccel kernels (stencil, B-splines, field evaluation, DOF kernels) in cunumpy.PyccelKernel, so they accept CuPy arrays. - Keep host-only metadata on NumPy: MPI/index bookkeeping in ddm (cart, partition, petsc) and fem.partitioning, Kronecker solver sizes, and index arithmetic with Python ints (compute_diag_len, math.prod). - Stage data for host-only libraries by array, not by global backend: LAPACK/SuperLU direct solvers, SciPy FFT, SciPy sparse products. - Fix calls that ran host kernels on device arrays: the second stencil2coo call in StencilMatrix.tosparse and the conjugate transpose. - Vectorize the construction of the 1D collocation matrices in the global projectors (element-wise indexing was a device round trip per entry: 334 s of a 348 s Derham setup on the GPU). - GMRES: take real scalars from CuPy views before modifying them. - Fix StencilMatrix._update_ghost_regions_serial: the ghost region is pads * shifts wide (wrong whenever shifts > 1, on both backends). - Tests: work with CuPy arrays; skip PETSc tests without petsc4py. Serial tests pass on both backends (core, ddm, fem, linalg). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Allow MPI on the CuPy backend and make it correct with device buffers (requires a CUDA-aware MPI library). - feectools.ddm.mpi no longer disables MPI when ARRAY_BACKEND=cupy; the segfaults it guarded against come from MPI libraries that are not CUDA-aware. - Call cunumpy.synchronize_for_mpi before every MPI call on device buffers: CuPy kernels run asynchronously and MPI does not know about CUDA streams, so a buffer still being written would be sent silently wrong. Covers the blocking, non-blocking and interface data exchangers, the Allreduce in StencilVectorSpace.inner and the Alltoallv calls of the parallel Kronecker solver. Requires cunumpy >= 0.3.0. - Fix CuPy incompatibilities reached only by the MPI tests: xp.dot/vdot on .flat iterators in StencilInterfaceMatrix._dot and the pure-Python inner product, and test_cart_1d assigning Python lists to CuPy arrays. - Add test_mpi_device.py: distributed results against global references, and that the exchangers synchronize before MPI. With 2 MPI ranks and a CUDA-aware Open MPI, all MPI tests in ddm and linalg pass on both backends; serial tests are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 1, 2026
Merged
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The top-level cunumpy.synchronize_for_mpi is deprecated in cunumpy 0.5 (removed in 0.6). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
1 task
max-models
added this pull request to stack #92
October 6, 2026 14:30
max-models
marked this pull request as ready for review
October 6, 2026 14:36
Merged
max-models
added a commit
to struphy-hub/struphy
that referenced
this pull request
Oct 7, 2026
**Solves the following issue(s):** PR 17 of the CUDA strategy (`CUDA_STRATEGY.md`), tracked in #650. No issue closed. Stack: #668 (PR 13) → #670 (PR 14) → #671 (PR 15) → #679 (PR 16) → **this PR**. ## Core changes The `feectools` submodule now points at the **top of the feectools CUDA stack**: `cuda-4-device-kernels` at `e4f4b2e` (struphy-hub/feectools#88), instead of `devel-tiny`. From this PR on, the struphy CUDA PRs run against a feectools stack set up like struphy's: each PR targets the previous branch and contains only its own step. | feectools PR | Branch → base | What | |---|---|---| | struphy-hub/feectools#90 (CUDA 1) | `cuda-development` → `devel-tiny` | feectools on the CuPy backend (contains #85), `devel-tiny` merged in, cunumpy 0.5 (`CUNUMPY_BACKEND`, `cunumpy>=0.5.0, <0.6`), feectools' own `CUDA_STRATEGY.md` | | struphy-hub/feectools#86 (CUDA 2) | `cuda-2-mpi-sync` → `cuda-development` | MPI with device buffers (`cunumpy.mpi.synchronize_for_mpi`) | | struphy-hub/feectools#87 (CUDA 3) | `cuda-3-device-binding` → `cuda-2-mpi-sync` | one GPU per MPI rank (`cunumpy.cuda.bind_local_device`) | | struphy-hub/feectools#88 (CUDA 4) | `cuda-4-device-kernels` → `cuda-3-device-binding` | stencil `dot`, `transpose`, `inner`, `axpy` as one folder per kernel (pyccel and CUDA side by side), with parity and CPU-emulation tests | That is the FEEC side a whole model needs on the GPU: field solves without host copies, and multi-rank, multi-GPU runs. The pointer moves again whenever the top of the feectools stack changes. **Before merging the struphy CUDA stack into `devel`:** 1. Merge the feectools stack into `devel-tiny` from the bottom up. 2. Release feectools. 3. Point the submodule and the `pyproject.toml` pin (`feectools>=0.3.0, <=0.3.0`) at that release again. Until then the `pr-feectools-submodule` CI check fails on this PR and the PRs above it, which is expected: it requires the submodule to be the latest `devel-tiny`. ## Documentation changes `CUDA_STRATEGY.md`: - New PR 17 entry. - Geometry evaluation for all analytic mappings becomes **PR 18**, and spline mappings become **PR 19**. - The *feectools* section explains how struphy follows the stack and what has to happen before the CUDA stack is merged. **Model-specific changes:** none. **Testing:** no code change in struphy. The feectools stack was tested on its own branches: - serial: 9456 passed; 6 failures in `test_cart_2d`/`test_cart_3d`, which also fail on `devel-tiny`; - 888 passed under `mpirun -n 2`; - CPU emulation: all 12 CUDA stencil kernels match pyccel. No GPU tests have been run. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Stefan Possanner <stefan.possanner@ipp.mpg.de>
3 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Solves the following issue(s):
CUDA 2 of the feectools CUDA strategy (
CUDA_STRATEGY.md, "CUDA 2 implementation notes"), tracked in struphy-hub/struphy#650. No issue closed.Stack: #90 → #86 (this PR) → #87 → #88.
Why
On the CuPy backend, the stencil data live on the GPU, but
feectools.ddm.mpiturned MPI off on that backend, so struphy could not run on more than one rank. With a CUDA-aware MPI library, device buffers can go to MPI directly. The segfaults the old check guarded against come from MPI libraries that are not CUDA-aware.CuPy kernels run asynchronously, and MPI does not know about CUDA streams. A buffer that is still being written would therefore be sent with wrong data, and no error.
Core changes
1. MPI on the CuPy backend
feectools.ddm.mpino longer turns MPI off on the CuPy backend.cunumpy.mpi.synchronize_for_mpiis called before every MPI call on device buffers. It does nothing on NumPy. The calls are in:AllreduceofStencilVectorSpace.inner;Alltoallvcalls of the parallel Kronecker solver.xp.dot/xp.vdoton.flatiterators inStencilInterfaceMatrix._dotand in the pure-Python inner product;test_cart_1dassigning Python lists to CuPy arrays.2. cunumpy 0.5 names
synchronize_for_mpiis imported fromcunumpy.mpi. The top-level name is deprecated in cunumpy 0.5 and removed in 0.6.cunumpy >= 0.5.0, < 0.6), so this PR no longer changespyproject.toml. Before, it setcunumpy>=0.3.0.3. Stack
cuda-development(CUDA 1 - Run feectools with the CuPy backend #90) is merged in, so the base of this PR is nowcuda-developmentinstead ofdevel-tiny. The conflicts inmpi.py(the removed check),stencil.py(imports) andpyproject.tomlwere resolved in favour of this step and the cunumpy 0.5 pin.Tests
linalg/tests/test_mpi_device.py(6 MPI tests):Run locally on macOS (no GPU, cunumpy 0.5.0, pyccel 2.2.1, Open MPI). After the changes, both runs used
-W error:cunumpy\.:DeprecationWarning.e8e30ef)pytest feectools -m "not mpi and not petsc"mpirun -n 2 pytest feectools -m "mpi and not petsc" --with-mpiThe 6 serial failures happen on
devel-tinytoo, so this PR does not cause them. They aretest_cart_2d/test_cart_3dwithMockComm, see #90. The higher numbers after come from the tests indevel-tiny.The GPU tests have not been run for this PR. The MPI tests on CuPy need two GPUs and a CUDA-aware MPI. The original commit (
ccb39f8) says all MPI tests inddmandlinalgpassed with 2 ranks and a CUDA-aware Open MPI on both backends. That has not been repeated after the merge.Before merging
CUNUMPY_BACKEND=cupy mpirun -n 2 pytest feectools -m mpi --with-mpi).Documentation changes:
CUDA_STRATEGY.md: new section "CUDA 2 implementation notes".🤖 Generated with Claude Code