Skip to content

CUDA 2 - MPI with device buffers - #86

Merged
max-models merged 8 commits into
cuda-developmentfrom
cuda-2-mpi-sync
Oct 7, 2026
Merged

max-models merged 8 commits into
cuda-developmentfrom
cuda-2-mpi-sync

Conversation

@max-models

@max-models max-models commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Solves the following issue(s):

CUDA 2 of the feectools CUDA strategy (CUDA_STRATEGY.md, "CUDA 2 implementation notes"), tracked in struphy-hub/struphy#650. No issue closed.

Stack: #90 → #86 (this PR) → #87 → #88.

Why

On the CuPy backend, the stencil data live on the GPU, but feectools.ddm.mpi turned MPI off on that backend, so struphy could not run on more than one rank. With a CUDA-aware MPI library, device buffers can go to MPI directly. The segfaults the old check guarded against come from MPI libraries that are not CUDA-aware.

CuPy kernels run asynchronously, and MPI does not know about CUDA streams. A buffer that is still being written would therefore be sent with wrong data, and no error.

Core changes

1. MPI on the CuPy backend

  • feectools.ddm.mpi no longer turns MPI off on the CuPy backend.
  • cunumpy.mpi.synchronize_for_mpi is called before every MPI call on device buffers. It does nothing on NumPy. The calls are in:
    • the blocking, non-blocking and interface data exchangers;
    • the Allreduce of StencilVectorSpace.inner;
    • the Alltoallv calls of the parallel Kronecker solver.
  • CuPy fixes that only the MPI tests reach:
    • xp.dot/xp.vdot on .flat iterators in StencilInterfaceMatrix._dot and in the pure-Python inner product;
    • test_cart_1d assigning Python lists to CuPy arrays.

2. cunumpy 0.5 names

  • synchronize_for_mpi is imported from cunumpy.mpi. The top-level name is deprecated in cunumpy 0.5 and removed in 0.6.
  • The requirement on cunumpy comes from CUDA 1 - Run feectools with the CuPy backend #90 (cunumpy >= 0.5.0, < 0.6), so this PR no longer changes pyproject.toml. Before, it set cunumpy>=0.3.0.

3. Stack

  • cuda-development (CUDA 1 - Run feectools with the CuPy backend #90) is merged in, so the base of this PR is now cuda-development instead of devel-tiny. The conflicts in mpi.py (the removed check), stencil.py (imports) and pyproject.toml were resolved in favour of this step and the cunumpy 0.5 pin.

Tests

  • New linalg/tests/test_mpi_device.py (6 MPI tests):
    • distributed results compared with global references. A distributed run compared with another distributed run would hide stale ghost regions, because every rank agrees on the same wrong answer;
    • the exchangers synchronize before they call MPI.

Run locally on macOS (no GPU, cunumpy 0.5.0, pyccel 2.2.1, Open MPI). After the changes, both runs used -W error:cunumpy\.:DeprecationWarning.

Tests Before (e8e30ef) After
serial: pytest feectools -m "not mpi and not petsc" 9408 passed, 6 failed, 4 skipped 9427 passed, 6 failed, 4 skipped
MPI: mpirun -n 2 pytest feectools -m "mpi and not petsc" --with-mpi 884 passed, 4 skipped 888 passed, 4 skipped

The 6 serial failures happen on devel-tiny too, so this PR does not cause them. They are test_cart_2d/test_cart_3d with MockComm, see #90. The higher numbers after come from the tests in devel-tiny.

The GPU tests have not been run for this PR. The MPI tests on CuPy need two GPUs and a CUDA-aware MPI. The original commit (ccb39f8) says all MPI tests in ddm and linalg passed with 2 ranks and a CUDA-aware Open MPI on both backends. That has not been repeated after the merge.

Before merging

  • Run the MPI tests on the CuPy backend with a CUDA-aware MPI (CUNUMPY_BACKEND=cupy mpirun -n 2 pytest feectools -m mpi --with-mpi).

Documentation changes:

CUDA_STRATEGY.md: new section "CUDA 2 implementation notes".

🤖 Generated with Claude Code

max-models and others added 2 commits September 30, 2026 23:53
Make feectools work when cunumpy's backend is CuPy, without any device
kernels yet: kernels that stay on the host still copy their arrays.

- Wrap all Pyccel kernels (stencil, B-splines, field evaluation, DOF
  kernels) in cunumpy.PyccelKernel, so they accept CuPy arrays.
- Keep host-only metadata on NumPy: MPI/index bookkeeping in ddm (cart,
  partition, petsc) and fem.partitioning, Kronecker solver sizes, and
  index arithmetic with Python ints (compute_diag_len, math.prod).
- Stage data for host-only libraries by array, not by global backend:
  LAPACK/SuperLU direct solvers, SciPy FFT, SciPy sparse products.
- Fix calls that ran host kernels on device arrays: the second
  stencil2coo call in StencilMatrix.tosparse and the conjugate transpose.
- Vectorize the construction of the 1D collocation matrices in the
  global projectors (element-wise indexing was a device round trip per
  entry: 334 s of a 348 s Derham setup on the GPU).
- GMRES: take real scalars from CuPy views before modifying them.
- Fix StencilMatrix._update_ghost_regions_serial: the ghost region is
  pads * shifts wide (wrong whenever shifts > 1, on both backends).
- Tests: work with CuPy arrays; skip PETSc tests without petsc4py.

Serial tests pass on both backends (core, ddm, fem, linalg).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Allow MPI on the CuPy backend and make it correct with device buffers
(requires a CUDA-aware MPI library).

- feectools.ddm.mpi no longer disables MPI when ARRAY_BACKEND=cupy; the
  segfaults it guarded against come from MPI libraries that are not
  CUDA-aware.
- Call cunumpy.synchronize_for_mpi before every MPI call on device
  buffers: CuPy kernels run asynchronously and MPI does not know about
  CUDA streams, so a buffer still being written would be sent silently
  wrong. Covers the blocking, non-blocking and interface data exchangers,
  the Allreduce in StencilVectorSpace.inner and the Alltoallv calls of the
  parallel Kronecker solver. Requires cunumpy >= 0.3.0.
- Fix CuPy incompatibilities reached only by the MPI tests: xp.dot/vdot
  on .flat iterators in StencilInterfaceMatrix._dot and the pure-Python
  inner product, and test_cart_1d assigning Python lists to CuPy arrays.
- Add test_mpi_device.py: distributed results against global references,
  and that the exchangers synchronize before MPI.

With 2 MPI ranks and a CUDA-aware Open MPI, all MPI tests in ddm and
linalg pass on both backends; serial tests are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@max-models
max-models changed the base branch from devel-tiny to cuda-development October 2, 2026 13:04
max-models and others added 3 commits October 6, 2026 14:55
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The top-level cunumpy.synchronize_for_mpi is deprecated in cunumpy 0.5 (removed in 0.6).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@max-models max-models changed the title Cuda 2 mpi sync CUDA 2 - MPI with device buffers Oct 6, 2026
@max-models
max-models added this pull request to stack #92 October 6, 2026 14:30
@max-models
max-models marked this pull request as ready for review October 6, 2026 14:36
@max-models
max-models merged commit 3d8c68f into devel-tiny Oct 7, 2026
9 checks passed
@max-models max-models mentioned this pull request Oct 7, 2026
max-models added a commit to struphy-hub/struphy that referenced this pull request Oct 7, 2026
**Solves the following issue(s):**

PR 17 of the CUDA strategy (`CUDA_STRATEGY.md`), tracked in #650. No
issue closed.

Stack: #668 (PR 13) → #670 (PR 14) → #671 (PR 15) → #679 (PR 16) →
**this PR**.

## Core changes

The `feectools` submodule now points at the **top of the feectools CUDA
stack**: `cuda-4-device-kernels` at `e4f4b2e`
(struphy-hub/feectools#88), instead of `devel-tiny`. From this PR on,
the struphy CUDA PRs run against a feectools stack set up like
struphy's: each PR targets the previous branch and contains only its own
step.

| feectools PR | Branch → base | What |
|---|---|---|
| struphy-hub/feectools#90 (CUDA 1) | `cuda-development` → `devel-tiny`
| feectools on the CuPy backend (contains #85), `devel-tiny` merged in,
cunumpy 0.5 (`CUNUMPY_BACKEND`, `cunumpy>=0.5.0, <0.6`), feectools' own
`CUDA_STRATEGY.md` |
| struphy-hub/feectools#86 (CUDA 2) | `cuda-2-mpi-sync` →
`cuda-development` | MPI with device buffers
(`cunumpy.mpi.synchronize_for_mpi`) |
| struphy-hub/feectools#87 (CUDA 3) | `cuda-3-device-binding` →
`cuda-2-mpi-sync` | one GPU per MPI rank
(`cunumpy.cuda.bind_local_device`) |
| struphy-hub/feectools#88 (CUDA 4) | `cuda-4-device-kernels` →
`cuda-3-device-binding` | stencil `dot`, `transpose`, `inner`, `axpy` as
one folder per kernel (pyccel and CUDA side by side), with parity and
CPU-emulation tests |

That is the FEEC side a whole model needs on the GPU: field solves
without host copies, and multi-rank, multi-GPU runs. The pointer moves
again whenever the top of the feectools stack changes.

**Before merging the struphy CUDA stack into `devel`:**
1. Merge the feectools stack into `devel-tiny` from the bottom up.
2. Release feectools.
3. Point the submodule and the `pyproject.toml` pin (`feectools>=0.3.0,
<=0.3.0`) at that release again.

Until then the `pr-feectools-submodule` CI check fails on this PR and
the PRs above it, which is expected: it requires the submodule to be the
latest `devel-tiny`.

## Documentation changes

`CUDA_STRATEGY.md`:
- New PR 17 entry.
- Geometry evaluation for all analytic mappings becomes **PR 18**, and
spline mappings become **PR 19**.
- The *feectools* section explains how struphy follows the stack and
what has to happen before the CUDA stack is merged.

**Model-specific changes:** none.

**Testing:** no code change in struphy. The feectools stack was tested
on its own branches:
- serial: 9456 passed; 6 failures in `test_cart_2d`/`test_cart_3d`,
which also fail on `devel-tiny`;
- 888 passed under `mpirun -n 2`;
- CPU emulation: all 12 CUDA stencil kernels match pyccel.

No GPU tests have been run.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Stefan Possanner <stefan.possanner@ipp.mpg.de>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant