Skip to content

CUDA 3 - One GPU per MPI rank - #87

Merged
max-models merged 6 commits into
cuda-2-mpi-syncfrom
cuda-3-device-binding
Oct 7, 2026
Merged

max-models merged 6 commits into
cuda-2-mpi-syncfrom
cuda-3-device-binding

Conversation

@max-models

@max-models max-models commented Sep 30, 2026 •

Copy link
Copy Markdown
Member

Solves the following issue(s):

CUDA 3 of the feectools CUDA strategy (CUDA_STRATEGY.md, "CUDA 3 implementation notes"), tracked in struphy-hub/struphy#650. No issue closed.

Stack: #90 → #86 → #87 (this PR) → #88.

Why

With MPI on device buffers (#86), every rank of a multi-GPU run should use its own GPU. Before, feectools always picked GPU 0. CUDA-aware MPI also requires the CUDA context to exist before MPI is initialized.

Core changes

1. One GPU per MPI rank

  • feectools.ddm.cart calls cunumpy.cuda.bind_local_device() when it is imported. Each process uses GPU local_rank % device_count, where local_rank is the rank within the node that the MPI launcher exports (cunumpy.mpi.local_rank). It does nothing on the NumPy backend.
  • Order: ddm/__init__.py imports cart first, and cart binds the device before it imports feectools.ddm.mpi, which starts MPI. So the GPU is chosen before MPI is initialized, whatever the user imports first.

2. cunumpy 0.5 names

  • bind_local_device is imported from cunumpy.cuda.
  • The test takes local_rank and device_count from cunumpy.mpi and cunumpy.cuda, and is skipped with cunumpy.kernel_testing.requires_cupy, like the other GPU tests.

3. Stack

Tests

  • New ddm/tests/test_device_binding.py: checks that the current device is local_rank % device_count after cart is imported. It runs on a GPU only.

Run locally on macOS (no GPU, cunumpy 0.5.0, pyccel 2.2.1, Open MPI). After the changes, both runs used -W error:cunumpy\.:DeprecationWarning.

Tests Before (3a11ad1) After
serial: pytest feectools -m "not mpi and not petsc" 9408 passed, 6 failed, 5 skipped 9427 passed, 6 failed, 5 skipped
MPI: mpirun -n 2 pytest feectools -m "mpi and not petsc" --with-mpi 884 passed, 4 skipped 888 passed, 4 skipped

The 6 serial failures happen on devel-tiny too, so this PR does not cause them (test_cart_2d/test_cart_3d, see #90). The extra skip is the GPU binding test.

The GPU tests have not been run for this PR.

Before merging

  • Run test_device_binding.py on a multi-GPU node with 2 ranks on the CuPy backend.

Documentation changes:

CUDA_STRATEGY.md: new section "CUDA 3 implementation notes".

🤖 Generated with Claude Code

Replace the initialization that always used GPU 0 by
cunumpy.bind_local_device(): each process uses GPU
local_rank % device_count, chosen from the node-local rank that the MPI
launcher exports, and its CUDA context is created before MPI is
initialized (as CUDA-aware MPI requires). No-op on the NumPy backend.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
max-models and others added 3 commits October 6, 2026 15:03
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The top-level cunumpy.bind_local_device, local_rank and device_count are deprecated in cunumpy 0.5
(removed in 0.6). The binding test imports them from cunumpy.cuda and cunumpy.mpi and is skipped with
cunumpy.kernel_testing.requires_cupy, like the other GPU tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@max-models
max-models changed the base branch from devel-tiny to cuda-2-mpi-sync October 6, 2026 13:58
@max-models max-models changed the title Cuda 3 device binding CUDA 3 - One GPU per MPI rank Oct 6, 2026
@max-models
max-models added this pull request to stack #92 October 6, 2026 14:30
@max-models
max-models marked this pull request as ready for review October 6, 2026 14:37
@max-models
max-models merged commit e07d5e9 into devel-tiny Oct 7, 2026
9 checks passed
@max-models max-models mentioned this pull request Oct 7, 2026
max-models added a commit to struphy-hub/struphy that referenced this pull request Oct 7, 2026
**Solves the following issue(s):**

PR 17 of the CUDA strategy (`CUDA_STRATEGY.md`), tracked in #650. No
issue closed.

Stack: #668 (PR 13) → #670 (PR 14) → #671 (PR 15) → #679 (PR 16) →
**this PR**.

## Core changes

The `feectools` submodule now points at the **top of the feectools CUDA
stack**: `cuda-4-device-kernels` at `e4f4b2e`
(struphy-hub/feectools#88), instead of `devel-tiny`. From this PR on,
the struphy CUDA PRs run against a feectools stack set up like
struphy's: each PR targets the previous branch and contains only its own
step.

| feectools PR | Branch → base | What |
|---|---|---|
| struphy-hub/feectools#90 (CUDA 1) | `cuda-development` → `devel-tiny`
| feectools on the CuPy backend (contains #85), `devel-tiny` merged in,
cunumpy 0.5 (`CUNUMPY_BACKEND`, `cunumpy>=0.5.0, <0.6`), feectools' own
`CUDA_STRATEGY.md` |
| struphy-hub/feectools#86 (CUDA 2) | `cuda-2-mpi-sync` →
`cuda-development` | MPI with device buffers
(`cunumpy.mpi.synchronize_for_mpi`) |
| struphy-hub/feectools#87 (CUDA 3) | `cuda-3-device-binding` →
`cuda-2-mpi-sync` | one GPU per MPI rank
(`cunumpy.cuda.bind_local_device`) |
| struphy-hub/feectools#88 (CUDA 4) | `cuda-4-device-kernels` →
`cuda-3-device-binding` | stencil `dot`, `transpose`, `inner`, `axpy` as
one folder per kernel (pyccel and CUDA side by side), with parity and
CPU-emulation tests |

That is the FEEC side a whole model needs on the GPU: field solves
without host copies, and multi-rank, multi-GPU runs. The pointer moves
again whenever the top of the feectools stack changes.

**Before merging the struphy CUDA stack into `devel`:**
1. Merge the feectools stack into `devel-tiny` from the bottom up.
2. Release feectools.
3. Point the submodule and the `pyproject.toml` pin (`feectools>=0.3.0,
<=0.3.0`) at that release again.

Until then the `pr-feectools-submodule` CI check fails on this PR and
the PRs above it, which is expected: it requires the submodule to be the
latest `devel-tiny`.

## Documentation changes

`CUDA_STRATEGY.md`:
- New PR 17 entry.
- Geometry evaluation for all analytic mappings becomes **PR 18**, and
spline mappings become **PR 19**.
- The *feectools* section explains how struphy follows the stack and
what has to happen before the CUDA stack is merged.

**Model-specific changes:** none.

**Testing:** no code change in struphy. The feectools stack was tested
on its own branches:
- serial: 9456 passed; 6 failures in `test_cart_2d`/`test_cart_3d`,
which also fail on `devel-tiny`;
- 888 passed under `mpirun -n 2`;
- CPU emulation: all 12 CUDA stencil kernels match pyccel.

No GPU tests have been run.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Stefan Possanner <stefan.possanner@ipp.mpg.de>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant