Repository navigation
CUDA 3 - One GPU per MPI rank - #87
Merged
Merged
Conversation
Replace the initialization that always used GPU 0 by cunumpy.bind_local_device(): each process uses GPU local_rank % device_count, chosen from the node-local rank that the MPI launcher exports, and its CUDA context is created before MPI is initialized (as CUDA-aware MPI requires). No-op on the NumPy backend. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 1, 2026
Merged
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The top-level cunumpy.bind_local_device, local_rank and device_count are deprecated in cunumpy 0.5 (removed in 0.6). The binding test imports them from cunumpy.cuda and cunumpy.mpi and is skipped with cunumpy.kernel_testing.requires_cupy, like the other GPU tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
1 task
max-models
added this pull request to stack #92
October 6, 2026 14:30
max-models
marked this pull request as ready for review
October 6, 2026 14:37
Merged
max-models
added a commit
to struphy-hub/struphy
that referenced
this pull request
Oct 7, 2026
**Solves the following issue(s):** PR 17 of the CUDA strategy (`CUDA_STRATEGY.md`), tracked in #650. No issue closed. Stack: #668 (PR 13) → #670 (PR 14) → #671 (PR 15) → #679 (PR 16) → **this PR**. ## Core changes The `feectools` submodule now points at the **top of the feectools CUDA stack**: `cuda-4-device-kernels` at `e4f4b2e` (struphy-hub/feectools#88), instead of `devel-tiny`. From this PR on, the struphy CUDA PRs run against a feectools stack set up like struphy's: each PR targets the previous branch and contains only its own step. | feectools PR | Branch → base | What | |---|---|---| | struphy-hub/feectools#90 (CUDA 1) | `cuda-development` → `devel-tiny` | feectools on the CuPy backend (contains #85), `devel-tiny` merged in, cunumpy 0.5 (`CUNUMPY_BACKEND`, `cunumpy>=0.5.0, <0.6`), feectools' own `CUDA_STRATEGY.md` | | struphy-hub/feectools#86 (CUDA 2) | `cuda-2-mpi-sync` → `cuda-development` | MPI with device buffers (`cunumpy.mpi.synchronize_for_mpi`) | | struphy-hub/feectools#87 (CUDA 3) | `cuda-3-device-binding` → `cuda-2-mpi-sync` | one GPU per MPI rank (`cunumpy.cuda.bind_local_device`) | | struphy-hub/feectools#88 (CUDA 4) | `cuda-4-device-kernels` → `cuda-3-device-binding` | stencil `dot`, `transpose`, `inner`, `axpy` as one folder per kernel (pyccel and CUDA side by side), with parity and CPU-emulation tests | That is the FEEC side a whole model needs on the GPU: field solves without host copies, and multi-rank, multi-GPU runs. The pointer moves again whenever the top of the feectools stack changes. **Before merging the struphy CUDA stack into `devel`:** 1. Merge the feectools stack into `devel-tiny` from the bottom up. 2. Release feectools. 3. Point the submodule and the `pyproject.toml` pin (`feectools>=0.3.0, <=0.3.0`) at that release again. Until then the `pr-feectools-submodule` CI check fails on this PR and the PRs above it, which is expected: it requires the submodule to be the latest `devel-tiny`. ## Documentation changes `CUDA_STRATEGY.md`: - New PR 17 entry. - Geometry evaluation for all analytic mappings becomes **PR 18**, and spline mappings become **PR 19**. - The *feectools* section explains how struphy follows the stack and what has to happen before the CUDA stack is merged. **Model-specific changes:** none. **Testing:** no code change in struphy. The feectools stack was tested on its own branches: - serial: 9456 passed; 6 failures in `test_cart_2d`/`test_cart_3d`, which also fail on `devel-tiny`; - 888 passed under `mpirun -n 2`; - CPU emulation: all 12 CUDA stencil kernels match pyccel. No GPU tests have been run. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Stefan Possanner <stefan.possanner@ipp.mpg.de>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Solves the following issue(s):
CUDA 3 of the feectools CUDA strategy (
CUDA_STRATEGY.md, "CUDA 3 implementation notes"), tracked in struphy-hub/struphy#650. No issue closed.Stack: #90 → #86 → #87 (this PR) → #88.
Why
With MPI on device buffers (#86), every rank of a multi-GPU run should use its own GPU. Before, feectools always picked GPU 0. CUDA-aware MPI also requires the CUDA context to exist before MPI is initialized.
Core changes
1. One GPU per MPI rank
feectools.ddm.cartcallscunumpy.cuda.bind_local_device()when it is imported. Each process uses GPUlocal_rank % device_count, wherelocal_rankis the rank within the node that the MPI launcher exports (cunumpy.mpi.local_rank). It does nothing on the NumPy backend.ddm/__init__.pyimportscartfirst, andcartbinds the device before it importsfeectools.ddm.mpi, which starts MPI. So the GPU is chosen before MPI is initialized, whatever the user imports first.2. cunumpy 0.5 names
bind_local_deviceis imported fromcunumpy.cuda.local_rankanddevice_countfromcunumpy.mpiandcunumpy.cuda, and is skipped withcunumpy.kernel_testing.requires_cupy, like the other GPU tests.3. Stack
cuda-2-mpi-sync(CUDA 2 - MPI with device buffers #86) is merged in, so the base of this PR is nowcuda-2-mpi-syncinstead ofdevel-tiny. This branch carried its own copies of steps 1 and 2 (0a55cbd,ccb39f8). Their trees are identical to Cuda 1 xp arrays #85 and CUDA 2 - MPI with device buffers #86, so the merge had no conflicts, and each change is in the branch only once.Tests
ddm/tests/test_device_binding.py: checks that the current device islocal_rank % device_countaftercartis imported. It runs on a GPU only.Run locally on macOS (no GPU, cunumpy 0.5.0, pyccel 2.2.1, Open MPI). After the changes, both runs used
-W error:cunumpy\.:DeprecationWarning.3a11ad1)pytest feectools -m "not mpi and not petsc"mpirun -n 2 pytest feectools -m "mpi and not petsc" --with-mpiThe 6 serial failures happen on
devel-tinytoo, so this PR does not cause them (test_cart_2d/test_cart_3d, see #90). The extra skip is the GPU binding test.The GPU tests have not been run for this PR.
Before merging
test_device_binding.pyon a multi-GPU node with 2 ranks on the CuPy backend.Documentation changes:
CUDA_STRATEGY.md: new section "CUDA 3 implementation notes".🤖 Generated with Claude Code