Repository navigation
CUDA 1 - Run feectools with the CuPy backend - #90
Merged
Merged
Conversation
* CUDA step 1: run feectools with the CuPy backend Make feectools work when cunumpy's backend is CuPy, without any device kernels yet: kernels that stay on the host still copy their arrays. - Wrap all Pyccel kernels (stencil, B-splines, field evaluation, DOF kernels) in cunumpy.PyccelKernel, so they accept CuPy arrays. - Keep host-only metadata on NumPy: MPI/index bookkeeping in ddm (cart, partition, petsc) and fem.partitioning, Kronecker solver sizes, and index arithmetic with Python ints (compute_diag_len, math.prod). - Stage data for host-only libraries by array, not by global backend: LAPACK/SuperLU direct solvers, SciPy FFT, SciPy sparse products. - Fix calls that ran host kernels on device arrays: the second stencil2coo call in StencilMatrix.tosparse and the conjugate transpose. - Vectorize the construction of the 1D collocation matrices in the global projectors (element-wise indexing was a device round trip per entry: 334 s of a 348 s Derham setup on the GPU). - GMRES: take real scalars from CuPy views before modifying them. - Fix StencilMatrix._update_ghost_regions_serial: the ghost region is pads * shifts wide (wrong whenever shifts > 1, on both backends). - Tests: work with CuPy arrays; skip PETSc tests without petsc4py. Serial tests pass on both backends (core, ddm, fem, linalg). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * run ci --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
9 of 11 tasks
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The top-level cunumpy.PyccelKernel is deprecated in cunumpy 0.5 (removed in 0.6). Import it from cunumpy.kernels, as struphy does, and require cunumpy >= 0.5.0, < 0.6 like struphy. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
cunumpy reads its backend from CUNUMPY_BACKEND; ARRAY_BACKEND is not used any more, so the check that turns MPI off on the CuPy backend (until CUDA 2) never triggered. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Goal, principles, the PR stack and the CUDA 1 implementation notes, in the style of struphy's CUDA_STRATEGY.md (struphy-hub/struphy#650). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 6, 2026
max-models
added this pull request to stack #92
October 6, 2026 14:30
max-models
marked this pull request as ready for review
October 6, 2026 14:36
spossann
reviewed
Oct 7, 2026
spossann
reviewed
Oct 7, 2026
Co-authored-by: Stefan Possanner <stefan.possanner@ipp.mpg.de>
Merged
max-models
added a commit
to struphy-hub/struphy
that referenced
this pull request
Oct 7, 2026
**Solves the following issue(s):** PR 17 of the CUDA strategy (`CUDA_STRATEGY.md`), tracked in #650. No issue closed. Stack: #668 (PR 13) → #670 (PR 14) → #671 (PR 15) → #679 (PR 16) → **this PR**. ## Core changes The `feectools` submodule now points at the **top of the feectools CUDA stack**: `cuda-4-device-kernels` at `e4f4b2e` (struphy-hub/feectools#88), instead of `devel-tiny`. From this PR on, the struphy CUDA PRs run against a feectools stack set up like struphy's: each PR targets the previous branch and contains only its own step. | feectools PR | Branch → base | What | |---|---|---| | struphy-hub/feectools#90 (CUDA 1) | `cuda-development` → `devel-tiny` | feectools on the CuPy backend (contains #85), `devel-tiny` merged in, cunumpy 0.5 (`CUNUMPY_BACKEND`, `cunumpy>=0.5.0, <0.6`), feectools' own `CUDA_STRATEGY.md` | | struphy-hub/feectools#86 (CUDA 2) | `cuda-2-mpi-sync` → `cuda-development` | MPI with device buffers (`cunumpy.mpi.synchronize_for_mpi`) | | struphy-hub/feectools#87 (CUDA 3) | `cuda-3-device-binding` → `cuda-2-mpi-sync` | one GPU per MPI rank (`cunumpy.cuda.bind_local_device`) | | struphy-hub/feectools#88 (CUDA 4) | `cuda-4-device-kernels` → `cuda-3-device-binding` | stencil `dot`, `transpose`, `inner`, `axpy` as one folder per kernel (pyccel and CUDA side by side), with parity and CPU-emulation tests | That is the FEEC side a whole model needs on the GPU: field solves without host copies, and multi-rank, multi-GPU runs. The pointer moves again whenever the top of the feectools stack changes. **Before merging the struphy CUDA stack into `devel`:** 1. Merge the feectools stack into `devel-tiny` from the bottom up. 2. Release feectools. 3. Point the submodule and the `pyproject.toml` pin (`feectools>=0.3.0, <=0.3.0`) at that release again. Until then the `pr-feectools-submodule` CI check fails on this PR and the PRs above it, which is expected: it requires the submodule to be the latest `devel-tiny`. ## Documentation changes `CUDA_STRATEGY.md`: - New PR 17 entry. - Geometry evaluation for all analytic mappings becomes **PR 18**, and spline mappings become **PR 19**. - The *feectools* section explains how struphy follows the stack and what has to happen before the CUDA stack is merged. **Model-specific changes:** none. **Testing:** no code change in struphy. The feectools stack was tested on its own branches: - serial: 9456 passed; 6 failures in `test_cart_2d`/`test_cart_3d`, which also fail on `devel-tiny`; - 888 passed under `mpirun -n 2`; - CPU emulation: all 12 CUDA stencil kernels match pyccel. No GPU tests have been run. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Stefan Possanner <stefan.possanner@ipp.mpg.de>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Solves the following issue(s):
CUDA 1 of the feectools CUDA strategy (
CUDA_STRATEGY.md, added in this PR), tracked in struphy-hub/struphy#650. No issue closed.Stack: #90 (this PR) → #86 → #87 → #88.
Each PR's base is the branch of the PR before it, so its diff on GitHub shows only its own step. This PR targets
devel-tiny. Merge bottom up, retargeting the next PR todevel-tinyafter each merge.Why
struphy runs on the GPU through cunumpy's CuPy backend, and its
Derhamholds feectools stencil vectors and matrices. This PR makes feectools work when the backend is CuPy, but adds no device kernels yet: the pyccel kernels still run on the host, and their arrays are copied there and back. Device MPI, GPU binding and device kernels follow in #86, #87 and #88.The CuPy changes were reviewed and merged in #85 (into
cuda-development). This PR brings them todevel-tiny, together with the updates below.Core changes
1. feectools on the CuPy backend (from #85)
cunumpy.kernels.PyccelKernel, so they accept CuPy arrays.ddm(cart,partition,petsc) andfem.partitioning, the Kronecker solver sizes, and index arithmetic with Python ints (compute_diag_len,math.prod).Derhamsetup on the GPU.stencil2coocall inStencilMatrix.tosparse, and the conjugate transpose.StencilMatrix._update_ghost_regions_serialnow uses a ghost regionpads * shiftswide. It was wrong on both backends whenevershifts > 1.2. Up to date with
devel-tinydevel-tiny(Generic LinearOperator.toarray() and tosparse() defaults #89, DomainDecomposition.coarsen, refine fix, ghost-sync fixes in axpy and essential BCs #91, version 0.3.0) is merged in. There were no conflicts.devel-tiny's version is kept.3. cunumpy 0.5, as in struphy
PyccelKernelis imported fromcunumpy.kernels. The top-level name is deprecated in cunumpy 0.5 and removed in 0.6.pyproject.tomlrequirescunumpy >= 0.5.0, < 0.6. Before, cunumpy was not pinned.feectools.ddm.mpireads cunumpy's backend variableCUNUMPY_BACKENDto turn MPI off on the CuPy backend. It used to readARRAY_BACKEND, which cunumpy no longer uses, so the check never triggered. CUDA 2 - MPI with device buffers #86 removes the check.Tests
Run locally on macOS (no GPU, cunumpy 0.5.0, pyccel 2.2.1, kernels compiled with
psydac-accelerate --language c). After the changes, both runs used-W error:cunumpy\.:DeprecationWarning, so a deprecated cunumpy name would fail.a9c8623)pytest feectools -m "not mpi and not petsc"devel-tiny)mpirun -n 2 pytest feectools -m "mpi and not petsc" --with-mpiThe 6 failures happen on
devel-tinytoo, so this PR does not cause them. They aretest_cart_2d/test_cart_3d(×6), which fail in the serial session with'NoneType' object has no attribute 'Shift'.The GPU tests have not been run for this PR (no GPU on the development machine).
Before merging
CUNUMPY_BACKEND=cupy).Documentation changes:
New
CUDA_STRATEGY.md, in the style of struphy's (section "feectools"). It has the goal, the principles (the backend decides, one folder per kernel, 1:1 correspondence, no silent CPU fallback), the PR stack table, the CUDA 1 implementation notes, the testing rules and the open questions. Each later PR of the stack extends it.🤖 Generated with Claude Code