Skip to content

CUDA 1 - Run feectools with the CuPy backend - #90

Merged
max-models merged 7 commits into
devel-tinyfrom
cuda-development
Oct 7, 2026
Merged

max-models merged 7 commits into
devel-tinyfrom
cuda-development

Conversation

@max-models

@max-models max-models commented Oct 2, 2026 •

Copy link
Copy Markdown
Member

Solves the following issue(s):

CUDA 1 of the feectools CUDA strategy (CUDA_STRATEGY.md, added in this PR), tracked in struphy-hub/struphy#650. No issue closed.

Stack: #90 (this PR) → #86 → #87 → #88.

Each PR's base is the branch of the PR before it, so its diff on GitHub shows only its own step. This PR targets devel-tiny. Merge bottom up, retargeting the next PR to devel-tiny after each merge.

Why

struphy runs on the GPU through cunumpy's CuPy backend, and its Derham holds feectools stencil vectors and matrices. This PR makes feectools work when the backend is CuPy, but adds no device kernels yet: the pyccel kernels still run on the host, and their arrays are copied there and back. Device MPI, GPU binding and device kernels follow in #86, #87 and #88.

The CuPy changes were reviewed and merged in #85 (into cuda-development). This PR brings them to devel-tiny, together with the updates below.

Core changes

1. feectools on the CuPy backend (from #85)

  • Kernels: the pyccel kernels (stencil, B-splines, field evaluation, DOF kernels) are wrapped in cunumpy.kernels.PyccelKernel, so they accept CuPy arrays.
  • Metadata on NumPy: host-only metadata stays on NumPy. This covers MPI and index bookkeeping in ddm (cart, partition, petsc) and fem.partitioning, the Kronecker solver sizes, and index arithmetic with Python ints (compute_diag_len, math.prod).
  • Host-only libraries: LAPACK/SuperLU, SciPy FFT and SciPy sparse products get host copies per array, instead of being chosen by the global backend.
  • Global projectors: the 1D collocation matrices are built vectorized. Indexing element by element cost one device round trip per entry, which was 334 s of a 348 s Derham setup on the GPU.
  • Fixes:
    • host kernels that were given device arrays: the second stencil2coo call in StencilMatrix.tosparse, and the conjugate transpose.
    • GMRES takes real scalars from CuPy views before it modifies them.
    • StencilMatrix._update_ghost_regions_serial now uses a ghost region pads * shifts wide. It was wrong on both backends whenever shifts > 1.

2. Up to date with devel-tiny

3. cunumpy 0.5, as in struphy

  • PyccelKernel is imported from cunumpy.kernels. The top-level name is deprecated in cunumpy 0.5 and removed in 0.6.
  • pyproject.toml requires cunumpy >= 0.5.0, < 0.6. Before, cunumpy was not pinned.
  • feectools.ddm.mpi reads cunumpy's backend variable CUNUMPY_BACKEND to turn MPI off on the CuPy backend. It used to read ARRAY_BACKEND, which cunumpy no longer uses, so the check never triggered. CUDA 2 - MPI with device buffers #86 removes the check.

Tests

Run locally on macOS (no GPU, cunumpy 0.5.0, pyccel 2.2.1, kernels compiled with psydac-accelerate --language c). After the changes, both runs used -W error:cunumpy\.:DeprecationWarning, so a deprecated cunumpy name would fail.

Tests Before (a9c8623) After
serial: pytest feectools -m "not mpi and not petsc" 9408 passed, 6 failed, 4 skipped 9427 passed, 6 failed, 4 skipped (+19 tests from devel-tiny)
MPI: mpirun -n 2 pytest feectools -m "mpi and not petsc" --with-mpi 878 passed, 4 skipped 882 passed, 4 skipped

The 6 failures happen on devel-tiny too, so this PR does not cause them. They are test_cart_2d/test_cart_3d (×6), which fail in the serial session with 'NoneType' object has no attribute 'Shift'.

The GPU tests have not been run for this PR (no GPU on the development machine).

Before merging

  • Run the serial tests on the CuPy backend on a GPU (CUNUMPY_BACKEND=cupy).

Documentation changes:

New CUDA_STRATEGY.md, in the style of struphy's (section "feectools"). It has the goal, the principles (the backend decides, one folder per kernel, 1:1 correspondence, no silent CPU fallback), the PR stack table, the CUDA 1 implementation notes, the testing rules and the open questions. Each later PR of the stack extends it.

🤖 Generated with Claude Code

* CUDA step 1: run feectools with the CuPy backend

Make feectools work when cunumpy's backend is CuPy, without any device
kernels yet: kernels that stay on the host still copy their arrays.

- Wrap all Pyccel kernels (stencil, B-splines, field evaluation, DOF
  kernels) in cunumpy.PyccelKernel, so they accept CuPy arrays.
- Keep host-only metadata on NumPy: MPI/index bookkeeping in ddm (cart,
  partition, petsc) and fem.partitioning, Kronecker solver sizes, and
  index arithmetic with Python ints (compute_diag_len, math.prod).
- Stage data for host-only libraries by array, not by global backend:
  LAPACK/SuperLU direct solvers, SciPy FFT, SciPy sparse products.
- Fix calls that ran host kernels on device arrays: the second
  stencil2coo call in StencilMatrix.tosparse and the conjugate transpose.
- Vectorize the construction of the 1D collocation matrices in the
  global projectors (element-wise indexing was a device round trip per
  entry: 334 s of a 348 s Derham setup on the GPU).
- GMRES: take real scalars from CuPy views before modifying them.
- Fix StencilMatrix._update_ghost_regions_serial: the ghost region is
  pads * shifts wide (wrong whenever shifts > 1, on both backends).
- Tests: work with CuPy arrays; skip PETSc tests without petsc4py.

Serial tests pass on both backends (core, ddm, fem, linalg).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* run ci

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
max-models and others added 4 commits October 6, 2026 14:43
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The top-level cunumpy.PyccelKernel is deprecated in cunumpy 0.5 (removed in 0.6). Import it from
cunumpy.kernels, as struphy does, and require cunumpy >= 0.5.0, < 0.6 like struphy.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
cunumpy reads its backend from CUNUMPY_BACKEND; ARRAY_BACKEND is not used any more, so the check that
turns MPI off on the CuPy backend (until CUDA 2) never triggered.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Goal, principles, the PR stack and the CUDA 1 implementation notes, in the style of struphy's
CUDA_STRATEGY.md (struphy-hub/struphy#650).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@max-models max-models changed the title CUDA support CUDA 1 - Run feectools with the CuPy backend Oct 6, 2026
@max-models
max-models added this pull request to stack #92 October 6, 2026 14:30
@max-models
max-models marked this pull request as ready for review October 6, 2026 14:36
Comment thread feectools/ddm/cart.py
Comment thread feectools/ddm/cart.py Outdated
max-models and others added 2 commits October 7, 2026 10:32
Co-authored-by: Stefan Possanner <stefan.possanner@ipp.mpg.de>
@max-models
max-models merged commit b8cc9b1 into devel-tiny Oct 7, 2026
9 checks passed
@max-models max-models mentioned this pull request Oct 7, 2026
max-models added a commit to struphy-hub/struphy that referenced this pull request Oct 7, 2026
**Solves the following issue(s):**

PR 17 of the CUDA strategy (`CUDA_STRATEGY.md`), tracked in #650. No
issue closed.

Stack: #668 (PR 13) → #670 (PR 14) → #671 (PR 15) → #679 (PR 16) →
**this PR**.

## Core changes

The `feectools` submodule now points at the **top of the feectools CUDA
stack**: `cuda-4-device-kernels` at `e4f4b2e`
(struphy-hub/feectools#88), instead of `devel-tiny`. From this PR on,
the struphy CUDA PRs run against a feectools stack set up like
struphy's: each PR targets the previous branch and contains only its own
step.

| feectools PR | Branch → base | What |
|---|---|---|
| struphy-hub/feectools#90 (CUDA 1) | `cuda-development` → `devel-tiny`
| feectools on the CuPy backend (contains #85), `devel-tiny` merged in,
cunumpy 0.5 (`CUNUMPY_BACKEND`, `cunumpy>=0.5.0, <0.6`), feectools' own
`CUDA_STRATEGY.md` |
| struphy-hub/feectools#86 (CUDA 2) | `cuda-2-mpi-sync` →
`cuda-development` | MPI with device buffers
(`cunumpy.mpi.synchronize_for_mpi`) |
| struphy-hub/feectools#87 (CUDA 3) | `cuda-3-device-binding` →
`cuda-2-mpi-sync` | one GPU per MPI rank
(`cunumpy.cuda.bind_local_device`) |
| struphy-hub/feectools#88 (CUDA 4) | `cuda-4-device-kernels` →
`cuda-3-device-binding` | stencil `dot`, `transpose`, `inner`, `axpy` as
one folder per kernel (pyccel and CUDA side by side), with parity and
CPU-emulation tests |

That is the FEEC side a whole model needs on the GPU: field solves
without host copies, and multi-rank, multi-GPU runs. The pointer moves
again whenever the top of the feectools stack changes.

**Before merging the struphy CUDA stack into `devel`:**
1. Merge the feectools stack into `devel-tiny` from the bottom up.
2. Release feectools.
3. Point the submodule and the `pyproject.toml` pin (`feectools>=0.3.0,
<=0.3.0`) at that release again.

Until then the `pr-feectools-submodule` CI check fails on this PR and
the PRs above it, which is expected: it requires the submodule to be the
latest `devel-tiny`.

## Documentation changes

`CUDA_STRATEGY.md`:
- New PR 17 entry.
- Geometry evaluation for all analytic mappings becomes **PR 18**, and
spline mappings become **PR 19**.
- The *feectools* section explains how struphy follows the stack and
what has to happen before the CUDA stack is merged.

**Model-specific changes:** none.

**Testing:** no code change in struphy. The feectools stack was tested
on its own branches:
- serial: 9456 passed; 6 failures in `test_cart_2d`/`test_cart_3d`,
which also fail on `devel-tiny`;
- 888 passed under `mpirun -n 2`;
- CPU emulation: all 12 CUDA stencil kernels match pyccel.

No GPU tests have been run.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Stefan Possanner <stefan.possanner@ipp.mpg.de>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants