Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
56 changes: 56 additions & 0 deletions .github/workflows/macos.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,56 @@
name: Tests (macOS)

on:
push:
branches:
- main
- devel
pull_request:
branches:
- main
- devel
workflow_dispatch:

jobs:
build:
# macos-latest is Apple silicon (arm64): checks the NumPy backend, the
# compiled host kernels and the MLX import on the platform of MetalKernel.
runs-on: macos-latest

strategy:
fail-fast: false
matrix:
python-version: ["3.10", "3.13"]

steps:
- name: Checkout code
uses: actions/checkout@v4

- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python-version }}
cache: pip
cache-dependency-path: pyproject.toml

# Pyccel is built from source here (no wheel for macOS arm64) and needs a
# Fortran compiler, which the macOS runners do not have.
- name: Install gfortran
run: |
brew install gcc
gfortran --version

- name: Install project
run: |
pip install --upgrade pip
pip install ".[test-compiled,metal]"

# Hosted macOS runners are virtual machines and usually have no Metal GPU,
# in which case the MetalKernel launch tests are skipped. This step shows
# which case this run is in.
- name: Report Metal availability
run: |
python -c "import platform, cunumpy as xp; print(platform.machine(), 'metal_available =', xp.kernels.metal_available())"

- name: Run tests
run: pytest . -rs
40 changes: 40 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,46 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

### Fixed
- Preserve Fortran order when copying CuPy arrays to the host through
`to_numpy`, `host_call`, `evaluate_on_host`, and `PyccelKernel` conversions.

### Added
- `kernel_testing.emulated_launches()` runs every `CudaKernel` launch in a block
on the CPU, on the host buffers of the fake CuPy arrays, so code that launches
kernels can be tested without a GPU. `emulate_cuda_kernel` now accepts struct
parameters (a mapping of field values, a `CudaStructValue`, or an object with
an attribute per field). `kernel_testing.host_buffer(array)` returns the NumPy
array behind a fake CuPy array.
- `CUNUMPY_REQUIRE_CUDA=1` makes the GPU markers of `kernel_testing` fail instead
of skipping: `requires_cupy`, the `cupy` run of the `backend` fixture (which
activates CuPy strictly) and `assert_kernels_agree`. New `kernel_testing.cuda_required()`.
- `count_transfers()` records `sync` events (the host waiting for the device):
`xp.synchronize()`, the waits of the MPI helpers and of the CUDA debug mode, and,
on the fake CuPy, scalar reads of device arrays. They are in `counter.syncs`
and the report but not in `total`; `assert_no_transfers(syncs=True)` rejects them.
The real CuPy's own `float(a)` cannot be observed from Python and is not counted.
- `xp.algorithms.compact_by_mask(mask, *arrays)` moves the masked rows of arrays to
the front, in place and in order, and returns their number.
- Kernel outputs: a name in `PyccelKernel(outputs=...)` also finds a positional
argument and an index a keyword argument, using the parameter names of the
function (or the new `parameters=`); a `Kernel` supplies those of its host
function. `xp.kernels.outputs_from_annotations()` reads the outputs from the
annotations (not `Final`, `const` or a scalar), and `Kernel.from_folder()` /
`KernelCatalog.from_package()` take `outputs=` (names, indices or
`"annotations"`). `assert_kernels_agree(outputs=...)` takes parameter names and
`"name.field"` to compare only some fields of a struct argument.
- `xp.kernels.MetalKernel` runs a Metal Shading Language kernel on the GPU of an
Apple silicon Mac through MLX (`pip install 'cunumpy[metal]'`). It takes and
fills NumPy arrays, is float32 only (`float64="cast"` computes float64 data in
float32), and its copies are counted by `xp.profiling.count_transfers()`.
`xp.kernels.metal_available()` tells whether it can run.
- `xp.host_call`, `xp.evaluate_on_host` and `xp.setup_on_host` run host-only code
(SciPy splines, file readers, external libraries) with arguments of either
backend: device arrays are copied to the host, the call runs on the NumPy
backend and array results are copied back. The copies are counted by
`xp.profiling.count_transfers()`.

### Changed
- The launcher detection and the serial MPI stand-in moved to the new package
[maybempi](https://github.com/max-models/maybempi), a dependency of cunumpy.
Expand Down
4 changes: 3 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@ never hide a NumPy name:

| Submodule | Contents |
|---|---|
| `xp.kernels` | `Kernel`, `KernelCatalog`, `PyccelKernel`, `CudaKernel`, host implementations, `fuse` |
| `xp.kernels` | `Kernel`, `KernelCatalog`, `PyccelKernel`, `CudaKernel`, `MetalKernel`, host implementations, `fuse` |
| `xp.arguments` | CUDA only: `CudaStruct`, `CudaStructArguments`, `CudaArguments` |
| `xp.cuda` | CUDA only: devices, streams, debug mode, CUDA headers |
| `xp.rng` | `random_streams`, `get_rng`, `philox_*` |
Expand All @@ -36,6 +36,8 @@ never hide a NumPy name:
| `cunumpy.kernel_testing` | pytest helpers for host/CUDA kernel pairs |

Everything except `xp.cuda`, `xp.arguments` and `CudaKernel` works on both backends.
`MetalKernel` runs Metal kernels on the GPU of an Apple silicon Mac (MLX, float32, NumPy arrays;
`pip install 'cunumpy[metal]'`, see [the guide](docs/source/kernels/metal-kernel.md)).

## Install

Expand Down
104 changes: 92 additions & 12 deletions docs/source/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,12 +39,12 @@ The submodules are named so that they do not hide a NumPy name (`rng`, not

| Submodule | Backends | Contents |
|---|---|---|
| `cunumpy` | both | NumPy/CuPy namespace, backend selection, array inspection and conversion, `synchronize`, `scipy`, `require_version` |
| `cunumpy.kernels` | both | `Kernel`, `KernelCatalog`, `PyccelKernel`, `CudaKernel`, `CudaKernelVariants`, host implementations, `as_kernel_array`, `kernel_output`, `fuse` |
| `cunumpy` | both | NumPy/CuPy namespace, backend selection, array inspection and conversion, `synchronize`, `host_call`, `evaluate_on_host`, `setup_on_host`, `scipy`, `require_version` |
| `cunumpy.kernels` | both | `Kernel`, `KernelCatalog`, `PyccelKernel`, `CudaKernel`, `CudaKernelVariants`, `MetalKernel`, `metal_available`, host implementations, `as_kernel_array`, `kernel_output`, `fuse` |
| `cunumpy.arguments` | CUDA only | `CudaArguments`, `CudaStruct`, `CudaStructArguments`, `CudaStructValue`, `write_cuda_header` |
| `cunumpy.cuda` | CUDA only | device selection and memory, `stream`, streams/events, `pin_memory`, debug mode, CUDA headers and source tools (`cuda_include_dir`, `parse_cuda_signature`) |
| `cunumpy.rng` | both | `random_streams`, `get_rng`, `philox_*` |
| `cunumpy.algorithms` | both | `morton_*`, `sort_by_key`, `segment_sum` |
| `cunumpy.algorithms` | both | `morton_*`, `sort_by_key`, `segment_sum`, `compact_by_mask` |
| `cunumpy.mpi` | both | `mpi_buffer`, CUDA-aware MPI detection, `local_rank`, `synchronize_for_mpi` |
| `cunumpy.profiling` | both | `timed_region`, `nvtx_range`, `count_transfers`, `assert_no_transfers` |
| `cunumpy.memory` | both | `HostStaging`, `DeviceMirror` |
Expand Down Expand Up @@ -206,6 +206,10 @@ Converts to a host-side NumPy array. CuPy arrays are copied from device to
host. Other array-like inputs are passed through `numpy.asarray`; NumPy arrays
may therefore be returned as-is rather than copied.

F-contiguous CuPy inputs keep F order on the host; other CuPy inputs become
C-contiguous. Arbitrary device strides are not preserved. See
[Array ordering and strides](guides/array-ordering.md).

### `to_cupy(array)`

Converts an array-like input to a CuPy array. Raises `ImportError` if CuPy or
Expand Down Expand Up @@ -267,6 +271,21 @@ keys, order, positions, charges = xp.algorithms.sort_by_key(keys, positions, cha
Returns `(keys[order], order, *(a[order] for a in arrays))`, `order` as
`int64`. Equal keys keep their order, so the result is reproducible.

### `algorithms.compact_by_mask(mask, *arrays)`

Moves the rows where the boolean `mask` is True to the front of every array, in
place and in their original order, and returns how many there are. Typical use:
keep the live particles at the front of the marker arrays.

```python
n = xp.algorithms.compact_by_mask(alive, markers, weights)
markers, weights = markers[:n], weights[:n]
```

The rows after the first `n` are unspecified. The count is needed on the host,
so on CuPy each call synchronizes once. The mask and the arrays must be on the
same backend.

## Count transfers

A transfer inside a time loop is the classic performance bug of a GPU port:
Expand All @@ -286,7 +305,7 @@ with xp.profiling.count_transfers() as counter:
assert counter.total == 0, counter.report()
```

Five kinds of events are recorded:
Six kinds of events are recorded:

* `to_host`: an actual device-to-host copy through conversion, mirror, staging,
serial MPI, or host kernel helpers;
Expand All @@ -298,7 +317,12 @@ Five kinds of events are recorded:
* `fallback`: a `Kernel` without CUDA kernel calling its host kernel on the
CuPy backend (`missing_cuda="fallback"`), one event per call, naming the
kernel. Physical copies and a `kernel_conversion` marker are recorded separately;
* `device_copy`: a device-only dtype/layout conversion through CuNumpy helpers.
* `device_copy`: a device-only dtype/layout conversion through CuNumpy helpers;
* `sync`: the host waited for the device: `xp.synchronize()`, the waits of the MPI
helpers (`synchronize_for_mpi()`, `mpi_buffer()` staging) and of the CUDA debug
mode, and, on the fake CuPy, a scalar read of a device array (`float(a)`,
`int(a)`, `bool(a)`, `a.item()`, `a.tolist()`). They are in `counter.syncs` and in
the report, but not in `total`.

Only real transfers count: `to_numpy()` of a NumPy array or `to_cupy()` of a
CuPy array records nothing. The counter has the attributes `to_host`,
Expand Down Expand Up @@ -327,16 +351,17 @@ not thread-safe.

**Limitation:** only transfers made through CuNumpy are seen. Raw
`cupy.ndarray.get()`, `cupy.asarray(numpy_array)`, `numpy.asarray(cupy_array)`,
`float(device_array)`, forwarded backend calls such as `xp.asarray()`, and
implicit conversions inside other libraries are
not counted. Use `nsys` (or CuPy's profiling hooks) to find those.
forwarded backend calls such as `xp.asarray()`, and implicit conversions inside
other libraries are not counted. Neither is `float(device_array)` (an implicit
sync) with the real CuPy, which cannot be observed from Python; the fake CuPy
counts it. Use `nsys` (or CuPy's profiling hooks) to find those.

### `profiling.assert_no_transfers()`
### `profiling.assert_no_transfers(*, syncs=False)`

Context manager that raises `AssertionError` with the counter's `report()` if
the block makes a host/device transfer or host fallback through CuNumpy.
Device-only dtype/layout conversions are allowed. It yields the `TransferCounter`
too. An exception raised inside the block propagates as it is:
Device-only dtype/layout conversions are allowed, and so are syncs unless
`syncs=True`. It yields the `TransferCounter` too. An exception raised inside the block propagates as it is:

```python
def test_time_step_stays_on_the_device():
Expand Down Expand Up @@ -808,8 +833,18 @@ CuPy arrays.
* `is_array`: predicate for host array values to convert back to CuPy. The
default is `isinstance(value, numpy.ndarray)`.
* `outputs`: sequence of arguments the kernel may write to. Entries are
positional indices or keyword names. If omitted, every converted array is
indices or parameter names; a name also finds a positional argument, and an
index a keyword argument, when the parameter names are known (from the
Python signature, or `parameters`). If omitted, every converted array is
copied back.
* `parameters`: the names of the positional parameters, or a function that
returns them, for compiled kernels without a Python signature. A `Kernel`
supplies the names of its host function.

`kernels.outputs_from_annotations(function)` returns the parameters a kernel may
write to, read from its annotations: everything that is not `Final`, `const` or a
scalar. Use it as `outputs=outputs_from_annotations(push)`, or pass
`outputs="annotations"` to `Kernel.from_folder()` / `KernelCatalog.from_package()`.

### Example and output declarations

Expand Down Expand Up @@ -857,6 +892,51 @@ lists) are converted back using `is_array`; dictionaries in return values are
not recursively converted. On the NumPy path, the original return value and
normal Python mutation and exception behavior are preserved.

## `kernels.MetalKernel`

A Metal Shading Language kernel for the GPU of an Apple silicon Mac, run with
[MLX](https://github.com/ml-explore/mlx) (`pip install 'cunumpy[metal]'`). It
takes NumPy arrays and fills the output arrays you pass, so no backend switch
is needed.

```python
import numpy as np
import cunumpy as xp

scale = xp.kernels.MetalKernel(
"uint i = thread_position_in_grid.x; y[i] = a[0] * x[i];",
inputs=["x", "a"],
outputs=["y"],
)
x = np.arange(8, dtype=np.float32)
y = np.empty_like(x)
scale(x, 2.0, out=y)
```

`MetalKernel(source, inputs, outputs, *, name="cunumpy_kernel", header="",
threadgroup=256, float64="error", atomic_outputs=False, init_value=None)`

- `source` is the body of the kernel function. MLX generates the signature: each
name in `inputs` and `outputs` is a pointer to the flat, row-major data of that
array, so `x[i]` is the flat index. `thread_position_in_grid` and the other
Metal attributes used in the body are added automatically. `header` goes before
the function (includes, defines, helper functions).
- Calling the kernel: `kernel(*inputs, out=array_or_arrays, n_threads=None,
template=None)`. `n_threads` is the total thread count (default: first axis of
the first output). `template` gives compile-time constants, e.g.
`template={"NSTEPS": 200}`. It returns the output array, or a tuple of them.
- Outputs are uninitialized: write every element, pass the old array as an input
too if the kernel reads it, or set `init_value`.
- **float32 only.** The Apple GPU has no float64: a float64 input or output
raises `TypeError`. With `float64="cast"` float64 data is computed in float32
(a push over 200 steps agreed with float64 to about 3e-5).
- Every call copies the inputs to MLX arrays and the results back, counted as
`to_device` and `to_host` transfers by `count_transfers()`. On a 4M-particle
push these copies were about 7 ms next to an 11.5 ms kernel.
- `xp.kernels.metal_available()` is True if MLX is installed and a Metal GPU is
present. Without them, calling a `MetalKernel` raises `ImportError` or
`RuntimeError`.

## `kernels.CudaKernel`

### Constructor
Expand Down
Loading
Loading