Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
56 changes: 56 additions & 0 deletions .github/workflows/macos.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,56 @@
name: Tests (macOS)

on:
push:
branches:
- main
- devel
pull_request:
branches:
- main
- devel
workflow_dispatch:

jobs:
build:
# macos-latest is Apple silicon (arm64): checks the NumPy backend, the
# compiled host kernels and the MLX import on the platform of MetalKernel.
runs-on: macos-latest

strategy:
fail-fast: false
matrix:
python-version: ["3.10", "3.13"]

steps:
- name: Checkout code
uses: actions/checkout@v4

- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python-version }}
cache: pip
cache-dependency-path: pyproject.toml

# Pyccel is built from source here (no wheel for macOS arm64) and needs a
# Fortran compiler, which the macOS runners do not have.
- name: Install gfortran
run: |
brew install gcc
gfortran --version

- name: Install project
run: |
pip install --upgrade pip
pip install ".[test-compiled,metal]"

# Hosted macOS runners are virtual machines and usually have no Metal GPU,
# in which case the MetalKernel launch tests are skipped. This step shows
# which case this run is in.
- name: Report Metal availability
run: |
python -c "import platform, cunumpy as xp; print(platform.machine(), 'metal_available =', xp.kernels.metal_available())"

- name: Run tests
run: pytest . -rs
82 changes: 82 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,88 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

## [0.6.3] - 2026-10-08

### Fixed
- The emulation of kernels with `__syncthreads` compiles with GCC on macOS.
- Preserve Fortran order when copying CuPy arrays to the host through
`to_numpy`, `host_call`, `evaluate_on_host`, and `PyccelKernel` conversions.

### Added
- `emulate_cuda_kernel` compiles each kernel once, into a shared library that
takes the arguments at run time and runs in the process (ctypes), on the arrays
themselves; later launches with any values and sizes reuse it. Libraries are
cached per process and on disk (`kernel_testing.emulation_cache_dir()`,
`CUNUMPY_EMULATION_CACHE`, default `~/.cache/cunumpy/emulation`).
`kernel_testing.compile_for_emulation(kernel)` builds one without a launch.
`__trap()` (a failed bounds check) still raises `RuntimeError`, but a
segmentation fault now ends the process.
- Inside `emulated_launches()`, `CudaKernel.compile()` (and so `recompile()`,
`CudaKernelVariants.compile_all()` and `KernelCatalog.compile_all()`) builds the
emulation library instead of compiling CUDA, and returns None. Code that compiles
its kernels before the time loop runs on the fake CuPy and still reports compile
errors there; patching `CudaKernel.compile` is no longer needed.
- The emulation compiles inline PTX (`asm(...)`, `asm volatile(...)`) as a trap, so
a kernel with an `asm("trap;")` branch needs no `-Dasm(x)=__trap()` option.
- `kernel_testing.fake_cupy_session()` runs a CuPy-backend program on the CPU
(CuPy backend on the fake CuPy, launches and compilation emulated);
`kernel_testing.device_backend_available()` and the marker
`requires_device_backend` (a GPU or the fake CuPy);
`kernel_testing.run_in_fake_cupy_subprocess(code)` runs code in a serial child
process on the fake CuPy (rank 0 only under MPI) and fails the test with the
signal or exit code and the end of the child's output.
- `profiling.TransferBudget` counts transfers per phase of a program
(`budget.phase(name)`, the decorator `budget.count(name)`, `start()`/`stop()`)
and checks a rule per phase (`require(phase, allow={"to_host": {"max_nbytes": 8}},
calls=n)`, `check()`, `report()`). `count_transfers(into=counter)` adds to an
existing counter (counted once when nested).
- `TransferEvent` has `blocking` (False for `to_host_async()` and
`HostStaging.copy()`) and `implicit` (True for the scalar reads of the fake CuPy).
- `xp.to_host_async(a)` copies a device scalar or small array to the host on a
separate stream without waiting; the returned `memory.HostCopy` has `ready()`
(never waits) and `result()`.
- `kernel_testing.emulated_launches()` runs every `CudaKernel` launch in a block
on the CPU, on the host buffers of the fake CuPy arrays, so code that launches
kernels can be tested without a GPU. `emulate_cuda_kernel` now accepts struct
parameters (a mapping of field values, a `CudaStructValue`, or an object with
an attribute per field). `kernel_testing.host_buffer(array)` returns the NumPy
array behind a fake CuPy array.
- `CUNUMPY_REQUIRE_CUDA=1` makes the GPU markers of `kernel_testing` fail instead
of skipping: `requires_cupy`, the `cupy` run of the `backend` fixture (which
activates CuPy strictly) and `assert_kernels_agree`. New `kernel_testing.cuda_required()`.
- `count_transfers()` records `sync` events (the host waiting for the device):
`xp.synchronize()`, the waits of the MPI helpers and of the CUDA debug mode, and,
on the fake CuPy, scalar reads of device arrays. They are in `counter.syncs`
and the report but not in `total`; `assert_no_transfers(syncs=True)` rejects them.
The real CuPy's own `float(a)` cannot be observed from Python and is not counted.
- `xp.algorithms.compact_by_mask(mask, *arrays, axis=0)` moves the masked entries
of arrays to the front along `axis`, in place and in order, and returns their
number; `axis=-1` compacts component-major `(ncomp, N)` and `(N,)` arrays together.
- `as_kernel_array(..., strided=True)` and `kernel_output(..., strided=True)` take a
NumPy array in C order with gaps (positive strides, each at least the extent of
the next axis, e.g. `storage[:, :n]`) unchanged for host kernels, instead of
copying it to a C-contiguous array; Pyccel's wrappers take such arrays.
- `CudaKernel(n_threads_from="last_axis")`: one thread per entry of the last axis of
the first array argument, for component-major `(ncomp, N)` marker arrays.
- Kernel outputs: a name in `PyccelKernel(outputs=...)` also finds a positional
argument and an index a keyword argument, using the parameter names of the
function (or the new `parameters=`); a `Kernel` supplies those of its host
function. `xp.kernels.outputs_from_annotations()` reads the outputs from the
annotations (not `Final`, `const` or a scalar), and `Kernel.from_folder()` /
`KernelCatalog.from_package()` take `outputs=` (names, indices or
`"annotations"`). `assert_kernels_agree(outputs=...)` takes parameter names and
`"name.field"` to compare only some fields of a struct argument.
- `xp.kernels.MetalKernel` runs a Metal Shading Language kernel on the GPU of an
Apple silicon Mac through MLX (`pip install 'cunumpy[metal]'`). It takes and
fills NumPy arrays, is float32 only (`float64="cast"` computes float64 data in
float32), and its copies are counted by `xp.profiling.count_transfers()`.
`xp.kernels.metal_available()` tells whether it can run.
- `xp.host_call`, `xp.evaluate_on_host` and `xp.setup_on_host` run host-only code
(SciPy splines, file readers, external libraries) with arguments of either
backend: device arrays are copied to the host, the call runs on the NumPy
backend and array results are copied back. The copies are counted by
`xp.profiling.count_transfers()`.

### Changed
- The launcher detection and the serial MPI stand-in moved to the new package
[maybempi](https://github.com/max-models/maybempi), a dependency of cunumpy.
Expand Down
4 changes: 3 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@ never hide a NumPy name:

| Submodule | Contents |
|---|---|
| `xp.kernels` | `Kernel`, `KernelCatalog`, `PyccelKernel`, `CudaKernel`, host implementations, `fuse` |
| `xp.kernels` | `Kernel`, `KernelCatalog`, `PyccelKernel`, `CudaKernel`, `MetalKernel`, host implementations, `fuse` |
| `xp.arguments` | CUDA only: `CudaStruct`, `CudaStructArguments`, `CudaArguments` |
| `xp.cuda` | CUDA only: devices, streams, debug mode, CUDA headers |
| `xp.rng` | `random_streams`, `get_rng`, `philox_*` |
Expand All @@ -36,6 +36,8 @@ never hide a NumPy name:
| `cunumpy.kernel_testing` | pytest helpers for host/CUDA kernel pairs |

Everything except `xp.cuda`, `xp.arguments` and `CudaKernel` works on both backends.
`MetalKernel` runs Metal kernels on the GPU of an Apple silicon Mac (MLX, float32, NumPy arrays;
`pip install 'cunumpy[metal]'`, see [the guide](docs/source/kernels/metal-kernel.md)).

## Install

Expand Down
Loading
Loading