Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
32 changes: 32 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,33 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
## [Unreleased]

### Changed
- The launcher detection and the serial MPI stand-in moved to the new package
[maybempi](https://github.com/max-models/maybempi), a dependency of cunumpy.
`xp.mpi` re-exports them (`get_mpi`, `launched_under_mpi`, `local_rank`,
`SerialMPI`, `SerialComm`, ...), plus the new `xp.mpi.is_serial`. The override
variable is now `MAYBEMPI=1`/`0`; `CUNUMPY_MPI` is no longer read.
- `CudaKernel` and `CudaKernelVariants` moved to `cunumpy.kernels`, next to
`Kernel` and `PyccelKernel`. The argument classes `CudaArguments`,
`CudaStruct`, `CudaStructArguments`, `CudaStructValue` and
`write_cuda_header` moved to the new `cunumpy.arguments`. `cunumpy.cuda`
keeps the device runtime and the CUDA source tools. The old
`cunumpy.cuda.<name>` names are removed.
- **Removed** the deprecated names that were kept for one release after the
0.5 reorganisation, with no replacement other than the submodules:
the top-level helpers (`xp.CudaKernel`, `xp.mpi_buffer`, `xp.fuse` as cunumpy's,
...; use `xp.kernels`, `xp.arguments`, `xp.cuda`, `xp.mpi`, `xp.rng`,
`xp.algorithms`, `xp.profiling`, `xp.memory` and `xp.petsc`), and the modules
`cunumpy.testing` (now `cunumpy.kernel_testing`), `cunumpy.kernel`,
`cunumpy.dispatch` and `cunumpy.cuda_kernel` (public names are in
`cunumpy.kernels` and `cunumpy.arguments`). Also removed the placeholder
`cunumpy.main`. `xp.fuse` is now CuPy's own `fuse`.
- **Removed** `kernels.KernelArguments`, `kernels.resolve_host_args`,
`kernels.PyccelStructArguments` and the `__host_args__()` protocol, with no
replacement. Kernels receive argument objects as they are. Write the host
argument class (e.g. pyccel) and a `CudaStructArguments` with the same
constructor, and let the owner of the arrays build the one for its backend.
`Kernel(dispatch="arrays")` now treats every object with `__cuda_args__()` as
a device argument.
- Rename `CUNUMPY_KERNEL_IMPLEMENTATION` to
`CUNUMPY_HOST_KERNEL_IMPLEMENTATION` to make its host-only scope explicit.
The former environment variable is no longer read; update job scripts.
Expand All @@ -17,6 +44,11 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
The former function names are removed without compatibility aliases.

### Added
- C-contiguous array views `CArray1D<T>` to `CArray4D<T>` in
`cunumpy/array_view.cuh`. They hold a pointer and shape only, so `a(i, j)` is
`data[i * shape[1] + j]`. As kernel parameters or struct fields, they reject
non-contiguous arrays and never copy them. `CudaStruct.from_signature` and
`from_pyccel_class` take `contiguous=True` or field names to generate them.
- Device implementation selection via `kernels.set_device_kernel_implementation`,
`get_device_kernel_implementation`, `use_device_kernel_implementation`, and
`CUNUMPY_DEVICE_KERNEL_IMPLEMENTATION`. Accept `"cuda"` or automatic selection
Expand Down
45 changes: 16 additions & 29 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,8 +24,9 @@ never hide a NumPy name:

| Submodule | Contents |
|---|---|
| `xp.cuda` | CUDA only: `CudaKernel`, `CudaStruct`, CUDA headers, devices, streams |
| `xp.kernels` | `Kernel`, `KernelCatalog`, `PyccelKernel`, host implementations, `fuse` |
| `xp.kernels` | `Kernel`, `KernelCatalog`, `PyccelKernel`, `CudaKernel`, host implementations, `fuse` |
| `xp.arguments` | CUDA only: `CudaStruct`, `CudaStructArguments`, `CudaArguments` |
| `xp.cuda` | CUDA only: devices, streams, debug mode, CUDA headers |
| `xp.rng` | `random_streams`, `get_rng`, `philox_*` |
| `xp.algorithms` | `morton_*`, `sort_by_key`, `cell_offsets`, `segment_boundaries`, `segment_sum`, `SegmentPlan` |
| `xp.mpi` | `mpi_buffer`, reusable `MPIStaging`, CUDA-aware MPI |
Expand All @@ -34,7 +35,7 @@ never hide a NumPy name:
| `xp.petsc` | `petsc_vec` |
| `cunumpy.kernel_testing` | pytest helpers for host/CUDA kernel pairs |

Everything except `xp.cuda` works on both backends.
Everything except `xp.cuda`, `xp.arguments` and `CudaKernel` works on both backends.

## Install

Expand Down Expand Up @@ -333,7 +334,7 @@ def axpy(a, x, y, n): # host version, e.g. compiled with Pyccel
y[i] += a * x[i]


kernel = xp.kernels.Kernel(axpy, xp.cuda.CudaKernel(AXPY, "axpy"))
kernel = xp.kernels.Kernel(axpy, xp.kernels.CudaKernel(AXPY, "axpy"))

with xp.use_backend("cupy"):
x = xp.arange(1000, dtype=xp.float64)
Expand All @@ -356,8 +357,8 @@ the matching memory layout) and packs values into it, which the kernel takes
as one parameter:

```python
Vec = xp.cuda.CudaStruct("Vec", [("data", "double*"), ("n", "int")])
scale = xp.cuda.CudaKernel(
Vec = xp.arguments.CudaStruct("Vec", [("data", "double*"), ("n", "int")])
scale = xp.kernels.CudaKernel(
Vec.declaration
+ r"""
extern "C" __global__ void scale(Vec v, double a) {
Expand All @@ -379,29 +380,15 @@ device copies. Call it once when the object is
built, not per kernel call; on the NumPy backend it raises, so host data is
never copied to the device implicitly.
When the host kernel takes such a group as one object too (e.g. a Pyccel class
holding NumPy arrays), give the group both forms with `KernelArguments`:
`__host_args__()` returns the object for the host kernel, `__cuda_args__()`
the flattened device arguments. `Kernel` and `PyccelKernel` resolve
`__host_args__()` on the host path and `CudaKernel` flattens `__cuda_args__()`
on the CUDA path, so the call site is the same on both backends and each form
can be built lazily on first access (a CPU run never builds device arguments):
holding NumPy arrays), write a CUDA class with the same constructor and
attributes (a `CudaStructArguments`, see below) and let the owner of the arrays
build the one for the active backend. CuNumpy passes argument objects through
as they are and never converts one form into the other:

```python
class ParticleArguments(xp.kernels.KernelArguments):
def __init__(self, markers):
self.markers = markers
self._host = None

def __host_args__(self):
if self._host is None:
self._host = MarkerArguments(self.markers) # Pyccel class
return self._host

def __cuda_args__(self):
return (self.markers, self.markers.shape[0])


kernel(particles.kernel_args, dt, n_threads=n) # host or CUDA kernel
args_class = CudaMarkerArguments if xp.is_gpu(markers) else MarkerArguments
particles.args_markers = args_class(markers, markers.shape[0])
kernel(particles.args_markers, dt) # host or CUDA kernel
```

Kernels ported from pyccel index arrays like `markers[ip, j]`, which needs
Expand All @@ -417,11 +404,11 @@ class MarkerArguments:
def __init__(self, markers: "float[:, :]", n_markers: int, valid: "bool[:]"): ...


MarkerArgs = xp.cuda.CudaStruct.from_signature(MarkerArguments.__init__, "MarkerArgs")
MarkerArgs = xp.arguments.CudaStruct.from_signature(MarkerArguments.__init__, "MarkerArgs")
MarkerArgs.to_header(
"marker_args.cuh"
) # Array2D<double> markers; long long n_markers; ...
push = xp.cuda.CudaKernel(
push = xp.kernels.CudaKernel(
r"""
#include "marker_args.cuh"
#include <cunumpy/index.cuh>
Expand Down
Loading
Loading