Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 14 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,21 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

### Changed
- Rename `CUNUMPY_KERNEL_IMPLEMENTATION` to
`CUNUMPY_HOST_KERNEL_IMPLEMENTATION` to make its host-only scope explicit.
The former environment variable is no longer read; update job scripts.
CUDA dispatch is unchanged.
- Rename the `kernels` selection functions to `set_host_kernel_implementation`,
`get_host_kernel_implementation`, and `use_host_kernel_implementation`.
The former function names are removed without compatibility aliases.

### Added
- Device implementation selection via `kernels.set_device_kernel_implementation`,
`get_device_kernel_implementation`, `use_device_kernel_implementation`, and
`CUNUMPY_DEVICE_KERNEL_IMPLEMENTATION`. Accept `"cuda"` or automatic selection
(`None`); explicit CUDA selection rejects missing CUDA kernels instead of
falling back to the host. Unsupported implementation names raise.
- CUDA launches infer thread counts from the first array by default, including
arrays in argument objects. 1D blocks use rows; multidimensional blocks use
matching leading shape axes. Explicit sizes and callbacks override inference.
Expand Down
40 changes: 33 additions & 7 deletions docs/source/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -1713,16 +1713,16 @@ files, as loaders by name (`{"numpy": lambda: push_numpy}`). Takes the options o
no host kernel module and `ModuleNotFoundError` if `package` is not a package.
`kernel.selected(device=True)` names the implementation for device arguments.

## `kernels.HostImplementations`, `kernels.set_kernel_implementation`
## `kernels.HostImplementations`, `kernels.set_host_kernel_implementation`

```python
host = xp.kernels.HostImplementations(
"push",
{"pyccel": load_compiled, "numpy": lambda: push_numpy, "python": lambda: push},
)
host(*args) # the default implementation
xp.kernels.set_kernel_implementation("numpy") # every kernel: like xp.set_backend
with xp.kernels.use_kernel_implementation("python"): # like xp.use_backend
xp.kernels.set_host_kernel_implementation("numpy") # every kernel: like xp.set_backend
with xp.kernels.use_host_kernel_implementation("python"): # like xp.use_backend
host(*args)
```

Expand All @@ -1733,14 +1733,40 @@ Loaded on first use; `available(name)` loads and reports, `get(name)` returns it
or raises `LookupError` (missing, or failed to load with the error as cause),
`errors` maps names to load errors, `names` lists them, `python` is the
uncompiled function, `build()` loads the default now. A call runs
`selected()`: the implementation set with `set_kernel_implementation(name)` (or
`use_kernel_implementation`, or the environment variable
`CUNUMPY_KERNEL_IMPLEMENTATION` read at import), which raises if the kernel
`selected()`: the implementation set with `set_host_kernel_implementation(name)` (or
`use_host_kernel_implementation`, or the environment variable
`CUNUMPY_HOST_KERNEL_IMPLEMENTATION` read at import), which raises if the kernel
lacks it or cannot load it, else the default: the first available of pyccel,
numba and NumPy, and else `"python"` with a `RuntimeWarning` (once).
`get_kernel_implementation()` reads the setting; `None` is the default. The
`get_host_kernel_implementation()` reads the setting; `None` is the default. The
setting is global, not per thread, and applies to host calls only.

## Device kernel implementation selection

```python
xp.kernels.set_device_kernel_implementation("cuda")
xp.kernels.get_device_kernel_implementation() # "cuda"
with xp.kernels.use_device_kernel_implementation(None):
push(positions, velocities, dt) # automatic selection
xp.kernels.set_device_kernel_implementation(None)
```

`xp.kernels.DEVICE_IMPLEMENTATIONS` is currently `("cuda",)`. The setter accepts
`"cuda"` or `None` (automatic selection); unsupported values raise `ValueError`
without changing the setting. The getter reports the requested setting, so
it returns `None` in automatic mode even when a call would use CUDA.
`CUNUMPY_DEVICE_KERNEL_IMPLEMENTATION` initializes the setting at import;
an unset or empty value means automatic selection. Environment values are
case-insensitive and stripped of whitespace; unsupported values fail at import.

Explicit `"cuda"` requires a CUDA implementation for device dispatch: missing
implementations raise `LookupError`, even with `missing_cuda="fallback"`.
Automatic mode preserves that per-kernel fallback policy. This affects
`Kernel` and `KernelCatalog` device dispatch; it does not switch the array
backend, alter host calls, or affect direct `CudaKernel`/CuPy RawKernel calls.
The context manager restores the previous setting even after an exception.
The setting is global, not per thread.

## `kernels.CompiledHostKernel`

```python
Expand Down
9 changes: 8 additions & 1 deletion docs/source/installation.md
Original file line number Diff line number Diff line change
Expand Up @@ -69,7 +69,8 @@ Tests that need a GPU are skipped automatically where CuPy is not functional.
| --- | --- |
| `CUNUMPY_BACKEND=cupy` | start with the CuPy backend instead of NumPy (read once, at import) |
| `CUNUMPY_CUDA_DEBUG=1` | enable [CUDA debug mode](kernels/debugging.md) for all kernels |
| `CUNUMPY_KERNEL_IMPLEMENTATION=numpy` | choose the host kernel implementation (read at import) |
| `CUNUMPY_HOST_KERNEL_IMPLEMENTATION=numpy` | choose the host kernel implementation (read at import) |
| `CUNUMPY_DEVICE_KERNEL_IMPLEMENTATION=cuda` | require CUDA for device kernel dispatch (read at import); unset allows the kernel's configured fallback |
| `CUNUMPY_MPI=1` / `0` | require MPI / use serial MPI regardless of launcher detection |
| `CUNUMPY_FAKE_CUPY=1` | install the strict CPU stand-in for CuPy for tests |
| `CUNUMPY_REQUIRE_CUDA=1` | require a real usable GPU when starting the test suite (CI guard) |
Expand All @@ -78,6 +79,12 @@ Use `CUNUMPY_BACKEND` in job scripts; the former `ARRAY_BACKEND` setting is no
longer read. Standard toolchain/device variables such as `CXX` and
`CUDA_VISIBLE_DEVICES` keep their standard meanings.

Use `CUNUMPY_HOST_KERNEL_IMPLEMENTATION` instead of the former
`CUNUMPY_KERNEL_IMPLEMENTATION`, which is no longer read. This selects host
implementations only; CUDA dispatch is unaffected. The Python functions
`set_host_kernel_implementation`, `get_host_kernel_implementation`, and
`use_host_kernel_implementation` provide runtime selection under `xp.kernels`.

MPI launchers also export node-local rank variables (`OMPI_COMM_WORLD_LOCAL_RANK`,
`SLURM_LOCALID`, ...), which `xp.mpi.local_rank()` reads to pick a GPU per process.

Expand Down
31 changes: 25 additions & 6 deletions docs/source/kernels/dispatch.md
Original file line number Diff line number Diff line change
Expand Up @@ -173,20 +173,39 @@ none is. Device arrays run the CUDA kernel. To choose, use the same pattern as
for the array backend:

```python
xp.kernels.set_kernel_implementation("numpy") # like xp.set_backend
with xp.kernels.use_kernel_implementation("numba"): # like xp.use_backend
xp.kernels.set_host_kernel_implementation("numpy") # like xp.set_backend
with xp.kernels.use_host_kernel_implementation("numba"): # like xp.use_backend
push(positions, velocities, dt)
xp.kernels.set_kernel_implementation(None) # back to the default
xp.kernels.set_host_kernel_implementation(None) # back to the default
```

or `CUNUMPY_KERNEL_IMPLEMENTATION=numpy` for a whole run (read at import, like
or `CUNUMPY_HOST_KERNEL_IMPLEMENTATION=numpy` for a whole run (read at import, like
`CUNUMPY_BACKEND`). A chosen implementation that a kernel does not have, or
cannot load, raises `LookupError` instead of running another one: a benchmark
of numba never silently measures NumPy. `kernel.implementations` lists the
implementations, `kernel.selected()` names the one a call with host arrays runs
now (`kernel.selected(device=True)` for device arrays), and
`kernel.host_kernel.kernel.errors` holds why an implementation failed to load.

### Device implementation selection

Currently CUDA is the only device implementation. To require it explicitly:

```python
xp.kernels.set_device_kernel_implementation("cuda")
assert xp.kernels.get_device_kernel_implementation() == "cuda"
with xp.kernels.use_device_kernel_implementation(None):
push(positions, velocities, dt) # automatic selection, normal fallback policy
xp.kernels.set_device_kernel_implementation(None) # restore automatic selection
```

`CUNUMPY_DEVICE_KERNEL_IMPLEMENTATION=cuda` sets the same choice at import.
Unsupported values raise `ValueError`. An explicit CUDA choice raises
`LookupError` for a missing device implementation, including kernels configured
with `missing_cuda="fallback"`. The default (`None`, or an unset/empty environment
variable) preserves that fallback policy. This setting controls device dispatch;
the array backend and host implementation are selected independently.

## Compiled Pyccel host kernels

By default the host kernel is the Python function itself, which is fine for
Expand Down Expand Up @@ -218,8 +237,8 @@ catalog = xp.kernels.KernelCatalog.from_package(
host implementations" above); `catalog["push"].host_kernel.kernel.available("pyccel")`
reports whether the compiled version builds, and `catalog["push"].selected()`
which version runs. To test the path of a machine without Pyccel, run the code
inside `with xp.kernels.use_kernel_implementation("numpy"):` (or set
`CUNUMPY_KERNEL_IMPLEMENTATION=numpy` for a whole run). Note that `epyccel` compiles again on every call; a
inside `with xp.kernels.use_host_kernel_implementation("numpy"):` (or set
`CUNUMPY_HOST_KERNEL_IMPLEMENTATION=numpy` for a whole run). Note that `epyccel` compiles again on every call; a
code that compiles at run time usually keeps the builds in an on-disk cache
keyed on the module source, so that only the first run after an edit compiles.

Expand Down
3 changes: 2 additions & 1 deletion src/cunumpy/LLM_GUIDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -78,7 +78,8 @@ https://max-models.github.io/cunumpy/ and in `docs/source/` of the repository.
| host arrays reach kernels while CuPy is active | `Kernel(..., dispatch="arrays")` / `from_package(..., dispatch="arrays")`: CUDA only for device arguments |
| one kernel folder declares its kernel in its own `__init__.py` | `kernel = xp.kernels.Kernel.from_folder(__name__, host_suffix="_pyccel", compile_host=..., dispatch="arrays")`; `<name>_numba.py`, `<name>_numpy.py` in the folder are further host implementations |
| bring a `dispatch="arrays"` kernel's arguments to the side of the main array | `xp.kernels.as_kernel_array(a, like=grid, dtype=float)`; outputs: `with xp.kernels.kernel_output(out, like=grid, dtype=float) as buf:` |
| choose the host implementation (pyccel/numba/numpy/python) | `xp.kernels.set_kernel_implementation("numpy")`, `with xp.kernels.use_kernel_implementation(...)`, `CUNUMPY_KERNEL_IMPLEMENTATION=numpy`; default: first available of pyccel, numba, numpy; `kernel.selected()` |
| choose the host implementation (pyccel/numba/numpy/python) | `xp.kernels.set_host_kernel_implementation("numpy")`, `with xp.kernels.use_host_kernel_implementation(...)`, `CUNUMPY_HOST_KERNEL_IMPLEMENTATION=numpy`; default: first available of pyccel, numba, numpy; `kernel.selected()` |
| require CUDA for device kernel dispatch | `xp.kernels.set_device_kernel_implementation("cuda")`, `get_device_kernel_implementation()`, `with xp.kernels.use_device_kernel_implementation(...)`, `CUNUMPY_DEVICE_KERNEL_IMPLEMENTATION=cuda`; default `None` preserves `missing_cuda` policy; explicit CUDA rejects host fallback |
| check host and CUDA kernels take the same parameters | `catalog.check_signatures()` (in a unit test) |
| test a CUDA kernel's arithmetic without a GPU | `cunumpy.kernel_testing.emulate_cuda_kernel(kernel, *numpy_args, n_threads=n)` (C++ compiler; shared memory and __syncthreads ok, no warp ops; `shared_mem=` for extern shared) |
| shared-memory budget of a block | `xp.cuda.max_shared_memory_per_block()` (48 KiB without a GPU) |
Expand Down
12 changes: 11 additions & 1 deletion src/cunumpy/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,17 @@
# moved to. They still resolve (with a DeprecationWarning) until cunumpy 0.6.
_MOVED = {
**dict.fromkeys(cuda.__all__, "cuda"),
**dict.fromkeys(kernels.__all__, "kernels"),
**dict.fromkeys(
(
name
for name in kernels.__all__
if not name.endswith(
("_host_kernel_implementation", "_device_kernel_implementation")
)
and name != "DEVICE_IMPLEMENTATIONS"
),
"kernels",
),
**dict.fromkeys(rng.__all__, "rng"),
**dict.fromkeys(algorithms.__all__, "algorithms"),
# the names of cunumpy.mpi that were at the top level (not the later ones)
Expand Down
16 changes: 13 additions & 3 deletions src/cunumpy/_dispatch.py
Original file line number Diff line number Diff line change
Expand Up @@ -39,6 +39,7 @@
CompiledHostKernel,
HostImplementations,
PyccelKernel,
get_device_kernel_implementation,
resolve_host_args,
)
from cunumpy._transfers import _ACTIVE as _COUNTERS
Expand Down Expand Up @@ -268,7 +269,7 @@ def from_folder(
* ``<name><cuda_suffix>``: the CUDA kernel.

The host implementations form a :class:`~cunumpy.kernels.HostImplementations`:
a call runs the one set with :func:`~cunumpy.kernels.set_kernel_implementation`,
a call runs the one set with :func:`~cunumpy.kernels.set_host_kernel_implementation`,
or by default the first available of pyccel, numba and NumPy. The
folder's own ``__init__.py`` can declare its kernel with this method, so
that the kernel is imported from where it is written::
Expand Down Expand Up @@ -516,10 +517,12 @@ def selected(self, device: bool = False) -> str:
"""The implementation a call with host (or `device`) arguments runs now.

For host arguments: the setting of
:func:`~cunumpy.kernels.set_kernel_implementation` or the default (loads it), or
:func:`~cunumpy.kernels.set_host_kernel_implementation` or the default (loads it), or
``"host"`` for a host kernel that is not a
:class:`~cunumpy.kernels.HostImplementations`. For device arguments ``"cuda"``,
or ``"host"`` if there is no CUDA kernel and ``missing_cuda="fallback"``.
Explicit CUDA selection rejects missing CUDA implementations with
``LookupError``, including when host fallback is configured.
Useful to check that a run does not use a slow path.
"""
if device:
Expand All @@ -536,6 +539,8 @@ def get_kernel(self) -> PyccelKernel | CudaKernel:

Raises
------
LookupError
On the CuPy backend, if CUDA is explicitly selected but missing.
NotImplementedError
On the CuPy backend, if there is no CUDA kernel and
``missing_cuda="raise"``.
Expand All @@ -545,9 +550,14 @@ def get_kernel(self) -> PyccelKernel | CudaKernel:
return self._device_kernel()

def _device_kernel(self) -> PyccelKernel | CudaKernel:
"""The CUDA kernel, or what ``missing_cuda`` says without one."""
"""Resolve the device selection, honoring fallback only in automatic mode."""
if self._cuda_kernel is not None:
return self._cuda_kernel
if get_device_kernel_implementation() == "cuda":
raise LookupError(
f"kernel {self._name!r} has no 'cuda' implementation "
"(explicitly selected device implementation)",
)
if self._missing_cuda == "raise":
expected = (
"" if self._cuda_path is None else f" (expected {self._cuda_path})"
Expand Down
65 changes: 56 additions & 9 deletions src/cunumpy/_kernel.py
Original file line number Diff line number Diff line change
Expand Up @@ -546,11 +546,11 @@ def _check_implementation(name: str | None) -> str | None:


_KERNEL_IMPLEMENTATION: str | None = _check_implementation(
os.environ.get("CUNUMPY_KERNEL_IMPLEMENTATION", "").strip().lower() or None,
os.environ.get("CUNUMPY_HOST_KERNEL_IMPLEMENTATION", "").strip().lower() or None,
)


def set_kernel_implementation(name: str | None) -> None:
def set_host_kernel_implementation(name: str | None) -> None:
"""Choose the host implementation every kernel runs, like :func:`set_backend`.

``"pyccel"``, ``"numba"``, ``"numpy"`` or ``"python"`` (the uncompiled
Expand All @@ -559,42 +559,89 @@ def set_kernel_implementation(name: str | None) -> None:
or whose chosen implementation is unavailable (e.g. pyccel failed to
compile), raises instead of running another one. CUDA kernels are not
affected: device arrays always run the CUDA version. The environment
variable ``CUNUMPY_KERNEL_IMPLEMENTATION`` (read when cunumpy is imported)
variable ``CUNUMPY_HOST_KERNEL_IMPLEMENTATION`` (read when cunumpy is imported)
sets it for a whole run.
"""
global _KERNEL_IMPLEMENTATION
_KERNEL_IMPLEMENTATION = _check_implementation(name)


def get_kernel_implementation() -> str | None:
"""The host implementation set with :func:`set_kernel_implementation`, or None."""
def get_host_kernel_implementation() -> str | None:
"""The host implementation set with :func:`set_host_kernel_implementation`, or None."""
return _KERNEL_IMPLEMENTATION


@contextmanager
def use_kernel_implementation(name: str | None) -> Iterator[None]:
def use_host_kernel_implementation(name: str | None) -> Iterator[None]:
"""Temporarily choose the host implementation, like :func:`use_backend`.

For tests and benchmarks, e.g. ``with xp.kernels.use_kernel_implementation("numpy"):``
For tests and benchmarks, e.g. ``with xp.kernels.use_host_kernel_implementation("numpy"):``
to run the code path of a machine without pyccel. The setting is global,
not per thread.
"""
global _KERNEL_IMPLEMENTATION
previous = _KERNEL_IMPLEMENTATION
set_kernel_implementation(name)
set_host_kernel_implementation(name)
try:
yield
finally:
_KERNEL_IMPLEMENTATION = previous


DEVICE_IMPLEMENTATIONS = ("cuda",)


def _check_device_implementation(name: str | None) -> str | None:
if name is not None and name not in DEVICE_IMPLEMENTATIONS:
raise ValueError(
f"device kernel implementation must be one of {DEVICE_IMPLEMENTATIONS} "
f"or None, got {name!r}",
)
return name


_DEVICE_KERNEL_IMPLEMENTATION = _check_device_implementation(
os.environ.get("CUNUMPY_DEVICE_KERNEL_IMPLEMENTATION", "").strip().lower() or None,
)


def set_device_kernel_implementation(name: str | None) -> None:
"""Choose the device implementation for dispatched kernels.

Currently only ``"cuda"`` is supported; ``None`` restores automatic selection.
Explicit CUDA selection raises if a kernel has no CUDA implementation, even
with ``missing_cuda="fallback"``. This does not switch the array backend or
affect host calls or direct CudaKernel calls. The import-time environment
variable ``CUNUMPY_DEVICE_KERNEL_IMPLEMENTATION`` initializes this setting.
"""
global _DEVICE_KERNEL_IMPLEMENTATION
_DEVICE_KERNEL_IMPLEMENTATION = _check_device_implementation(name)


def get_device_kernel_implementation() -> str | None:
"""Return the requested device implementation, or None for automatic selection."""
return _DEVICE_KERNEL_IMPLEMENTATION


@contextmanager
def use_device_kernel_implementation(name: str | None) -> Iterator[None]:
"""Temporarily choose the device implementation; global, not per thread."""
global _DEVICE_KERNEL_IMPLEMENTATION
previous = _DEVICE_KERNEL_IMPLEMENTATION
set_device_kernel_implementation(name)
try:
yield
finally:
_DEVICE_KERNEL_IMPLEMENTATION = previous


class HostImplementations:
"""The host implementations of one kernel, run by name or by the default rule.

Each implementation is loaded on first use (a pyccel build, an import) and
may be unavailable (no compiler, numba not installed); a failed load is
remembered with its exception. A call runs the implementation chosen with
:func:`set_kernel_implementation`, which must exist and load, or else the
:func:`set_host_kernel_implementation`, which must exist and load, or else the
default: the first available of ``"pyccel"``, ``"numba"`` and ``"numpy"``,
and as a last resort ``"python"``, with a warning (correct, but slow).
Built by :meth:`Kernel.from_folder` from the files of a kernel folder.
Expand Down
4 changes: 1 addition & 3 deletions src/cunumpy/_mpi.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,9 +11,7 @@
import array_api_compat
import array_api_compat.numpy as np

from cunumpy._mpi_serial import (
_LOCAL_RANK_VARIABLES, # noqa: F401 - re-exported
)
from cunumpy._mpi_serial import _LOCAL_RANK_VARIABLES # noqa: F401 - re-exported
from cunumpy._transfers import _ACTIVE as _COUNTERS
from cunumpy._transfers import _describe, _nbytes, _record
from cunumpy.xp import array_backend, cupy_available, to_numpy
Expand Down
Loading
Loading