Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,9 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- Support for Python 3.8 and 3.9 (both end-of-life); `cunumpy` now requires Python 3.10 or newer.

### Changed
- `Kernel(..., dispatch="arrays")` calls the host kernel directly for host arguments, without `PyccelKernel`'s check for device arrays to convert, which converted nothing but ran while CuPy was the active backend (unless the `PyccelKernel` was built with `use_cupy=True`).
- `Kernel.check_signature()` (and `KernelCatalog.check_signatures()`) also compares the parameters of every host implementation (numba, NumPy; nothing is compiled) and of a `CompiledHostKernel`'s fallback with those of the host kernel.
- `KernelCatalog.from_package(..., compile_host=..., host_fallback=...)` builds `HostImplementations` instead of `CompiledHostKernel`s: `host_fallback` becomes the `"numpy"` implementation, and `<name>_numba.py`/`<name>_numpy.py` files in a kernel folder are picked up.
- `CudaKernel` pointer and array view parameters and `CudaStruct` pointer and view fields reject an array that lives on another CUDA device than the current one (`ValueError`; `cupy.RawKernel` would read the foreign address silently). Checked only when CuPy is imported and the array reports a `device`.
- `assert_kernels_agree(..., n_threads=...)` also takes a function of the argument tuple, and `n_threads` may be omitted when the CUDA kernel has `n_threads_from`.
- NumPy integer scalars passed to integer kernel parameters (and struct fields) are checked by value, like Python ints: `np.int64(5)` fits an `int` parameter; before, they raised `TypeError` unless their dtype was exactly the parameter's.
Expand All @@ -26,6 +29,10 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- `KernelCatalog.from_package(..., include_dirs=None)` is now an explicit keyword; by default the source root of the top-level package (the directory containing it) is an include directory of every CUDA kernel, in addition to the kernel's own folder, so kernels can `#include "my_pkg/common.cuh"`.

### Added
- `Kernel.from_folder(package, ...)`: the kernel of one kernel folder, with the options of `KernelCatalog.from_package` (`host_suffix`, `compile_host`, `dispatch`, `include_dirs`, CUDA options...). Every version in the folder is an implementation: `<name><host_suffix>.py` (`"pyccel"`, compiled with `compile_host`, and `"python"`, uncompiled), `<name>_numba.py`, `<name>_numpy.py` and `<name>_cuda.cu`. A folder's own `__init__.py` can declare `kernel = xp.Kernel.from_folder(__name__, ...)`, so code imports the kernel from where it is written; `from_package` now builds each kernel with it. `Kernel.implementations` lists them, `Kernel.selected(device=False)` tells which one a call runs.
- `xp.HostImplementations` and `xp.HOST_IMPLEMENTATIONS`: the host implementations of a kernel, loaded on first use; a call runs the default, the first available of pyccel, numba and NumPy (the uncompiled Python version, with a warning, if none is), or the one chosen with `xp.set_kernel_implementation(name)` / `with xp.use_kernel_implementation(name):` / `CUNUMPY_KERNEL_IMPLEMENTATION=name` (read at import), like `set_backend`/`use_backend`/`ARRAY_BACKEND`; a chosen implementation that a kernel lacks or cannot load raises instead of running another. `xp.get_kernel_implementation()` reads the setting.
- `xp.as_kernel_array(value, like, dtype=None)` and `xp.kernel_output(out, like, dtype=None)`: bring the arguments of a `dispatch="arrays"` kernel to the side of the main array `like` (CuPy or NumPy, C-contiguous, `dtype`), without a copy when they already fit; `kernel_output` yields the buffer the kernel writes and copies it back into `out` if it had to be converted.
- `CompiledHostKernel.fallback` returns the fallback.
- `cunumpy/random.cuh` and `xp.philox_uniform`, `philox_uniform2`, `philox_normal`, `philox_normal2`, `philox4x32_10`: counter-based random numbers (Philox4x32-10, passing the Random123 known-answer tests) as a pure function of `(seed, stream, counter)`, the same in a kernel and on the host (uniform numbers bit for bit, normal numbers up to the last bits of the math functions), so kernels that draw random numbers can be compared with their host versions.
- `CudaKernel(..., n_threads_from="first_array")`: one thread per row of the first array argument when a launch gives neither `n_threads` nor `grid`.
- `CudaKernel` launches with `shared_mem` above 48 KiB set the kernel's `max_dynamic_shared_size_bytes` once, up to the device's opt-in limit, and raise `ValueError` beyond it.
Expand Down
77 changes: 73 additions & 4 deletions docs/source/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -288,6 +288,23 @@ class DeviceParticles(xp.CudaArguments):
super().__init__(self.markers, self.degree, self.markers.shape[0])
```

### `as_kernel_array(value, like, dtype=None)`, `kernel_output(out, like, dtype=None)`

For the arguments of a `Kernel` with `dispatch="arrays"`, whose choice follows
the arrays. `as_kernel_array` returns `value` on the side of `like` (a CuPy
array if `like` is one, a NumPy array otherwise), C-contiguous and with `dtype`
(any if `None`): `value` itself if it already is such an array, else one copy,
moved across if needed (counted by `count_transfers()`). `kernel_output` is a
context manager yielding the buffer for an array the kernel writes: `out`
itself if `as_kernel_array` takes it unchanged, else a converted copy whose
contents are written into `out` (on its own side) when the block ends without
an error.

```python
with xp.kernel_output(result, like=field, dtype=float) as buffer:
gather(xp.as_kernel_array(positions, like=field, dtype=float), field, buffer)
```

## Random numbers and dtype

### `get_rng(seed=None)`
Expand Down Expand Up @@ -1350,11 +1367,12 @@ array, or a device-only argument object: one with `__cuda_args__()` but no
`__host_args__()`, such as a `CudaArguments` or a struct value), else the host
kernel, whatever the backend. Use `"arrays"` in codes that hand host arrays to
kernels while CuPy is active (diagnostics, MPI staging, CPU fallbacks): those
calls then run the host kernel instead of failing in the CUDA argument checks.
calls then run the host kernel instead of failing in the CUDA argument checks,
calling the host function directly (no device-array conversion).
`missing_cuda` applies to device arguments without a CUDA kernel.

`kernel.check_signature()` checks that the host and CUDA kernels take the same
parameters in the same order (the names of the Python host function, or of the
`kernel.check_signature()` checks that the host and CUDA kernels (and the
fallback of a `CompiledHostKernel`) take the same parameters in the same order (the names of the Python host function, or of the
uncompiled Python version of a `CompiledHostKernel`, against the parsed
`__global__` signature) and raises `ValueError` showing both lists otherwise.
It does nothing without a CUDA kernel, with `check_signature=False`, or when
Expand Down Expand Up @@ -1471,6 +1489,56 @@ that have a CUDA kernel, for a parametrised parity test (see "Testing
utilities"). `KernelCatalog(kernels)` and `catalog.register(kernel, name=None)`
build a catalog by hand.

## `Kernel.from_folder`

```python
# my_sim/kernels/push/__init__.py
kernel = xp.Kernel.from_folder(__name__, host_suffix="_pyccel", dispatch="arrays",
compile_host=compile_kernels)
kernel.implementations # ("pyccel", "numpy", "python", "cuda")
kernel.selected() # "pyccel": what a call with host arrays runs now
```

The kernel of one kernel folder `package` (its dotted name, `__name__` in its
`__init__.py`). Each version of `<name>` in the folder is an implementation:
`<name><host_suffix>.py` gives `"pyccel"` (compiled with `compile_host` on
first use, as it is without one) and `"python"` (uncompiled), `<name>_numba.py`
gives `"numba"`, `<name>_numpy.py` gives `"numpy"`, and `<name><cuda_suffix>` the
CUDA kernel. `extra_implementations` adds host implementations that are not
files, as loaders by name (`{"numpy": lambda: push_numpy}`). Takes the options of
`KernelCatalog.from_package` for one kernel: `host_suffix`, `cuda_suffix`,
`test_args_suffix`, `check_name_length`, `missing_cuda`, `host_options`,
`include_dirs`, `dispatch`, `compile_host` and CUDA options such as
`block_size` or `n_threads_from`. Raises `FileNotFoundError` if the folder has
no host kernel module and `ModuleNotFoundError` if `package` is not a package.
`kernel.selected(device=True)` names the implementation for device arguments.

## `HostImplementations`, `set_kernel_implementation`

```python
host = xp.HostImplementations("push", {"pyccel": load_compiled, "numpy": lambda: push_numpy,
"python": lambda: push})
host(*args) # the default implementation
xp.set_kernel_implementation("numpy") # every kernel: like xp.set_backend
with xp.use_kernel_implementation("python"): # like xp.use_backend
host(*args)
```

The host implementations of one kernel (names from `xp.HOST_IMPLEMENTATIONS`:
`"pyccel"`, `"numba"`, `"numpy"`, `"python"`; `"python"` is required), each
given as a loader that returns the function or raises if it is unavailable.
Loaded on first use; `available(name)` loads and reports, `get(name)` returns it
or raises `LookupError` (missing, or failed to load with the error as cause),
`errors` maps names to load errors, `names` lists them, `python` is the
uncompiled function, `build()` loads the default now. A call runs
`selected()`: the implementation set with `set_kernel_implementation(name)` (or
`use_kernel_implementation`, or the environment variable
`CUNUMPY_KERNEL_IMPLEMENTATION` read at import), which raises if the kernel
lacks it or cannot load it, else the default: the first available of pyccel,
numba and NumPy, and else `"python"` with a `RuntimeWarning` (once).
`get_kernel_implementation()` reads the setting; `None` is the default. The
setting is global, not per thread, and applies to host calls only.

## `CompiledHostKernel`

```python
Expand All @@ -1487,7 +1555,8 @@ on-disk cache, or one that imports modules compiled ahead of time with the
the uncompiled Python function with a `RuntimeWarning`. `kernel.compiled`
builds and reports whether that worked (so that callers can choose another
path), `kernel.error` is the exception of a failed build, `kernel.python` the
uncompiled function and `kernel.build()` compiles now.
uncompiled function, `kernel.fallback` the fallback and `kernel.build()`
compiles now.

## Testing utilities

Expand Down
92 changes: 87 additions & 5 deletions docs/source/kernels/dispatch.md
Original file line number Diff line number Diff line change
Expand Up @@ -123,6 +123,69 @@ A catalog is a read-only mapping: `catalog["push"]`, `"push" in catalog`,
`len(catalog)`, iteration over names. `KernelCatalog()` plus
`catalog.register(kernel)` builds one by hand.

### One folder, declared in its own `__init__.py`

`Kernel.from_folder()` builds the kernel of one folder, with the same options
as `from_package` (`from_package` calls it for every folder). The folder's
`__init__.py` can then declare its kernel, and code imports the kernel from
where it is written:

```text
my_sim/kernels/push/
├── __init__.py # kernel = xp.Kernel.from_folder(__name__, ...)
├── push_pyccel.py # "pyccel" (compiled with compile_host) and "python" (as is)
├── push_numba.py # "numba": def push(...) decorated with numba.njit
├── push_numpy.py # "numpy": vectorized
└── push_cuda.cu # the CUDA kernel
```

```python
# my_sim/kernels/push/__init__.py
import cunumpy as xp

kernel = xp.Kernel.from_folder(
__name__,
host_suffix="_pyccel",
compile_host=compile_kernels, # see below
dispatch="arrays",
n_threads_from="first_array",
)
```

```python
# anywhere in the code
from my_sim.kernels.push import kernel as push

push(positions, velocities, dt) # CUDA for CuPy arrays, host kernel otherwise
```

Per-kernel options (`missing_cuda`, launch defaults) then sit next to the
kernel instead of in mappings keyed by name.

### Several host implementations

Every file of the folder is one implementation; only the host side has a
choice. A call with host arrays runs the first available of `"pyccel"`,
`"numba"` and `"numpy"` (an implementation is unavailable if it fails to
compile or import), and the uncompiled `"python"` version, with a warning, if
none is. Device arrays run the CUDA kernel. To choose, use the same pattern as
for the array backend:

```python
xp.set_kernel_implementation("numpy") # like xp.set_backend
with xp.use_kernel_implementation("numba"): # like xp.use_backend
push(positions, velocities, dt)
xp.set_kernel_implementation(None) # back to the default
```

or `CUNUMPY_KERNEL_IMPLEMENTATION=numpy` for a whole run (read at import, like
`ARRAY_BACKEND`). A chosen implementation that a kernel does not have, or
cannot load, raises `LookupError` instead of running another one: a benchmark
of numba never silently measures NumPy. `kernel.implementations` lists the
implementations, `kernel.selected()` names the one a call with host arrays runs
now (`kernel.selected(device=True)` for device arrays), and
`kernel.host_kernel.kernel.errors` holds why an implementation failed to load.

## Compiled Pyccel host kernels

By default the host kernel is the Python function itself, which is fine for
Expand Down Expand Up @@ -150,9 +213,12 @@ catalog = xp.KernelCatalog.from_package(
)
```

Each host kernel is a `cunumpy.CompiledHostKernel`:
`catalog["push"].host_kernel.kernel.compiled` reports whether the compiled
version is available. Note that `epyccel` compiles again on every call; a
`host_fallback` becomes each kernel's `"numpy"` implementation (see "Several
host implementations" above); `catalog["push"].host_kernel.kernel.available("pyccel")`
reports whether the compiled version builds, and `catalog["push"].selected()`
which version runs. To test the path of a machine without Pyccel, run the code
inside `with xp.use_kernel_implementation("numpy"):` (or set
`CUNUMPY_KERNEL_IMPLEMENTATION=numpy` for a whole run). Note that `epyccel` compiles again on every call; a
code that compiles at run time usually keeps the builds in an on-disk cache
keyed on the module source, so that only the first run after an edit compiles.

Expand All @@ -176,7 +242,22 @@ catalog["gather"](host_positions, host_field, host_result) # host kernel, also

The CUDA kernel runs if any top-level argument is on the GPU: a CuPy array, or
a device-only argument object (`CudaArguments`, a struct value). A
`KernelArguments` object has both forms and does not decide on its own.
`KernelArguments` object has both forms and does not decide on its own. Host
arguments go to the host function directly, without conversion.

The arguments of one call must then all be on one side, with the dtype and
layout the kernels take. `xp.as_kernel_array(value, like, dtype)` brings an
input to the side of the main array `like` (no copy if it already fits), and
`xp.kernel_output(out, like, dtype)` gives the buffer for an output, copied
back into `out` after the block if it had to be converted:

```python
from functools import partial

convert = partial(xp.as_kernel_array, like=field, dtype=float)
with xp.kernel_output(result, like=field, dtype=float) as buffer:
catalog["gather"](convert(positions), convert(field), buffer, n_threads=n)
```

## Same parameters on both sides

Expand All @@ -191,7 +272,8 @@ def test_kernel_signatures():
It compares the parameter names of each host function (for a compiled host
kernel, its Python source, or the `__pyccel__/<module>.pyi` stub that pyccel
writes next to a compiled extension module) with the parsed `__global__`
signature, and lists every kernel that differs, e.g.
signature, and the parameters of a compiled host kernel's fallback with the host
function's, and lists every kernel that differs, e.g.
`kernel 'push': the host kernel takes (x, v, dt, n), the CUDA kernel (x, v, n, dt)`.

## Porting status and setup
Expand Down
6 changes: 6 additions & 0 deletions src/cunumpy/LLM_GUIDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,6 +68,9 @@ https://max-models.github.io/cunumpy/ and in `docs/source/` of the repository.
| many kernels in a package, ported incrementally | `xp.KernelCatalog.from_package(__name__, missing_cuda="fallback")` |
| host kernels compiled at first call (your compile function), NumPy fallback | `from_package(..., host_suffix="_pyccel", compile_host=my_compile, host_fallback={...})` -> `xp.CompiledHostKernel` |
| host arrays reach kernels while CuPy is active | `Kernel(..., dispatch="arrays")` / `from_package(..., dispatch="arrays")`: CUDA only for device arguments |
| one kernel folder declares its kernel in its own `__init__.py` | `kernel = xp.Kernel.from_folder(__name__, host_suffix="_pyccel", compile_host=..., dispatch="arrays")`; `<name>_numba.py`, `<name>_numpy.py` in the folder are further host implementations |
| bring a `dispatch="arrays"` kernel's arguments to the side of the main array | `xp.as_kernel_array(a, like=grid, dtype=float)`; outputs: `with xp.kernel_output(out, like=grid, dtype=float) as buf:` |
| choose the host implementation (pyccel/numba/numpy/python) | `xp.set_kernel_implementation("numpy")`, `with xp.use_kernel_implementation(...)`, `CUNUMPY_KERNEL_IMPLEMENTATION=numpy`; default: first available of pyccel, numba, numpy; `kernel.selected()` |
| check host and CUDA kernels take the same parameters | `catalog.check_signatures()` (in a unit test) |
| test a CUDA kernel's arithmetic without a GPU | `cunumpy.testing.emulate_cuda_kernel(kernel, *numpy_args, n_threads=n)` (C++ compiler; shared memory and __syncthreads ok, no warp ops; `shared_mem=` for extern shared) |
| shared-memory budget of a block | `xp.max_shared_memory_per_block()` (48 KiB without a GPU) |
Expand Down Expand Up @@ -244,6 +247,9 @@ k.get_kernel(); k.compile(); k.has_cuda
catalog = xp.KernelCatalog.from_package(__name__, *, host_suffix="_kernels",
cuda_suffix="_cuda.cu", missing_cuda="raise", host_options=None,
include_dirs=None, **cuda_options)
kernel = xp.Kernel.from_folder(__name__, *, host_suffix="_kernels", compile_host=None,
dispatch="backend", extra_implementations=None, **same options as from_package)
kernel.implementations; kernel.selected(device=False)
catalog["push"]; catalog.summary(); catalog.with_cuda; catalog.without_cuda
catalog.compile_all(jobs=1); catalog.parity_cases(); catalog.register(kernel)
```
Expand Down
Loading
Loading