Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- `xp.parse_cuda_signature(source, name)` and `xp.CudaParameter`: Parse the parameters of a `__global__` function.
- `xp.Kernel`: A host kernel (`PyccelKernel`) and its CUDA counterpart, calling the one matching the active backend. Without a CUDA kernel on the CuPy backend it raises `NotImplementedError` (`missing_cuda="raise"`, default) or falls back to the host kernel with host copies (`missing_cuda="fallback"`).
- `xp.KernelCatalog`: Read-only mapping of `Kernel`s; `KernelCatalog.from_package()` collects them from a package with one folder per kernel (`name/name_kernels.py`, `name/name_cuda.cu`); `without_cuda` lists the kernels still to port.
- `xp.CudaStructArguments`: Base class for argument objects that are passed to CUDA kernels as one C struct (the class form of `CudaStruct`). A subclass sets `struct_name` and `fields`, stores each field as an attribute and calls `pack()`; the `CudaStruct` is built once per class (`cls.struct`), `__cuda_args__()` returns the packed struct, and copies and unpickled objects are packed again from their own arrays.
- `xp.CudaStruct` and `xp.CudaStructValue`: C structs passed to CUDA kernels by value. A `CudaStruct` is defined once from `(field, C type)` pairs; it provides the C `declaration`, the NumPy `dtype` with the C memory layout, and packs values (device arrays as addresses, scalars checked and cast) into a `CudaStructValue` that is passed as one kernel argument. `CudaKernel(..., structs=[...])` checks struct parameters and that a struct definition in the source matches.
- `CudaKernel(..., template_args=...)`: Instantiate C++ function templates (e.g. `template_args=(np.float64, 3)` for `name<double, 3>`); the template parameters are substituted into the checked signature.
- `xp.CudaKernelVariants`: Creates and caches one `CudaKernel` per variant key for generated kernel sources (e.g. per dimension and dtype); `compile_all()` compiles given and existing variants.
Expand Down
37 changes: 37 additions & 0 deletions docs/source/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -1034,6 +1034,43 @@ def test_pusher_args_header_is_up_to_date():
Kernels created with `structs=[MarkerArgs, ...]` also check a definition in
their own source against the Python definition (`check_source`).

## `CudaStructArguments`

```python
class MarkerArguments(xp.CudaStructArguments):
struct_name = "MarkerArgs"
fields = (("markers", "Array2D<double>"), ("valid", "bool*"), ("n_markers", "int"))

def __init__(self, markers, valid):
self.markers = markers
self.valid = valid
self.n_markers = markers.shape[0]
self.pack()

push = xp.CudaKernel(source, "push", structs=[MarkerArguments.struct])
push(MarkerArguments(markers, valid), dt, n_threads=markers.shape[0])
```

Base class for argument objects that are one C struct: the class form of
`CudaStruct`. A subclass sets `struct_name` and `fields` (`(field name, C type)`
pairs, as for `CudaStruct`), stores every field as an attribute of the same
name, and calls `pack()` at the end of its constructor. The `CudaStruct` is
built once per subclass when the class is defined and is the class attribute
`struct` (for `structs=[...]`, `declaration`, `to_header()`,
`write_cuda_header()`); invalid field types raise when the class is defined.

* `pack()` packs the field attributes, with the checks of `CudaStruct`
(C-contiguous CuPy arrays of the declared dtype, range-checked scalars). A
field without an attribute raises `AttributeError`. Call it again after
replacing an array attribute: the packed struct holds device addresses.
* `packed` is the packed struct (`numpy.void`); `__cuda_args__()` returns
`(packed,)`, so a `CudaKernel` receives the struct.
* Copies (`copy.copy`, `copy.deepcopy`) and unpickled objects are packed again
from their own arrays; the packed struct is not part of the pickled state.
* A subclass that sets neither `struct_name` nor `fields` is an intermediate
base class (its instances cannot be packed); setting only one raises
`TypeError`. Subclasses of a complete class inherit its struct.

## `CudaArguments`

```python
Expand Down
44 changes: 44 additions & 0 deletions docs/source/kernels/arguments.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@ argument:
| `CudaArguments` | (not used) | several parameters, flattened | grouping device arguments only |
| `KernelArguments` | one object (`__host_args__()`) | several parameters (`__cuda_args__()`) | the host kernel takes an argument class, the CUDA kernel flat parameters |
| `CudaStruct` | (not used) | one C struct parameter | many fields; one definition shared by all CUDA kernels |
| `CudaStructArguments` | (not used) | one C struct parameter | the struct as a class: an object with attributes, built once and reused |

They combine: a `KernelArguments` object can return a `CudaStruct` value from
`__cuda_args__()`.
Expand Down Expand Up @@ -134,6 +135,49 @@ field back; `Particles.dtype` is the NumPy structured dtype of the layout.
the kernel source matches the Python definition, so a hand-edited header that
drifted raises a `ValueError` instead of reading fields at wrong offsets.

### `CudaStructArguments`: the struct as a class

When the device arguments are an object of their own (built once per particle
species, domain or grid, and kept next to the host argument object), subclass
`CudaStructArguments`. The class declares the struct, the instance holds the
field values as attributes and is passed to kernels as it is:

```python
class CudaMarkerArguments(xp.CudaStructArguments):
struct_name = "MarkerArgs"
fields = (
("markers", "double*"),
("valid_mks", "bool*"),
("n_markers", "int"),
("n_cols", "int"),
("weight_idx", "int"),
)

def __init__(self, markers, valid_mks, weight_idx):
self.markers = xp.as_device_array(markers, np.float64, ndim=2, name="markers")
self.valid_mks = xp.as_device_array(valid_mks, np.bool_, ndim=1, name="valid_mks")
self.n_markers, self.n_cols = self.markers.shape
self.weight_idx = weight_idx
self.pack()


xp.write_cuda_header("kernels/marker_args.cuh", [CudaMarkerArguments.struct])
push = xp.CudaKernel.from_file("kernels/push_cuda.cu", structs=[CudaMarkerArguments.struct])

args = CudaMarkerArguments(markers, valid_mks, weight_idx=6)
push(args, dt, n_threads=args.n_markers)
```

* The `CudaStruct` is built when the class is defined (`CudaMarkerArguments.struct`),
so a bad field type raises at import, not at the first launch.
* `pack()` checks every field like `CudaStruct` does. Call it again after
replacing an array attribute, e.g. after the marker array was resized.
* Copies and unpickled objects are packed again from their own arrays, so a
`deepcopy` never points at the device memory of the original.
* When the same call site must also reach a host kernel, pair it with the host
argument object in a `KernelArguments`: `__host_args__()` returns the host
object, `__cuda_args__()` returns `cuda_args.__cuda_args__()`.

### Generate the struct from the host argument class

If the host kernels already use an annotated argument class (Pyccel style), the
Expand Down
9 changes: 8 additions & 1 deletion src/cunumpy/LLM_GUIDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -66,7 +66,7 @@ https://max-models.github.io/cunumpy/ and in `docs/source/` of the repository.
| launch a hand-written CUDA C kernel | `xp.CudaKernel(source, "name")` / `CudaKernel.from_file(path)` |
| host kernel + CUDA port, chosen by backend | `xp.Kernel(host_fn, cuda_kernel_or_None)` |
| many kernels in a package, ported incrementally | `xp.KernelCatalog.from_package(__name__, missing_cuda="fallback")` |
| group arrays/scalars into one kernel argument | `xp.CudaArguments` (device only), `xp.KernelArguments` (host object + device tuple), `xp.CudaStruct` (C struct) |
| group arrays/scalars into one kernel argument | `xp.CudaArguments` (device only), `xp.KernelArguments` (host object + device tuple), `xp.CudaStruct` (C struct), `xp.CudaStructArguments` (C struct as a class) |
| CUDA struct from a Pyccel argument class | `xp.CudaStruct.from_signature(Cls.__init__, "Name")`, `xp.write_cuda_header(...)` |
| kernel writes into a host buffer owned by another library | `xp.DeviceMirror(host_array)` + `<cunumpy/atomic.cuh>` |
| N-D indexing in CUDA, non-contiguous arrays | `Array1D<T>`..`Array3D<T>` params from `<cunumpy/array_view.cuh>` |
Expand Down Expand Up @@ -242,6 +242,13 @@ S = xp.CudaStruct("S", [("x", "double*"), ("n", "long long"), ("a", "Array2D<dou
S.declaration; S.dtype; S.to_header(path); value = S(x=..., n=..., a=...)
S = xp.CudaStruct.from_signature(Cls.__init__, "S", int_type="long long")
xp.write_cuda_header("args.cuh", [S1, S2])

class A(xp.CudaStructArguments): # the struct as a class; A.struct is the CudaStruct
struct_name = "A"
fields = (("x", "double*"), ("n", "int"))
def __init__(self, x):
self.x, self.n = x, x.shape[0]
self.pack() # again after replacing an array; copies repack
xp.CudaKernel(S.declaration + src, "k", structs=[S])
xp.resolve_host_args(args, kwargs)
```
Expand Down
2 changes: 2 additions & 0 deletions src/cunumpy/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,7 @@
CudaKernelVariants,
CudaParameter,
CudaStruct,
CudaStructArguments,
CudaStructValue,
ctype_of,
cuda_include_dir,
Expand Down Expand Up @@ -77,6 +78,7 @@
"CudaKernelVariants",
"CudaParameter",
"CudaStruct",
"CudaStructArguments",
"CudaStructValue",
"DeviceMirror",
"Kernel",
Expand Down
1 change: 1 addition & 0 deletions src/cunumpy/__init__.pyi
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,7 @@ from .cuda_kernel import CudaKernel as CudaKernel
from .cuda_kernel import CudaKernelVariants as CudaKernelVariants
from .cuda_kernel import CudaParameter as CudaParameter
from .cuda_kernel import CudaStruct as CudaStruct
from .cuda_kernel import CudaStructArguments as CudaStructArguments
from .cuda_kernel import CudaStructValue as CudaStructValue
from .cuda_kernel import ctype_of as ctype_of
from .cuda_kernel import cuda_include_dir as cuda_include_dir
Expand Down
129 changes: 129 additions & 0 deletions src/cunumpy/cuda_kernel.py
Original file line number Diff line number Diff line change
Expand Up @@ -67,6 +67,7 @@ class is the one definition of the arguments.
"CudaKernelVariants",
"CudaParameter",
"CudaStruct",
"CudaStructArguments",
"CudaStructValue",
"ctype_of",
"cuda_include_dir",
Expand Down Expand Up @@ -1151,6 +1152,134 @@ def __cuda_args__(self) -> tuple[np.void]:
return (self._packed,)


class CudaStructArguments(CudaArguments):
"""Base class for argument objects that are passed to CUDA kernels as one C struct.

The class form of :class:`CudaStruct`: a subclass names the struct
(:attr:`struct_name`) and lists its fields (:attr:`fields`), stores every
field as an attribute of the same name, and calls :meth:`pack` at the end
of its constructor. The object is then passed to a :class:`CudaKernel` as
it is and arrives as the packed struct.

The :class:`CudaStruct` is built once per subclass, when the class is
defined, and is available as :attr:`struct` (for ``CudaKernel(...,
structs=[...])``, :attr:`CudaStruct.declaration` and
:meth:`CudaStruct.to_header`). Packing checks every field like a kernel
argument: pointers need C-contiguous CuPy arrays of the declared dtype
(never copied), scalars are range-checked and cast.

The packed struct holds device addresses. Call :meth:`pack` again after
replacing an array attribute. Copies (``copy.copy``, ``copy.deepcopy``)
and unpickled objects are packed again from their own arrays.

A subclass that sets neither :attr:`struct_name` nor :attr:`fields` is an
intermediate base class; a subclass that sets only one of them raises
``TypeError``.

Attributes
----------
struct_name : str
Name of the C struct type.
fields : Sequence[tuple[str, str]]
``(field name, C type)`` pairs, in declaration order, as for
:class:`CudaStruct`.
struct : CudaStruct
The struct type, built from :attr:`struct_name` and :attr:`fields`.

Examples
--------
>>> class MarkerArguments(CudaStructArguments):
... struct_name = "MarkerArgs"
... fields = (
... ("markers", "Array2D<double>"),
... ("valid", "bool*"),
... ("n_markers", "int"),
... )
...
... def __init__(self, markers, valid):
... self.markers = markers
... self.valid = valid
... self.n_markers = markers.shape[0]
... self.pack()
>>> print(MarkerArguments.struct.declaration)
struct MarkerArgs {
Array2D<double> markers;
bool* valid;
int n_markers;
};
>>> push = CudaKernel(source, "push", structs=[MarkerArguments.struct]) # doctest: +SKIP
>>> push(MarkerArguments(markers, valid), 0.1, n_threads=markers.shape[0]) # doctest: +SKIP
"""

struct_name: str
fields: Sequence[tuple[str, str]]
struct: CudaStruct

def __init_subclass__(cls, **kwargs: Any) -> None:
super().__init_subclass__(**kwargs)
has_name = "struct_name" in cls.__dict__
has_fields = "fields" in cls.__dict__
if not has_name and not has_fields:
return # an intermediate base class, or a subclass of a complete one
if not (has_name and has_fields):
raise TypeError(
f"{cls.__qualname__} must define both struct_name and fields"
)
cls.struct = CudaStruct(cls.struct_name, cls.fields)

def pack(self) -> None:
"""Pack the field attributes into the struct.

Called at the end of the constructor, and again after an array
attribute has been replaced.

Raises
------
TypeError
If the class does not define a struct, or a field value has the
wrong type (see :meth:`CudaStruct.__call__`).
AttributeError
If a field has no attribute of the same name.
"""
struct = getattr(type(self), "struct", None)
if struct is None:
raise TypeError(
f"{type(self).__qualname__} does not define struct_name and fields"
)
values = {}
for field in struct.fields:
try:
values[field.name] = getattr(self, field.name)
except AttributeError:
raise AttributeError(
f"{type(self).__qualname__} has no attribute {field.name!r} "
f"for the field of struct {struct.name}"
) from None
self._struct_value = struct(**values)

@property
def packed(self) -> np.void:
"""The packed struct, with the memory layout of the C struct; packs on first use."""
if self.__dict__.get("_struct_value") is None:
self.pack()
return self._struct_value.packed

def __cuda_args__(self) -> tuple[np.void]:
"""The packed struct, as the one kernel argument this object stands for."""
return (self.packed,)

def __getstate__(self) -> dict[str, Any]:
# the packed struct holds the device addresses of the original arrays,
# which a copy (or another process) does not share
state = self.__dict__.copy()
state.pop("_struct_value", None)
return state

def __setstate__(self, state: dict[str, Any]) -> None:
self.__dict__.update(state)
self.pack()


def _as_shape(value: int | Sequence[int], what: str) -> tuple[int, ...]:
shape = (value,) if isinstance(value, (int, np.integer)) else tuple(value)
if not 1 <= len(shape) <= 3:
Expand Down
Loading
Loading