diff --git a/CHANGELOG.md b/CHANGELOG.md index cfd0d1f..4f96452 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -25,6 +25,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - `xp.parse_cuda_signature(source, name)` and `xp.CudaParameter`: Parse the parameters of a `__global__` function. - `xp.Kernel`: A host kernel (`PyccelKernel`) and its CUDA counterpart, calling the one matching the active backend. Without a CUDA kernel on the CuPy backend it raises `NotImplementedError` (`missing_cuda="raise"`, default) or falls back to the host kernel with host copies (`missing_cuda="fallback"`). - `xp.KernelCatalog`: Read-only mapping of `Kernel`s; `KernelCatalog.from_package()` collects them from a package with one folder per kernel (`name/name_kernels.py`, `name/name_cuda.cu`); `without_cuda` lists the kernels still to port. +- `xp.CudaStructArguments`: Base class for argument objects that are passed to CUDA kernels as one C struct (the class form of `CudaStruct`). A subclass sets `struct_name` and `fields`, stores each field as an attribute and calls `pack()`; the `CudaStruct` is built once per class (`cls.struct`), `__cuda_args__()` returns the packed struct, and copies and unpickled objects are packed again from their own arrays. - `xp.CudaStruct` and `xp.CudaStructValue`: C structs passed to CUDA kernels by value. A `CudaStruct` is defined once from `(field, C type)` pairs; it provides the C `declaration`, the NumPy `dtype` with the C memory layout, and packs values (device arrays as addresses, scalars checked and cast) into a `CudaStructValue` that is passed as one kernel argument. `CudaKernel(..., structs=[...])` checks struct parameters and that a struct definition in the source matches. - `CudaKernel(..., template_args=...)`: Instantiate C++ function templates (e.g. `template_args=(np.float64, 3)` for `name`); the template parameters are substituted into the checked signature. - `xp.CudaKernelVariants`: Creates and caches one `CudaKernel` per variant key for generated kernel sources (e.g. per dimension and dtype); `compile_all()` compiles given and existing variants. diff --git a/docs/source/api.md b/docs/source/api.md index 2305a91..4ab5e54 100644 --- a/docs/source/api.md +++ b/docs/source/api.md @@ -1034,6 +1034,43 @@ def test_pusher_args_header_is_up_to_date(): Kernels created with `structs=[MarkerArgs, ...]` also check a definition in their own source against the Python definition (`check_source`). +## `CudaStructArguments` + +```python +class MarkerArguments(xp.CudaStructArguments): + struct_name = "MarkerArgs" + fields = (("markers", "Array2D"), ("valid", "bool*"), ("n_markers", "int")) + + def __init__(self, markers, valid): + self.markers = markers + self.valid = valid + self.n_markers = markers.shape[0] + self.pack() + +push = xp.CudaKernel(source, "push", structs=[MarkerArguments.struct]) +push(MarkerArguments(markers, valid), dt, n_threads=markers.shape[0]) +``` + +Base class for argument objects that are one C struct: the class form of +`CudaStruct`. A subclass sets `struct_name` and `fields` (`(field name, C type)` +pairs, as for `CudaStruct`), stores every field as an attribute of the same +name, and calls `pack()` at the end of its constructor. The `CudaStruct` is +built once per subclass when the class is defined and is the class attribute +`struct` (for `structs=[...]`, `declaration`, `to_header()`, +`write_cuda_header()`); invalid field types raise when the class is defined. + +* `pack()` packs the field attributes, with the checks of `CudaStruct` + (C-contiguous CuPy arrays of the declared dtype, range-checked scalars). A + field without an attribute raises `AttributeError`. Call it again after + replacing an array attribute: the packed struct holds device addresses. +* `packed` is the packed struct (`numpy.void`); `__cuda_args__()` returns + `(packed,)`, so a `CudaKernel` receives the struct. +* Copies (`copy.copy`, `copy.deepcopy`) and unpickled objects are packed again + from their own arrays; the packed struct is not part of the pickled state. +* A subclass that sets neither `struct_name` nor `fields` is an intermediate + base class (its instances cannot be packed); setting only one raises + `TypeError`. Subclasses of a complete class inherit its struct. + ## `CudaArguments` ```python diff --git a/docs/source/kernels/arguments.md b/docs/source/kernels/arguments.md index 6dce885..973ecdf 100644 --- a/docs/source/kernels/arguments.md +++ b/docs/source/kernels/arguments.md @@ -11,6 +11,7 @@ argument: | `CudaArguments` | (not used) | several parameters, flattened | grouping device arguments only | | `KernelArguments` | one object (`__host_args__()`) | several parameters (`__cuda_args__()`) | the host kernel takes an argument class, the CUDA kernel flat parameters | | `CudaStruct` | (not used) | one C struct parameter | many fields; one definition shared by all CUDA kernels | +| `CudaStructArguments` | (not used) | one C struct parameter | the struct as a class: an object with attributes, built once and reused | They combine: a `KernelArguments` object can return a `CudaStruct` value from `__cuda_args__()`. @@ -134,6 +135,49 @@ field back; `Particles.dtype` is the NumPy structured dtype of the layout. the kernel source matches the Python definition, so a hand-edited header that drifted raises a `ValueError` instead of reading fields at wrong offsets. +### `CudaStructArguments`: the struct as a class + +When the device arguments are an object of their own (built once per particle +species, domain or grid, and kept next to the host argument object), subclass +`CudaStructArguments`. The class declares the struct, the instance holds the +field values as attributes and is passed to kernels as it is: + +```python +class CudaMarkerArguments(xp.CudaStructArguments): + struct_name = "MarkerArgs" + fields = ( + ("markers", "double*"), + ("valid_mks", "bool*"), + ("n_markers", "int"), + ("n_cols", "int"), + ("weight_idx", "int"), + ) + + def __init__(self, markers, valid_mks, weight_idx): + self.markers = xp.as_device_array(markers, np.float64, ndim=2, name="markers") + self.valid_mks = xp.as_device_array(valid_mks, np.bool_, ndim=1, name="valid_mks") + self.n_markers, self.n_cols = self.markers.shape + self.weight_idx = weight_idx + self.pack() + + +xp.write_cuda_header("kernels/marker_args.cuh", [CudaMarkerArguments.struct]) +push = xp.CudaKernel.from_file("kernels/push_cuda.cu", structs=[CudaMarkerArguments.struct]) + +args = CudaMarkerArguments(markers, valid_mks, weight_idx=6) +push(args, dt, n_threads=args.n_markers) +``` + +* The `CudaStruct` is built when the class is defined (`CudaMarkerArguments.struct`), + so a bad field type raises at import, not at the first launch. +* `pack()` checks every field like `CudaStruct` does. Call it again after + replacing an array attribute, e.g. after the marker array was resized. +* Copies and unpickled objects are packed again from their own arrays, so a + `deepcopy` never points at the device memory of the original. +* When the same call site must also reach a host kernel, pair it with the host + argument object in a `KernelArguments`: `__host_args__()` returns the host + object, `__cuda_args__()` returns `cuda_args.__cuda_args__()`. + ### Generate the struct from the host argument class If the host kernels already use an annotated argument class (Pyccel style), the diff --git a/src/cunumpy/LLM_GUIDE.md b/src/cunumpy/LLM_GUIDE.md index f1fe3ba..95b50f2 100644 --- a/src/cunumpy/LLM_GUIDE.md +++ b/src/cunumpy/LLM_GUIDE.md @@ -66,7 +66,7 @@ https://max-models.github.io/cunumpy/ and in `docs/source/` of the repository. | launch a hand-written CUDA C kernel | `xp.CudaKernel(source, "name")` / `CudaKernel.from_file(path)` | | host kernel + CUDA port, chosen by backend | `xp.Kernel(host_fn, cuda_kernel_or_None)` | | many kernels in a package, ported incrementally | `xp.KernelCatalog.from_package(__name__, missing_cuda="fallback")` | -| group arrays/scalars into one kernel argument | `xp.CudaArguments` (device only), `xp.KernelArguments` (host object + device tuple), `xp.CudaStruct` (C struct) | +| group arrays/scalars into one kernel argument | `xp.CudaArguments` (device only), `xp.KernelArguments` (host object + device tuple), `xp.CudaStruct` (C struct), `xp.CudaStructArguments` (C struct as a class) | | CUDA struct from a Pyccel argument class | `xp.CudaStruct.from_signature(Cls.__init__, "Name")`, `xp.write_cuda_header(...)` | | kernel writes into a host buffer owned by another library | `xp.DeviceMirror(host_array)` + `` | | N-D indexing in CUDA, non-contiguous arrays | `Array1D`..`Array3D` params from `` | @@ -242,6 +242,13 @@ S = xp.CudaStruct("S", [("x", "double*"), ("n", "long long"), ("a", "Array2D tuple[np.void]: return (self._packed,) +class CudaStructArguments(CudaArguments): + """Base class for argument objects that are passed to CUDA kernels as one C struct. + + The class form of :class:`CudaStruct`: a subclass names the struct + (:attr:`struct_name`) and lists its fields (:attr:`fields`), stores every + field as an attribute of the same name, and calls :meth:`pack` at the end + of its constructor. The object is then passed to a :class:`CudaKernel` as + it is and arrives as the packed struct. + + The :class:`CudaStruct` is built once per subclass, when the class is + defined, and is available as :attr:`struct` (for ``CudaKernel(..., + structs=[...])``, :attr:`CudaStruct.declaration` and + :meth:`CudaStruct.to_header`). Packing checks every field like a kernel + argument: pointers need C-contiguous CuPy arrays of the declared dtype + (never copied), scalars are range-checked and cast. + + The packed struct holds device addresses. Call :meth:`pack` again after + replacing an array attribute. Copies (``copy.copy``, ``copy.deepcopy``) + and unpickled objects are packed again from their own arrays. + + A subclass that sets neither :attr:`struct_name` nor :attr:`fields` is an + intermediate base class; a subclass that sets only one of them raises + ``TypeError``. + + Attributes + ---------- + struct_name : str + Name of the C struct type. + fields : Sequence[tuple[str, str]] + ``(field name, C type)`` pairs, in declaration order, as for + :class:`CudaStruct`. + struct : CudaStruct + The struct type, built from :attr:`struct_name` and :attr:`fields`. + + Examples + -------- + >>> class MarkerArguments(CudaStructArguments): + ... struct_name = "MarkerArgs" + ... fields = ( + ... ("markers", "Array2D"), + ... ("valid", "bool*"), + ... ("n_markers", "int"), + ... ) + ... + ... def __init__(self, markers, valid): + ... self.markers = markers + ... self.valid = valid + ... self.n_markers = markers.shape[0] + ... self.pack() + >>> print(MarkerArguments.struct.declaration) + struct MarkerArgs { + Array2D markers; + bool* valid; + int n_markers; + }; + >>> push = CudaKernel(source, "push", structs=[MarkerArguments.struct]) # doctest: +SKIP + >>> push(MarkerArguments(markers, valid), 0.1, n_threads=markers.shape[0]) # doctest: +SKIP + """ + + struct_name: str + fields: Sequence[tuple[str, str]] + struct: CudaStruct + + def __init_subclass__(cls, **kwargs: Any) -> None: + super().__init_subclass__(**kwargs) + has_name = "struct_name" in cls.__dict__ + has_fields = "fields" in cls.__dict__ + if not has_name and not has_fields: + return # an intermediate base class, or a subclass of a complete one + if not (has_name and has_fields): + raise TypeError( + f"{cls.__qualname__} must define both struct_name and fields" + ) + cls.struct = CudaStruct(cls.struct_name, cls.fields) + + def pack(self) -> None: + """Pack the field attributes into the struct. + + Called at the end of the constructor, and again after an array + attribute has been replaced. + + Raises + ------ + TypeError + If the class does not define a struct, or a field value has the + wrong type (see :meth:`CudaStruct.__call__`). + AttributeError + If a field has no attribute of the same name. + """ + struct = getattr(type(self), "struct", None) + if struct is None: + raise TypeError( + f"{type(self).__qualname__} does not define struct_name and fields" + ) + values = {} + for field in struct.fields: + try: + values[field.name] = getattr(self, field.name) + except AttributeError: + raise AttributeError( + f"{type(self).__qualname__} has no attribute {field.name!r} " + f"for the field of struct {struct.name}" + ) from None + self._struct_value = struct(**values) + + @property + def packed(self) -> np.void: + """The packed struct, with the memory layout of the C struct; packs on first use.""" + if self.__dict__.get("_struct_value") is None: + self.pack() + return self._struct_value.packed + + def __cuda_args__(self) -> tuple[np.void]: + """The packed struct, as the one kernel argument this object stands for.""" + return (self.packed,) + + def __getstate__(self) -> dict[str, Any]: + # the packed struct holds the device addresses of the original arrays, + # which a copy (or another process) does not share + state = self.__dict__.copy() + state.pop("_struct_value", None) + return state + + def __setstate__(self, state: dict[str, Any]) -> None: + self.__dict__.update(state) + self.pack() + + def _as_shape(value: int | Sequence[int], what: str) -> tuple[int, ...]: shape = (value,) if isinstance(value, (int, np.integer)) else tuple(value) if not 1 <= len(shape) <= 3: diff --git a/tests/unit/test_cuda_kernel.py b/tests/unit/test_cuda_kernel.py index 4876ae2..592e13b 100644 --- a/tests/unit/test_cuda_kernel.py +++ b/tests/unit/test_cuda_kernel.py @@ -1601,3 +1601,174 @@ def test_as_device_array_checks_ndim(): as_device_array(x, np.float64, ndim=1, name="x") with pytest.raises(ValueError, match="device argument must have 2"): as_device_array((1, 2, 3), np.int32, ndim=2) + + +# --------------------------------------------------------------------------- +# struct argument classes +# --------------------------------------------------------------------------- + + +class ParticleArguments(xp.CudaStructArguments): + """The class form of PARTICLES.""" + + struct_name = "Particles" + fields = tuple( + (f.name, f"{f.ctype}{'*' if f.pointer else ''}") for f in PARTICLES.fields + ) + + def __init__(self, x, alive, ids, charge=2.0, weight=0.5): + self.x = x + self.n = x.shape[0] + self.charge = charge + self.alive = alive + self.ids = ids + self.weight = weight + self.pack() + + +def _particle_arguments(ptr=0xABC0, n=3, **kwargs): + return ParticleArguments( + FakeDeviceArray(np.float64, ptr=ptr, shape=(n,)), + FakeDeviceArray(np.bool_, shape=(n,)), + FakeDeviceArray(np.int64, shape=(2,)), + **kwargs, + ) + + +def test_struct_arguments_define_the_struct_once(): + struct = ParticleArguments.struct + assert isinstance(struct, CudaStruct) and struct.name == "Particles" + assert struct.declaration == PARTICLES.declaration + assert struct.dtype == PARTICLES.dtype + assert ParticleArguments.struct is _particle_arguments().struct # one per class + + +def test_struct_arguments_pack_their_attributes(): + args = _particle_arguments(ptr=0xF00, charge=3) + assert isinstance(args, CudaArguments) + assert args.packed["x"] == 0xF00 and args.packed["n"] == 3 + assert args.packed["charge"] == 3.0 and args.packed["weight"] == 0.5 + assert args.__cuda_args__() == (args.packed,) + + +def test_struct_arguments_check_their_fields(): + with pytest.raises(TypeError, match="must have dtype float64"): + ParticleArguments( + FakeDeviceArray(np.float32), + FakeDeviceArray(np.bool_), + FakeDeviceArray(np.int64), + ) + with pytest.raises(TypeError, match="must be a CuPy array"): + ParticleArguments( + np.zeros(3), FakeDeviceArray(np.bool_), FakeDeviceArray(np.int64) + ) + args = _particle_arguments() + args.n = 2.5 # a float for the int field n + with pytest.raises(TypeError, match="int n"): + args.pack() + with pytest.raises(OverflowError): + _particle_arguments(n=2**31) + + +def test_struct_arguments_need_every_field_attribute(): + class Incomplete(xp.CudaStructArguments): + struct_name = "Incomplete" + fields = (("x", "double*"), ("n", "int")) + + def __init__(self, x): + self.x = x + self.pack() + + with pytest.raises( + AttributeError, match="no attribute 'n' for the field of struct Incomplete" + ): + Incomplete(FakeDeviceArray(np.float64)) + + +def test_struct_arguments_class_definition(): + with pytest.raises(TypeError, match="must define both struct_name and fields"): + + class OnlyName(xp.CudaStructArguments): + struct_name = "OnlyName" + + with pytest.raises(ValueError, match="unsupported type"): + + class BadField(xp.CudaStructArguments): + struct_name = "BadField" + fields = (("a", "Other"),) + + class Base(xp.CudaStructArguments): # intermediate base: no struct + def __init__(self): + self.pack() + + with pytest.raises(TypeError, match="does not define struct_name and fields"): + Base() + + class Derived(ParticleArguments): # inherits the struct of its parent + pass + + assert Derived.struct is ParticleArguments.struct + + +def test_struct_arguments_repack_after_replacing_an_array(): + args = _particle_arguments(ptr=0x100) + args.x = FakeDeviceArray(np.float64, ptr=0x200, shape=(3,)) + assert args.packed["x"] == 0x100 # still the old address + args.pack() + assert args.packed["x"] == 0x200 + + +def test_struct_arguments_are_packed_again_when_copied(): + import copy + import pickle + + args = _particle_arguments(ptr=0x100) + shallow = copy.copy(args) + assert shallow.packed["x"] == 0x100 and shallow.packed is not args.packed + + deep = copy.deepcopy(args) + deep_x = deep.x + assert deep_x is not args.x and deep.packed["x"] == deep_x.data.ptr + + restored = pickle.loads(pickle.dumps(args)) + assert "_struct_value" in vars(restored) + assert restored.packed["charge"] == 2.0 + assert "_struct_value" not in args.__getstate__() + + +def test_struct_arguments_as_kernel_arguments(): + kernel = CudaKernel(PUSH_SOURCE, "push", structs=[ParticleArguments.struct]) + args = _particle_arguments() + out, size = FakeDeviceArray(np.float64), FakeDeviceArray(np.uint64) + packed, dt, _, _ = kernel.prepare_args(args, 1, out, size) + assert packed is args.packed and type(dt) is np.float64 + + class Other(xp.CudaStructArguments): + struct_name = "Other" + fields = (("x", "double*"),) + + def __init__(self, x): + self.x = x + self.pack() + + with pytest.raises(TypeError, match="must be a value of struct Particles"): + kernel.prepare_args(Other(FakeDeviceArray(np.float64)), 1.0, out, size) + + +def test_struct_arguments_on_gpu(): + _skip_without_cupy() + import cupy as cp + + n = 300 + alive = cp.ones(n, dtype=bool) + alive[::2] = False + args = ParticleArguments( + cp.zeros(n), alive, cp.array([7, 42], dtype=cp.int64), weight=1.5 + ) + out, size = cp.zeros(4), cp.zeros(1, dtype=cp.uint64) + CudaKernel(PUSH_SOURCE, "push", structs=[ParticleArguments.struct])( + args, 0.5, out, size, n_threads=n + ) + assert int(size.get()[0]) == ParticleArguments.struct.dtype.itemsize + assert out.get().tolist() == [n, 2.0, 42.0, 1.5] + assert cp.all(args.x[1::2] == 1.0) and cp.all(args.x[::2] == 0.0)