diff --git a/CHANGELOG.md b/CHANGELOG.md index 4f96452..f41ad41 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -7,17 +7,50 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [Unreleased] +### Fixed +- `CudaKernel`'s header hash (`-DCUNUMPY_INCLUDE_HASH`) now covers the headers shipped with cunumpy (`cunumpy/atomic.cuh`, `reduce.cuh`, ...), also when included in angle brackets. Before, an upgrade of cunumpy that changed one of them left CuPy's kernel cache serving the kernel compiled with the old header. `resolve_includes(..., angle_dirs=...)` tracks angle-bracket includes found in the given directories. +- `xp.testing.assert_kernels_agree` reads `CudaStructArguments` objects and struct values through their struct fields, so their arrays get the same names as the attributes of the host argument object (before, arrays behind properties were named after the private attribute holding the owner, and the comparison failed with "do not have the same array arguments"). + ### Removed - Support for Python 3.8 and 3.9 (both end-of-life); `cunumpy` now requires Python 3.10 or newer. ### Changed +- `CudaKernel` pointer and array view parameters and `CudaStruct` pointer and view fields reject an array that lives on another CUDA device than the current one (`ValueError`; `cupy.RawKernel` would read the foreign address silently). Checked only when CuPy is imported and the array reports a `device`. +- `assert_kernels_agree(..., n_threads=...)` also takes a function of the argument tuple, and `n_threads` may be omitted when the CUDA kernel has `n_threads_from`. +- NumPy integer scalars passed to integer kernel parameters (and struct fields) are checked by value, like Python ints: `np.int64(5)` fits an `int` parameter; before, they raised `TypeError` unless their dtype was exactly the parameter's. - Python 3.14 is supported. +- The `test` extra installs SciPy, so the `xp.scipy` tests run in CI instead of being skipped. - `KernelCatalog.compile_all(jobs=1)` and `CudaKernelVariants.compile_all(keys, jobs=1)`: With `jobs > 1` the CUDA kernels are compiled in threads (NVRTC releases the GIL); `jobs=None` uses the number of CPUs. All kernels are compiled even if one fails, and the first error is raised afterwards. - CI now tests every supported Python version (3.10, 3.11, 3.12, 3.13 and 3.14) instead of 3.8/3.10/3.13. - `CudaKernel.compile()` passes `compile_options()` to CuPy: the given `options` plus `-DCUNUMPY_INCLUDE_HASH=0x` when the source includes header files, so CuPy's kernel cache (keyed on source and options only) is invalidated when an included header changes. `options` still returns the options as given. - `KernelCatalog.from_package(..., include_dirs=None)` is now an explicit keyword; by default the source root of the top-level package (the directory containing it) is an include directory of every CUDA kernel, in addition to the kernel's own folder, so kernels can `#include "my_pkg/common.cuh"`. ### Added +- `cunumpy/random.cuh` and `xp.philox_uniform`, `philox_uniform2`, `philox_normal`, `philox_normal2`, `philox4x32_10`: counter-based random numbers (Philox4x32-10, passing the Random123 known-answer tests) as a pure function of `(seed, stream, counter)`, the same in a kernel and on the host (uniform numbers bit for bit, normal numbers up to the last bits of the math functions), so kernels that draw random numbers can be compared with their host versions. +- `CudaKernel(..., n_threads_from="first_array")`: one thread per row of the first array argument when a launch gives neither `n_threads` nor `grid`. +- `CudaKernel` launches with `shared_mem` above 48 KiB set the kernel's `max_dynamic_shared_size_bytes` once, up to the device's opt-in limit, and raise `ValueError` beyond it. +- `emulate_cuda_kernel` emulates block shared memory (`__shared__`, `extern __shared__` with `shared_mem=`) and `__syncthreads` (the threads of a block run as coroutines and meet at every barrier), so per-block deposits, shared-memory reductions and tiled kernels can be tested without a GPU; warp intrinsics, now also in included headers, are still refused. +- `xp.HostStaging(shape, dtype, buffers=2)` and `StagedCopy`: copy device arrays to page-locked host buffers in the background (snapshot on the device, then an asynchronous copy on a separate stream), for output that overlaps the next time steps; stale results raise. +- Documentation: a "Particle codes" guide with tested recipes (removing and sorting markers, sort-then-reduce deposits, marker exchange between MPI ranks, random numbers per particle, background output, CUDA graphs, block sizes). +- `xp.PyccelStructArguments`: `CudaStructArguments` with a host form. A subclass names its pyccel-compiled (or any) host argument class in `host_class` and the attributes passed to it in `host_fields`; `__host_args__()` builds that object once and again when an attribute was replaced, so one object is passed to a `Kernel` on both backends (the pyccel class itself cannot inherit from anything). On the CuPy backend `__host_args__()` raises unless `host_copies = True`, which builds the host object from host copies for read-only evaluations (results are not copied back). Objects holding host arrays are copied and pickled without packing; `has_device_arrays()` tells whether the struct can be packed. +- `CudaStruct.from_pyccel_class(source, class_name, name=None, *, int_type, scalar_names, exclude, attribute_names)`: build a struct from the annotated `__init__` of a class in a `.py` file (or source string) by parsing it with `ast`, for classes whose module is compiled by pyccel (the compiled class has no Python signature). Fields are named after the attributes the parameters are stored in (`self. = `); `exclude` drops parameters (e.g. scratch arrays). +- `CudaKernel(..., n_threads_from=callable)` (also a settable property): a launch that gives neither `n_threads` nor `grid` takes `n_threads_from(args)`, e.g. `lambda args: args[2].n_markers`; `Kernel.__call__` and `assert_kernels_agree` accept a missing `n_threads` then. +- `CudaKernel(..., check_finite=True)` (also a settable property): after every launch (synchronized) the floating-point and complex arrays among the arguments, including the array fields of struct argument objects, are scanned, and a NaN or inf raises `RuntimeError` naming the kernel and the argument. For debugging a kernel that produces NaN; costs a synchronization and a pass over the arrays per launch. +- `cunumpy.testing.device_function_kernel` accepts struct parameters, by value (`DomainArgs d`) or by const reference (`const DomainArgs& d`), for the structs given in `structs=`; they are passed through to every thread unchanged, so device helpers that take argument structs can be tested from Python. +- Test arguments next to the kernel: `KernelCatalog.from_package(..., test_args_suffix="_test_args")` records `/_test_args.py` as `Kernel.test_args_module` (imported on first access as `Kernel.test_args`; `Kernel(..., test_args=...)` sets it by hand). The module defines `make_args(backend, seed)` and `N_THREADS` (an integer, a tuple, or a function of the argument tuple) or `GRID`, optionally `BLOCK`, `RTOL`, `ATOL`, `N_CALLS`, `OUTPUTS`, `SEED` (`cunumpy.testing.TEST_ARGS_SETTINGS`). `cunumpy.testing.parity_cases(catalog)` gives one `pytest.param` per kernel with a CUDA version, kernels without a test-arguments module marked `skip` with a reason naming the missing file, and `check_parity(kernel, **overrides)` runs `assert_kernels_agree` with the module's settings. A ported package's parity test is then one parametrized test, and a developer who adds a kernel adds a file, not test code. +- `KernelCatalog.from_package(..., check_name_length=True)` warns about a kernel whose module name does not fit pyccel's Fortran wrapper module `bind_c__kernels` into Fortran's 63-character limit (`cunumpy.dispatch.FORTRAN_NAME_LIMIT`): with the default suffix a kernel name has at most 48 characters. +- `Kernel.host_parameters()` reads the parameter names of a pyccel-compiled host kernel from the `__pyccel__/.pyi` stub pyccel writes next to the extension module, so `Kernel.check_signature()` and `KernelCatalog.check_signatures()` also check compiled kernels. +- The fake CuPy (`cunumpy._fake_cupy`): a strict host stand-in for CuPy for CI machines without a GPU, installed with the environment variable `CUNUMPY_FAKE_CUPY=1` (read when cunumpy is imported) or `cunumpy.testing.install_fake_cupy()`. Its arrays live in host memory but are not NumPy arrays (`numpy.asarray` raises, as for real CuPy arrays), reject host arrays and lists as CuPy does, have `data.ptr`, `device` and `__cuda_array_interface__`, so argument objects, struct packing, `as_device_array`, `count_transfers` and backend branches run on the CuPy code path on the CPU; CUDA kernels cannot run (`NotImplementedError`). `cunumpy.testing.fake_cupy_active()` tells; `requires_cupy` and `assert_kernels_agree` skip while it is active. +- `xp.mpi_buffer(array, *, send=True, recv=False, cuda_aware=None)`: context manager yielding the buffer to hand to MPI: a host array unchanged; a device array unchanged (after `synchronize_for_mpi`) when MPI is CUDA-aware; otherwise a pinned host staging copy, copied from the device before the block (`send`) and back after it (`recv`), counted by `count_transfers()`. `cuda_aware=None` uses the answer recorded by `mpi_is_cuda_aware()` (which now remembers its result) or `xp.set_mpi_cuda_aware()`; `xp.get_mpi_cuda_aware()` reads it. One MPI call site for both backends and both kinds of MPI builds. +- `xp.segment_sum(values, keys, n_segments)`: `out[k] = sum(values[keys == k])` on either backend (`bincount` per column), the reduction step of a sort-then-reduce accumulation; negative keys drop the value. +- `xp.require_version("0.4.0")`: raises `ImportError` if the installed cunumpy is older. +- `Kernel(..., dispatch="arrays")` and `KernelCatalog.from_package(..., dispatch="arrays")`: choose the CUDA kernel when an argument lives on the GPU (a CuPy array or a device-only argument object) and the host kernel otherwise, whatever the backend, for codes that hand host arrays to kernels while CuPy is active. The default, `dispatch="backend"`, is unchanged. +- `xp.CompiledHostKernel`: a host kernel compiled on its first call by a compile function the caller provides (cunumpy does not compile anything itself, e.g. a wrapper around `pyccel.epyccel`), falling back to a given callable (or, with a warning, to the uncompiled Python function) when compilation fails. `KernelCatalog.from_package(..., compile_host=..., host_fallback=...)` uses it for every host kernel. +- `Kernel.check_signature()`, `Kernel.host_parameters()` and `KernelCatalog.check_signatures()`: check that a host kernel and its CUDA kernel take the same parameters in the same order, listing every kernel that differs. +- `cunumpy.testing.emulate_cuda_kernel(kernel, *args, n_threads=...)` (in `cunumpy.emulation`) and `emulation_compiler()`: run a CUDA kernel on the CPU, one thread after another, on NumPy arrays (any strides, written back), with the kernel compiled as C++ and the CUDA built-ins stubbed, so CPU-only CI can check the arithmetic of ported kernels. Kernels using shared memory, `__syncthreads` or warp intrinsics are refused. +- `Array4D` in `cunumpy/array_view.cuh` and as a kernel parameter, struct field and `from_signature` annotation (`"float[:, :, :, :]"`), e.g. for a 3D grid of vector components. +- `xp.max_shared_memory_per_block(device=None, *, opt_in=False)` and `xp.DEFAULT_SHARED_MEMORY_PER_BLOCK`: the shared memory a block may use on a device (48 KiB without a GPU). +- `xp.random_streams` (`cunumpy.random_streams.RandomStreams`): one seeded generator per process and backend, with the stream `(seed, rank)` for MPI runs, a choice of NumPy bit generator, per-component generators and draw functions that work with NumPy and CuPy generators. - Documentation restructured into getting started, user guide (backends, backend-agnostic code, data movement, devices, MPI, profiling), kernel porting guides (`PyccelKernel`, `CudaKernel`, `Kernel`/`KernelCatalog`, argument objects, accumulation, debugging, testing), worked examples, best practices and troubleshooting pages. - `cunumpy/LLM_GUIDE.md`: a self-contained guide to the API and its rules for AI coding assistants, shipped as package data and rendered in the documentation. - `xp.CudaKernel`: Wraps a CUDA C kernel (`cupy.RawKernel`, compiled lazily with NVRTC) so it can be called with the same arguments as the host kernel it mirrors, plus `n_threads`. The `extern "C" __global__` signature is parsed once and every call is checked against it: argument count, array dtypes (host arrays raise, they are never copied), and scalars (Python scalars are cast to the declared C types with range checks; lossy or mismatching scalars raise instead of reaching the kernel as silently wrong values). Supports `block_size`, NVRTC `options`, `include_dirs`, `shared_mem`, `stream`, `CudaKernel.from_file()` (`_cuda.cu`), `compile()` and `prepare_args()`; `check_signature=False` skips the checks. @@ -25,7 +58,13 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - `xp.parse_cuda_signature(source, name)` and `xp.CudaParameter`: Parse the parameters of a `__global__` function. - `xp.Kernel`: A host kernel (`PyccelKernel`) and its CUDA counterpart, calling the one matching the active backend. Without a CUDA kernel on the CuPy backend it raises `NotImplementedError` (`missing_cuda="raise"`, default) or falls back to the host kernel with host copies (`missing_cuda="fallback"`). - `xp.KernelCatalog`: Read-only mapping of `Kernel`s; `KernelCatalog.from_package()` collects them from a package with one folder per kernel (`name/name_kernels.py`, `name/name_cuda.cu`); `without_cuda` lists the kernels still to port. -- `xp.CudaStructArguments`: Base class for argument objects that are passed to CUDA kernels as one C struct (the class form of `CudaStruct`). A subclass sets `struct_name` and `fields`, stores each field as an attribute and calls `pack()`; the `CudaStruct` is built once per class (`cls.struct`), `__cuda_args__()` returns the packed struct, and copies and unpickled objects are packed again from their own arrays. +- `xp.CudaStructArguments`: Base class for argument objects that are passed to CUDA kernels as one C struct (the class form of `CudaStruct`). A subclass sets `struct_name` and `fields`, stores each field as an attribute and calls `pack()`; the `CudaStruct` is built once per class (`cls.struct`), and `__cuda_args__()` returns the packed struct. The struct is packed again automatically at the next use when a field attribute changed (another array address, shape or strides, or another scalar value), so fields can be properties that follow an owner's resized arrays; copies and unpickled objects are packed again from their own arrays. +- Debug mode no longer synchronizes after a launch while the stream is being captured into a CUDA graph, which would invalidate the capture. +- `xp.scipy`: SciPy for the active backend, `scipy` on NumPy and `cupyx.scipy` on CuPy, resolved at every access (`fft`, `fftpack`, `interpolate`, `linalg`, `ndimage`, `signal`, `sparse`, `sparse.csgraph`, `sparse.linalg`, `spatial`, `special`, `stats`). A name missing on the active backend raises `AttributeError` naming the backend; `available(name)` checks without raising; `resolve()` returns the module. SciPy stays an optional dependency. +- `xp.fuse`: Decorator that compiles an elementwise function into one kernel with `cupy.fuse` when it is called with CuPy arrays (with the CuPy backend active while tracing) and calls it unchanged otherwise. +- `xp.petsc_vec(array, comm=None)`: A `petsc4py` vector sharing the memory of a NumPy or CuPy array through DLPack (a CUDA or HIP vector for a CuPy array), never copying; raises instead of falling back to a host copy when petsc4py has no GPU support, or when the dtype or layout would need a copy. +- `cunumpy/reduce.cuh`: Warp and block reductions for CUDA kernels (`cunumpy_warp_sum/min/max`, `cunumpy_block_sum/min/max`, and `cunumpy_block_sum_to(out, v)` with one atomic add per block), for in-kernel diagnostics and per-block accumulation. +- `CudaStruct.verify_layout(include=None, *, include_dirs=(), options=())` and `CudaStruct.layout_source()`: Compile and run a one-thread kernel that reports `sizeof`, `alignof` and the field offsets of the struct as the CUDA compiler lays it out, and raise `ValueError` if they differ from the NumPy `dtype` that values are packed into; with `include` the struct is defined by a header instead of `declaration`. - `xp.CudaStruct` and `xp.CudaStructValue`: C structs passed to CUDA kernels by value. A `CudaStruct` is defined once from `(field, C type)` pairs; it provides the C `declaration`, the NumPy `dtype` with the C memory layout, and packs values (device arrays as addresses, scalars checked and cast) into a `CudaStructValue` that is passed as one kernel argument. `CudaKernel(..., structs=[...])` checks struct parameters and that a struct definition in the source matches. - `CudaKernel(..., template_args=...)`: Instantiate C++ function templates (e.g. `template_args=(np.float64, 3)` for `name`); the template parameters are substituted into the checked signature. - `xp.CudaKernelVariants`: Creates and caches one `CudaKernel` per variant key for generated kernel sources (e.g. per dimension and dtype); `compile_all()` compiles given and existing variants. diff --git a/README.md b/README.md index 90a9000..d485f4c 100644 --- a/README.md +++ b/README.md @@ -377,7 +377,7 @@ kernel(particles.kernel_args, dt, n_threads=n) # host or CUDA kernel Kernels ported from pyccel index arrays like `markers[ip, j]`, which needs shapes and strides rather than bare pointers. The shipped header `cunumpy/array_view.cuh` (found by every `CudaKernel`) provides the strided -views `Array1D` to `Array3D`; a parameter or struct field of that type +views `Array1D` to `Array4D`; a parameter or struct field of that type takes a CuPy array, contiguous or not, and indexes `a(i, j)`. The struct can be generated from the annotations of the pyccel argument class, so the Python class is the one definition, and written to a header that a test keeps in sync: diff --git a/docs/source/api.md b/docs/source/api.md index 4ab5e54..06dee24 100644 --- a/docs/source/api.md +++ b/docs/source/api.md @@ -27,6 +27,14 @@ NumPy and CuPy are not interchangeable for every function or object. A function that needs to follow an input array's location should use `get_array_module(array)` instead of assuming the global backend matches it. +## Version + +### `require_version(minimum)` + +Raises `ImportError` if the installed cunumpy is older than `minimum` +(`xp.require_version("0.4.0")`). Only the numeric parts are compared; nothing +is checked when the version is unknown (not installed as a package). + ## Backend selection ### `set_backend(backend)` @@ -172,6 +180,15 @@ assert xp.get_array_backend(normalized) == xp.get_backend() Each conversion returns a suitable array; it does not change the active backend or mutate the source. +### `segment_sum(values, keys, n_segments)` + +`out[k] = sum(values[i] for keys[i] == k)` on the backend of `keys`, with +`bincount` under the hood: the reduction step of a sort-then-reduce +accumulation. `values` has shape `(n,)` or `(n, m)` (columns summed +separately); a negative key drops the value; keys must be smaller than +`n_segments`. The result keeps a floating-point or complex dtype and is +`float64` otherwise. + ## Count transfers A transfer inside a time loop is the classic performance bug of a GPU port: @@ -285,6 +302,33 @@ samples = rng.uniform(size=100) The generator APIs are similar, but seeds do not guarantee identical random sequences across NumPy and CuPy. +### `random_streams` + +```python +xp.random_streams.seed(42, rank=comm.Get_rank(), bit_generator="PCG64") +v = xp.random_streams.normal(0.0, v_th, (n, 3)) +rng = xp.random_streams.generator() # numpy or cupy Generator +own = xp.random_streams.make_generator(seed) # a component's own generator +``` + +One seeded random generator per process and backend, for reproducible MPI +runs: after `seed(value, rank)` every draw of the process comes from the stream +`(value, rank)` (a NumPy `SeedSequence` with the rank as spawn key), so the same +seed and number of ranks give the same results and the ranks' streams are +independent. `seed(None)` (or no seed) seeds from the operating system. +`bit_generator` selects the NumPy bit generator (`MT19937`, `PCG64`, +`PCG64DXSM`, `Philox`, `SFC64`); the CuPy generator uses CuPy's default. +`seed` also seeds NumPy's global state (and CuPy's on the CuPy backend) for code +that calls `np.random.*` directly. + +`generator(backend=None)` returns the process generator of a backend (the +active one by default), created on first use. `make_generator(seed=None, +backend=None)` returns a separate generator for a component with a seed of its +own, and the process generator otherwise. `random`, `standard_normal`, +`normal` and `uniform` draw from the process generator (or `rng=`); `normal` +and `uniform` fall back to `standard_normal` and `random` for CuPy generators +without those methods. `xp.RandomStreams()` makes an independent instance. + ### `default_float_dtype()` Returns the active backend module's `float64` dtype object. Pass it to array @@ -418,6 +462,22 @@ comm.Sendrecv(send_buffer, dest, recvbuf=recv_buffer, source=source) No synchronization is needed after MPI returns: kernels launched afterwards see the received data. +### `mpi_buffer(array, *, send=True, recv=False, cuda_aware=None)` + +Context manager yielding the buffer to pass to MPI for `array`: a host array +unchanged; a device array unchanged (after `synchronize_for_mpi`) when MPI is +CUDA-aware; otherwise a pinned host staging buffer, filled from the device +before the block (`send`) and copied back after it (`recv`), both counted by +`count_transfers()`. `cuda_aware=None` uses the answer recorded by +`mpi_is_cuda_aware()` or `set_mpi_cuda_aware()`; without one, a device array +raises `RuntimeError`. + +### `set_mpi_cuda_aware(value)`, `get_mpi_cuda_aware()` + +Record (or read) whether MPI can take device buffers, for `mpi_buffer()`. +`mpi_is_cuda_aware()` records its own result; set it by hand when the answer +is known otherwise, or `None` to forget it. + ### `memory_info()` Returns `(free_bytes, total_bytes)` reported by the CUDA runtime for the @@ -431,10 +491,19 @@ It is a no-op on NumPy. It does not release blocks still referenced by live arrays. CuPy normally caches freed allocations for reuse, so cached memory does not necessarily indicate a leak. +### `max_shared_memory_per_block(device=None, *, opt_in=False)` + +The bytes of shared memory a block may use on the current (or given) device, +from the device attributes, e.g. to decide whether a per-block copy of a grid +fits. `opt_in=True` gives the larger limit of newer GPUs, which a kernel uses +only after setting `max_dynamic_shared_size_bytes` on its compiled +`cupy.RawKernel`. Without CuPy it returns `DEFAULT_SHARED_MEMORY_PER_BLOCK` +(48 KiB, which every CUDA device provides). + ### `cuda_include_dir()` Returns the directory (as `str`) of the CUDA headers shipped with CuNumpy, -`cunumpy/array_view.cuh`, `cunumpy/atomic.cuh` and `cunumpy/index.cuh`. `CudaKernel` adds it to its NVRTC options as +`cunumpy/array_view.cuh`, `cunumpy/atomic.cuh`, `cunumpy/index.cuh` and `cunumpy/reduce.cuh`. `CudaKernel` adds it to its NVRTC options as `-I` automatically (and only once), so kernel sources can write `#include ` without configuration. Use it to pass the same headers to other compilers. @@ -694,19 +763,24 @@ kernel.compile_options() # options + ('-DCUNUMPY_INCLUDE_HASH=0x3f9a...',) `#include "name"`, recursively, each once in order of first inclusion. A name is looked up relative to the including file (`source_dir` for the kernel source, the header's own directory for nested includes), then in - `include_dirs` in order, like NVRTC does. System headers in angle brackets - and includes that cannot be found are ignored (NVRTC reports the latter). - Recomputed at every access, so it follows the files on disk. + `include_dirs` in order, then in cunumpy's header directory, like NVRTC + does. cunumpy's shipped headers are tracked also when included in angle + brackets (`#include `), so upgrading cunumpy with a + changed header recompiles the kernels that use it. Other angle-bracket + (system) headers and includes that cannot be found are ignored (NVRTC + reports the latter). Recomputed at every access, so it follows the files on + disk. * `compile_options()`: the options passed to CuPy at compile time: `options` plus `-DCUNUMPY_INCLUDE_HASH=0x` if the source includes any header, where the hash covers the contents of `included_headers` (not their paths). A changed header gives another define, hence another cache entry. Sources - without quoted includes never touch the file system. + without includes never touch the file system. The two building blocks are available on their own: -* `xp.resolve_includes(source, include_dirs=(), *, base_dir=None)`: the - resolved header paths of a source, as a list. +* `xp.resolve_includes(source, include_dirs=(), *, base_dir=None, + angle_dirs=())`: the resolved header paths of a source, as a list; + `angle_dirs` are searched last and also for `#include `. * `xp.include_hash(paths)`: the first 16 hex digits of the SHA-256 digest of the contents of the files, in order. @@ -726,7 +800,20 @@ shape is given either by `n_threads` or by `grid`: * `grid`: number of blocks per dimension, instead of `n_threads`. * `block`: block shape for this call, instead of `block_size`. * `shared_mem`: dynamic shared memory per block in bytes, for - `extern __shared__` arrays. + `extern __shared__` arrays. Above 48 KiB (the limit every device has) the + compiled kernel's `max_dynamic_shared_size_bytes` is raised to `shared_mem` + once, up to the device's opt-in limit (`xp.max_shared_memory_per_block( + opt_in=True)`); a larger request raises `ValueError` before the launch. + +Without `n_threads` and `grid`, a launch uses `n_threads_from(args)` if the +kernel has one (constructor argument and settable property): a function of the +argument tuple, or `"first_array"` for the length of the first array argument +(its first axis), i.e. one thread per marker: + +```python +push = xp.CudaKernel.from_file("push_cuda.cu", n_threads_from="first_array") +push(positions, velocities, e_field, dt) # n_threads = positions.shape[0] +``` Nothing is launched if the grid has a zero dimension (e.g. `n_threads=0`). `kernel.launch_shape(n_threads=None, *, grid=None, block=None)` returns the @@ -778,7 +865,9 @@ The arguments are prepared by `kernel.prepare_args(*args)`: and integer parameters; anything else raises `TypeError`; * NumPy scalars are passed as they are if their dtype matches, cast if the cast is safe (e.g. `np.float32` into `double`), and raise `TypeError` - otherwise (e.g. `np.float64` into `float`). + otherwise (e.g. `np.float64` into `float`). NumPy integer scalars are + checked by value, like Python ints: `np.int64(5)` fits an `int` + parameter, `np.int64(2**31)` raises `OverflowError`. This matters because `cupy.RawKernel` reads each argument with the size declared in the signature and does not check types: an integer passed to a @@ -819,8 +908,9 @@ cunumpy ships CUDA headers that every `CudaKernel` finds automatically; `xp.cuda_include_dir()` returns their directory (a `str`) for other compilers (`-I`). -`cunumpy/array_view.cuh` defines the strided views `Array1D`, `Array2D` -and `Array3D`: `T* data`, `long long shape[ndim]`, `long long +`cunumpy/array_view.cuh` defines the strided views `Array1D` to +`Array4D` (4D e.g. for a 3D grid of vector components `(nx, ny, nz, +ncomp)`): `T* data`, `long long shape[ndim]`, `long long strides[ndim]` (in elements, not bytes), `operator()(i, j, ...)` returning a reference to the element, and `size()`. A kernel indexes `a(i, j)` like the pyccel kernel it is ported from indexes `a[i, j]`, without hand-passed sizes. @@ -956,7 +1046,7 @@ changes one definition instead of every kernel signature. `CudaStruct(name, fields)` takes the fields as `(name, C type)` pairs; scalar fields, pointers to the scalar types above (or `void*`), and array views -`Array1D` to `Array3D` of those scalar types (see "CUDA headers and +`Array1D` to `Array4D` of those scalar types (see "CUDA headers and array views") are supported. * `declaration`: the C definition of the struct, to put in the CUDA source @@ -967,6 +1057,15 @@ array views") are supported. * `fields`: the parsed fields (`CudaParameter` tuples). * `check_source(source)`: raises `ValueError` if `source` defines the struct with other fields; a kernel created with `structs=[...]` does this check. +* `verify_layout(include=None, *, include_dirs=(), options=())` (needs CuPy): + compiles and runs a one-thread kernel that reports `sizeof`, `alignof` and + every field offset as the CUDA compiler lays the struct out, and raises + `ValueError` listing the differences from `dtype`. Returns the measured + layout as a dict. With `include` (a header file name or `#include` line, + found in `include_dirs`) the struct is defined by that header instead of + `declaration`, which checks a hand-written or generated header. Call it once + per struct in a GPU test, and on every new platform (e.g. ROCm). + `layout_source(include=None)` returns the kernel source. * Calling the struct with keyword arguments, one per field, packs the values: pointer fields take C-contiguous CuPy arrays of the declared dtype (never copied), array view fields take CuPy arrays of the declared dtype and number @@ -1008,6 +1107,18 @@ scalar types such as `np.float32` -> `float`, and an array `"float[:, :]"` -> types, e.g. `{"float": "float"}` for single precision. A parameter without annotation, or with an annotation that cannot be mapped, raises `ValueError`. +### Structs from a pyccel source file + +`CudaStruct.from_pyccel_class(source, class_name, name=None, *, int_type="long long", +scalar_names=None, exclude=(), attribute_names=True)` builds a struct from the +annotated `__init__` of a class in a `.py` file (or a source string), parsed +with `ast` and never imported, for classes whose module is compiled by pyccel. +Fields are named after the attributes the parameters are stored in +(`self. = `; `attribute_names=False` keeps the parameter +names), parameters in `exclude` are skipped, and the mappings are those of +`from_signature`. Raises `ValueError` if the class or its `__init__` is missing +or an annotation cannot be mapped. + ### Generating headers ```python @@ -1061,8 +1172,14 @@ built once per subclass when the class is defined and is the class attribute * `pack()` packs the field attributes, with the checks of `CudaStruct` (C-contiguous CuPy arrays of the declared dtype, range-checked scalars). A - field without an attribute raises `AttributeError`. Call it again after - replacing an array attribute: the packed struct holds device addresses. + field without an attribute raises `AttributeError`. +* The packed struct always matches the current attributes: at every use + (`packed`, `__cuda_args__()`, so at every launch) the device address, shape + and strides of each array field and the value of each scalar field are + compared with what was packed, and the struct is packed again if anything + changed. Fields may be properties that read an owner's current arrays, so a + resized array is picked up at the next launch. Calling `pack()` again is + never needed. * `packed` is the packed struct (`numpy.void`); `__cuda_args__()` returns `(packed,)`, so a `CudaKernel` receives the struct. * Copies (`copy.copy`, `copy.deepcopy`) and unpickled objects are packed again @@ -1071,6 +1188,25 @@ built once per subclass when the class is defined and is the class attribute base class (its instances cannot be packed); setting only one raises `TypeError`. Subclasses of a complete class inherit its struct. +## `PyccelStructArguments` + +`CudaStructArguments` with a host form (the `KernelArguments` protocol). Class +attributes, besides `struct_name` and `fields`: + +* `host_class`: the class of the host argument object, e.g. the + pyccel-compiled class (which cannot inherit from anything). +* `host_fields`: the attributes passed to `host_class(...)`, positionally and + in this order; by default the struct fields. +* `host_copies`: whether `__host_args__()` may build the host object from host + copies of device arrays (default `False`: it raises on the CuPy backend). + Copies are counted by `count_transfers()`; results are not copied back. + +`__host_args__()` builds `host_class(*host_fields)` once, and again when one of +the attributes was replaced (array identity or address, scalar value). +`__cuda_args__()` is the packed struct. `has_device_arrays()` tells whether the +array fields are device arrays; objects holding host arrays are copied and +pickled without packing, and the host object is never pickled. + ## `CudaArguments` ```python @@ -1174,6 +1310,7 @@ xp.Kernel( missing_cuda="raise", cuda_path=None, host_options=None, + dispatch="backend", ) ``` @@ -1206,8 +1343,38 @@ arrays inside application objects, and `outputs` limits the copies back to the device. Passing `host_options` together with a `PyccelKernel` raises `ValueError`; configure that `PyccelKernel` directly. +`dispatch` decides which kernel a call runs. `"backend"` (default): the CUDA +kernel on the CuPy backend, the host kernel on the NumPy backend. +`"arrays"`: the CUDA kernel if any top-level argument lives on the GPU (a CuPy +array, or a device-only argument object: one with `__cuda_args__()` but no +`__host_args__()`, such as a `CudaArguments` or a struct value), else the host +kernel, whatever the backend. Use `"arrays"` in codes that hand host arrays to +kernels while CuPy is active (diagnostics, MPI staging, CPU fallbacks): those +calls then run the host kernel instead of failing in the CUDA argument checks. +`missing_cuda` applies to device arguments without a CUDA kernel. + +`kernel.check_signature()` checks that the host and CUDA kernels take the same +parameters in the same order (the names of the Python host function, or of the +uncompiled Python version of a `CompiledHostKernel`, against the parsed +`__global__` signature) and raises `ValueError` showing both lists otherwise. +It does nothing without a CUDA kernel, with `check_signature=False`, or when +the host kernel has no Python signature (a compiled function). +`kernel.host_parameters()` returns the host names, or None. + Properties: `name`, `host_kernel`, `cuda_kernel`, `has_cuda`, `missing_cuda`, -`cuda_path`. +`cuda_path`, `dispatch`. + +### Test arguments and compiled host kernels + +* `Kernel(..., test_args="pkg.push.push_test_args")`, + `Kernel.test_args_module` and `Kernel.test_args` (the module, imported on + first access): the test-arguments module of the kernel, set by + `KernelCatalog.from_package()`; see `check_parity` below. +* `Kernel.host_parameters()` falls back to the `__pyccel__/.pyi` stub + for a pyccel-compiled host function, so `check_signature()` works for + compiled kernels. +* `Kernel.__call__` needs no `n_threads` when the CUDA kernel has + `n_threads_from`. ## `KernelCatalog` @@ -1220,6 +1387,9 @@ catalog = xp.KernelCatalog.from_package( missing_cuda="raise", host_options=None, include_dirs=None, + dispatch="backend", + compile_host=None, + host_fallback=None, **cuda_options, ) kernel = catalog["push"] @@ -1252,9 +1422,36 @@ my_kernels/ `my_pkg.kernels` can `#include "my_pkg/common.cuh"`. Headers found this way take part in the compile cache key, see "Included headers and the compile cache" under `CudaKernel`. +* `dispatch`: passed on to every `Kernel` (`"backend"` or `"arrays"`). +* `compile_host`: your function that compiles a host kernel module (cunumpy + does not compile anything itself), e.g. a wrapper around `pyccel.epyccel` + with a cache. Each host kernel is then a `CompiledHostKernel` (below), + compiled on its first call. Without it, the plain Python functions are + called. +* `host_fallback`: for kernels whose compilation fails, a callable with the + same arguments (e.g. a vectorized NumPy version), given as a mapping from + names or a function of the name. Without one, a failed compilation runs the + uncompiled Python function, with a warning. * `cuda_options`: passed on to `CudaKernel.from_file`, e.g. `block_size` or `structs`. +A Pyccel package with the layout `/_pyccel.py` and +`/_cuda.cu`, compiled host kernels and dispatch by argument: + +```python +catalog = xp.KernelCatalog.from_package( + __name__, + host_suffix="_pyccel", + dispatch="arrays", + compile_host=my_pkg.compile_kernels, # e.g. pyccel.epyccel with a cache + host_fallback=NUMPY_VERSIONS, # {"gather": gather_numpy, ...} +) +``` + +`catalog.check_signatures()` runs `check_signature()` on every kernel and +raises one `ValueError` listing every kernel whose host and CUDA parameters +differ; call it in a unit test of a ported package. + `catalog.without_cuda` lists the kernels still to port, `catalog.with_cuda` the ported ones. `catalog.summary()` (also `str(catalog)`) is one line on the porting status, e.g. for a `--status` command: @@ -1274,6 +1471,24 @@ that have a CUDA kernel, for a parametrised parity test (see "Testing utilities"). `KernelCatalog(kernels)` and `catalog.register(kernel, name=None)` build a catalog by hand. +## `CompiledHostKernel` + +```python +kernel = xp.CompiledHostKernel(my_kernels_module, "push", compiler, fallback=push_numpy) +kernel(*args) +``` + +A host kernel compiled on its first call, for `KernelCatalog.from_package(..., +compile_host=...)` or by hand. cunumpy does not compile anything itself: +`compiler(module)` is your function returning the compiled form of the module +(with the same function names), e.g. a wrapper around `pyccel.epyccel` with an +on-disk cache, or one that imports modules compiled ahead of time with the +`pyccel` command. If compilation fails, the kernel calls `fallback`, or else +the uncompiled Python function with a `RuntimeWarning`. `kernel.compiled` +builds and reports whether that worked (so that callers can choose another +path), `kernel.error` is the exception of a failed build, `kernel.python` the +uncompiled function and `kernel.build()` compiles now. + ## Testing utilities ```python @@ -1281,6 +1496,7 @@ from cunumpy.testing import ( BACKENDS, assert_kernels_agree, device_function_kernel, + emulate_cuda_kernel, requires_cupy, ) ``` @@ -1431,6 +1647,88 @@ pointer return type, an unsupported parameter type, or a parameter named like `out_param` or `n_threads_param` raises `ValueError` (rename the generated parameter in that case). +### `parity_cases(catalog)`, `check_parity(kernel, **overrides)` + +`parity_cases(catalog)` returns one `pytest.param(kernel, id=name)` per kernel +of `catalog.parity_cases()`; a kernel without a test-arguments module is +marked `skip` with a reason naming the missing `_test_args.py`. +`check_parity(kernel)` runs `assert_kernels_agree` with the module's +`make_args` and the settings of `TEST_ARGS_SETTINGS` (`N_THREADS`, `GRID`, +`BLOCK`, `RTOL`, `ATOL`, `N_CALLS`, `OUTPUTS`, `SEED`), overridden by keyword +arguments. Raises `ValueError` without a module and `TypeError` without a +callable `make_args`. + +### `install_fake_cupy()`, `fake_cupy_active()` + +Install the fake CuPy of `cunumpy._fake_cupy` (a strict host stand-in for CuPy +for machines without a GPU; also installed by `CUNUMPY_FAKE_CUPY=1` when +cunumpy is imported), and tell whether it is active. `requires_cupy` and +`assert_kernels_agree` skip while it is. `install_fake_cupy()` raises if the +real CuPy was imported already or cunumpy already checked for CuPy. + +### `emulate_cuda_kernel(kernel, *args, n_threads=None, grid=None, block=None, compiler=None, options=(), shared_mem=0)` + +```python +x, y = rng.random(1000), np.zeros(1000) +emulate_cuda_kernel(axpy, 2.0, x, y, 1000, n_threads=1000) +np.testing.assert_allclose(y, 2.0 * x, rtol=1e-15) +``` + +Runs a `CudaKernel` on the CPU, serially, as if it were launched with `args`, +so that CI without a GPU can compare a kernel with its host version. The +kernel source is compiled as C++ (C++17, `CXX` or `c++`; see +`emulation_compiler()`) with the CUDA built-ins replaced: `threadIdx`, +`blockIdx`, `blockDim`, `gridDim`, atomics (`atomicAdd`, `atomicMin`, ..., +plain operations), `__ldg`, `rsqrt`, `__trap` (aborts). The shipped headers +and the kernel's include directories and `-D` options apply. Then the kernel is +called once per thread, for every block and thread index of the launch shape. + +Arguments follow the signature, with NumPy arrays in place of CuPy arrays: +pointer and view parameters (`Array1D` to `Array4D`) take arrays of the +declared dtype (and ndim), passed as contiguous copies, so any strides work, +and written back into the given arrays; scalars are checked and cast like in a +launch. Like NVRTC by default, the compiler may fuse `a * b + c` into an FMA, +so compare with NumPy using a tolerance of a few ulp, or pass +`options=("-ffp-contract=off",)` for NumPy's rounding. + +Block shared memory and `__syncthreads` are emulated: `__shared__` variables +are one copy per block (blocks run one after another), `extern __shared__` +arrays point into a buffer of `shared_mem` bytes, and in a kernel that calls +`__syncthreads` the threads of a block run as coroutines (POSIX `ucontext`, +each with its own stack), so that every thread reaches a barrier before any +thread continues past it. Per-block deposits, shared-memory reductions and +tiled kernels work. + +Not emulated: concurrency between barriers (races and atomic ordering never +show), warp intrinsics, struct parameters and complex scalars. A kernel, or a +header it includes, using `__syncwarp`, warp shuffles or votes raises +`NotImplementedError` (serial threads would give wrong results); a kernel that +does not compile, or crashes (an out-of-bounds +index with `-DCUNUMPY_BOUNDS_CHECK`, `__trap()`), raises `RuntimeError` with +the compiler or program output. + +## `HostStaging` + +```python +staging = xp.HostStaging(rho.shape, rho.dtype, buffers=2) +copy = staging.copy(rho) # returns at once; rho may be overwritten +... +if copy.ready(): + h5file["rho"] = copy.result() # a NumPy array +``` + +Copies device arrays to page-locked host buffers in the background, so output +overlaps the next time steps. `copy(array)` snapshots the array on the device +(on the current stream, after the kernels that wrote it) and copies the +snapshot to the next of `buffers` pinned host buffers on its own stream. It +waits only if that buffer's previous copy has not finished, so the program +runs at most `buffers` copies ahead. The returned `StagedCopy` has `ready()` +(never waits) and `result()` (waits, returns the host buffer, valid until the +buffer is reused `buffers` copies later; a stale result raises +`RuntimeError`). `staging.synchronize()` waits for all copies. Host arrays and +the NumPy backend copy at once. Device copies are counted by +`count_transfers()`. The arrays must have the staging shape and dtype. + ## `DeviceMirror` ```python @@ -1492,6 +1790,143 @@ The indexed helpers (also for `float`) address C-contiguous arrays of shape for `double` from compute capability 6.0 (sm_60) on; older devices use a compare-and-swap loop. +### `cunumpy/random.cuh` and `philox_uniform` + +Counter-based random numbers (Philox4x32-10, as in Random123 and cuRAND): a +pure function of a key and a counter, with no generator state, so each thread +draws from `(seed, stream, counter)`, e.g. `(seed, particle id, step)`, and +the host computes the same numbers: + +```c +#include + +cunumpy_u32x4 cunumpy_philox4x32_10(cunumpy_u32x4 ctr, unsigned int key0, unsigned int key1); +double cunumpy_uniform(seed, stream, counter); // [0, 1), 53 bits +void cunumpy_uniform2(seed, stream, counter, &u0, &u1); // two from one call +double cunumpy_normal(seed, stream, counter); // Box-Muller +void cunumpy_normal2(seed, stream, counter, &z0, &z1); +``` + +(`seed`, `stream` and `counter` are `unsigned long long`.) + +```python +ids = xp.arange(n, dtype=xp.uint64) +u0, u1 = xp.philox_uniform2(seed, ids, step) # == cunumpy_uniform2 in thread i +z0, z1 = xp.philox_normal2(seed, ids, step) +words = xp.philox4x32_10(counter_words, key0, key1) # the raw generator +``` + +`xp.philox_uniform`, `philox_uniform2`, `philox_normal`, `philox_normal2` and +`philox4x32_10` broadcast their arguments and return NumPy or CuPy arrays, +matching the inputs. The uniform numbers equal the kernel's bit for bit; the +normal numbers can differ in the last bits (`log`, `sqrt`, `sin`, `cos` on the +GPU are not the host's). The generator passes the Random123 known-answer +tests. Use a different `counter` for every random decision of a step. + +### `cunumpy/reduce.cuh` + +Warp- and block-level reductions for hand-written kernels: in-kernel +diagnostics (energy, momentum, total charge, the maximum velocity for a CFL +check) and combining values in a block before one atomic write. + +```c +#include + +T cunumpy_warp_sum(T v); T cunumpy_warp_min(T v); T cunumpy_warp_max(T v); +T cunumpy_block_sum(T v); T cunumpy_block_min(T v); T cunumpy_block_max(T v); +void cunumpy_block_sum_to(T* out, T v); // *out += block sum, one atomic per block +int cunumpy_block_thread(); // linear thread index in a 1D-3D block +int cunumpy_block_threads(); // threads per block +``` + +`T` is `int`, `unsigned`, `long long`, `unsigned long long`, `float` or +`double` (`block_sum_to`: `double`, `float`, `int`, `unsigned long long`). Every +thread gets the result. Rules: every thread of the block calls block +functions and all 32 lanes call warp functions (no early `return`; threads +without a value pass the identity, e.g. `0.0` for a sum); the block size is a +multiple of 32. The block functions use 32 values of static shared memory per +type and may be called several times in a kernel. + +```c +extern "C" __global__ void kinetic_energy(const double* v, long long n, + double mass, double* energy) { + long long i = blockIdx.x * (long long)blockDim.x + threadIdx.x; + double e = i < n ? 0.5 * mass * v[i] * v[i] : 0.0; + cunumpy_block_sum_to(energy, e); // zero *energy before the launch +} +``` + +## `scipy` + +```python +A = xp.scipy.sparse.csr_matrix((data, (rows, cols)), shape=(n, n)) +x, info = xp.scipy.sparse.linalg.cg(A, b) +rho_k = xp.scipy.fft.rfftn(rho) +``` + +SciPy for the active backend: `scipy` on NumPy, `cupyx.scipy` on CuPy. The +forwarded subpackages are those `cupyx.scipy` has (`SUBMODULES`): `fft`, +`fftpack`, `interpolate`, `linalg`, `ndimage`, `signal`, `sparse`, +`sparse.csgraph`, `sparse.linalg`, `spatial`, `special`, `stats`. Names are +looked up at every access, so a backend switch takes effect immediately +(Python caches the imports). Nothing is imported until a name is used; SciPy +is not a dependency of cunumpy. + +* A name missing on the active backend (`cupyx.scipy` covers part of SciPy) + raises `AttributeError` naming the backend. Keyword arguments can differ + too (SciPy's `cg(..., rtol=)` is `tol=` in CuPy). +* `xp.scipy.special.available("erfcx")` checks a name without raising. +* `xp.scipy.sparse.linalg.resolve()` returns the module itself. +* A missing SciPy (NumPy backend) or CuPy raises `ImportError` with the module + it needs. + +Sparse matrices assembled on the host move to the device once, with the +constructor of the device type: `xp.scipy.sparse.csr_matrix(host_matrix)` on +the CuPy backend copies a SciPy matrix; `matrix.get()` copies back. + +## `fuse(function=None, *, kernel_name=None)` + +```python +@xp.fuse +def pressure(rho, T, gamma): + return (gamma - 1.0) * rho * T +``` + +Compiles an elementwise function into one kernel with `cupy.fuse` when it is +called with a CuPy array (positional or keyword), with the CuPy backend active +while it is traced, so `xp.exp` etc. resolve to CuPy ufuncs; other calls run +the function as it is. The fused kernel is created on first use and reused. +`kernel_name` names it in profilers (default: the function name). The function +must be elementwise in the sense of `cupy.fuse`: arithmetic, comparisons, +ufuncs, `xp.where`, and supported reductions as the last operation; no Python +control flow on array values or indexing. Test the CuPy path: a function +`cupy.fuse` cannot trace raises at its first call with CuPy arrays. + +## `petsc_vec(array, comm=None)` + +```python +b_vec = xp.petsc_vec(b) # b: NumPy or CuPy array, shared, never copied +x_vec = xp.petsc_vec(x) +xp.synchronize() +ksp.solve(b_vec, x_vec) # PETSc writes into x +xp.synchronize() +``` + +A `petsc4py.PETSc.Vec` that shares the memory of a C-contiguous array of +`PETSc.ScalarType`, through DLPack: a `seq`/`mpi` vector for a NumPy array, a +`seqcuda`/`mpicuda` (or HIP) vector for a CuPy array. The vector keeps a +reference to the array. With several processes, `array` is this process's +part of the vector (`comm`, default `COMM_WORLD`). + +* Another dtype raises `TypeError` and a non-contiguous array `ValueError` + (both would need a copy). +* A CuPy array with a petsc4py built without CUDA/HIP raises `RuntimeError` + instead of PETSc working on a host copy. +* Synchronize between CuPy and PETSc work on the same memory (they may use + different streams). +* For the solve to stay on the GPU, the matrix must be a GPU type too + (`aijcusparse`, or `-mat_type aijcusparse -vec_type cuda`). + ## Version `xp.__version__` is the installed package version. When package metadata is diff --git a/docs/source/best-practices.md b/docs/source/best-practices.md index 4ba86a5..9ba9873 100644 --- a/docs/source/best-practices.md +++ b/docs/source/best-practices.md @@ -37,7 +37,12 @@ A condensed checklist. Each item links to the guide with the reasoning. kernels by profile order, switch to `"raise"` when done. ([Porting kernels](kernels/overview.md)) * Keep the CUDA kernel's argument list identical to the host kernel's; put both - in one folder. + in one folder, and test it with `catalog.check_signatures()`. +* Compile Pyccel host kernels ahead of time with the `pyccel` command, or at run + time with `from_package(..., compile_host=...)`, and keep a NumPy version as + `host_fallback`. +* Use `dispatch="arrays"` if host arrays reach kernels while CuPy is active. +* Ship `.cu`/`.cuh` files as package data. * Declare `outputs` on host kernels so the fallback copies back only what was written, and never forget an argument that is written. * Pass device arrays to `CudaKernel`; build argument objects once with @@ -58,6 +63,8 @@ A condensed checklist. Each item links to the guide with the reasoning. everywhere and use the GPU where there is one. ([Testing kernels](kernels/testing.md)) * One `assert_kernels_agree` test over `catalog.parity_cases()`. +* `emulate_cuda_kernel` tests, so CPU-only CI checks the CUDA arithmetic. +* Seed `xp.random_streams` with `(seed, rank)` and draw only from it. * Debug crashes with `CUNUMPY_CUDA_DEBUG=1`, then `compute-sanitizer`. ([Debugging](kernels/debugging.md)) * Time with `timed_region()`, profile with `nvtx_range()` and `nsys`. diff --git a/docs/source/examples/particle_recipes.py b/docs/source/examples/particle_recipes.py new file mode 100644 index 0000000..52847eb --- /dev/null +++ b/docs/source/examples/particle_recipes.py @@ -0,0 +1,100 @@ +"""Recipes for particle codes with cunumpy: the code of the "Particle codes" page. + +Every function runs on NumPy and on CuPy arrays (``xp`` follows the active +backend); ``tests/unit/test_particle_recipes.py`` runs them. +""" + +from __future__ import annotations + +import numpy as np + +import cunumpy as xp + + +def remove_dead(markers, alive): + """Drop the markers whose ``alive`` flag is False (a new, compact array).""" + return markers[alive] + + +def compact_in_place(markers, alive): + """Move the live markers to the front of a preallocated buffer; return their count. + + ``markers[:n]`` are the live markers afterwards, in their original order, + and the rows behind them are free for injection. The right-hand side is a + copy (boolean indexing), so source and destination may overlap. + """ + n_alive = int(alive.sum()) + markers[:n_alive] = markers[alive] + return n_alive + + +def sort_by_cell(positions, lower, cell_size, n_cells): + """Order the markers by cell; return the order, the cells and the cell offsets. + + After ``markers = markers[order]`` the markers of cell ``c`` are + ``markers[offsets[c]:offsets[c + 1]]``. The sort is stable, so markers + keep their relative order within a cell (reproducible results). + """ + cell = xp.floor((positions - lower) / cell_size).astype(xp.int64) + cell = xp.clip(cell, 0, n_cells - 1) + order = xp.argsort(cell, kind="stable") + counts = xp.bincount(cell, minlength=n_cells) + offsets = xp.zeros(n_cells + 1, dtype=xp.int64) + offsets[1:] = xp.cumsum(counts) + return order, cell[order], offsets + + +def deposit_nearest_cell(cell, weights, n_cells): + """Sum the weights per cell without atomics: the sort-then-reduce deposit.""" + return xp.segment_sum(weights, cell, n_cells) + + +def pack_for_ranks(markers, destination, n_ranks): + """Group the markers by destination rank; return the send buffer and the counts. + + ``destination[i]`` is the rank marker ``i`` moves to (its own rank to stay). + The markers for rank ``r`` are the rows ``displacements[r]`` to + ``displacements[r] + counts[r]`` of the buffer. + """ + order = xp.argsort(destination, kind="stable") + counts = xp.bincount(destination, minlength=n_ranks) + return xp.ascontiguousarray(markers[order]), counts + + +def exchange(comm, markers, destination): + """Send every marker to its destination rank (``MPI_Alltoallv``); return the received ones. + + Works for host and device arrays, with or without CUDA-aware MPI + (``xp.mpi_buffer`` stages device buffers through the host when needed). + """ + n_ranks = comm.Get_size() + width = markers.shape[1] + sendbuf, counts = pack_for_ranks(markers, destination, n_ranks) + send_counts = np.asarray(xp.to_numpy(counts), dtype=np.int64) + recv_counts = np.empty(n_ranks, dtype=np.int64) + comm.Alltoall(send_counts, recv_counts) # how many rows come from each rank + received = xp.empty((int(recv_counts.sum()), width), dtype=markers.dtype) + + def displacements(counts): + return np.concatenate([[0], np.cumsum(counts)[:-1]]) + + with ( + xp.mpi_buffer(sendbuf) as send, + xp.mpi_buffer(received, send=False, recv=True) as recv, + ): + comm.Alltoallv( + [send, send_counts * width, displacements(send_counts) * width, None], + [recv, recv_counts * width, displacements(recv_counts) * width, None], + ) + return received + + +def thermal_velocities(seed, particle_ids, step, v_th): + """Maxwellian velocities from (seed, particle id, step): no generator state. + + The same numbers as ``cunumpy_normal2(seed, id, step, ...)`` in a kernel + (up to the last bits of the math functions), whatever the order of the + particles or the number of ranks. + """ + z0, z1 = xp.philox_normal2(seed, particle_ids, step) + return v_th * z0, v_th * z1 diff --git a/docs/source/guides/mpi.md b/docs/source/guides/mpi.md index 10c440e..c6211ea 100644 --- a/docs/source/guides/mpi.md +++ b/docs/source/guides/mpi.md @@ -95,6 +95,47 @@ columns are not (copy them with `xp.ascontiguousarray()` first). No synchronization is needed after MPI returns; kernels launched afterwards see the received data. +### 5. One call site for both MPI builds: `mpi_buffer` + +Code that must also run with an MPI library that is not CUDA-aware (or on +the NumPy backend) stages device buffers through the host. `mpi_buffer()` +does the right thing for each case, so the MPI call is written once: + +```python +xp.mpi_is_cuda_aware(comm) # once at startup; the answer is remembered + +with xp.mpi_buffer(send_r) as sendbuf, xp.mpi_buffer(recv_l, send=False, recv=True) as recvbuf: + comm.Sendrecv(sendbuf, dest=right, recvbuf=recvbuf, source=left) +``` + +A host array is yielded as it is. A device array is yielded as it is (after +`synchronize_for_mpi`) when MPI is CUDA-aware, and otherwise replaced by a +pinned host copy: filled from the device before the block when `send=True`, +copied back into the device array after the block when `recv=True`. The +copies are counted by `count_transfers()`, so a GPU run with a plain MPI +build is visible in the transfer report. Without a recorded answer (no +`mpi_is_cuda_aware()` call and no `set_mpi_cuda_aware()`), a device array +raises instead of guessing. + +## Reproducible random numbers + +Each rank needs its own random stream, and a run is reproducible only if every +draw comes from a seeded generator. Seed `xp.random_streams` once, after MPI is +initialized, and draw from it everywhere: + +```python +xp.random_streams.seed(config.seed, rank=comm.Get_rank()) + +positions = xp.random_streams.random((n, 3)) +velocities = xp.random_streams.normal(0.0, v_th, (n, 3)) +rng = xp.random_streams.generator() # for other distributions +``` + +Rank `r` draws the stream `(seed, r)`; the same seed and number of ranks give +the same results, and different ranks never share numbers. Avoid unseeded +generators (`np.random.default_rng()` without a seed) anywhere in the time loop: +one of them makes every run different. + ## Launching ```bash diff --git a/docs/source/guides/particle-codes.md b/docs/source/guides/particle-codes.md new file mode 100644 index 0000000..584ecbc --- /dev/null +++ b/docs/source/guides/particle-codes.md @@ -0,0 +1,186 @@ +# Particle codes + +Recipes for the parts of a particle-in-cell (or any particle) code around the +kernels: removing and sorting markers, depositing without atomics, moving +markers between MPI ranks, random numbers per particle, output, and running a +time step as a CUDA graph. The NumPy recipes are the functions of +`docs/source/examples/particle_recipes.py`, which the test suite runs; each +works on NumPy and CuPy arrays alike. + +## Remove dead markers + +Markers leave the domain, hit a wall or are absorbed. Keep a boolean `alive` +array and compact: + +```{literalinclude} ../examples/particle_recipes.py +:pyobject: remove_dead +``` + +To avoid reallocating, keep the markers in a preallocated buffer with a count +of live rows, and move the live ones to the front; the free rows behind them +take injected markers: + +```{literalinclude} ../examples/particle_recipes.py +:pyobject: compact_in_place +``` + +`alive.sum()` on the GPU is a reduction plus one small transfer (the count); +do it once per step, not per species and kernel. + +## Sort markers by cell + +Sorting by cell makes deposits and gathers read and write memory in order, +which is often worth more than the sort costs, and gives each cell a +contiguous range of markers, which binary collisions and per-cell diagnostics +need: + +```{literalinclude} ../examples/particle_recipes.py +:pyobject: sort_by_cell +``` + +Apply `order` to every per-marker array (positions, velocities, weights, ids), +e.g. by keeping them as columns of one `(n, k)` array. A stable sort keeps the +result independent of how the markers were ordered before. + +## Deposit without atomics: sort, then reduce + +With markers sorted by cell, a nearest-cell deposit is a segmented sum, +deterministic on both backends: + +```{literalinclude} ../examples/particle_recipes.py +:pyobject: deposit_nearest_cell +``` + +For linear (cloud-in-cell) weights, deposit each of the two (2D: four, 3D: +eight) neighbours with its weight, one `segment_sum` each. + +The two other GPU strategies, as kernels: + +* **Global atomics** (`cunumpy_atomic_add` from ``): the + simplest; slow when many threads hit the same cells, and the summation order + varies between runs (results differ in the last bits). +* **Per-block shared memory**: each block deposits into a copy of the grid in + shared memory, then adds it to the global grid once per cell. Fast for small + grids. Check the size with `xp.max_shared_memory_per_block()` and pass + `shared_mem=` at the launch; above 48 KiB the kernel is set up for the larger + limit automatically. + +```c +extern "C" __global__ +void deposit(Array1D x, Array1D w, Array1D rho, + double lower, double dx, int nx) { + extern __shared__ double block_rho[]; + for (int k = threadIdx.x; k < nx; k += blockDim.x) block_rho[k] = 0.0; + __syncthreads(); + long long i = blockIdx.x * (long long)blockDim.x + threadIdx.x; + if (i < x.shape[0]) { /* cunumpy_atomic_add(&block_rho[cell], ...) */ } + __syncthreads(); + for (int k = threadIdx.x; k < nx; k += blockDim.x) + cunumpy_atomic_add(&rho(k), block_rho[k]); +} +``` + +`cunumpy.testing.emulate_cuda_kernel(..., shared_mem=8 * nx)` runs such a +kernel on the CPU, barriers included, so it can be checked against the host +version without a GPU. + +## Move markers between MPI ranks + +With a spatial decomposition, markers that leave a rank's subdomain move to the +rank that owns their new position: compute each marker's destination rank, +group the markers by destination, exchange the counts, then the markers: + +```{literalinclude} ../examples/particle_recipes.py +:pyobject: pack_for_ranks +``` + +```{literalinclude} ../examples/particle_recipes.py +:pyobject: exchange +``` + +`xp.mpi_buffer` hands device arrays to MPI directly when it is CUDA-aware and +stages them through host memory otherwise (see [MPI](mpi.md)). Only the counts +are host arrays. Remove the markers that left with the compaction above, and +append the received ones. + +## Random numbers per particle + +A time loop is reproducible only if every random number comes from a seed. +Counter-based random numbers make that independent of the order of the +markers, the launch shape and the number of ranks: particle `id` at step +`step` always draws the same numbers. On the host: + +```{literalinclude} ../examples/particle_recipes.py +:pyobject: thermal_velocities +``` + +and in a kernel, with ``, the same numbers (uniforms +bit for bit): + +```c +#include +double z0, z1; +cunumpy_normal2(seed, particle_id, step, &z0, &z1); +v[i] = v_th * z0; +``` + +Use a different counter for every random decision of a step (e.g. +`4 * step + 0` for injection, `4 * step + 1` for collisions) so that they are +independent. For draws that need not be per particle (e.g. a collision +operator's own sampling), `xp.random_streams` gives one seeded generator per +rank. + +## Write output without stalling the GPU + +```python +staging = xp.HostStaging(rho.shape, rho.dtype) # once +pending = [] +for step in range(n_steps): + advance() + if step % output_every == 0: + pending.append((step, staging.copy(rho))) # returns at once + while pending and pending[0][1].ready(): + s, copy = pending.pop(0) + h5file[f"rho/{s}"] = copy.result() +``` + +`copy()` snapshots the array on the device and copies the snapshot to pinned +host memory on its own stream, so the next steps overwrite `rho` while the copy +runs. See `HostStaging` in the [API reference](../api.md). + +## Run a time step as a CUDA graph + +A step of a small simulation is dozens of short kernels, and the launch +overhead (a few microseconds each) can dominate. CuPy can capture the launches +of a step on a stream once and replay them: + +```python +stream = cp.cuda.Stream(non_blocking=True) +catalog.compile_all() # compile before capturing: no compilation in a graph +with stream: + stream.begin_capture() + step_kernels() # the kernel launches of one step + graph = stream.end_capture() +for _ in range(n_steps): + graph.launch(stream) +stream.synchronize() +``` + +A graph replays exactly the captured launches: the same arrays (by address), +the same scalar arguments and launch shapes. So allocate every array before +capturing, keep scalars that change per step in device arrays, and keep host +synchronization (`.get()`, `float(x)`, `alive.sum()` read on the host, MPI) +outside the captured part. Debug mode skips its synchronization while a stream +is capturing. + +## Choose the block size from the compiled kernel + +```python +raw = kernel.compile() # the cupy.RawKernel +raw.num_regs, raw.max_threads_per_block, raw.shared_size_bytes +``` + +A kernel that uses many registers per thread cannot run 1024 threads per block; +`max_threads_per_block` is the limit for this kernel on this device. Start +with 128 or 256 threads per block, then time a few sizes on the target GPU +(`xp.timed_region`) for the kernels that dominate a step. diff --git a/docs/source/guides/portable-code.md b/docs/source/guides/portable-code.md index abc66ca..02e04e2 100644 --- a/docs/source/guides/portable-code.md +++ b/docs/source/guides/portable-code.md @@ -130,8 +130,9 @@ of the installed CuPy version. The common differences: device array on CuPy. Keep it as an array, or call `float()` and accept the synchronization. * SciPy functions accept only NumPy arrays; CuPy has its own `cupyx.scipy`. - Convert with `to_numpy()` at that boundary, or branch on - `xp.get_array_backend(a)`. + Use `xp.scipy`, which is the one for the active backend (see + [Solvers and fluid updates](solvers.md)); for functions `cupyx.scipy` lacks, + convert with `to_numpy()` at that boundary. * Plotting, HDF5 and most I/O libraries need host arrays: `to_numpy()` first. When a function genuinely needs a different implementation per backend, branch diff --git a/docs/source/guides/solvers.md b/docs/source/guides/solvers.md new file mode 100644 index 0000000..058e7e3 --- /dev/null +++ b/docs/source/guides/solvers.md @@ -0,0 +1,141 @@ +# Solvers and fluid updates + +Particle pushes and deposits are only part of a plasma code. Field solves need +sparse matrices, iterative solvers and FFTs; fluid (MHD) updates are long +chains of pointwise operations; diagnostics reduce over all cells or +particles. This page covers the three tools for that: `xp.scipy`, `xp.fuse` +and `xp.petsc_vec`, plus in-kernel reductions with `cunumpy/reduce.cuh`. + +## SciPy on both backends: `xp.scipy` + +`xp.scipy` is SciPy on the NumPy backend and `cupyx.scipy` on the CuPy +backend, so solver code is written once: + +```python +import numpy as np + +import cunumpy as xp + + +def poisson_1d(rho, dx): + """Solve -phi'' = rho with phi = 0 at both ends (second-order FD).""" + n = rho.size + main = xp.full(n, 2.0 / dx**2) + off = xp.full(n - 1, -1.0 / dx**2) + A = xp.scipy.sparse.diags([off, main, off], [-1, 0, 1], format="csr") + phi, info = xp.scipy.sparse.linalg.cg(A, rho) + assert info == 0 + return phi + + +def periodic_poisson(rho, length): + """Solve -phi'' = rho on a periodic grid with FFTs.""" + n = rho.size + k = 2 * np.pi * xp.fft.rfftfreq(n, d=length / n) + rho_k = xp.scipy.fft.rfft(rho) + phi_k = xp.where(k == 0, 0.0, rho_k / xp.where(k == 0, 1.0, k**2)) + return xp.scipy.fft.irfft(phi_k, n) +``` + +* The subpackages forwarded are the ones `cupyx.scipy` has: `fft`, `linalg`, + `ndimage`, `signal`, `sparse`, `sparse.linalg`, `sparse.csgraph`, + `spatial`, `special`, `stats`, `interpolate`, `fftpack`. +* `cupyx.scipy` has only part of SciPy. A missing name raises + `AttributeError` saying which backend lacks it; check with + `xp.scipy.special.available("erfcx")` where a fallback is possible. +* Keyword arguments can differ between the two: SciPy's `cg` takes `rtol` + (since SciPy 1.12), CuPy's takes `tol`. Pass tolerances only after checking + both, or through a small wrapper. +* Assemble a sparse matrix once and move it to the device once: + `A = xp.scipy.sparse.csr_matrix(host_matrix)` on the CuPy backend copies a + SciPy matrix to the device. Do not rebuild matrices in the time loop. +* Special functions for distribution functions and cross sections + (`erf`, `erfc`, Bessel functions `i0`, `i1`, `k0`, ...) are in + `xp.scipy.special` on both backends. + +## Fused elementwise updates: `xp.fuse` + +On the GPU, `(gamma - 1) * (E - 0.5 * rho * u**2)` runs as five kernels, each +reading and writing a full temporary array. `xp.fuse` turns the whole +function into one kernel with `cupy.fuse` when it is called with CuPy arrays, +and calls it unchanged with NumPy arrays: + +```python +@xp.fuse +def pressure(rho, mom, energy, gamma): + u = mom / rho + return (gamma - 1.0) * (energy - 0.5 * rho * u * u) + + +@xp.fuse +def maxwellian(v, n, u, v_th): + return n / (xp.sqrt(2.0 * np.pi) * v_th) * xp.exp(-0.5 * ((v - u) / v_th) ** 2) + + +p = pressure(rho, mom, energy, 5.0 / 3.0) +``` + +Memory-bound chains like these typically get several times faster. The +function must be elementwise: arithmetic, comparisons, ufuncs (`xp.exp`, +`xp.sqrt`, ...), `xp.where`; no `if` on array values, no indexing. Test the +CuPy path (`cunumpy.testing.BACKENDS`): `cupy.fuse` reports a function it +cannot trace at the first call with CuPy arrays. + +## PETSc without copies: `xp.petsc_vec` + +When the field solve goes through PETSc, `xp.petsc_vec(array)` wraps an +array as a PETSc vector that uses the array's memory, a CUDA vector for a +CuPy array. The deposit writes into `rho`, PETSc reads it and writes `phi`, +the gather reads `phi`, and no data leaves the GPU: + +```python +from petsc4py import PETSc + +rho = xp.zeros(n_local) +phi = xp.zeros(n_local) +rho_vec, phi_vec = xp.petsc_vec(rho), xp.petsc_vec(phi) + +A = assemble_laplacian() # a PETSc Mat +A.setType("aijcusparse") # keep the matrix on the GPU too +ksp = PETSc.KSP().create() +ksp.setOperators(A) + +for step in range(n_steps): + deposit(rho) + xp.synchronize() # CuPy finished writing rho + ksp.solve(rho_vec, phi_vec) + xp.synchronize() # PETSc finished writing phi + gather(phi) +``` + +This needs a petsc4py built with CUDA (or HIP) support; otherwise +`petsc_vec` raises for CuPy arrays rather than let PETSc work on a host copy. +The array must be C-contiguous and of `PETSc.ScalarType`, since anything else +would need a copy. With MPI, each process passes its local part and the +vector's communicator. + +## Reductions inside kernels: `cunumpy/reduce.cuh` + +Outside kernels, `xp.sum` and friends are the right tool. Inside a +hand-written kernel, for a diagnostic computed in the same pass as the push, +or to combine values in a block before one atomic write, include the header: + +```c +#include + +extern "C" __global__ void push_and_energy(double* x, double* v, const double* E, + long long n, double qm_dt, double dt, + double half_m, double* energy) { + long long i = blockIdx.x * (long long)blockDim.x + threadIdx.x; + double e = 0.0; + if (i < n) { + v[i] += qm_dt * E[i]; + x[i] += dt * v[i]; + e = half_m * v[i] * v[i]; + } + cunumpy_block_sum_to(energy, e); // every thread calls it: no early return +} +``` + +The block size must be a multiple of 32, and every thread must reach the +reduction. See the API reference for the warp and block functions. diff --git a/docs/source/index.md b/docs/source/index.md index 54d3585..fc72edf 100644 --- a/docs/source/index.md +++ b/docs/source/index.md @@ -49,6 +49,8 @@ guides/portable-code guides/data-movement guides/gpu-devices guides/mpi +guides/particle-codes +guides/solvers guides/profiling array-api-compat ``` diff --git a/docs/source/kernels/accumulation.md b/docs/source/kernels/accumulation.md index 4ccad9d..f4135b9 100644 --- a/docs/source/kernels/accumulation.md +++ b/docs/source/kernels/accumulation.md @@ -99,6 +99,24 @@ rho.to_host() # rho_host now holds the charge density, on both backends The same lines run on the NumPy backend, where `rho.device is rho_host`, the host kernel writes into it directly, and `to_host()` does nothing. +## Sort-then-reduce: `segment_sum` + +The alternative to atomics: bin the particles (sort or compute a cell key per +particle), compute each particle's contribution into an array, and sum per +cell. The last step is `xp.segment_sum(values, keys, n_segments)`, on either +backend: + +```python +cell = ix + nx * (iy + ny * iz) # (n_particles,), -1 for outside +weights = compute_weights(markers) # (n_particles, 8), one per corner +rho_cells = xp.segment_sum(weights, cell, nx * ny * nz) # (n_cells, 8) +``` + +Negative keys drop the value; a 2D `values` is summed column by column. Measure +both strategies on a real case before choosing: atomics are simpler and often +fast enough, sort-then-reduce is deterministic (the summation order does not +depend on thread scheduling). + ## Guidelines * Create the mirror once and keep it with the owner of the buffer. Allocation @@ -111,5 +129,7 @@ host kernel writes into it directly, and `to_host()` does nothing. contents; then call `to_device()` first if the host side changed. * Atomics on one hot cell serialize. If most particles hit few cells, consider sorting particles by cell or accumulating per block in shared memory first. + For a single value (a total charge, an energy), `cunumpy_block_sum_to` from + `cunumpy/reduce.cuh` makes one atomic add per block instead of one per thread. * Floating-point atomics make the summation order non-deterministic. Results differ between runs in the last bits; compare with a tolerance in tests. diff --git a/docs/source/kernels/arguments.md b/docs/source/kernels/arguments.md index 973ecdf..ce3e163 100644 --- a/docs/source/kernels/arguments.md +++ b/docs/source/kernels/arguments.md @@ -121,7 +121,7 @@ push(value, 0.1, n_threads=x.size) ``` Fields may be scalars, pointers to scalar types (or `void*`), and array views -`Array1D` to `Array3D`. Packing checks every field like a kernel +`Array1D` to `Array4D`. Packing checks every field like a kernel argument: pointers need C-contiguous CuPy arrays of the declared dtype, scalars are range-checked and cast. Adding a field means editing the one Python definition; kernels that use the struct pick it up. @@ -170,14 +170,95 @@ push(args, dt, n_threads=args.n_markers) * The `CudaStruct` is built when the class is defined (`CudaMarkerArguments.struct`), so a bad field type raises at import, not at the first launch. -* `pack()` checks every field like `CudaStruct` does. Call it again after - replacing an array attribute, e.g. after the marker array was resized. +* `pack()` checks every field like `CudaStruct` does. +* The struct is packed again automatically, at the next launch, when a field + attribute changed: another array (address, shape or strides) or another + scalar value. Make the fields properties when the arrays belong to another + object that replaces them, e.g. a particle container that resizes its marker + array: + + ```python + class CudaMarkerArguments(xp.CudaStructArguments): + struct_name = "MarkerArgs" + fields = (("markers", "Array2D"), ("n_markers", "int")) + + def __init__(self, particles): + self._particles = particles + self.pack() + + @property + def markers(self): + return self._particles.markers + + @property + def n_markers(self): + return self._particles.markers.shape[0] + ``` + + Without the properties, the object keeps the old array alive and kernels + keep working on it: assign the new array to the attribute instead. * Copies and unpickled objects are packed again from their own arrays, so a `deepcopy` never points at the device memory of the original. * When the same call site must also reach a host kernel, pair it with the host argument object in a `KernelArguments`: `__host_args__()` returns the host object, `__cuda_args__()` returns `cuda_args.__cuda_args__()`. +### `PyccelStructArguments`: a pyccel host class and a struct + +When the host kernels take a pyccel-compiled argument class, that class cannot +inherit from `CudaStructArguments` (or anything else). `PyccelStructArguments` +holds it instead: the same object is passed to a `Kernel` on both backends, and +arrives as the pyccel object on the host path and as the struct on the device +path: + +```python +from my_sim.kernel_arguments import pusher_args_kernels # compiled by pyccel + + +class MarkerArguments(xp.PyccelStructArguments): + struct_name = "MarkerArgs" + fields = (("markers", "Array2D"), ("Np", "long long"), ("n_markers", "int")) + host_class = pusher_args_kernels.MarkerArguments + host_fields = ("markers", "Np") # its constructor arguments, in order + + def __init__(self, markers, Np): + self.markers = markers # NumPy or CuPy, whatever the owner has + self.Np = Np + self.n_markers = markers.shape[0] + if self.has_device_arrays(): + self.pack() # fail early on a bad device array + + +args = MarkerArguments(particles.markers, Np) +push(args, dt, n_threads=args.n_markers) # Kernel: same call on both backends +``` + +* `__host_args__()` builds `host_class(*host_fields)` once and again when one of + those attributes was replaced (a resized array, a changed scalar). The host + object is not pickled; a copy or an unpickled object rebuilds it. +* On the CuPy backend the attributes are device arrays, and there is no host + form: `__host_args__()` raises. Set `host_copies = True` on a class whose host + kernels only *read* the arrays (an evaluation, not a push): the host object + is then built from host copies, which `count_transfers()` reports, and what + the host kernel writes is not copied back. +* Objects holding host arrays are copied and pickled without packing; the + struct is only built from device arrays. + +### Check the layout against the compiler + +Values are packed with the NumPy dtype of the struct. If the compiler lays the +struct out differently (a hand-edited header, a `#pragma pack`, another +compiler such as hiprtc), kernels read fields at the wrong offsets without any +error. One GPU test per struct catches it: + +```python +def test_marker_args_layout(): + CudaMarkerArguments.struct.verify_layout() # the Python declaration + CudaMarkerArguments.struct.verify_layout( # the committed header + "marker_args.cuh", include_dirs=["kernels"] + ) +``` + ### Generate the struct from the host argument class If the host kernels already use an annotated argument class (Pyccel style), the @@ -208,6 +289,19 @@ C types, `"float[:, :]"` to `Array2D` (1 to 3 dimensions). `Final[...]` and `const` are ignored. `scalar_names={"float": "float"}` switches to single precision. Parameters without a mappable annotation raise `ValueError`. +If the module of the argument class is compiled by pyccel, importing it gives +the compiled class, whose `__init__` has no Python signature. +`CudaStruct.from_pyccel_class(path, class_name, name)` parses the `.py` source +with `ast` instead (never importing it), names each field after the attribute +the parameter is stored in (`self.first_init_idx = first_pusher_idx` gives a +field `first_init_idx`), and skips the parameters in `exclude=`: + +```python +MarkerArgs = xp.CudaStruct.from_pyccel_class( + "my_sim/kernel_arguments/pusher_args_kernels.py", "MarkerArguments", "MarkerArgs" +) +``` + Array view fields accept non-contiguous arrays, and the kernel indexes them like the host kernel does: diff --git a/docs/source/kernels/cuda-kernel.md b/docs/source/kernels/cuda-kernel.md index 195789a..b109f6f 100644 --- a/docs/source/kernels/cuda-kernel.md +++ b/docs/source/kernels/cuda-kernel.md @@ -53,8 +53,8 @@ every call: | `double*`, `const int*`, ... | C-contiguous CuPy array of exactly that dtype | NumPy arrays (`TypeError`, never copied), other dtypes, non-contiguous views | | `void*` | C-contiguous CuPy array of any dtype | host arrays, views | | `double`, `float`, `complex` | Python `int`/`float`, NumPy scalars that cast safely | strings, arrays, unsafe casts (`np.float64` into `float`) | -| `int`, `long long`, `size_t`, `int64_t`, ... | Python `int` in range, `bool`, matching NumPy integers | out-of-range values (`OverflowError`), floats | -| `Array1D` ... `Array3D` | CuPy array of dtype `T` and that ndim, contiguous or not | wrong dtype or ndim | +| `int`, `long long`, `size_t`, `int64_t`, ... | Python `int` and NumPy integers whose value is in range, `bool` | out-of-range values (`OverflowError`), floats | +| `Array1D` ... `Array4D` | CuPy array of dtype `T` and that ndim, contiguous or not | wrong dtype or ndim | | a `CudaStruct` type | a value of that struct | anything else | C types map to NumPy dtypes as on 64-bit Linux: `int` is `int32`, `long` and @@ -154,8 +154,9 @@ Pass extra include directories with `include_dirs=[...]` and NVRTC flags with | Header | Provides | | --- | --- | | `` | `CUNUMPY_THREAD_1D(i, n)`, `_2D`, `_3D`, `CUNUMPY_GRID_STRIDE_1D(i, n)` | -| `` | strided views `Array1D`, `Array2D`, `Array3D` | +| `` | strided views `Array1D` to `Array4D` | | `` | `cunumpy_atomic_add` and indexed 2D/3D variants, see [Accumulation kernels](accumulation.md) | +| `` | counter-based random numbers `cunumpy_uniform(seed, stream, counter)`, `cunumpy_normal2(...)`, equal to `xp.philox_uniform` on the host | `xp.cuda_include_dir()` returns their directory for use with other compilers. diff --git a/docs/source/kernels/debugging.md b/docs/source/kernels/debugging.md index e952e69..1b702ff 100644 --- a/docs/source/kernels/debugging.md +++ b/docs/source/kernels/debugging.md @@ -26,11 +26,14 @@ In debug mode a `CudaKernel`: * is compiled with `-lineinfo` (source lines for `compute-sanitizer` and profilers) and `-DCUNUMPY_BOUNDS_CHECK`, which turns on bounds checks in - `Array1D`/`Array2D`/`Array3D` views (an out-of-bounds index prints the index + `Array1D` to `Array4D` views (an out-of-bounds index prints the index and shape, then traps); * synchronizes after every launch, so a failure raises at the launch that caused it, as a `RuntimeError` naming the kernel and its grid and block, with - the CUDA error chained. + the CUDA error chained. While the stream is being captured into a CUDA graph + (`stream.begin_capture()`), the synchronization is skipped, since it would + invalidate the capture; errors of captured kernels surface when the graph is + launched, so synchronize after `graph.launch()` to see them. ```python with xp.cuda_debug(): @@ -79,6 +82,21 @@ that differs (see [Testing kernels](testing.md)). For `__device__` helpers, `device_function_kernel()` exposes a single function to Python so it can be compared value by value with its host version. +## NaN hunting: `check_finite` + +A kernel that reads a wrong index usually produces a NaN or inf that surfaces +many steps later. With `check_finite`, every launch is synchronized and the +floating-point arrays among its arguments (including the array fields of +struct argument objects) are scanned afterwards: + +```python +catalog["push"].cuda_kernel.check_finite = True +# RuntimeError: kernel 'push' left a NaN or inf in argument 0.markers (dtype float64, shape (1000, 8)) +``` + +It costs a synchronization and a pass over the arrays per launch, so switch it +on for the kernel under suspicion, not in production. + ## Common causes | Symptom | Likely cause | diff --git a/docs/source/kernels/dispatch.md b/docs/source/kernels/dispatch.md index 43d541f..5ea0151 100644 --- a/docs/source/kernels/dispatch.md +++ b/docs/source/kernels/dispatch.md @@ -123,6 +123,77 @@ A catalog is a read-only mapping: `catalog["push"]`, `"push" in catalog`, `len(catalog)`, iteration over names. `KernelCatalog()` plus `catalog.register(kernel)` builds one by hand. +## Compiled Pyccel host kernels + +By default the host kernel is the Python function itself, which is fine for +NumPy-vectorized code but slow for the loops of a Pyccel kernel. cunumpy does +not compile Pyccel code, but `from_package` can call your compile function on +each host kernel module, on the kernel's first call, and fall back to another +version when that fails (no Pyccel, no compiler): + +```python +import cunumpy as xp +import pyccel + +from my_sim.kernels.numpy_versions import NUMPY_VERSIONS # {"push": push_numpy, ...} + + +def compile_kernels(module): + return pyccel.epyccel(module, language="c") # add a cache in real code + + +catalog = xp.KernelCatalog.from_package( + __name__, + host_suffix="_pyccel", # push/push_pyccel.py next to push/push_cuda.cu + compile_host=compile_kernels, + host_fallback=NUMPY_VERSIONS, +) +``` + +Each host kernel is a `cunumpy.CompiledHostKernel`: +`catalog["push"].host_kernel.kernel.compiled` reports whether the compiled +version is available. Note that `epyccel` compiles again on every call; a +code that compiles at run time usually keeps the builds in an on-disk cache +keyed on the module source, so that only the first run after an edit compiles. + +Packages that compile their kernels at install time with the `pyccel` command +do not need `compile_host`: the compiled modules are imported like the Python +ones. + +## Choosing the kernel by where the arrays are + +`Kernel` picks the CUDA kernel on the CuPy backend. Real codes also hand host +arrays to kernels while CuPy is active: a diagnostic on a copy, a buffer staged +for MPI, a path that has no GPU version yet. With `dispatch="arrays"` such calls +run the host kernel: + +```python +catalog = xp.KernelCatalog.from_package(__name__, dispatch="arrays") + +catalog["gather"](positions, field, result, n_threads=n) # CUDA if positions are CuPy +catalog["gather"](host_positions, host_field, host_result) # host kernel, also on CuPy +``` + +The CUDA kernel runs if any top-level argument is on the GPU: a CuPy array, or +a device-only argument object (`CudaArguments`, a struct value). A +`KernelArguments` object has both forms and does not decide on its own. + +## Same parameters on both sides + +The call site is the same for both kernels only if they take the same +parameters in the same order. Check it once, in a test: + +```python +def test_kernel_signatures(): + catalog.check_signatures() +``` + +It compares the parameter names of each host function (for a compiled host +kernel, its Python source, or the `__pyccel__/.pyi` stub that pyccel +writes next to a compiled extension module) with the parsed `__global__` +signature, and lists every kernel that differs, e.g. +`kernel 'push': the host kernel takes (x, v, dt, n), the CUDA kernel (x, v, n, dt)`. + ## Porting status and setup ```python @@ -146,6 +217,23 @@ All kernels are compiled even if one fails; the first error is raised afterwards. On later runs CuPy loads the binaries from its disk cache, so this is fast. +`from_package()` also warns about kernel names that pyccel's Fortran backend +cannot compile: the wrapper module `bind_c__kernels` must fit Fortran's +63-character limit, so a kernel name has at most 48 characters +(`check_name_length=False` silences it, e.g. for the C backend). + +## Launch size from the arguments + +When the number of threads is a property of an argument (the marker count of +a particle argument object), set it once instead of at every call site: + +```python +catalog["push"].cuda_kernel.n_threads_from = lambda args: args[0].n_markers +push(args_markers, dt) # no n_threads needed, on either backend +``` + +An explicit `n_threads` or `grid` still wins. + ## Recommended project layout * One folder per kernel, the host kernel and its CUDA port side by side, so a @@ -155,6 +243,11 @@ is fast. `device_function_kernel` ([Testing kernels](testing.md)). * Generated struct headers (`write_cuda_header`) committed next to the kernels, with a test that they are up to date. -* One parametrised parity test over `catalog.parity_cases()`. +* One parametrised parity test over `catalog.parity_cases()`, and + `catalog.check_signatures()` in a test. +* Declare the `.cu` and `.cuh` files as package data (e.g. + `[tool.setuptools.package-data] my_sim = ["kernels/*/*_cuda.cu", + "kernels/*.cuh"]`); otherwise wheels contain the host kernels but no CUDA + sources, and every kernel looks unported after `pip install`. * `missing_cuda="fallback"` while porting, `"raise"` once the time loop is fully ported. diff --git a/docs/source/kernels/testing.md b/docs/source/kernels/testing.md index 145a168..d7ff0e3 100644 --- a/docs/source/kernels/testing.md +++ b/docs/source/kernels/testing.md @@ -98,7 +98,10 @@ Things to know: * **Which arguments are compared**: `outputs=(2,)` selects them by index; otherwise the host kernel's declared `outputs` are used, and if there are none, every array argument. Arrays held by argument objects (one level deep, - e.g. a `CudaArguments` object or a list) are compared too. + e.g. a `CudaArguments` object or a list) are compared too. A + `CudaStructArguments` object is read through its struct fields, so its + arrays get the names of the host argument object's attributes + (`argument 0.markers`), also when the fields are properties. * **Tolerances**: the default `rtol=1e-12` suits deterministic kernels. Kernels with atomics or a different summation order need looser tolerances, for example `rtol=1e-10, atol=1e-14`. @@ -124,6 +127,85 @@ def test_parity(name, kernel): ported kernel is tested as soon as its `.cu` file exists (it needs an entry in `MAKE_ARGS`, which fails loudly with a `KeyError` if forgotten). +### Test arguments next to the kernel + +Instead of one `make_args` per kernel in the test file, each kernel folder can +hold `_test_args.py`: + +```python +# my_sim/kernels/push/push_test_args.py +import numpy as np + +import cunumpy as xp + +N_THREADS = 1000 # or GRID; also BLOCK, RTOL, ATOL, N_CALLS, OUTPUTS, SEED + + +def make_args(backend, seed): + x = xp.to_cunumpy(np.random.default_rng(seed).random(1000)) + return (x, 2.0, x.size) +``` + +`KernelCatalog.from_package()` records these modules (imported only when a +test asks for them), and the parity test of the whole package becomes + +```python +from cunumpy.testing import check_parity, parity_cases + + +@pytest.mark.parametrize("kernel", parity_cases(catalog)) +def test_parity(kernel): + check_parity(kernel) +``` + +A kernel with a CUDA version but no test-arguments module shows up as skipped, +with the name of the missing file in the reason, so the report lists what is +left to do. `N_THREADS` may be a function of the argument tuple +(`lambda args: args[0].shape[0]`), or be omitted when the CUDA kernel has +`n_threads_from`. Keyword arguments of `check_parity` override the module. + +## Test CUDA kernels without a GPU: `emulate_cuda_kernel` + +On a CPU-only CI runner the parity tests are skipped, so nothing checks the +CUDA kernels' index and weight arithmetic. `emulate_cuda_kernel` runs a kernel +on the CPU instead: the source is compiled as C++ with the CUDA built-ins +replaced, and the kernel is called once per thread, one thread after another. +Compare it with the host kernel: + +```python +import numpy as np +import pytest + +from cunumpy.testing import emulate_cuda_kernel, emulation_compiler + +from my_sim.kernels import catalog + +pytestmark = pytest.mark.skipif(emulation_compiler() is None, reason="no C++ compiler") + + +def test_gather_cuda_arithmetic(): + rng = np.random.default_rng(0) + positions, field = rng.random((500, 2)), rng.normal(size=(17, 9, 2)) + expected, result = np.zeros((500, 2)), np.zeros((500, 2)) + catalog["gather"].host_kernel(positions, field, expected, 0.0, 0.0, 0.06, 0.11, 17, 9) + emulate_cuda_kernel( + catalog["gather"].cuda_kernel, + positions, field, result, 0.0, 0.0, 0.06, 0.11, 17, 9, + n_threads=500, + ) + np.testing.assert_allclose(result, expected, rtol=1e-12, atol=1e-14) +``` + +Arrays are passed as NumPy arrays (any strides) and written back; scalars are +checked like in a launch. It catches wrong indices, clamping, periodic wrapping +and weights, i.e. most porting bugs of gather, scatter and push kernels. Block +shared memory and `__syncthreads` are emulated (pass `shared_mem=` for +`extern __shared__` arrays), so per-block deposits and shared-memory reductions +are covered too. It does not emulate concurrency between barriers or warp +intrinsics; kernels using the latter are refused with `NotImplementedError`, so +those still need a GPU run. The compiler may fuse multiply-adds as NVRTC does, +so compare with a tolerance of a few ulp. + ## Test `__device__` helpers: `device_function_kernel` Helpers such as B-spline evaluation or coordinate maps are `__device__` @@ -166,10 +248,46 @@ def test_find_span_matches_host(): In the generated kernel, pointer parameters are passed unchanged to every thread (shared data), scalar parameters become per-thread arrays, the return -value of thread `i` goes to `out[i]`, and `n` is the number of elements. Extra +value of thread `i` goes to `out[i]`, and `n` is the number of elements. A +struct parameter, by value or by `const` reference (`const DomainArgs& d`), +is passed through unchanged as well; give the struct types in `structs=` +and pass a `CudaStructArguments` object or a packed value, so helpers that +take the argument structs of the kernels are tested the same way. Extra keyword arguments (`include_dirs`, `includes`, `block_size`) go to the `CudaKernel`. +## Run the CuPy code paths without a GPU: the fake CuPy + +`emulate_cuda_kernel` covers the kernels. Everything around them (argument +objects built from device arrays, struct packing, `as_device_array`, +transfer counting, the backend branches of a simulation) runs only with CuPy +present. For CI machines without a GPU, cunumpy ships a strict stand-in: + +```bash +CUNUMPY_FAKE_CUPY=1 ARRAY_BACKEND=cupy pytest tests/ +``` + +Its arrays live in host memory but are not NumPy arrays: `numpy.asarray(a)` +raises (as it does for real CuPy arrays, so a compiled host kernel rejects +them), CuPy functions reject NumPy arrays and lists, mixing the two raises, +reductions return 0-d arrays, and arrays have `data.ptr`, `device` and +`__cuda_array_interface__`. Kernels cannot run: `RawKernel` raises +`NotImplementedError`, `requires_cupy` skips and `assert_kernels_agree` +skips while the fake is active (`cunumpy.testing.fake_cupy_active()`). It +can also be installed from code, before the first backend use: + +```python +# conftest.py +from cunumpy.testing import install_fake_cupy + +install_fake_cupy() +``` + +Most host/device bugs (a NumPy array reaching a device argument object, a +device array reaching SciPy or MPI, a missing `xp.asarray`) show up this way +long before the code reaches a GPU. The fake is never installed when the +real CuPy is importable. + ## Test that a step stays on the device ```python @@ -197,7 +315,12 @@ offsets. ## CI setup * Run the suite on a normal CPU runner: everything on NumPy runs, GPU cases - are reported as skipped. + are reported as skipped, and `emulate_cuda_kernel` tests check the CUDA + kernels' arithmetic (the runner needs a C++ compiler, which Linux images + have). +* Run it a second time with `CUNUMPY_FAKE_CUPY=1 ARRAY_BACKEND=cupy`, so the + CuPy code paths are exercised on the CPU runner too (kernel launches are + skipped). * Run the same suite on a GPU runner, optionally with `CUNUMPY_CUDA_DEBUG=1` so kernels are bounds-checked and errors are attributed to the right launch. * Every so often, run the GPU suite under `compute-sanitizer` (see [Debugging diff --git a/docs/source/troubleshooting.md b/docs/source/troubleshooting.md index 245ae20..afbfa10 100644 --- a/docs/source/troubleshooting.md +++ b/docs/source/troubleshooting.md @@ -70,10 +70,24 @@ launch, then use `compute-sanitizer` ([Debugging CUDA kernels](kernels/debugging.md)). Restart the process afterwards; the CUDA context is unusable. -**A change in a `.cuh` header is ignored.** Headers included with quotes are -hashed into the compile options and trigger recompilation; headers included -with angle brackets are not. Also, a kernel object compiles once per process: -restart the process after editing. +**A change in a `.cuh` header is ignored.** Headers included with quotes, and +cunumpy's own headers in either form, are hashed into the compile options and +trigger recompilation; other headers included with angle brackets are not. +Also, a kernel object compiles once per process: restart the process after +editing. + +**After `pip install`, every kernel has no CUDA version.** The `.cu` files are +not package data, so the wheel contains only the Python files. Declare them, +e.g. `[tool.setuptools.package-data] my_sim = ["kernels/*/*_cuda.cu", +"kernels/*.cuh"]`, and check the wheel's contents. + +**`TypeError` for a CUDA kernel called with host arrays on the CuPy backend.** +The code hands host arrays to a kernel while CuPy is active. Create the kernels +with `dispatch="arrays"` so such calls run the host kernel. + +**`ValueError: ... the host kernel takes (...), the CUDA kernel (...)`.** From +`check_signature()`: the two versions of a kernel take different parameters, +so one of them reads its arguments in the wrong order. Make the lists equal. **The first time step is much slower.** Kernels are compiled on first call. Compile at setup with `catalog.compile_all()` or `kernel.compile()`; later runs diff --git a/pyproject.toml b/pyproject.toml index ba99ffd..43a8bda 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -43,7 +43,7 @@ optional-dependencies.docs = [ "sphinx", "sphinx-book-theme", ] -optional-dependencies.test = [ "coverage", "pytest" ] +optional-dependencies.test = [ "coverage", "pytest", "scipy" ] optional-dependencies.test-compiled = [ "cunumpy[test]", "pyccel" ] urls."Source" = "https://github.com/max-models/cunumpy" @@ -55,3 +55,7 @@ cunumpy = [ "py.typed", "*.pyi", "LLM_GUIDE.md", "cuda/include/cunumpy/*.cuh" ] [tool.isort] profile = "black" + +[tool.ruff] +# agent worktrees and scratch scripts are local copies, not part of the project +extend-exclude = [ ".claude" ] diff --git a/src/cunumpy/LLM_GUIDE.md b/src/cunumpy/LLM_GUIDE.md index 95b50f2..72a80cf 100644 --- a/src/cunumpy/LLM_GUIDE.md +++ b/src/cunumpy/LLM_GUIDE.md @@ -66,10 +66,24 @@ https://max-models.github.io/cunumpy/ and in `docs/source/` of the repository. | launch a hand-written CUDA C kernel | `xp.CudaKernel(source, "name")` / `CudaKernel.from_file(path)` | | host kernel + CUDA port, chosen by backend | `xp.Kernel(host_fn, cuda_kernel_or_None)` | | many kernels in a package, ported incrementally | `xp.KernelCatalog.from_package(__name__, missing_cuda="fallback")` | +| host kernels compiled at first call (your compile function), NumPy fallback | `from_package(..., host_suffix="_pyccel", compile_host=my_compile, host_fallback={...})` -> `xp.CompiledHostKernel` | +| host arrays reach kernels while CuPy is active | `Kernel(..., dispatch="arrays")` / `from_package(..., dispatch="arrays")`: CUDA only for device arguments | +| check host and CUDA kernels take the same parameters | `catalog.check_signatures()` (in a unit test) | +| test a CUDA kernel's arithmetic without a GPU | `cunumpy.testing.emulate_cuda_kernel(kernel, *numpy_args, n_threads=n)` (C++ compiler; shared memory and __syncthreads ok, no warp ops; `shared_mem=` for extern shared) | +| shared-memory budget of a block | `xp.max_shared_memory_per_block()` (48 KiB without a GPU) | +| random numbers inside a kernel, equal on the host | `#include `: `cunumpy_uniform(seed, particle_id, step)`; host: `xp.philox_uniform(seed, ids, step)` | +| one thread per marker without passing n_threads | `CudaKernel(..., n_threads_from="first_array")` | +| copy device arrays to the host for output without stalling | `xp.HostStaging(shape, dtype)`: `c = staging.copy(a)` ... `c.result()` | +| PIC recipes (compaction, sort by cell, MPI exchange, graphs) | docs guide "Particle codes" | +| reproducible random numbers per MPI rank | `xp.random_streams.seed(seed, rank=rank)`, then `xp.random_streams.normal(...)` / `.generator()` | | group arrays/scalars into one kernel argument | `xp.CudaArguments` (device only), `xp.KernelArguments` (host object + device tuple), `xp.CudaStruct` (C struct), `xp.CudaStructArguments` (C struct as a class) | | CUDA struct from a Pyccel argument class | `xp.CudaStruct.from_signature(Cls.__init__, "Name")`, `xp.write_cuda_header(...)` | +| SciPy (sparse, sparse.linalg, fft, special, ndimage, ...) on either backend | `xp.scipy..` (SciPy or `cupyx.scipy`); `xp.scipy.special.available(name)` | +| chain of elementwise operations as one GPU kernel | `@xp.fuse` (`cupy.fuse` for CuPy arrays, plain call otherwise) | +| PETSc solve on device arrays without copies | `xp.petsc_vec(array)` (CUDA/HIP petsc4py for CuPy arrays); `xp.synchronize()` around PETSc calls | +| reduction inside a CUDA kernel (energy, max velocity) | ``: `cunumpy_block_sum_to(out, v)`, `cunumpy_block_min/max`, `cunumpy_warp_sum` | | kernel writes into a host buffer owned by another library | `xp.DeviceMirror(host_array)` + `` | -| N-D indexing in CUDA, non-contiguous arrays | `Array1D`..`Array3D` params from `` | +| N-D indexing in CUDA, non-contiguous arrays | `Array1D`..`Array4D` params from `` | | one MPI rank per GPU | `bind_local_device()` → `from mpi4py import MPI` → `require_cuda_aware_mpi()` → `synchronize_for_mpi(...)` before each call | | timing GPU code | `with xp.timed_region("name") as t:` → `t.elapsed` | | profiler markers | `xp.nvtx_range("name")` (context manager or decorator) | @@ -113,6 +127,17 @@ with xp.assert_no_transfers(): ... # AssertionError with report if anything c Only transfers through cunumpy are counted (not raw `cupy.asarray`, `.get()`, `float(device_scalar)`, or `DeviceMirror.to_host()/to_device()`). +MPI, accumulation and versions: + +```python +xp.mpi_is_cuda_aware(comm) # collective, once at startup; remembered +with xp.mpi_buffer(a) as buf: comm.Send(buf, ...) # host array, CUDA-aware device +with xp.mpi_buffer(a, send=False, recv=True) as buf: ... # array, or pinned staging copy +xp.set_mpi_cuda_aware(True | False | None), xp.get_mpi_cuda_aware() +xp.segment_sum(values, keys, n_segments) # out[k] = sum(values[keys == k]); keys < 0 dropped +xp.require_version("0.4.0") # ImportError if cunumpy is older +``` + Random numbers and dtypes: ```python @@ -190,7 +215,8 @@ xp.cuda_include_dir() `double`=float64, `complex`=complex128, `bool`=bool; `int64_t`, `size_t` etc. supported. * Scalars: Python `int`/`float`/`bool` are cast with range checks - (`OverflowError`); NumPy scalars must match or cast safely. + (`OverflowError`); NumPy integers are checked by value too (`np.int64(5)` + fits `int`); other NumPy scalars must match or cast safely. * Pointer params: C-contiguous CuPy arrays, exact dtype (`void*` any). `ArrayND` params: CuPy arrays of dtype T and ndim N, any strides. * Compiled lazily with NVRTC on first call; cached on disk by CuPy; quoted @@ -201,7 +227,7 @@ Shipped CUDA headers (always on the include path): ```c #include // CUNUMPY_THREAD_1D(i, n) /_2D/_3D, CUNUMPY_GRID_STRIDE_1D(i, n) {...} -#include // Array1D..Array3D: data, shape[], strides[] (elements), a(i, j), size() +#include // Array1D..Array4D: data, shape[], strides[] (elements), a(i, j), size() #include // cunumpy_atomic_add(double*|float*, v), _2d(data, n1, i, j, v), _3d(...) ``` @@ -240,6 +266,7 @@ class Args(xp.KernelArguments): # one object, host form + device form S = xp.CudaStruct("S", [("x", "double*"), ("n", "long long"), ("a", "Array2D")]) S.declaration; S.dtype; S.to_header(path); value = S(x=..., n=..., a=...) +S.verify_layout() # GPU test: compiler layout == S.dtype (also verify_layout("hdr.cuh")) S = xp.CudaStruct.from_signature(Cls.__init__, "S", int_type="long long") xp.write_cuda_header("args.cuh", [S1, S2]) @@ -248,14 +275,15 @@ class A(xp.CudaStructArguments): # the struct as a class; A.struct is th fields = (("x", "double*"), ("n", "int")) def __init__(self, x): self.x, self.n = x, x.shape[0] - self.pack() # again after replacing an array; copies repack + self.pack() # repacks itself when a field changes; copies repack xp.CudaKernel(S.declaration + src, "k", structs=[S]) xp.resolve_host_args(args, kwargs) ``` Only top-level arguments are resolved. Cache both forms lazily and invalidate them when the underlying arrays are replaced. A packed struct holds device -addresses: re-pack after replacing an array. +addresses: re-pack after replacing an array (a `CudaStructArguments` does this +itself at the next launch; make its fields properties to follow an owner's arrays). `DeviceMirror`: @@ -294,6 +322,24 @@ assert_kernels_agree(kernel, make_args, *, n_threads=None, grid=None, block=None k = device_function_kernel(header_source, "int f(const double* t, int p, double x)") k(t, p_array, x_array, out, n, n_threads=n) # scalars become per-thread arrays +k = device_function_kernel(src, "double g(const DomainArgs& d, double x)", structs=[DomainArgs]) + +# /_test_args.py: make_args(backend, seed) + N_THREADS (or GRID), RTOL, ... +from cunumpy.testing import parity_cases, check_parity +@pytest.mark.parametrize("kernel", parity_cases(catalog)) # skip-marked if no test args +def test_parity(kernel): check_parity(kernel) + +# without a GPU: CUNUMPY_FAKE_CUPY=1 ARRAY_BACKEND=cupy pytest (fake CuPy: strict host +# stand-in, no kernel launches; fake_cupy_active(); requires_cupy skips) + +# argument classes with a pyccel host class: one object on both backends +class MarkerArguments(xp.PyccelStructArguments): + struct_name = "MarkerArgs"; fields = (("markers", "Array2D"), ("Np", "long long")) + host_class = pusher_args_kernels.MarkerArguments # pyccel class; cannot inherit + host_fields = ("markers", "Np") # its constructor args, in order +MarkerArgs = xp.CudaStruct.from_pyccel_class("pusher_args_kernels.py", "MarkerArguments", "MarkerArgs") +kernel.n_threads_from = lambda args: args[0].n_markers # launch size from an argument +kernel.check_finite = True # NaN/inf after each launch (debug) ``` ## Canonical patterns diff --git a/src/cunumpy/__init__.py b/src/cunumpy/__init__.py index 04d8b39..4aa6692 100644 --- a/src/cunumpy/__init__.py +++ b/src/cunumpy/__init__.py @@ -1,4 +1,5 @@ # cunumpy/__init__.py +import re as _re from importlib.metadata import PackageNotFoundError, version from . import xp @@ -11,6 +12,7 @@ CudaStruct, CudaStructArguments, CudaStructValue, + PyccelStructArguments, ctype_of, cuda_include_dir, cuda_kernel_names, @@ -20,8 +22,20 @@ write_cuda_header, ) from .dispatch import Kernel, KernelCatalog -from .kernel import KernelArguments, PyccelKernel, resolve_host_args +from .fusion import fuse +from .kernel import CompiledHostKernel, KernelArguments, PyccelKernel, resolve_host_args from .mirror import DeviceMirror +from .petsc import petsc_vec +from .philox import ( + philox4x32_10, + philox_normal, + philox_normal2, + philox_uniform, + philox_uniform2, +) +from .random_streams import RandomStreams, random_streams +from .scipy_backend import scipy +from .staging import HostStaging, StagedCopy from .transfers import ( TransferCounter, TransferEvent, @@ -29,6 +43,7 @@ count_transfers, ) from .xp import ( + DEFAULT_SHARED_MEMORY_PER_BLOCK, Timing, as_device_array, assert_same_backend, @@ -42,20 +57,25 @@ get_array_module, get_backend, get_cuda_debug, + get_mpi_cuda_aware, get_rng, is_cpu, is_gpu, local_rank, + max_shared_memory_per_block, memory_info, + mpi_buffer, mpi_is_cuda_aware, nvtx_range, pin_memory, require_cuda_aware_mpi, same_backend, + segment_sum, set_backend, set_cuda_debug, set_device, set_device_for_rank, + set_mpi_cuda_aware, stream, synchronize, synchronize_for_mpi, @@ -71,8 +91,38 @@ except PackageNotFoundError: __version__ = "0.0.0+unknown" + +def _version_key(text: str) -> tuple[int, ...]: + return tuple(int(part) for part in _re.findall(r"\d+", text.split("+")[0])[:3]) + + +def require_version(minimum: str) -> None: + """Raise ``ImportError`` if this cunumpy is older than `minimum`. + + For projects that depend on a feature of a given release, as a clearer + error than an ``AttributeError`` later:: + + import cunumpy as xp + + xp.require_version("0.4.0") + + Only the numeric part of the versions is compared (``0.4.0`` and + ``0.4.0.dev1`` compare equal). Nothing is checked when the installed + version is unknown (cunumpy not installed as a package). + """ + if __version__.startswith("0.0.0+unknown"): + return + if _version_key(__version__) < _version_key(minimum): + raise ImportError( + f"cunumpy {minimum} or newer is required, but {__version__} is " + "installed: pip install --upgrade cunumpy" + ) + + __all__ = [ "DEBUG_OPTIONS", + "DEFAULT_SHARED_MEMORY_PER_BLOCK", + "CompiledHostKernel", "CudaArguments", "CudaKernel", "CudaKernelVariants", @@ -81,10 +131,14 @@ "CudaStructArguments", "CudaStructValue", "DeviceMirror", + "HostStaging", "Kernel", "KernelArguments", "KernelCatalog", "PyccelKernel", + "PyccelStructArguments", + "RandomStreams", + "StagedCopy", "Timing", "TransferCounter", "TransferEvent", @@ -103,29 +157,44 @@ "default_float_dtype", "device_count", "free_memory", + "fuse", "get_array_backend", "get_array_module", "get_backend", "get_cuda_debug", + "get_mpi_cuda_aware", "get_rng", "include_hash", "is_cpu", "is_gpu", "local_rank", + "max_shared_memory_per_block", "memory_info", + "mpi_buffer", "mpi_is_cuda_aware", "numpy_backend", "nvtx_range", "parse_cuda_signature", + "petsc_vec", + "philox4x32_10", + "philox_normal", + "philox_normal2", + "philox_uniform", + "philox_uniform2", "pin_memory", + "random_streams", "require_cuda_aware_mpi", + "require_version", "resolve_host_args", "resolve_includes", "same_backend", + "scipy", + "segment_sum", "set_backend", "set_cuda_debug", "set_device", "set_device_for_rank", + "set_mpi_cuda_aware", "stream", "synchronize", "synchronize_for_mpi", diff --git a/src/cunumpy/__init__.pyi b/src/cunumpy/__init__.pyi index cf24ab4..f54f7ad 100644 --- a/src/cunumpy/__init__.pyi +++ b/src/cunumpy/__init__.pyi @@ -16,14 +16,28 @@ from .cuda_kernel import CudaParameter as CudaParameter from .cuda_kernel import CudaStruct as CudaStruct from .cuda_kernel import CudaStructArguments as CudaStructArguments from .cuda_kernel import CudaStructValue as CudaStructValue +from .cuda_kernel import PyccelStructArguments as PyccelStructArguments from .cuda_kernel import ctype_of as ctype_of from .cuda_kernel import cuda_include_dir as cuda_include_dir from .cuda_kernel import parse_cuda_signature as parse_cuda_signature from .cuda_kernel import write_cuda_header as write_cuda_header from .dispatch import Kernel as Kernel from .dispatch import KernelCatalog as KernelCatalog +from .fusion import fuse as fuse +from .kernel import CompiledHostKernel as CompiledHostKernel from .kernel import PyccelKernel as PyccelKernel from .mirror import DeviceMirror as DeviceMirror +from .petsc import petsc_vec as petsc_vec +from .philox import philox4x32_10 as philox4x32_10 +from .philox import philox_normal as philox_normal +from .philox import philox_normal2 as philox_normal2 +from .philox import philox_uniform as philox_uniform +from .philox import philox_uniform2 as philox_uniform2 +from .random_streams import RandomStreams as RandomStreams +from .random_streams import random_streams as random_streams +from .scipy_backend import scipy as scipy +from .staging import HostStaging as HostStaging +from .staging import StagedCopy as StagedCopy from .transfers import TransferCounter as TransferCounter from .transfers import TransferEvent as TransferEvent @@ -46,9 +60,25 @@ def set_device_for_rank(rank: int, devices_per_node: int | None = ...) -> int: . def local_rank() -> int: ... def bind_local_device() -> int | None: ... def synchronize_for_mpi(*arrays: Any) -> None: ... +def mpi_is_cuda_aware(comm: Any = ..., *, method: str = ...) -> bool: ... +def require_cuda_aware_mpi(comm: Any = ...) -> None: ... +def set_mpi_cuda_aware(value: bool | None) -> None: ... +def get_mpi_cuda_aware() -> bool | None: ... +@contextmanager +def mpi_buffer( + array: Any, *, send: bool = ..., recv: bool = ..., cuda_aware: bool | None = ... +) -> Generator[Any]: ... +def segment_sum(values: Any, keys: Any, n_segments: int) -> Any: ... +def require_version(minimum: str) -> None: ... def device_count() -> int: ... def memory_info() -> tuple[int, int] | None: ... def free_memory() -> None: ... +def max_shared_memory_per_block( + device: int | None = ..., *, opt_in: bool = ... +) -> int: ... + +DEFAULT_SHARED_MEMORY_PER_BLOCK: int + def pin_memory(array: Any) -> Any: ... @contextmanager def stream() -> Generator[Any]: ... diff --git a/src/cunumpy/_fake_cupy.py b/src/cunumpy/_fake_cupy.py new file mode 100644 index 0000000..d8eb8e5 --- /dev/null +++ b/src/cunumpy/_fake_cupy.py @@ -0,0 +1,97 @@ +"""A strict stand-in for CuPy on machines without a GPU, for tests only. + +With the real CuPy absent, the CuPy code paths of a program (argument objects +built from device arrays, host/device conversions, MPI buffers, backend +branches) cannot run in CI at all. This module installs a fake ``cupy`` +package whose arrays live in host memory but behave like CuPy arrays where it +matters for finding host/device bugs: + +* ``cupy.ndarray`` is **not** a ``numpy.ndarray``; ``numpy.asarray(a)`` raises + (use ``.get()``), so compiled host kernels reject these arrays exactly as + they reject real CuPy arrays; +* ``cupy.`` rejects NumPy arrays and Python lists as array arguments, + like CuPy does; mixing CuPy and NumPy arrays in arithmetic raises; +* reductions and scalar indexing return 0-d arrays, not Python scalars; +* arrays have ``data.ptr``, ``device`` and ``__cuda_array_interface__``, so + :class:`~cunumpy.CudaStruct` packing and the argument checks of + :class:`~cunumpy.CudaKernel` work; +* CUDA kernels cannot run: ``RawKernel`` and friends raise + ``NotImplementedError`` when called, and :func:`cunumpy.testing.requires_cupy` + skips tests while the fake is active. + +Activate it before CuPy or cunumpy's backend is first used, either with the +environment variable ``CUNUMPY_FAKE_CUPY=1`` (read when cunumpy is imported) +or by calling :func:`install` first thing in a test session (e.g. in +``conftest.py``). Then ``ARRAY_BACKEND=cupy`` or ``xp.set_backend("cupy")`` +selects the fake like the real thing. + +Never installed with a real CuPy present (``install`` raises), never shipped +as a top-level ``cupy`` package. +""" + +from __future__ import annotations + +import sys +import types +from pathlib import Path + +__all__ = ["install", "is_active", "uninstall"] + +_IMPLEMENTATION = Path(__file__).with_name("_fake_cupy_impl.py") +_SUBMODULES = ("cupy.cuda", "cupy.cuda.device", "cupy.cuda.runtime", "cupy.linalg") + + +def is_active() -> bool: + """Whether the fake CuPy is the ``cupy`` module of this process.""" + module = sys.modules.get("cupy") + return bool(getattr(module, "__cunumpy_fake__", False)) + + +def install() -> types.ModuleType: + """Install the fake ``cupy`` package into ``sys.modules`` and return it. + + Idempotent. Raises ``RuntimeError`` if the real CuPy has been imported + already, and ``RuntimeError`` if cunumpy has already decided that CuPy is + unavailable (import cunumpy after installing, or set the environment + variable instead). + + Returns + ------- + types.ModuleType + The fake ``cupy`` module. + """ + existing = sys.modules.get("cupy") + if existing is not None: + if getattr(existing, "__cunumpy_fake__", False): + return existing + raise RuntimeError( + "the real CuPy is already imported; the fake CuPy must be installed " + "before CuPy (set CUNUMPY_FAKE_CUPY=1 or call install() first)" + ) + xp = sys.modules.get("cunumpy.xp") + if xp is not None and getattr(xp, "_CUPY_AVAILABLE_CACHE", None) is not None: + raise RuntimeError( + "cunumpy has already checked for CuPy; install the fake CuPy before " + "the first backend use (set CUNUMPY_FAKE_CUPY=1 in the environment)" + ) + module = types.ModuleType("cupy") + module.__file__ = str(_IMPLEMENTATION) + module.__cunumpy_fake__ = True + code = compile(_IMPLEMENTATION.read_text(), str(_IMPLEMENTATION), "exec") + sys.modules["cupy"] = module + try: + exec(code, module.__dict__) # noqa: S102 - our own file, classes named cupy.* + except BaseException: + uninstall() + raise + return module + + +def uninstall() -> None: + """Remove the fake ``cupy`` package from ``sys.modules`` (no-op otherwise).""" + if not is_active(): + return + for name in [n for n in sys.modules if n == "cupy" or n.startswith("cupy.")]: + del sys.modules[name] + for name in [n for n in sys.modules if n.startswith("array_api_compat.cupy")]: + del sys.modules[name] diff --git a/src/cunumpy/_fake_cupy_impl.py b/src/cunumpy/_fake_cupy_impl.py new file mode 100644 index 0000000..7226868 --- /dev/null +++ b/src/cunumpy/_fake_cupy_impl.py @@ -0,0 +1,513 @@ +# The body of the fake ``cupy`` module, see cunumpy/_fake_cupy.py. It is executed +# into a fresh module named ``cupy`` (so classes defined here are ``cupy.ndarray`` +# etc.), never imported as cunumpy._fake_cupy_impl. + +import builtins +import operator +import sys as _sys +import types + +import numpy as _np + +__version__ = "0.0.0+cunumpy-fake" +__cunumpy_fake__ = True + +_HOST = _np.ndarray + + +def _err(obj, where=""): + return TypeError( + f"Unsupported type {type(obj)}{where} (fake CuPy: host arrays/lists are " + "not accepted)" + ) + + +class ndarray: + __slots__ = ("__weakref__", "_a") + __array_priority__ = 100 + + def __init__(self, a): + assert isinstance(a, _HOST) + object.__setattr__(self, "_a", a) + + # --- conversions --------------------------------------------------------- + def __array__(self, *args, **kwargs): + raise TypeError( + "Implicit conversion to a NumPy array is not allowed. Please use " + "`.get()` to construct a NumPy array explicitly." + ) + + def get(self, *args, **kwargs): + return self._a.copy() + + def set(self, arr, *args, **kwargs): + self._a[...] = arr + + def __reduce__(self): + return (_from_host, (self._a.copy(),)) + + @property + def __cuda_array_interface__(self): + a = self._a + return { + "shape": a.shape, + "typestr": a.dtype.str, + "data": (a.ctypes.data, False), + "version": 3, + "strides": None if a.flags.c_contiguous else a.strides, + } + + @property + def data(self): + return types.SimpleNamespace(ptr=self._a.ctypes.data) + + @property + def device(self): + return cuda.Device(0) + + # --- attributes ---------------------------------------------------------- + def __getattr__(self, name): + # no __array_interface__ etc.: NumPy must not see the host buffer + if name.startswith("__"): + raise AttributeError(name) + attr = getattr(self._a, name) + if callable(attr): + return _wrap_callable(attr, strict=True, name=f"ndarray.{name}") + return _wrap(attr) + + def __setattr__(self, name, value): + setattr(self._a, name, _unwrap(value)) + + # --- python protocol ----------------------------------------------------- + def __len__(self): + return len(self._a) + + def __iter__(self): + for x in self._a: + yield _wrap(x) + + def __repr__(self): + return "fakecupy." + repr(self._a) + + def __str__(self): + return str(self._a) + + def __format__(self, spec): + return format(self._a.item() if self._a.ndim == 0 else self._a, spec) + + def __bool__(self): + return bool(self._a) + + def __int__(self): + return int(self._a) + + def __float__(self): + return float(self._a) + + def __complex__(self): + return complex(self._a) + + def __index__(self): + return operator.index(self._a.item() if self._a.ndim == 0 else self._a) + + __hash__ = None + + def __contains__(self, x): + return _unwrap(x) in self._a + + def __getitem__(self, key): + return _wrap(self._a[_unwrap(key)]) + + def __setitem__(self, key, value): + self._a[_unwrap(key)] = _unwrap(value) + + def __copy__(self): + return ndarray(self._a.copy()) + + def __deepcopy__(self, memo): + return ndarray(self._a.copy()) + + # --- numpy protocols ----------------------------------------------------- + def __array_ufunc__(self, ufunc, method, *inputs, **kwargs): + _strict(inputs, ufunc.__name__) + out = kwargs.get("out") + res = getattr(ufunc, method)(*_unwrap(inputs), **_unwrap(kwargs)) + if out is not None: + return out[0] if isinstance(out, tuple) and len(out) == 1 else out + return _wrap(res) + + def __array_function__(self, func, types_, args, kwargs): + mine = globals().get(func.__name__) + if mine is None: + mine = __getattr__(func.__name__) + return mine(*args, **kwargs) + + +def _from_host(a): + # fresh, aligned allocation like a device array restored from a pickle + return ndarray(_np.array(a, copy=True, order="K")) + + +def _wrap(x): + if isinstance(x, _HOST): + return ndarray(x) + if isinstance(x, _np.generic) and not isinstance( + x, (_np.str_, _np.bytes_, _np.void) + ): + return ndarray(_np.asarray(x)) + if isinstance(x, tuple): + return tuple(_wrap(i) for i in x) + if isinstance(x, list): + return [_wrap(i) for i in x] + return x + + +def _unwrap(x): + if isinstance(x, ndarray): + return x._a + if isinstance(x, tuple): + return tuple(_unwrap(i) for i in x) + if isinstance(x, list): + return [_unwrap(i) for i in x] + if isinstance(x, dict): + return {k: _unwrap(v) for k, v in x.items()} + return x + + +def _strict(args, where, first_is_data=False): + """Raise on NumPy arrays anywhere (top level or inside tuples/lists). + + 0-d NumPy arrays pass, like NumPy scalars: CuPy treats them as scalars. + With `first_is_data`, a list or tuple as first argument raises too. + """ + for i, a in enumerate(args): + if isinstance(a, _HOST) and a.ndim > 0: + raise _err(a, f" in {where}") + if isinstance(a, (list, tuple)): + if i == 0 and first_is_data: + raise _err(a, f" as array argument of {where}") + _strict(a, where) + + +# binary / unary operators ------------------------------------------------------ +def _binop(op): + def f(self, other): + if isinstance(other, _HOST) and other.ndim > 0: + raise _err(other, f" in operator {op.__name__}") + return _wrap(op(self._a, _unwrap(other))) + + return f + + +def _rbinop(op): + def f(self, other): + if isinstance(other, _HOST) and other.ndim > 0: + raise _err(other, f" in operator {op.__name__}") + return _wrap(op(_unwrap(other), self._a)) + + return f + + +def _unop(op): + def f(self): + return _wrap(op(self._a)) + + return f + + +def _ibinop(op): + def f(self, other): + if isinstance(other, _HOST) and other.ndim > 0: + raise _err(other, f" in operator {op.__name__}") + op(self._a, _unwrap(other)) + return self + + return f + + +for _name, _op, _iop in ( + ("add", operator.add, operator.iadd), + ("sub", operator.sub, operator.isub), + ("mul", operator.mul, operator.imul), + ("truediv", operator.truediv, operator.itruediv), + ("floordiv", operator.floordiv, operator.ifloordiv), + ("mod", operator.mod, operator.imod), + ("pow", operator.pow, operator.ipow), + ("matmul", operator.matmul, operator.imatmul), + ("and", operator.and_, operator.iand), + ("or", operator.or_, operator.ior), + ("xor", operator.xor, operator.ixor), + ("lshift", operator.lshift, operator.ilshift), + ("rshift", operator.rshift, operator.irshift), +): + setattr(ndarray, f"__{_name}__", _binop(_op)) + setattr(ndarray, f"__r{_name}__", _rbinop(_op)) + setattr(ndarray, f"__i{_name}__", _ibinop(_iop)) +for _name, _op in ( + ("lt", operator.lt), + ("le", operator.le), + ("gt", operator.gt), + ("ge", operator.ge), + ("eq", operator.eq), + ("ne", operator.ne), +): + setattr(ndarray, f"__{_name}__", _binop(_op)) +for _name, _op in ( + ("neg", operator.neg), + ("pos", operator.pos), + ("abs", operator.abs), + ("invert", operator.invert), +): + setattr(ndarray, f"__{_name}__", _unop(_op)) + + +# module functions -------------------------------------------------------------- +_CREATION = { + "array", "asarray", "asanyarray", "ascontiguousarray", "asfortranarray", + "zeros", "ones", "empty", "full", "arange", "linspace", "logspace", "eye", + "identity", "indices", "fromfunction", "frombuffer", "diag", "tri", + "geomspace", "copy", +} # fmt: skip +_SEQUENCE = { + "concatenate", "stack", "hstack", "vstack", "dstack", "column_stack", + "row_stack", "block", "meshgrid", "broadcast_arrays", "ix_", + "ravel_multi_index", "atleast_1d", "atleast_2d", "atleast_3d", "einsum", + "result_type", +} # fmt: skip +# Python-level CuPy routines that start with cupy.asanyarray(a), hence accept +# host arrays (one copy) +_ASANYARRAY = { + "diff", "ediff1d", "gradient", "trapezoid", "unique", "flip", "rot90", + "roll", "cumsum", "cumprod", +} # fmt: skip +_PASSTHROUGH = { + "dtype", "issubdtype", "promote_types", "can_cast", "finfo", "iinfo", + "isscalar", "result_type", "broadcast_shapes", "get_printoptions", + "set_printoptions", "printoptions", "errstate", "seterr", "geterr", + "iterable", "ndim", "shape", +} # fmt: skip + + +def _wrap_callable(f, strict, name, first_is_data=False): + def w(*args, **kwargs): + if strict: + _strict(args, name, first_is_data) + _strict(tuple(kwargs.values()), name) + res = f(*_unwrap(args), **_unwrap(kwargs)) + # like CuPy: no new array object when nothing was copied + # (e.g. ascontiguousarray of a contiguous array) + for a in args[:1]: + if isinstance(a, ndarray) and isinstance(res, _HOST): + same = res is a._a or ( + res.base is a._a + and res.shape == a._a.shape + and res.strides == a._a.strides + and res.dtype == a._a.dtype + ) + if same: + return a + return _wrap(res) + + w.__name__ = getattr(f, "__name__", name) + w.__doc__ = getattr(f, "__doc__", None) + return w + + +def _module_attr(name, src=_np, prefix="cupy"): + attr = getattr(src, name) + if isinstance(attr, (type, _np.dtype)) or not callable(attr): + return attr + if name in _PASSTHROUGH: + return lambda *a, **k: attr(*_unwrap(a), **_unwrap(k)) + if name in _CREATION or name in _ASANYARRAY: + return _wrap_callable(attr, strict=False, name=f"{prefix}.{name}") + if name in _SEQUENCE: + return _wrap_callable(attr, strict=True, name=f"{prefix}.{name}") + return _wrap_callable( + attr, strict=True, name=f"{prefix}.{name}", first_is_data=True + ) + + +def asarray(a, dtype=None, order=None, **kwargs): + return ndarray( + _np.array( + _unwrap(a), dtype=dtype, order=order or "K", copy=kwargs.pop("copy", None) + ) + ) + + +def array(a, dtype=None, copy=True, order="K", ndmin=0, **kwargs): + return ndarray( + _np.array(_unwrap(a), dtype=dtype, copy=copy, order=order, ndmin=ndmin) + ) + + +def asnumpy(a, *args, **kwargs): + return a.get() if isinstance(a, ndarray) else _np.asarray(a) + + +def is_available(): + return True + + +def get_array_module(*args): + if builtins.any(isinstance(a, ndarray) for a in args): + return _sys.modules["cupy"] + return _np + + +def fuse(*args, **kwargs): + """``cupy.fuse``: a plain decorator here (the function runs unfused).""" + if args and callable(args[0]) and len(args) == 1 and not kwargs: + return args[0] + return lambda f: f + + +class _Pool: + def free_all_blocks(self): + pass + + def used_bytes(self): + return 0 + + def total_bytes(self): + return 0 + + +def get_default_memory_pool(): + return _Pool() + + +def get_default_pinned_memory_pool(): + return _Pool() + + +class _NoGPU: + def __init__(self, *args, **kwargs): + pass + + def __call__(self, *args, **kwargs): + raise NotImplementedError("fake CuPy cannot run CUDA kernels") + + def get_function(self, name): + return _NoGPU() + + +RawKernel = RawModule = ElementwiseKernel = ReductionKernel = _NoGPU + + +class _Device: + def __init__(self, id=0): + self.id = 0 if id is None else id + + def __enter__(self): + return self + + def __exit__(self, *exc): + return False + + def use(self): + pass + + def synchronize(self): + pass + + @property + def attributes(self): + return {"MaxSharedMemoryPerBlock": 48 * 1024} + + +class _Stream(_Device): + def __init__(self, *args, **kwargs): + super().__init__() + + def __enter__(self): + return self + + +_Stream.null = _Stream() + + +def _alloc_pinned_memory(nbytes): + raise NotImplementedError("fake CuPy has no pinned memory") + + +cuda = types.ModuleType("cupy.cuda") +cuda.Device = _Device +cuda.Stream = _Stream +cuda.get_current_stream = lambda: _Stream.null +cuda.alloc_pinned_memory = _alloc_pinned_memory +cuda.device = types.ModuleType("cupy.cuda.device") +cuda.device.Device = _Device +cuda.runtime = types.ModuleType("cupy.cuda.runtime") +cuda.runtime.getDevice = lambda: 0 +cuda.runtime.setDevice = lambda device: None +cuda.runtime.getDeviceCount = lambda: 1 +cuda.runtime.memGetInfo = lambda: (1 << 34, 1 << 34) +cuda.runtime.CUDARuntimeError = RuntimeError +cuda.driver = types.SimpleNamespace(CUDADriverError=RuntimeError) + + +def _submodule(name, src): + m = types.ModuleType(f"cupy.{name}") + m.__getattr__ = lambda attr: _module_attr(attr, src, prefix=f"cupy.{name}") + m.__all__ = [n for n in dir(src) if not n.startswith("_")] + return m + + +linalg = _submodule("linalg", _np.linalg) +fft = _submodule("fft", _np.fft) + + +class _RNGProxy: + def __init__(self, obj): + self._obj = obj + + def __getattr__(self, name): + attr = getattr(self._obj, name) + if callable(attr): + return _wrap_callable(attr, strict=False, name=f"cupy.random.{name}") + return attr + + +random = types.ModuleType("cupy.random") +random.default_rng = lambda seed=None: _RNGProxy(_np.random.default_rng(_unwrap(seed))) +random.__getattr__ = lambda attr: _wrap_callable( + getattr(_np.random, attr), strict=False, name=f"cupy.random.{attr}" +) + +for _m in (cuda, cuda.device, cuda.runtime, linalg, fft, random): + _sys.modules[_m.__name__] = _m + +bool_ = _np.bool_ +_SKIP = { + "ndarray", "array", "asarray", "linalg", "fft", "random", "cuda", "testing", + "ma", "char", "rec", "lib", "polynomial", "ctypeslib", "emath", "typing", + "exceptions", "dtypes", "strings", "f2py", "version", "matlib", +} # fmt: skip +__all__ = sorted( # noqa: PLE0605 - the NumPy namespace plus the CuPy extras + {n for n in dir(_np) if not n.startswith("_") and n not in _SKIP} + | { + "ndarray", + "array", + "asarray", + "asnumpy", + "cuda", + "linalg", + "fft", + "random", + "is_available", + "bool_", + "fuse", + } +) + + +def __getattr__(name): + if name.startswith("__"): + raise AttributeError(name) + value = _module_attr(name) + # never shadow bool, int, all, ... used inside this module + if not hasattr(builtins, name): + globals()[name] = value + return value diff --git a/src/cunumpy/cuda/include/cunumpy/array_view.cuh b/src/cunumpy/cuda/include/cunumpy/array_view.cuh index af3efb9..c1ebe6a 100644 --- a/src/cunumpy/cuda/include/cunumpy/array_view.cuh +++ b/src/cunumpy/cuda/include/cunumpy/array_view.cuh @@ -1,6 +1,6 @@ // Strided array views for CUDA kernels, passed by value from Python. // -// Array1D, Array2D and Array3D describe a (possibly non-contiguous) +// Array1D to Array4D describe a (possibly non-contiguous) // device array the way NumPy/CuPy do: a data pointer, a shape and strides. // Strides are in ELEMENTS, not bytes, so that `a(i, j)` is // `data[i * strides[0] + j * strides[1]]`. Elements are accessed with @@ -9,7 +9,7 @@ // // The views are created on the Python side by cunumpy (a kernel parameter or a // CudaStruct field of type `Array2D` takes a CuPy array). They are -// passed by value, so the memory layout must be exactly, for ndim = 1, 2, 3: +// passed by value, so the memory layout must be exactly, for ndim = 1 to 4: // // T* data; // 8 bytes // long long shape[ndim]; // ndim * 8 bytes @@ -95,10 +95,33 @@ struct Array3D { } }; +// A 4D view, e.g. a 3D grid of vector components (nx, ny, nz, ncomp). +template +struct Array4D { + T* data; + long long shape[4]; + long long strides[4]; + + __device__ __forceinline__ T& operator()(long long i, long long j, + long long k, long long l) const { + CUNUMPY_CHECK_INDEX(i, 0, shape[0]); + CUNUMPY_CHECK_INDEX(j, 1, shape[1]); + CUNUMPY_CHECK_INDEX(k, 2, shape[2]); + CUNUMPY_CHECK_INDEX(l, 3, shape[3]); + return data[i * strides[0] + j * strides[1] + k * strides[2] + l * strides[3]]; + } + + // Number of elements. + __device__ __forceinline__ long long size() const { + return shape[0] * shape[1] * shape[2] * shape[3]; + } +}; + // The layout the Python side packs: pointer, shape, strides, 8-byte aligned. static_assert(sizeof(Array1D) == 24, "unexpected Array1D layout"); static_assert(sizeof(Array2D) == 40, "unexpected Array2D layout"); static_assert(sizeof(Array3D) == 56, "unexpected Array3D layout"); +static_assert(sizeof(Array4D) == 72, "unexpected Array4D layout"); static_assert(sizeof(Array1D) == 24, "unexpected Array1D layout"); static_assert(alignof(Array2D) == 8, "unexpected Array2D alignment"); diff --git a/src/cunumpy/cuda/include/cunumpy/random.cuh b/src/cunumpy/cuda/include/cunumpy/random.cuh new file mode 100644 index 0000000..3aebf35 --- /dev/null +++ b/src/cunumpy/cuda/include/cunumpy/random.cuh @@ -0,0 +1,127 @@ +// cunumpy/random.cuh: counter-based random numbers (Philox4x32-10) for kernels. +// +// A counter-based generator has no state: the numbers are a pure function of +// a key (the seed) and a counter. Each particle draws its own numbers from +// (seed, stream = particle id, counter = draw index), so a kernel needs no +// per-thread generator state, the result does not depend on the launch shape +// or the order of the threads, and the host computes exactly the same numbers +// (cunumpy.philox_uniform / philox_normal), so a kernel that samples random +// numbers can be compared with its host version. +// +// #include +// +// extern "C" __global__ void thermalize(double* v, long long n, +// unsigned long long seed, +// unsigned long long step, double v_th) { +// long long i = blockIdx.x * (long long)blockDim.x + threadIdx.x; +// if (i >= n) return; +// double z0, z1; +// cunumpy_normal2(seed, (unsigned long long)i, step, &z0, &z1); +// v[i] = v_th * z0; +// } +// +// Philox4x32-10 is the generator of Salmon et al., "Parallel random numbers: +// as easy as 1, 2, 3" (SC'11), as in Random123 and cuRAND. One call maps a +// 128-bit counter and a 64-bit key to 128 random bits: +// counter = (counter low, counter high, stream low, stream high), key = seed. +// The uniform numbers are bit-identical on host and device. The normal numbers +// (Box-Muller: log, sqrt, sin, cos) can differ in the last bits, since the GPU's +// math functions are not the host's. +// +// Everything is plain integer arithmetic (no CUDA vector types or intrinsics), +// so the header also compiles as C++ (see cunumpy.testing.emulate_cuda_kernel). + +#ifndef CUNUMPY_RANDOM_CUH +#define CUNUMPY_RANDOM_CUH + +struct cunumpy_u32x4 { + unsigned int v[4]; +}; + +#define CUNUMPY_PHILOX_M0 0xD2511F53u +#define CUNUMPY_PHILOX_M1 0xCD9E8D57u +#define CUNUMPY_PHILOX_W0 0x9E3779B9u +#define CUNUMPY_PHILOX_W1 0xBB67AE85u + +// Philox4x32-10: 128 random bits from a 128-bit counter and a 64-bit key. +__device__ __forceinline__ cunumpy_u32x4 cunumpy_philox4x32_10( + cunumpy_u32x4 ctr, unsigned int key0, unsigned int key1) +{ + for (int round = 0; round < 10; ++round) { + const unsigned long long p0 = (unsigned long long)CUNUMPY_PHILOX_M0 * ctr.v[0]; + const unsigned long long p1 = (unsigned long long)CUNUMPY_PHILOX_M1 * ctr.v[2]; + const unsigned int hi0 = (unsigned int)(p0 >> 32), lo0 = (unsigned int)p0; + const unsigned int hi1 = (unsigned int)(p1 >> 32), lo1 = (unsigned int)p1; + cunumpy_u32x4 next; + next.v[0] = hi1 ^ ctr.v[1] ^ key0; + next.v[1] = lo1; + next.v[2] = hi0 ^ ctr.v[3] ^ key1; + next.v[3] = lo0; + ctr = next; + key0 += CUNUMPY_PHILOX_W0; + key1 += CUNUMPY_PHILOX_W1; + } + return ctr; +} + +// The 128 random bits for (seed, stream, counter). +__device__ __forceinline__ cunumpy_u32x4 cunumpy_random_bits( + unsigned long long seed, unsigned long long stream, unsigned long long counter) +{ + cunumpy_u32x4 ctr; + ctr.v[0] = (unsigned int)counter; + ctr.v[1] = (unsigned int)(counter >> 32); + ctr.v[2] = (unsigned int)stream; + ctr.v[3] = (unsigned int)(stream >> 32); + return cunumpy_philox4x32_10(ctr, (unsigned int)seed, (unsigned int)(seed >> 32)); +} + +// A double in [0, 1) with 53 random bits, from two 32-bit words. +__device__ __forceinline__ double cunumpy_u64_to_uniform(unsigned int hi, unsigned int lo) +{ + const unsigned long long bits = ((unsigned long long)hi << 32) | lo; + return (double)(bits >> 11) * (1.0 / 9007199254740992.0); // 2^-53 +} + +// Two uniform doubles in [0, 1) for (seed, stream, counter). +__device__ __forceinline__ void cunumpy_uniform2( + unsigned long long seed, unsigned long long stream, unsigned long long counter, + double* u0, double* u1) +{ + const cunumpy_u32x4 r = cunumpy_random_bits(seed, stream, counter); + *u0 = cunumpy_u64_to_uniform(r.v[0], r.v[1]); + *u1 = cunumpy_u64_to_uniform(r.v[2], r.v[3]); +} + +// One uniform double in [0, 1): the first of cunumpy_uniform2. +__device__ __forceinline__ double cunumpy_uniform( + unsigned long long seed, unsigned long long stream, unsigned long long counter) +{ + const cunumpy_u32x4 r = cunumpy_random_bits(seed, stream, counter); + return cunumpy_u64_to_uniform(r.v[0], r.v[1]); +} + +// Two independent standard normal doubles for (seed, stream, counter), by the +// Box-Muller transform of cunumpy_uniform2 (1 - u0 is in (0, 1], so log is finite). +__device__ __forceinline__ void cunumpy_normal2( + unsigned long long seed, unsigned long long stream, unsigned long long counter, + double* z0, double* z1) +{ + double u0, u1; + cunumpy_uniform2(seed, stream, counter, &u0, &u1); + const double radius = sqrt(-2.0 * log(1.0 - u0)); + const double angle = 6.283185307179586 * u1; // 2 pi + *z0 = radius * cos(angle); + *z1 = radius * sin(angle); +} + +// One standard normal double: the first of cunumpy_normal2. +__device__ __forceinline__ double cunumpy_normal( + unsigned long long seed, unsigned long long stream, unsigned long long counter) +{ + double z0, z1; + cunumpy_normal2(seed, stream, counter, &z0, &z1); + return z0; +} + +#endif // CUNUMPY_RANDOM_CUH diff --git a/src/cunumpy/cuda/include/cunumpy/reduce.cuh b/src/cunumpy/cuda/include/cunumpy/reduce.cuh new file mode 100644 index 0000000..2371fa1 --- /dev/null +++ b/src/cunumpy/cuda/include/cunumpy/reduce.cuh @@ -0,0 +1,154 @@ +// cunumpy/reduce.cuh: warp- and block-level reductions for CUDA kernels. +// +// Diagnostics computed inside a kernel (kinetic energy, momentum, total +// charge, a maximum velocity for the CFL condition) and accumulation that +// combines values in a block before writing to global memory all need the +// same reduction: shuffles within each warp, then shared memory across the +// warps of a block. These helpers implement it once: +// +// #include +// +// extern "C" __global__ void kinetic_energy(const double* v, long long n, +// double mass, double* energy) { +// long long i = blockIdx.x * (long long)blockDim.x + threadIdx.x; +// double e = i < n ? 0.5 * mass * v[i] * v[i] : 0.0; // no early return +// cunumpy_block_sum_to(energy, e); // one atomic add per block +// } +// +// Rules: +// * Every thread of the block must call a block function, and all 32 lanes of +// the warp a warp function: do not return early. Threads without a value +// pass the identity (0 for a sum, the largest value for a minimum, ...). +// * The number of threads per block must be a multiple of 32 (CudaKernel's +// default block size is 128). +// * The result is returned to every thread (warp functions: every lane). +// * Supported types are those of the shuffle intrinsics: int, unsigned, +// long long, unsigned long long, float, double. +// +// The warp size is 32, as on all NVIDIA GPUs. +// +// The header directory is added to every CudaKernel's NVRTC options. + +#ifndef CUNUMPY_REDUCE_CUH +#define CUNUMPY_REDUCE_CUH + +#include "cunumpy/atomic.cuh" + +#define CUNUMPY_WARP_SIZE 32 +#define CUNUMPY_FULL_WARP_MASK 0xffffffffu + +// The binary operations; fill() gives the value of the lanes of the last +// step that have no partial result of their own. +struct cunumpy_sum_op { + template + __device__ __forceinline__ T operator()(T a, T b) const { return a + b; } + template + __device__ __forceinline__ static T fill(const T*) { return T(0); } +}; + +struct cunumpy_min_op { + template + __device__ __forceinline__ T operator()(T a, T b) const { return b < a ? b : a; } + template + __device__ __forceinline__ static T fill(const T* partial) { return partial[0]; } +}; + +struct cunumpy_max_op { + template + __device__ __forceinline__ T operator()(T a, T b) const { return a < b ? b : a; } + template + __device__ __forceinline__ static T fill(const T* partial) { return partial[0]; } +}; + +// Linear index of the thread in its block, and the number of threads per block +// (for 1D to 3D blocks). +__device__ __forceinline__ int cunumpy_block_thread() +{ + return threadIdx.x + blockDim.x * (threadIdx.y + blockDim.y * threadIdx.z); +} + +__device__ __forceinline__ int cunumpy_block_threads() +{ + return blockDim.x * blockDim.y * blockDim.z; +} + +// Reduce v over the 32 lanes of the warp; every lane gets the result. +template +__device__ __forceinline__ T cunumpy_warp_reduce(T v, Op op) +{ + for (int mask = CUNUMPY_WARP_SIZE / 2; mask > 0; mask /= 2) { + v = op(v, __shfl_xor_sync(CUNUMPY_FULL_WARP_MASK, v, mask)); + } + return v; +} + +template +__device__ __forceinline__ T cunumpy_warp_sum(T v) { return cunumpy_warp_reduce(v, cunumpy_sum_op()); } + +template +__device__ __forceinline__ T cunumpy_warp_min(T v) { return cunumpy_warp_reduce(v, cunumpy_min_op()); } + +template +__device__ __forceinline__ T cunumpy_warp_max(T v) { return cunumpy_warp_reduce(v, cunumpy_max_op()); } + +// Reduce v over all threads of the block; every thread gets the result. +// Uses 32 values of static shared memory per type and synchronizes the block +// (__syncthreads) three times; it may be called several times in a kernel. +template +__device__ T cunumpy_block_reduce(T v, Op op) +{ + __shared__ T partial[CUNUMPY_WARP_SIZE]; + const int thread = cunumpy_block_thread(); + const int lane = thread % CUNUMPY_WARP_SIZE; + const int warp = thread / CUNUMPY_WARP_SIZE; + const int n_warps = (cunumpy_block_threads() + CUNUMPY_WARP_SIZE - 1) / CUNUMPY_WARP_SIZE; + + v = cunumpy_warp_reduce(v, op); + __syncthreads(); // a previous call may still be reading partial + if (lane == 0) partial[warp] = v; + __syncthreads(); + if (warp == 0) { + v = lane < n_warps ? partial[lane] : Op::fill(partial); + v = cunumpy_warp_reduce(v, op); + if (lane == 0) partial[0] = v; + } + __syncthreads(); + return partial[0]; +} + +template +__device__ T cunumpy_block_sum(T v) { return cunumpy_block_reduce(v, cunumpy_sum_op()); } + +template +__device__ T cunumpy_block_min(T v) { return cunumpy_block_reduce(v, cunumpy_min_op()); } + +template +__device__ T cunumpy_block_max(T v) { return cunumpy_block_reduce(v, cunumpy_max_op()); } + +// *out += sum of v over the block, with one atomic add per block (thread 0). +// For a sum over the whole grid, zero *out before the launch. +__device__ __forceinline__ void cunumpy_block_sum_to(double* out, double v) +{ + const double total = cunumpy_block_sum(v); + if (cunumpy_block_thread() == 0) cunumpy_atomic_add(out, total); +} + +__device__ __forceinline__ void cunumpy_block_sum_to(float* out, float v) +{ + const float total = cunumpy_block_sum(v); + if (cunumpy_block_thread() == 0) cunumpy_atomic_add(out, total); +} + +__device__ __forceinline__ void cunumpy_block_sum_to(int* out, int v) +{ + const int total = cunumpy_block_sum(v); + if (cunumpy_block_thread() == 0) atomicAdd(out, total); +} + +__device__ __forceinline__ void cunumpy_block_sum_to(unsigned long long* out, unsigned long long v) +{ + const unsigned long long total = cunumpy_block_sum(v); + if (cunumpy_block_thread() == 0) atomicAdd(out, total); +} + +#endif // CUNUMPY_REDUCE_CUH diff --git a/src/cunumpy/cuda_kernel.py b/src/cunumpy/cuda_kernel.py index 4a4c65f..6b6a8a3 100644 --- a/src/cunumpy/cuda_kernel.py +++ b/src/cunumpy/cuda_kernel.py @@ -46,11 +46,13 @@ class is the one definition of the arguments. from __future__ import annotations +import ast import hashlib import inspect import math import os import re +import sys import typing from collections.abc import Callable, Hashable, Iterable, Iterator, Mapping, Sequence from concurrent.futures import ThreadPoolExecutor @@ -69,6 +71,7 @@ class is the one definition of the arguments. "CudaStruct", "CudaStructArguments", "CudaStructValue", + "PyccelStructArguments", "ctype_of", "cuda_include_dir", "cuda_kernel_names", @@ -90,10 +93,12 @@ def cuda_include_dir() -> str: """The directory of the CUDA headers shipped with cunumpy. :class:`CudaKernel` adds it to the include path automatically, so kernels - can ``#include "cunumpy/array_view.cuh"`` (strided ``Array1D``, - ``Array2D``, ``Array3D`` views passed by value) and + can ``#include "cunumpy/array_view.cuh"`` (strided ``Array1D`` + to ``Array4D`` views passed by value) and ``#include "cunumpy/index.cuh"`` (thread-index and grid-stride macros such - as ``CUNUMPY_THREAD_1D(i, n)``). Pass it as ``-I`` to other compilers. + as ``CUNUMPY_THREAD_1D(i, n)``), ``#include "cunumpy/atomic.cuh"`` (atomic + adds) and ``#include "cunumpy/reduce.cuh"`` (warp and block reductions). + Pass it as ``-I`` to other compilers. """ return str(_CUDA_INCLUDE_DIR) @@ -154,7 +159,7 @@ class CudaParameter(NamedTuple): The struct type, for a struct passed by value. view_ndim : int | None The number of dimensions, for an array view (``Array1D`` to - ``Array3D``, see :func:`cuda_include_dir`) passed by value. + ``Array4D``, see :func:`cuda_include_dir`) passed by value. """ name: str @@ -226,10 +231,10 @@ class CudaParameter(NamedTuple): _QUALIFIERS = {"const", "volatile", "__restrict__", "__restrict", "restrict"} _COMPLEX = re.compile(r"(?:(?:thrust|cuda::std)::)?complex\s*<\s*(float|double)\s*>") -# Array1D to Array3D (cunumpy/array_view.cuh), T a scalar type of _CTYPES -_VIEW = re.compile(r"\bArray([123])D\s*<((?:[^<>]|complex<[^<>]*>)+?)>") +# Array1D to Array4D (cunumpy/array_view.cuh), T a scalar type of _CTYPES +_VIEW = re.compile(r"\bArray([1234])D\s*<((?:[^<>]|complex<[^<>]*>)+?)>") _TOKEN = re.compile( - r"Array[123]D<[^<>]*(?:<[^<>]*>[^<>]*)?>|complex<(?:float|double)>" + r"Array[1234]D<[^<>]*(?:<[^<>]*>[^<>]*)?>|complex<(?:float|double)>" r"|[A-Za-z_]\w*|\*|\[\s*\]" ) @@ -271,12 +276,23 @@ def _strip_comments(source: str) -> str: # ``#include "name"``: quoted includes are the project's own headers. Angle -# bracket includes are system headers and are not tracked. -_QUOTED_INCLUDE = re.compile(r'^[ \t]*#[ \t]*include[ \t]*"([^"\n]+)"', re.MULTILINE) +# bracket includes are system headers and are not tracked, except in the +# directories given as ``angle_dirs`` (cunumpy's shipped headers). +_INCLUDE = re.compile( + r'^[ \t]*#[ \t]*include[ \t]*(?:"([^"\n]+)"|<([^>\n]+)>)', re.MULTILINE +) + + +def _includes(source: str) -> list[tuple[str, bool]]: + """``(name, quoted)`` of every ``#include`` in `source`, in order.""" + return [ + (quoted or angle, bool(quoted)) + for quoted, angle in _INCLUDE.findall(_strip_comments(source)) + ] def _quoted_includes(source: str) -> list[str]: - return _QUOTED_INCLUDE.findall(_strip_comments(source)) + return [name for name, quoted in _includes(source) if quoted] def resolve_includes( @@ -284,15 +300,17 @@ def resolve_includes( include_dirs: Iterable[str | Path] = (), *, base_dir: str | Path | None = None, + angle_dirs: Iterable[str | Path] = (), ) -> list[Path]: """The header files a CUDA source includes, recursively. Scans `source` (comments removed) for ``#include "name"`` and resolves each name like NVRTC does: relative to `base_dir` (the directory of the - including file), then in `include_dirs`, in order. Found headers are - scanned in turn, relative to their own directory. Includes in angle - brackets (system headers) and includes that cannot be found are ignored; - NVRTC reports the latter when the kernel is compiled. + including file), then in `include_dirs`, then in `angle_dirs`, in order. + Found headers are scanned in turn, relative to their own directory. + Includes in angle brackets (``#include ``) are system headers and + ignored, unless they are found in `angle_dirs`. Includes that cannot be + found are ignored; NVRTC reports them when the kernel is compiled. Parameters ---------- @@ -303,22 +321,32 @@ def resolve_includes( base_dir : str | Path | None Directory of the file `source` was read from, searched first; None if the source is not from a file. + angle_dirs : Iterable[str | Path] + Directories whose headers are tracked also when included in angle + brackets, searched last; :class:`CudaKernel` passes + :func:`cuda_include_dir`, so that ``#include `` + is tracked. Returns ------- list[Path] The resolved header files, each once, in order of first inclusion - (depth first). Empty if the source has no quoted includes; the file + (depth first). Empty if the source has no includes to track; the file system is not touched in that case. """ + angle = tuple(Path(d) for d in angle_dirs) dirs = tuple(Path(d) for d in include_dirs) + dirs += tuple(d for d in angle if d not in dirs) found: list[Path] = [] seen: set[Path] = set() def visit(code: str, directory: Path | None) -> None: - for name in _quoted_includes(code): - candidates = [directory / name] if directory is not None else [] - candidates += [d / name for d in dirs] + for name, quoted in _includes(code): + if quoted: + candidates = [directory / name] if directory is not None else [] + candidates += [d / name for d in dirs] + else: + candidates = [d / name for d in angle] for candidate in candidates: if candidate.is_file(): path = candidate.resolve() @@ -572,8 +600,23 @@ def _describe(param: CudaParameter, index: int) -> str: return f"argument {index} ({ctype} {param.name})" +def _current_device_id() -> int | None: + """The id of the current CUDA device, or None if CuPy has not been imported. + + CuPy is looked up in ``sys.modules`` and never imported here: without it + there are no device arrays whose device could be checked. + """ + cp = sys.modules.get("cupy") + if cp is None: + return None + try: + return int(cp.cuda.runtime.getDevice()) + except Exception: # noqa: BLE001 - no device check without a working runtime + return None + + def _check_device_array(param: CudaParameter, index: int, value: Any) -> None: - """Raise unless `value` is a device array of the declared dtype.""" + """Raise unless `value` is a device array of the declared dtype on the current device.""" # checked on the class: on the instance, CuPy builds the whole interface dict if not hasattr(type(value), "__cuda_array_interface__") and not hasattr( value, "__cuda_array_interface__" @@ -587,6 +630,17 @@ def _check_device_array(param: CudaParameter, index: int, value: Any) -> None: f"{_describe(param, index)} must have dtype {param.dtype}, got " f"{value.dtype}" ) + # an array on another GPU: the kernel would read a foreign address, which + # neither cupy.RawKernel nor the struct packing notices + device_id = getattr(getattr(value, "device", None), "id", None) + if device_id is not None: + current = _current_device_id() + if current is not None and device_id != current: + raise ValueError( + f"{_describe(param, index)} is on CUDA device {device_id}, but the " + f"current device is {current}; kernels only take arrays of the " + "current device (see cunumpy.bind_local_device)" + ) def _pointer_checker(param: CudaParameter, index: int) -> Callable[[Any], Any]: @@ -621,8 +675,7 @@ def check(value: Any) -> Any: _check_device_array(param, index, value) if value.ndim != ndim: raise TypeError( - f"{_describe(param, index)} must be a {ndim}D array, got " - f"{value.ndim}D" + f"{_describe(param, index)} must be a {ndim}D array, got {value.ndim}D" ) itemsize = value.dtype.itemsize strides = [s // itemsize for s in value.strides] @@ -677,6 +730,9 @@ def check(value: Any) -> Any: if isinstance(value, np.generic): if value.dtype == dtype: return value + if isinstance(value, np.integer) and kind in "iuf": + # by value, like a Python int: np.int64(5) fits an int parameter + return cast_int(int(value)) if np.can_cast(value.dtype, dtype, casting="safe"): return scalar_type(value) raise TypeError( @@ -782,8 +838,8 @@ def _pyccel_ctype(annotation: Any, scalars: Mapping[str, str], what: str) -> str ctype = scalars[scalar] if ndim == 0: return ctype - if ndim > 3: - raise ValueError(f"{what}: arrays have at most 3 dimensions, got {ndim}") + if ndim > 4: + raise ValueError(f"{what}: arrays have at most 4 dimensions, got {ndim}") return f"Array{ndim}D<{ctype}>" @@ -853,6 +909,59 @@ def write_cuda_header( return source +def _read_python_source(source: str | Path) -> tuple[str, str]: + """`(text, description)` of a Python source given as a path or as code.""" + if isinstance(source, Path) or ( + "\n" not in source and source.strip().endswith(".py") + ): + path = Path(source) + return path.read_text(), str(path) + return str(source), "" + + +def _find_init(tree: ast.Module, class_name: str, where: str) -> ast.FunctionDef: + """The ``__init__`` function of the class `class_name` in a parsed module.""" + for node in tree.body: + if isinstance(node, ast.ClassDef) and node.name == class_name: + for item in node.body: + if isinstance(item, ast.FunctionDef) and item.name == "__init__": + return item + raise ValueError(f"class {class_name!r} in {where} has no __init__") + raise ValueError(f"no class {class_name!r} in {where}") + + +def _stored_parameters(init: ast.FunctionDef) -> dict[str, str]: + """``{parameter: attribute}`` for the ``self. = `` statements.""" + stored: dict[str, str] = {} + for node in ast.walk(init): + targets: list[ast.expr] = [] + if isinstance(node, ast.Assign): + targets, value = list(node.targets), node.value + elif isinstance(node, ast.AnnAssign) and node.value is not None: + targets, value = [node.target], node.value + else: + continue + if not isinstance(value, ast.Name): + continue + for target in targets: + if ( + isinstance(target, ast.Attribute) + and isinstance(target.value, ast.Name) + and target.value.id == "self" + ): + stored.setdefault(value.id, target.attr) + return stored + + +def _annotation_text(annotation: ast.expr) -> str: + """The pyccel-style annotation string of a parsed annotation.""" + if isinstance(annotation, ast.Constant) and isinstance(annotation.value, str): + return annotation.value + text = ast.unparse(annotation) + # Final['float[:]'] -> float[:] (the quotes come from the string annotation) + return text.replace("'", "").replace('"', "") + + class CudaStruct: """A C struct type passed to CUDA kernels by value. @@ -876,7 +985,7 @@ class CudaStruct: ``(field name, C type)`` pairs, in order, e.g. ``("x", "double*")`` or ``("n", "int")``. Scalar fields, pointers to the scalar types of :func:`ctype_of` (or ``void*``), and array views ``Array1D`` to - ``Array3D`` of those scalar types (from ``cunumpy/array_view.cuh``, + ``Array4D`` of those scalar types (from ``cunumpy/array_view.cuh``, packed as pointer, shape and strides in elements) are supported. Examples @@ -984,6 +1093,82 @@ def from_signature( fields.append((param.name, _pyccel_ctype(param.annotation, scalars, what))) return cls(name, fields) + @classmethod + def from_pyccel_class( + cls, + source: str | Path, + class_name: str, + name: str | None = None, + *, + int_type: str = "long long", + scalar_names: Mapping[str, str] | None = None, + exclude: Sequence[str] = (), + attribute_names: bool = True, + ) -> CudaStruct: + """Build the struct from the ``__init__`` of a class in a Python source file. + + Like :meth:`from_signature`, but for a class whose module is compiled + by pyccel: importing such a module gives the compiled class, whose + ``__init__`` has no Python signature. The ``.py`` source is parsed + with :mod:`ast` instead and is never imported or executed. + + One field per parameter of ``__init__`` (``self`` skipped), in order. + A parameter that ``__init__`` stores as ``self. = `` + gives a field named after the attribute (`attribute_names`), so that + the struct members are the attribute names the host kernels use, also + when the constructor parameter is called differently. + + Parameters + ---------- + source : str | Path + Path of the ``.py`` file, or the source code itself (a string that + contains a newline or does not end with ``.py``). + class_name : str + Name of the class in the source. + name : str | None + Name of the struct type in C; `class_name` by default. + int_type, scalar_names + As for :meth:`from_signature`. + exclude : Sequence[str] + Parameter or attribute names that do not become fields, e.g. + scratch arrays the host class allocates for itself. + attribute_names : bool + Name the fields after the attributes the parameters are stored in + (default); False keeps the parameter names. + + Raises + ------ + ValueError + The class or its ``__init__`` is not found, or an annotation + cannot be mapped (see :meth:`from_signature`). + + Examples + -------- + >>> MarkerArgs = CudaStruct.from_pyccel_class( + ... "kernel_arguments/pusher_args_kernels.py", "MarkerArguments", "MarkerArgs" + ... ) # doctest: +SKIP + """ + text, where = _read_python_source(source) + init = _find_init(ast.parse(text, filename=where), class_name, where) + stored = _stored_parameters(init) if attribute_names else {} + scalars = dict(_PYCCEL_SCALARS, int=int_type) + if scalar_names: + scalars.update(scalar_names) + excluded = set(exclude) + fields = [] + for arg in init.args.posonlyargs + init.args.args + init.args.kwonlyargs: + if arg.arg == "self": + continue + field = stored.get(arg.arg, arg.arg) + if arg.arg in excluded or field in excluded: + continue + what = f"parameter {arg.arg!r} of {class_name}.__init__ in {where}" + if arg.annotation is None: + raise ValueError(f"{what} has no type annotation") + annotation = _annotation_text(arg.annotation) + fields.append((field, _pyccel_ctype(annotation, scalars, what))) + return cls(class_name if name is None else name, fields) + def __repr__(self) -> str: return f"CudaStruct({self._name!r}, {len(self._fields)} fields)" @@ -1090,6 +1275,115 @@ def check_source(self, source: str) -> None: f"match its CudaStruct:\n{self.declaration}" ) + def layout_source(self, include: str | None = None) -> str: + """The CUDA source of the kernel used by :meth:`verify_layout`. + + The kernel ``cunumpy_layout_(unsigned long long* out)`` writes + ``sizeof``, ``alignof`` and the offset of every field (in field order) + of the struct, as the compiler lays it out. + + Parameters + ---------- + include : str | None + Header that defines the struct, as a file name or an ``#include`` + line; by default the struct is defined by :attr:`declaration`. + """ + if include is None: + lines = ['#include "cunumpy/array_view.cuh"'] if self.has_views else [] + lines.append(self.declaration) + else: + line = include.strip() + lines = [line if line.startswith("#") else f'#include "{line}"'] + offsets = "\n".join( + f" out[{i + 2}] = (unsigned long long)((const char*)&s.{f.name}" + " - (const char*)&s);" + for i, f in enumerate(self._fields) + ) + lines.append( + f'extern "C" __global__ void cunumpy_layout_{self._name}(' + "unsigned long long* out) {\n" + f" {self._name} s;\n" + f" out[0] = sizeof({self._name});\n" + f" out[1] = alignof({self._name});\n" + f"{offsets}\n" + "}\n" + ) + return "\n".join(lines) + + def verify_layout( + self, + include: str | None = None, + *, + include_dirs: Sequence[str | Path] = (), + options: Sequence[str] = (), + ) -> dict[str, int]: + """Check the struct layout of the CUDA compiler against :attr:`dtype`. + + Compiles and runs a one-thread kernel (:meth:`layout_source`) that + reports the size, alignment and field offsets of the struct, and + compares them with the NumPy dtype that values are packed into. A + difference means every kernel taking the struct reads some fields at + the wrong place, without any error. Run it once per struct in a GPU + test, especially with a hand-written or generated header (`include`) + and on a new compiler or platform (e.g. ROCm). + + Parameters + ---------- + include : str | None + Header that defines the struct (file name or ``#include`` line), + found in `include_dirs`; by default :attr:`declaration` is compiled. + include_dirs : Sequence[str | Path] + Directories searched for `include`. + options : Sequence[str] + Additional compiler options. + + Returns + ------- + dict[str, int] + ``"sizeof"``, ``"alignof"`` and the offset of each field, by name. + + Raises + ------ + ValueError + If the compiled layout differs from :attr:`dtype` (the message lists + each difference). + RuntimeError + If CuPy is not available. + """ + from .xp import cupy_available + + if not cupy_available(): + raise RuntimeError("verify_layout() compiles a CUDA kernel and needs CuPy") + import cupy as cp + + kernel = CudaKernel( + self.layout_source(include), + f"cunumpy_layout_{self._name}", + include_dirs=include_dirs, + options=options, + structs=[self], + block_size=1, + ) + out = cp.zeros(len(self._fields) + 2, dtype=cp.uint64) + kernel(out, n_threads=1) + measured = [int(v) for v in out.get()] + layout = {"sizeof": measured[0], "alignof": measured[1]} + layout.update({f.name: n for f, n in zip(self._fields, measured[2:])}) + + expected = {"sizeof": self._dtype.itemsize, "alignof": self._dtype.alignment} + expected.update({f.name: self._dtype.fields[f.name][1] for f in self._fields}) + differences = [ + f"{key}: compiler {layout[key]}, dtype {expected[key]}" + for key in expected + if layout[key] != expected[key] + ] + if differences: + raise ValueError( + f"the compiled layout of struct {self._name!r} differs from its " + "CudaStruct dtype:\n " + "\n ".join(differences) + ) + return layout + def __call__(self, **values: Any) -> CudaStructValue: """Pack values into the struct. @@ -1168,9 +1462,15 @@ class CudaStructArguments(CudaArguments): argument: pointers need C-contiguous CuPy arrays of the declared dtype (never copied), scalars are range-checked and cast. - The packed struct holds device addresses. Call :meth:`pack` again after - replacing an array attribute. Copies (``copy.copy``, ``copy.deepcopy``) - and unpickled objects are packed again from their own arrays. + The packed struct always reflects the current field attributes: at every + use (:attr:`packed`, :meth:`__cuda_args__`, so at every kernel launch) the + device address, shape and strides of every array field and the value of + every scalar field are compared with those that were packed, and the + struct is packed again if any changed. A field may therefore be a + property that reads the owner's current array, so that resizing the + owner's arrays never leaves the struct pointing at freed device memory + (see the second example). Copies (``copy.copy``, ``copy.deepcopy``) and + unpickled objects are packed again from their own arrays. A subclass that sets neither :attr:`struct_name` nor :attr:`fields` is an intermediate base class; a subclass that sets only one of them raises @@ -1209,6 +1509,25 @@ class CudaStructArguments(CudaArguments): }; >>> push = CudaKernel(source, "push", structs=[MarkerArguments.struct]) # doctest: +SKIP >>> push(MarkerArguments(markers, valid), 0.1, n_threads=markers.shape[0]) # doctest: +SKIP + + Fields as properties follow the arrays of an owner object, also after the + owner replaced them (e.g. when it resized its marker array): + + >>> class ParticleArguments(CudaStructArguments): + ... struct_name = "ParticleArgs" + ... fields = (("markers", "Array2D"), ("n_markers", "int")) + ... + ... def __init__(self, particles): + ... self._particles = particles + ... self.pack() + ... + ... @property + ... def markers(self): + ... return self._particles.markers + ... + ... @property + ... def n_markers(self): + ... return self._particles.markers.shape[0] """ struct_name: str @@ -1230,8 +1549,10 @@ def __init_subclass__(cls, **kwargs: Any) -> None: def pack(self) -> None: """Pack the field attributes into the struct. - Called at the end of the constructor, and again after an array - attribute has been replaced. + Called at the end of the constructor, so that invalid field values + raise there. Afterwards the struct is packed again automatically when + a field changes (see the class documentation); calling this method + again is never needed, but harmless. Raises ------ @@ -1246,6 +1567,12 @@ def pack(self) -> None: raise TypeError( f"{type(self).__qualname__} does not define struct_name and fields" ) + values = self._field_values(struct) + self._struct_value = struct(**values) + self._packed_state = _field_state(struct, values) + + def _field_values(self, struct: CudaStruct) -> dict[str, Any]: + """The current value of every field attribute, by field name.""" values = {} for field in struct.fields: try: @@ -1255,13 +1582,24 @@ def pack(self) -> None: f"{type(self).__qualname__} has no attribute {field.name!r} " f"for the field of struct {struct.name}" ) from None - self._struct_value = struct(**values) + return values @property def packed(self) -> np.void: - """The packed struct, with the memory layout of the C struct; packs on first use.""" - if self.__dict__.get("_struct_value") is None: + """The packed struct, with the memory layout of the C struct. + + Packed on first use, and again whenever a field attribute changed + since the last packing (a different array, or a different scalar). + """ + value = self.__dict__.get("_struct_value") + if value is None: self.pack() + return self._struct_value.packed + struct = value.struct + values = self._field_values(struct) + if _field_state(struct, values) != self._packed_state: + self._struct_value = struct(**values) + self._packed_state = _field_state(struct, values) return self._struct_value.packed def __cuda_args__(self) -> tuple[np.void]: @@ -1273,6 +1611,7 @@ def __getstate__(self) -> dict[str, Any]: # which a copy (or another process) does not share state = self.__dict__.copy() state.pop("_struct_value", None) + state.pop("_packed_state", None) return state def __setstate__(self, state: dict[str, Any]) -> None: @@ -1280,6 +1619,186 @@ def __setstate__(self, state: dict[str, Any]) -> None: self.pack() +def _field_state(struct: CudaStruct, values: Mapping[str, Any]) -> tuple[Any, ...]: + """What the packed struct depends on, to detect changed field attributes. + + For an array field the device address, shape and strides (the address + alone for a pointer field); for a scalar field its value. A value that is + neither (e.g. a host array in a pointer field) is identified by its id, + so replacing it triggers a repack, which then raises the type error. + """ + state = [] + for field in struct.fields: + value = values[field.name] + if field.pointer or field.view_ndim is not None: + ptr = getattr(getattr(value, "data", None), "ptr", None) + if ptr is None: + state.append(("id", id(value))) + elif field.pointer: + state.append(ptr) + else: + state.append((ptr, tuple(value.shape), tuple(value.strides))) + elif isinstance(value, (bool, int, float, complex, np.generic)): + state.append((type(value), value)) + else: + state.append(("id", id(value))) + return tuple(state) + + +def _is_device_array(value: Any) -> bool: + return hasattr(type(value), "__cuda_array_interface__") or hasattr( + value, "__cuda_array_interface__" + ) + + +def _host_state(value: Any) -> Any: + """What a host argument object depends on, to detect replaced values.""" + if _is_device_array(value) or isinstance(value, np.ndarray): + ptr = getattr(getattr(value, "data", None), "ptr", None) + if ptr is None: + ptr = getattr(getattr(value, "ctypes", None), "data", None) + return (id(value), ptr, tuple(getattr(value, "shape", ()))) + if isinstance(value, (bool, int, float, complex, str, np.generic)): + return (type(value), value) + return ("id", id(value)) + + +class PyccelStructArguments(CudaStructArguments): + """Argument object with a pyccel host class and a C struct for CUDA kernels. + + The :class:`~cunumpy.KernelArguments` form of :class:`CudaStructArguments`: + the same object is passed to a :class:`~cunumpy.Kernel` on both backends. + On the device path it arrives as the packed struct (``__cuda_args__()``); + on the host path, ``__host_args__()`` builds an instance of + :attr:`host_class` (typically a pyccel-compiled argument class, which + cannot inherit from anything) from the attributes named in + :attr:`host_fields`, once, and again when one of them was replaced. + + A subclass sets :attr:`struct_name`, :attr:`fields` and :attr:`host_class`, + stores every field as an attribute (NumPy or CuPy arrays, whatever the + owner has) and, on the CuPy backend, calls :meth:`pack` at the end of its + constructor so that invalid arrays raise there (:meth:`has_device_arrays` + tells). Objects holding host arrays are copied and pickled without + packing; the struct is only built from device arrays. + + On the CuPy backend there is no host form: the arrays are device arrays, + and a host kernel would have to copy them. ``__host_args__()`` raises + then, unless :attr:`host_copies` is True, in which case the host object is + built from host copies (counted by :func:`~cunumpy.count_transfers`) and + what the host kernel writes is **not** copied back; use it for read-only + evaluations only. + + Attributes + ---------- + host_class : type + The class of the host argument object, e.g. the pyccel class. + host_fields : Sequence[str] | None + The attributes passed to ``host_class(...)``, positionally and in this + order; by default the struct fields in declaration order. + host_copies : bool + Whether ``__host_args__()`` may copy device arrays to the host + (default False). + + Examples + -------- + >>> from my_kernels import pusher_args_kernels # pyccel-compiled module + >>> class MarkerArguments(PyccelStructArguments): + ... struct_name = "MarkerArgs" + ... fields = (("markers", "Array2D"), ("Np", "long long")) + ... host_class = pusher_args_kernels.MarkerArguments + ... + ... def __init__(self, markers, Np): + ... self.markers = markers + ... self.Np = Np + ... if xp.is_gpu(markers): + ... self.pack() + >>> push(MarkerArguments(markers, Np), dt, n_threads=markers.shape[0]) # doctest: +SKIP + """ + + host_class: type | None = None + host_fields: Sequence[str] | None = None + host_copies: bool = False + + def _host_field_names(self) -> tuple[str, ...]: + if self.host_fields is not None: + return tuple(self.host_fields) + struct = getattr(type(self), "struct", None) + if struct is None: + raise TypeError( + f"{type(self).__qualname__} does not define struct_name and fields" + ) + return tuple(field.name for field in struct.fields) + + def __host_args__(self) -> Any: + """The host argument object, built from the current attributes.""" + host_class = self.host_class + if host_class is None: + raise TypeError( + f"{type(self).__qualname__}.host_class is not set: the class of " + "the host argument object (e.g. the pyccel class) is required" + ) + names = self._host_field_names() + values = [getattr(self, name) for name in names] + state = tuple(_host_state(value) for value in values) + if ( + self.__dict__.get("_host_value") is not None + and self.__dict__.get("_host_state") == state + ): + return self._host_value + host_values = [] + for name, value in zip(names, values): + if _is_device_array(value): + if not self.host_copies: + raise RuntimeError( + f"{type(self).__qualname__}.{name} is a device array: there " + "is no host form on the CuPy backend. Call the CUDA kernel, " + "or set host_copies = True for a read-only host evaluation " + "from host copies" + ) + from .xp import to_numpy + + value = to_numpy(value) + host_values.append(value) + self._host_value = host_class(*host_values) + self._host_state = state + return self._host_value + + def has_device_arrays(self) -> bool: + """Whether any array field is a device array (then the struct can be packed).""" + struct = getattr(type(self), "struct", None) + if struct is None: + return False + return any( + _is_device_array(getattr(self, field.name, None)) + for field in struct.fields + if field.pointer or field.view_ndim is not None + ) + + def __getstate__(self) -> dict[str, Any]: + # the host object is rebuilt from the copied or restored attributes + state = super().__getstate__() + state.pop("_host_value", None) + state.pop("_host_state", None) + return state + + def __setstate__(self, state: dict[str, Any]) -> None: + # on the NumPy backend the fields are host arrays: nothing to pack + self.__dict__.update(state) + if self.has_device_arrays(): + self.pack() + + +def _first_array_length(args: tuple[Any, ...]) -> int: + """``n_threads_from="first_array"``: the first axis of the first array argument.""" + for arg in args: + shape = getattr(arg, "shape", None) + if shape is not None and len(shape) > 0 and hasattr(arg, "dtype"): + return int(shape[0]) + raise TypeError( + "n_threads_from='first_array' needs an array argument; pass n_threads" + ) + + def _as_shape(value: int | Sequence[int], what: str) -> tuple[int, ...]: shape = (value,) if isinstance(value, (int, np.integer)) else tuple(value) if not 1 <= len(shape) <= 3: @@ -1287,6 +1806,33 @@ def _as_shape(value: int | Sequence[int], what: str) -> tuple[int, ...]: return tuple(int(n) for n in shape) +def _device_arrays_in(args: Sequence[Any]) -> Iterator[tuple[str, Any]]: + """``(label, array)`` for every device array among kernel arguments. + + Arrays are found directly, in the fields of :class:`CudaStructArguments` + objects and struct values, and in the values of other ``__cuda_args__()`` + objects. + """ + for index, arg in enumerate(args): + label = f"argument {index}" + if _is_device_array(arg): + yield label, arg + elif isinstance(arg, CudaStructArguments) and hasattr(type(arg), "struct"): + for field in arg.struct.fields: + value = getattr(arg, field.name, None) + if _is_device_array(value): + yield f"{label}.{field.name}", value + elif isinstance(arg, CudaStructValue): + for field in arg.struct.fields: + value = arg[field.name] + if _is_device_array(value): + yield f"{label}.{field.name}", value + elif hasattr(arg, "__cuda_args__"): + for j, value in enumerate(arg.__cuda_args__()): + if _is_device_array(value): + yield f"{label}[{j}]", value + + class CudaKernel: """A CUDA C kernel, compiled with NVRTC through CuPy. @@ -1307,8 +1853,9 @@ class CudaKernel: The headers shipped with cunumpy (:func:`cuda_include_dir`) are always found at compile time (see :meth:`compile_options`): ``#include "cunumpy/array_view.cuh"`` gives the ``Array1D`` to - ``Array3D`` views, ``#include "cunumpy/index.cuh"`` the thread-index - macros, ``#include "cunumpy/atomic.cuh"`` atomic adds. + ``Array4D`` views, ``#include "cunumpy/index.cuh"`` the thread-index + macros, ``#include "cunumpy/atomic.cuh"`` atomic adds, + ``#include "cunumpy/reduce.cuh"`` warp and block reductions. source_dir : str | Path | None Directory the source was read from (set by :meth:`from_file`), where ``#include "..."`` files are looked up first. @@ -1338,7 +1885,9 @@ class CudaKernel: options, but not on the files pulled in by ``#include "..."``. At compile time the headers are resolved (:attr:`included_headers`) and a define with the hash of their contents is added to the options - (:meth:`compile_options`), so editing a header recompiles the kernel. + (:meth:`compile_options`), so editing a header recompiles the kernel; this + includes cunumpy's own headers (``#include ``), so an + upgrade that changes them recompiles too. Examples -------- @@ -1364,8 +1913,14 @@ def __init__( template_args: Sequence[Any] | None = None, check_signature: bool = True, debug: bool | None = None, + n_threads_from: Callable[[tuple[Any, ...]], Any] | str | None = None, + check_finite: bool = False, ) -> None: self._block = self._check_block(_as_shape(block_size, "block_size")) + self.n_threads_from = n_threads_from + # dynamic shared memory the compiled kernel is set up for (see __call__) + self._shared_mem_opt_in = 48 * 1024 + self._check_finite = bool(check_finite) self._debug = None if debug is None else bool(debug) self._source = source self._name = name @@ -1538,15 +2093,22 @@ def source_dir(self) -> Path | None: @property def included_headers(self) -> tuple[Path, ...]: - """The header files the source includes with ``#include "..."``. - - Resolved recursively in `source_dir` and `include_dirs` at every access - (see :func:`resolve_includes`), so the result follows the files on - disk. Empty if the source has no quoted includes. + """The header files the source includes, recursively. + + Quoted includes (``#include "..."``) are resolved in `source_dir`, + `include_dirs` and cunumpy's header directory + (:func:`cuda_include_dir`); cunumpy's shipped headers are tracked + also when included in angle brackets (``#include ``). + Resolved at every access (see :func:`resolve_includes`), so the result + follows the files on disk. Empty if the source includes no project or + cunumpy headers. """ return tuple( resolve_includes( - self._source, self._include_dirs, base_dir=self._source_dir + self._source, + self._include_dirs, + base_dir=self._source_dir, + angle_dirs=(cuda_include_dir(),), ) ) @@ -1561,7 +2123,8 @@ def compile_options(self) -> tuple[str, ...]: it), and, if the source includes header files, ``-DCUNUMPY_INCLUDE_HASH=0x`` with the hash of the contents of :attr:`included_headers` (see :func:`include_hash`). CuPy keys its - kernel cache on the options, so a changed header means a recompile. + kernel cache on the options, so a changed header means a recompile, + also for a header shipped with cunumpy that changed in an upgrade. """ options = self._options # cunumpy's own headers (, ...) are always found @@ -1580,6 +2143,47 @@ def structs(self) -> tuple[CudaStruct, ...]: """Struct types passed to the kernel by value.""" return self._structs + @property + def n_threads_from(self) -> Callable[[tuple[Any, ...]], Any] | None: + """Default launch size: a function of the positional arguments, or None. + + Called with the tuple of arguments of a launch that gives neither + `n_threads` nor `grid`, and returns `n_threads` (an integer or a + tuple), e.g. ``lambda args: args[2].n_markers`` for a kernel whose + third argument is a struct argument object with the marker count. + Set it to ``"first_array"`` for the most common case, one thread per + row: the length of the first array argument (its first axis). + Settable, also on the ``cuda_kernel`` of a :class:`~cunumpy.Kernel`. + """ + return self._n_threads_from + + @n_threads_from.setter + def n_threads_from( + self, value: Callable[[tuple[Any, ...]], Any] | str | None + ) -> None: + if value == "first_array": + value = _first_array_length + if value is not None and not callable(value): + raise TypeError("n_threads_from must be callable, 'first_array' or None") + self._n_threads_from = value + + @property + def check_finite(self) -> bool: + """Whether every launch checks the floating-point arrays for NaN or inf. + + After the launch (synchronized), every floating-point or complex array + among the arguments, including the array fields of struct argument + objects, is scanned, and a non-finite value raises ``RuntimeError`` + naming the kernel and the argument. Costs a synchronization and one + pass over the arrays per launch; for debugging (e.g. a pusher writing + NaN velocities), not for production. Settable. + """ + return self._check_finite + + @check_finite.setter + def check_finite(self, value: bool) -> None: + self._check_finite = bool(value) + @property def template_args(self) -> tuple[Any, ...] | None: """Template arguments, or None if the kernel is not a template.""" @@ -1752,6 +2356,8 @@ def __call__( mode, such an error surfaces at a later synchronization (a ``.get()``, an MPI call, ...), not necessarily in this kernel. """ + if n_threads is None and grid is None and self._n_threads_from is not None: + n_threads = self._n_threads_from(args) grid_shape, block_shape = self.launch_shape(n_threads, grid=grid, block=block) if shared_mem < 0: raise ValueError(f"shared_mem must be non-negative, got {shared_mem}") @@ -1760,32 +2366,94 @@ def __call__( return kernel = self.compile() + if shared_mem > self._shared_mem_opt_in: + self._opt_in_shared_memory(kernel, shared_mem) debug = self.debug_active() with stream if stream is not None else nullcontext(): kernel(grid_shape, block_shape, values, shared_mem=shared_mem) - if debug: + if debug or self._check_finite: self._synchronize_after_launch(stream, grid_shape, block_shape) + if self._check_finite: + self._check_finite_arrays(args) + + def _opt_in_shared_memory(self, kernel: Any, shared_mem: int) -> None: + """Allow `shared_mem` bytes of dynamic shared memory above the default limit. + + Up to :data:`~cunumpy.DEFAULT_SHARED_MEMORY_PER_BLOCK` (48 KiB) every + device launches without setup. Above it, newer GPUs need the kernel + attribute ``max_dynamic_shared_size_bytes``; it is set once (and again + for a larger request) up to the device's opt-in limit. + """ + from .xp import DEFAULT_SHARED_MEMORY_PER_BLOCK, max_shared_memory_per_block + + if shared_mem <= DEFAULT_SHARED_MEMORY_PER_BLOCK: + return + limit = max_shared_memory_per_block(opt_in=True) + if shared_mem > limit: + raise ValueError( + f"kernel {self.expression!r}: shared_mem={shared_mem} bytes exceeds " + f"the {limit} bytes a block may use on this device" + ) + kernel.max_dynamic_shared_size_bytes = shared_mem + self._shared_mem_opt_in = shared_mem + + def _check_finite_arrays(self, args: tuple[Any, ...]) -> None: + """Raise if a floating-point array among `args` holds NaN or inf.""" + import cupy as cp + + for label, array in _device_arrays_in(args): + kind = getattr(array.dtype, "kind", "") + if kind not in "fc": + continue + if not bool(cp.isfinite(array).all()): + raise RuntimeError( + f"kernel {self.expression!r} left a NaN or inf in {label} " + f"(dtype {array.dtype}, shape {tuple(array.shape)})" + ) def _synchronize_after_launch( self, stream: Any, grid: tuple[int, ...], block: tuple[int, ...] ) -> None: - """Wait for the launch and re-raise a CUDA error naming this kernel.""" - import cupy as cp + """Wait for the launch and re-raise a CUDA error naming this kernel. + Skipped while the stream is being captured into a CUDA graph: the + launch is only recorded then, and synchronizing would invalidate the + capture. Errors of the captured kernels surface when the graph is + launched (in debug mode, synchronize after ``graph.launch()``). + """ + if stream is None: + import cupy as cp + + stream = cp.cuda.get_current_stream() + if _is_capturing(stream): + return try: - if stream is None: - stream = cp.cuda.get_current_stream() stream.synchronize() - except ( - cp.cuda.runtime.CUDARuntimeError, - cp.cuda.driver.CUDADriverError, - ) as error: + except Exception as error: + import cupy as cp + + if not isinstance( + error, + (cp.cuda.runtime.CUDARuntimeError, cp.cuda.driver.CUDADriverError), + ): + raise raise RuntimeError( f"CUDA error after launching kernel {self.expression!r} with " f"grid {grid} and block {block}: {error}" ) from error +def _is_capturing(stream: Any) -> bool: + """Whether `stream` is being captured into a CUDA graph.""" + is_capturing = getattr(stream, "is_capturing", None) + if is_capturing is None: + return False + try: + return bool(is_capturing()) + except Exception: # noqa: BLE001 - e.g. the legacy null stream cannot capture + return False + + class CudaKernelVariants: """Kernels generated per variant (e.g. per dtype and dimension), compiled once. diff --git a/src/cunumpy/dispatch.py b/src/cunumpy/dispatch.py index f3ba201..37742f3 100644 --- a/src/cunumpy/dispatch.py +++ b/src/cunumpy/dispatch.py @@ -1,10 +1,11 @@ -"""Pairs of host and CUDA kernels, chosen by the active backend. +"""Pairs of host and CUDA kernels, chosen by the backend or by the arguments. A :class:`Kernel` holds a host kernel (a :class:`~cunumpy.PyccelKernel`, e.g. a Pyccel-compiled function) and, optionally, its 1:1 corresponding CUDA kernel (:class:`~cunumpy.CudaKernel`). It calls the host kernel on the NumPy backend and -the CUDA kernel on the CuPy backend, so a code base can port its kernels to CUDA -one by one. +the CUDA kernel on the CuPy backend (or, with ``dispatch="arrays"``, the CUDA +kernel for device arguments and the host kernel for host arguments), so a code +base can port its kernels to CUDA one by one. :class:`KernelCatalog` collects such pairs from a package laid out with one folder per kernel:: @@ -20,14 +21,20 @@ from __future__ import annotations +import ast import importlib +import inspect +import sys import warnings from collections.abc import Callable, Iterator, Mapping, Sequence from pathlib import Path +from types import ModuleType from typing import Any +import array_api_compat + from .cuda_kernel import CudaKernel, _compile_in_threads -from .kernel import PyccelKernel, resolve_host_args +from .kernel import CompiledHostKernel, PyccelKernel, resolve_host_args from .transfers import _ACTIVE as _COUNTERS from .transfers import _record from .xp import get_backend @@ -35,6 +42,61 @@ __all__ = ["Kernel", "KernelCatalog"] _MISSING_CUDA = ("raise", "fallback") +_DISPATCH = ("backend", "arrays") + +# substituted in tests that have no GPU +_is_device_array = array_api_compat.is_cupy_array + + +def _on_device(arg: Any) -> bool: + """Whether a kernel argument lives on the GPU. + + A CuPy array, or a device-only argument object (one with ``__cuda_args__`` + but no ``__host_args__``, e.g. a :class:`~cunumpy.CudaArguments` or a struct + value). A :class:`~cunumpy.KernelArguments` object has both forms and does + not decide. + """ + if _is_device_array(arg): + return True + kind = type(arg) + return callable(getattr(kind, "__cuda_args__", None)) and not callable( + getattr(kind, "__host_args__", None) + ) + + +#: Longest name a Fortran compiler accepts. pyccel names the wrapper module of a +#: host kernel module ``bind_c_``, so a kernel module name longer than +#: 63 - len("bind_c_") characters cannot be compiled with the Fortran backend. +FORTRAN_NAME_LIMIT = 63 + + +def _pyccel_stub_parameters(function: Any) -> list[str] | None: + """Parameter names of a pyccel-compiled function, from its ``.pyi`` stub. + + pyccel writes ``__pyccel__/.pyi`` next to the compiled extension + module; the compiled function itself has no Python signature. Returns + None if the stub or the function is not found. + """ + module = getattr(function, "__self__", None) # extension functions: the module + if not isinstance(module, ModuleType): + name = getattr(function, "__module__", None) + module = sys.modules.get(name) if isinstance(name, str) else None + file = getattr(module, "__file__", None) + func_name = getattr(function, "__name__", None) + if not file or not func_name: + return None + path = Path(file) + stub = path.parent / "__pyccel__" / (path.name.split(".")[0] + ".pyi") + if not stub.is_file(): + return None + try: + tree = ast.parse(stub.read_text(), filename=str(stub)) + except SyntaxError: + return None + for node in tree.body: + if isinstance(node, ast.FunctionDef) and node.name == func_name: + return [a.arg for a in node.args.posonlyargs + node.args.args] + return None def _source_root(package: str) -> Path: @@ -70,6 +132,16 @@ class Kernel: plain callable `host_kernel`, e.g. ``{"object_modules": ("my_pkg.",), "outputs": (2,)}``; they matter for the fallback on the CuPy backend. Not allowed if `host_kernel` already is a ``PyccelKernel``. + dispatch : {"backend", "arrays"} + How a call chooses between the two kernels. ``"backend"`` (default): + the CUDA kernel on the CuPy backend, the host kernel on the NumPy + backend. ``"arrays"``: the CUDA kernel if any top-level argument lives + on the GPU (a CuPy array, or a device-only argument object such as a + :class:`~cunumpy.CudaArguments` or a struct value), else the host + kernel, whatever the backend. Use ``"arrays"`` when a code deliberately + hands host arrays to kernels while CuPy is active (diagnostics, MPI + staging, CPU fallbacks): the host kernel then runs on the host arrays + instead of the CUDA kernel rejecting them. Notes ----- @@ -91,7 +163,12 @@ def __init__( missing_cuda: str = "raise", cuda_path: str | Path | None = None, host_options: Mapping[str, Any] | None = None, + dispatch: str = "backend", + test_args: str | None = None, ) -> None: + if dispatch not in _DISPATCH: + raise ValueError(f"dispatch must be one of {_DISPATCH}, got {dispatch!r}") + self._dispatch = dispatch if missing_cuda not in _MISSING_CUDA: raise ValueError( f"missing_cuda must be one of {_MISSING_CUDA}, got {missing_cuda!r}" @@ -113,6 +190,8 @@ def __init__( self._name = name if name is not None else host_kernel.name self._missing_cuda = missing_cuda self._cuda_path = None if cuda_path is None else Path(cuda_path) + self._test_args_module = test_args + self._test_args: ModuleType | None = None self._warned = False def __repr__(self) -> str: @@ -154,6 +233,74 @@ def cuda_path(self) -> Path | None: """Where the CUDA kernel is expected, if known.""" return self._cuda_path + @property + def dispatch(self) -> str: + """How calls choose a kernel: ``"backend"`` or ``"arrays"``.""" + return self._dispatch + + @property + def test_args_module(self) -> str | None: + """Dotted name of the module with the test arguments of this kernel, or None. + + Set by :meth:`KernelCatalog.from_package` for a kernel folder that + contains ``_test_args.py``; see :func:`cunumpy.testing.check_parity`. + """ + return self._test_args_module + + @property + def test_args(self) -> ModuleType | None: + """The test-arguments module, imported on first access, or None.""" + if self._test_args is None and self._test_args_module is not None: + self._test_args = importlib.import_module(self._test_args_module) + return self._test_args + + def host_parameters(self) -> list[str] | None: + """The parameter names of the host kernel, or None if they are unknown. + + Read from the Python function (for a :class:`~cunumpy.CompiledHostKernel`, + its uncompiled Python version), or, for a pyccel-compiled function + without a Python signature, from the ``__pyccel__/.pyi`` stub + pyccel writes next to the extension module. None if neither is + available. + """ + function = self._host_kernel.kernel + if isinstance(function, CompiledHostKernel): + function = function.python + try: + signature = inspect.signature(function) + except (TypeError, ValueError): + return _pyccel_stub_parameters(function) + positional = ( + inspect.Parameter.POSITIONAL_ONLY, + inspect.Parameter.POSITIONAL_OR_KEYWORD, + ) + return [p.name for p in signature.parameters.values() if p.kind in positional] + + def check_signature(self) -> None: + """Check that the host and CUDA kernels take the same parameters, in order. + + Compares the parameter names of the host function with those of the + parsed ``__global__`` signature. Nothing is checked without a CUDA + kernel, without a parsed CUDA signature (``check_signature=False``) or + without a Python host signature. + + Raises + ------ + ValueError + If the names or their order differ; the message shows both lists. + """ + if self._cuda_kernel is None or self._cuda_kernel.signature is None: + return + host = self.host_parameters() + if host is None: + return + cuda = [param.name for param in self._cuda_kernel.signature] + if host != cuda: + raise ValueError( + f"kernel {self._name!r}: the host kernel takes ({', '.join(host)}), " + f"the CUDA kernel ({', '.join(cuda)})" + ) + def get_kernel(self) -> PyccelKernel | CudaKernel: """The kernel for the active backend. @@ -167,6 +314,10 @@ def get_kernel(self) -> PyccelKernel | CudaKernel: """ if get_backend() != "cupy": return self._host_kernel + return self._device_kernel() + + def _device_kernel(self) -> PyccelKernel | CudaKernel: + """The CUDA kernel, or what ``missing_cuda`` says without one.""" if self._cuda_kernel is not None: return self._cuda_kernel if self._missing_cuda == "raise": @@ -222,9 +373,14 @@ def __call__( `n_threads` (or `grid`) is required when the CUDA kernel is called. Ignored by the host kernel. """ - kernel = self.get_kernel() + if self._dispatch == "arrays": + on_device = any(_on_device(arg) for arg in args) + kernel = self._device_kernel() if on_device else self._host_kernel + else: + on_device = get_backend() == "cupy" + kernel = self.get_kernel() if kernel is self._host_kernel: - if _COUNTERS and self._cuda_kernel is None and get_backend() == "cupy": + if _COUNTERS and self._cuda_kernel is None and on_device: _record( "fallback", f"Kernel {self._name!r} has no CUDA kernel: host kernel " @@ -232,10 +388,10 @@ def __call__( ) args, _ = resolve_host_args(args) return kernel(*args) - if n_threads is None and grid is None: + if n_threads is None and grid is None and kernel.n_threads_from is None: raise ValueError( f"{self._name}: n_threads is required to launch the CUDA kernel " - "(or pass grid)" + "(or pass grid, or set cuda_kernel.n_threads_from)" ) return kernel( *args, @@ -271,11 +427,20 @@ def from_package( *, host_suffix: str = "_kernels", cuda_suffix: str = "_cuda.cu", + test_args_suffix: str | None = "_test_args", + check_name_length: bool = True, missing_cuda: str = "raise", host_options: ( Mapping[str, Any] | Callable[[str], Mapping[str, Any]] | None ) = None, include_dirs: Sequence[str | Path] | None = None, + dispatch: str = "backend", + compile_host: Callable[[Any], Any] | None = None, + host_fallback: ( + Mapping[str, Callable[..., Any]] + | Callable[[str], Callable[..., Any] | None] + | None + ) = None, **cuda_options: Any, ) -> KernelCatalog: """Collect the kernels of a package with one folder per kernel. @@ -296,6 +461,16 @@ def from_package( Module name suffix of the host kernels. cuda_suffix : str File name suffix of the CUDA kernels. + test_args_suffix : str | None + Module name suffix of the test arguments: ``.py`` + in the kernel's folder, if present, is recorded as + :attr:`Kernel.test_args_module` (imported only when a test asks for + it, see :func:`cunumpy.testing.check_parity`). None disables this. + check_name_length : bool + Warn about a kernel whose module name is too long for the Fortran + backend of pyccel: the wrapper module ``bind_c_`` + must fit :data:`FORTRAN_NAME_LIMIT` (63) characters, so with the + default suffix a kernel name has at most 48 characters. missing_cuda : {"raise", "fallback"} Passed on to every :class:`Kernel`. host_options : Mapping | Callable[[str], Mapping] | None @@ -307,6 +482,23 @@ def from_package( kernel's own folder. By default the source root of the top-level package (the directory containing it), so that a kernel in ``my_pkg.kernels`` can ``#include "my_pkg/common.cuh"``. + dispatch : {"backend", "arrays"} + Passed on to every :class:`Kernel`: choose the CUDA kernel by the + active backend or by where the arguments live. + compile_host : Callable | None + Compiles a host kernel module, e.g. + a function wrapping ``pyccel.epyccel`` (cunumpy does not compile + anything itself). Each host kernel then is a + :class:`~cunumpy.CompiledHostKernel`: compiled on its first + call, falling back to + `host_fallback`, or to the uncompiled Python function with a + warning, if compilation fails. By default the Python function is + called as it is. + host_fallback : Mapping | Callable | None + The fallback per kernel name, for a failed compilation (e.g. a + vectorized NumPy implementation with the same arguments): a mapping + from names to callables, or a function of the name returning a + callable or None. **cuda_options Passed on to :meth:`CudaKernel.from_file`, e.g. ``block_size`` or ``structs``. @@ -320,7 +512,32 @@ def from_package( name = folder.name if not (folder / f"{name}{host_suffix}.py").is_file(): continue + wrapper = f"bind_c_{name}{host_suffix}" + if check_name_length and len(wrapper) > FORTRAN_NAME_LIMIT: + warnings.warn( + f"kernel {name!r}: the module name {name}{host_suffix} gives the " + f"pyccel Fortran wrapper module {wrapper!r} ({len(wrapper)} " + f"characters), longer than Fortran's limit of " + f"{FORTRAN_NAME_LIMIT}; shorten the kernel name to at most " + f"{FORTRAN_NAME_LIMIT - len('bind_c_') - len(host_suffix)} " + "characters or compile with the C backend", + stacklevel=2, + ) + test_args = None + if ( + test_args_suffix is not None + and (folder / f"{name}{test_args_suffix}.py").is_file() + ): + test_args = f"{package}.{name}.{name}{test_args_suffix}" module = importlib.import_module(f"{package}.{name}.{name}{host_suffix}") + host: Callable[..., Any] = getattr(module, name) + if compile_host is not None: + fallback = ( + host_fallback(name) + if callable(host_fallback) + else (host_fallback or {}).get(name) + ) + host = CompiledHostKernel(module, name, compile_host, fallback) cuda_path = folder / f"{name}{cuda_suffix}" cuda_kernel = ( CudaKernel.from_file(cuda_path, name, **cuda_options) @@ -328,7 +545,7 @@ def from_package( else None ) kernels[name] = Kernel( - getattr(module, name), + host, cuda_kernel, name=name, missing_cuda=missing_cuda, @@ -336,6 +553,8 @@ def from_package( host_options=( host_options(name) if callable(host_options) else host_options ), + dispatch=dispatch, + test_args=test_args, ) return cls(kernels) @@ -414,6 +633,30 @@ def test_parity(name, kernel): (name, kernel) for name, kernel in self._kernels.items() if kernel.has_cuda ] + def check_signatures(self) -> None: + """Check every kernel's host and CUDA parameters (:meth:`Kernel.check_signature`). + + A cheap test for a ported package: call it once, e.g. in a unit test, + to catch a CUDA kernel whose parameters are missing, extra or in + another order than those of its host kernel. + + Raises + ------ + ValueError + Listing every kernel whose two signatures differ. + """ + problems = [] + for kernel in self._kernels.values(): + try: + kernel.check_signature() + except ValueError as error: + problems.append(str(error)) + if problems: + raise ValueError( + "host and CUDA kernels take different parameters:\n " + + "\n ".join(problems) + ) + def compile_all(self, jobs: int | None = 1) -> list[str]: """Compile all CUDA kernels now, e.g. at setup instead of in the first step. diff --git a/src/cunumpy/emulation.py b/src/cunumpy/emulation.py new file mode 100644 index 0000000..453a7b7 --- /dev/null +++ b/src/cunumpy/emulation.py @@ -0,0 +1,436 @@ +"""Run a CUDA kernel on the CPU, one thread after another, for tests without a GPU. + +A ported kernel is usually checked against its host version on a GPU +(:func:`cunumpy.testing.assert_kernels_agree`). Without one, CI cannot run that +check, and the kernel's index and weight arithmetic go untested. +:func:`emulate_cuda_kernel` closes that gap: it compiles the kernel source as +C++ with the CUDA built-ins replaced by plain C++ (``threadIdx``, ``blockIdx``, +``blockDim``, ``gridDim``, ``atomicAdd``, ...), calls the kernel once per +thread, serially, on copies of the NumPy arguments, and copies the arrays back, +so the call looks like a launch:: + + from cunumpy.testing import emulate_cuda_kernel + + y = np.zeros(1000) + emulate_cuda_kernel(axpy, 2.0, x, y, 1000, n_threads=1000) + np.testing.assert_allclose(y, 2.0 * x) + +Arguments follow the kernel signature: NumPy arrays for pointer and array view +parameters (``Array1D`` to ``Array4D``; any strides, they are passed as +contiguous copies), Python or NumPy scalars for scalar parameters (cast and +checked like in a launch). Arrays are written back into the given arrays. + +Block shared memory and ``__syncthreads`` are emulated: ``__shared__`` +variables (and ``extern __shared__`` arrays, sized by ``shared_mem``) are one +copy per block, and in a kernel that calls ``__syncthreads`` the threads of a +block run as coroutines (POSIX ``ucontext``, each with its own stack): every +thread runs to its next barrier before any thread continues past it, as on a +GPU. Per-block deposits, shared-memory reductions and tiled kernels therefore +work. + +What it does not emulate: concurrency between the barriers (threads run one +after another, so atomics are plain additions and races never show), warp +intrinsics (``__shfl_*``, ``__syncwarp``, ``__ballot_sync``, ...; a kernel or an +included header using them is refused), structs and ``CudaArguments`` objects +(not supported), and ````. Use a GPU for those. + +Floating point: like NVRTC (``--fmad=true`` by default) the C++ compiler may +fuse ``a * b + c`` into one fused multiply-add, so results can differ from +NumPy's in the last bit; compare with a tolerance of a few ulp, or pass +``options=("-ffp-contract=off",)`` for NumPy's rounding. + +It needs a C++17 compiler: ``CXX`` from the environment, else ``c++``. +""" + +from __future__ import annotations + +import os +import re +import shutil +import subprocess +import tempfile +from collections.abc import Sequence +from pathlib import Path +from typing import Any + +import numpy as np + +from .cuda_kernel import ( + CudaKernel, + CudaParameter, + _scalar_checker, + _strip_comments, + cuda_include_dir, +) + +__all__ = ["emulate_cuda_kernel", "emulation_compiler"] + +# CUDA constructs that serial emulation would get wrong +_UNSUPPORTED = { + r"__syncwarp": "warp synchronization (__syncwarp)", + r"__shfl\w*": "warp shuffles", + r"__ballot_sync|__any_sync|__all_sync": "warp votes", +} + +_STUBS = r""" +// --- cunumpy emulation: CUDA built-ins as plain C++, one thread at a time --- +#include +#include +#include +#include +#include +#define __device__ +#define __global__ +#define __host__ +#define __forceinline__ inline +#define __noinline__ +#define __restrict__ __restrict +#define __launch_bounds__(...) +// block shared memory: one copy for the block that runs (blocks run one after +// another); `extern __shared__` arrays point into a buffer of shared_mem bytes +#define __shared__ static +static void* cunumpy_dynamic_shared = 0; +struct cunumpy_dim3 { unsigned int x, y, z; }; +static cunumpy_dim3 threadIdx, blockIdx, blockDim, gridDim; +static const int warpSize = 32; +inline void __trap() { abort(); } +template inline T __ldg(const T* p) { return *p; } +inline double rsqrt(double v) { return 1.0 / sqrt(v); } +inline float rsqrtf(float v) { return 1.0f / sqrtf(v); } +using std::min; +using std::max; +template inline T cunumpy_atomic_op_add(T* p, T v) { T o = *p; *p += v; return o; } +inline double atomicAdd(double* p, double v) { return cunumpy_atomic_op_add(p, v); } +inline float atomicAdd(float* p, float v) { return cunumpy_atomic_op_add(p, v); } +inline int atomicAdd(int* p, int v) { return cunumpy_atomic_op_add(p, v); } +inline unsigned int atomicAdd(unsigned int* p, unsigned int v) { return cunumpy_atomic_op_add(p, v); } +inline unsigned long long atomicAdd(unsigned long long* p, unsigned long long v) { return cunumpy_atomic_op_add(p, v); } +inline int atomicSub(int* p, int v) { int o = *p; *p -= v; return o; } +inline unsigned int atomicSub(unsigned int* p, unsigned int v) { unsigned int o = *p; *p -= v; return o; } +template inline T atomicExch(T* p, T v) { T o = *p; *p = v; return o; } +template inline T atomicMin(T* p, T v) { T o = *p; if (v < o) *p = v; return o; } +template inline T atomicMax(T* p, T v) { T o = *p; if (v > o) *p = v; return o; } +inline unsigned long long atomicCAS(unsigned long long* p, unsigned long long c, unsigned long long v) { unsigned long long o = *p; if (o == c) *p = v; return o; } +inline int atomicCAS(int* p, int c, int v) { int o = *p; if (o == c) *p = v; return o; } +inline long long __double_as_longlong(double v) { long long r; memcpy(&r, &v, sizeof r); return r; } +inline double __longlong_as_double(long long v) { double r; memcpy(&r, &v, sizeof r); return r; } +static void cunumpy_read(const char* path, void* data, size_t bytes) { + FILE* f = fopen(path, "rb"); + if (!f || fread(data, 1, bytes, f) != bytes) { perror(path); exit(2); } + fclose(f); +} +static void cunumpy_write(const char* path, const void* data, size_t bytes) { + FILE* f = fopen(path, "wb"); + if (!f || fwrite(data, 1, bytes, f) != bytes) { perror(path); exit(2); } + fclose(f); +} +// --------------------------------------------------------------------------- +""" + +_COROUTINES = r""" +// --- __syncthreads: the threads of a block are coroutines (ucontext); each +// runs until the next barrier, then the scheduler resumes the next one --- +#include +struct cunumpy_thread { ucontext_t ctx; char* stack; int done; }; +static ucontext_t cunumpy_scheduler; +static cunumpy_thread* cunumpy_threads = 0; +static unsigned int cunumpy_current = 0; +#define CUNUMPY_STACK_BYTES (256 * 1024) +inline void __syncthreads() { + swapcontext(&cunumpy_threads[cunumpy_current].ctx, &cunumpy_scheduler); +} +""" + +_MAIN_SERIAL = r""" +{globals} +int main() {{ +{inits} + cunumpy_dynamic_shared = calloc({shared_mem} + 16, 1); + gridDim = {{{gx}u, {gy}u, {gz}u}}; + blockDim = {{{bx}u, {by}u, {bz}u}}; + for (unsigned int bz = 0; bz < gridDim.z; ++bz) + for (unsigned int by = 0; by < gridDim.y; ++by) + for (unsigned int bx = 0; bx < gridDim.x; ++bx) + for (unsigned int tz = 0; tz < blockDim.z; ++tz) + for (unsigned int ty = 0; ty < blockDim.y; ++ty) + for (unsigned int tx = 0; tx < blockDim.x; ++tx) {{ + blockIdx = {{bx, by, bz}}; + threadIdx = {{tx, ty, tz}}; + {call}; + }} +{writes} + return 0; +}} +""" + +_MAIN_COROUTINES = r""" +{globals} +static void cunumpy_entry() {{ + {call}; + cunumpy_threads[cunumpy_current].done = 1; +}} +int main() {{ +{inits} + cunumpy_dynamic_shared = calloc({shared_mem} + 16, 1); + gridDim = {{{gx}u, {gy}u, {gz}u}}; + blockDim = {{{bx}u, {by}u, {bz}u}}; + const unsigned int n = blockDim.x * blockDim.y * blockDim.z; + cunumpy_threads = (cunumpy_thread*)calloc(n, sizeof(cunumpy_thread)); + for (unsigned int t = 0; t < n; ++t) + cunumpy_threads[t].stack = (char*)malloc(CUNUMPY_STACK_BYTES); + for (unsigned int bz = 0; bz < gridDim.z; ++bz) + for (unsigned int by = 0; by < gridDim.y; ++by) + for (unsigned int bx = 0; bx < gridDim.x; ++bx) {{ + blockIdx = {{bx, by, bz}}; + for (unsigned int t = 0; t < n; ++t) {{ + cunumpy_thread* th = &cunumpy_threads[t]; + getcontext(&th->ctx); + th->ctx.uc_stack.ss_sp = th->stack; + th->ctx.uc_stack.ss_size = CUNUMPY_STACK_BYTES; + th->ctx.uc_link = &cunumpy_scheduler; + makecontext(&th->ctx, (void (*)())cunumpy_entry, 0); + th->done = 0; + }} + // rounds: every thread runs to its next __syncthreads (or its end) + // before any thread continues past it + for (int running = 1; running;) {{ + running = 0; + for (unsigned int t = 0; t < n; ++t) {{ + if (cunumpy_threads[t].done) continue; + cunumpy_current = t; + threadIdx = {{t % blockDim.x, (t / blockDim.x) % blockDim.y, + t / (blockDim.x * blockDim.y)}}; + swapcontext(&cunumpy_scheduler, &cunumpy_threads[t].ctx); + if (!cunumpy_threads[t].done) running = 1; + }} + }} + }} +{writes} + return 0; +}} +""" + + +def emulation_compiler() -> str | None: + """The C++ compiler :func:`emulate_cuda_kernel` uses, or None if there is none.""" + compiler = os.environ.get("CXX") + if compiler and shutil.which(compiler): + return shutil.which(compiler) + return shutil.which("c++") + + +def _scalar_literal(param: CudaParameter, index: int, value: Any) -> str: + """A C++ literal of `value`, checked and cast like a scalar kernel argument.""" + cast = _scalar_checker(param, index)(value) + kind = np.dtype(param.dtype).kind + if kind == "b": + return "true" if cast else "false" + if kind in "iu": + return f"(({param.ctype}){int(cast)}{'ULL' if kind == 'u' else 'LL'})" + if kind == "f": + return f"(({param.ctype}){float(cast).hex()})" # exact + raise NotImplementedError(f"emulation does not support {param.ctype} scalars") + + +def _code(kernel: CudaKernel) -> str: + """The kernel source and the headers it includes, without comments.""" + texts = [kernel.source] + for header in kernel.included_headers: + try: + texts.append(Path(header).read_text(errors="replace")) + except OSError: + pass + return _strip_comments("\n".join(texts)) + + +_EXTERN_SHARED = re.compile( + r"extern\s+__shared__\s+(?:__align__\(\s*\d+\s*\)\s+)?([\w:<>\s]+?)\s+(\w+)\s*\[\s*\]\s*;" +) + + +def _check_supported(kernel: CudaKernel) -> None: + code = _code(kernel) + for pattern, what in _UNSUPPORTED.items(): + if re.search(pattern, code): + raise NotImplementedError( + f"kernel {kernel.name!r} uses {what}, which emulation cannot run " + "correctly one thread at a time; test it on a GPU" + ) + + +def emulate_cuda_kernel( + kernel: CudaKernel, + *args: Any, + n_threads: int | Sequence[int] | None = None, + grid: int | Sequence[int] | None = None, + block: int | Sequence[int] | None = None, + compiler: str | None = None, + options: Sequence[str] = (), + shared_mem: int = 0, +) -> None: + """Run `kernel` on the CPU, serially, as if it were launched with `args`. + + Parameters + ---------- + kernel : CudaKernel + The kernel; its signature must be parsed (the default). + *args + The kernel arguments with NumPy arrays in place of CuPy arrays. Arrays + are updated in place with what the kernel wrote. + n_threads, grid, block + Launch shape, as for :meth:`CudaKernel.__call__`. + compiler : str | None + C++ compiler; by default :func:`emulation_compiler`. + options : Sequence[str] + Additional compiler options, e.g. ``("-DMY_FLAG=1",)``. ``-D`` options + of the kernel are passed on as well. + shared_mem : int + Dynamic shared memory per block in bytes, for ``extern __shared__`` + arrays, as in a launch. + + Raises + ------ + NotImplementedError + If the kernel (or a header it includes) uses warp intrinsics, or the + kernel has struct parameters or complex scalars. + TypeError + If an argument does not match its parameter (dtype, dimensions, a + scalar that does not fit). + RuntimeError + If there is no C++ compiler, or the kernel does not compile or crashes. + """ + if kernel.signature is None: + raise TypeError( + "emulation needs a parsed kernel signature (check_signature=True)" + ) + _check_supported(kernel) + params = kernel.signature + if len(args) != len(params): + raise TypeError( + f"kernel {kernel.name!r} takes {len(params)} arguments, got {len(args)}" + ) + compiler = compiler or emulation_compiler() + if compiler is None: + raise RuntimeError("emulation needs a C++ compiler (set CXX or install c++)") + grid_shape, block_shape = kernel.launch_shape(n_threads, grid=grid, block=block) + grid_shape = tuple(grid_shape) + (1,) * (3 - len(grid_shape)) + block_shape = tuple(block_shape) + (1,) * (3 - len(block_shape)) + + with tempfile.TemporaryDirectory(prefix="cunumpy-emulation-") as tmp: + tmp_path = Path(tmp) + globals_, inits, call_args, writes, arrays = [], [], [], [], [] + for i, (param, value) in enumerate(zip(params, args)): + name = f"cunumpy_arg{i}" + if param.struct is not None: + raise NotImplementedError( + "emulation does not support struct parameters" + ) + if param.pointer or param.view_ndim is not None: + if not isinstance(value, np.ndarray): + raise TypeError( + f"argument {i} ({param.name}) must be a NumPy array, got " + f"{type(value).__name__}" + ) + if param.dtype is not None and value.dtype != param.dtype: + raise TypeError( + f"argument {i} ({param.name}) must have dtype " + f"{np.dtype(param.dtype)}, got {value.dtype}" + ) + if param.view_ndim is not None and value.ndim != param.view_ndim: + raise TypeError( + f"argument {i} ({param.name}) must be a {param.view_ndim}D " + f"array, got {value.ndim}D" + ) + buffer = np.ascontiguousarray(value) + path = tmp_path / f"{name}.bin" + buffer.tofile(path) + ctype = param.ctype if param.dtype is not None else "unsigned char" + element = ( + re.match(r"Array\dD<(.*)>", ctype).group(1) + if param.view_ndim is not None + else ctype + ) + globals_.append(f"static {element}* {name};") + inits.append( + f" {name} = ({element}*)malloc({max(buffer.nbytes, 1)});\n" + f' cunumpy_read("{path}", {name}, {buffer.nbytes});' + ) + if param.view_ndim is not None: + shape = ", ".join(f"{n}LL" for n in buffer.shape) + strides = ", ".join( + f"{s // buffer.itemsize}LL" for s in buffer.strides + ) + globals_.append(f"static {ctype} {name}_view;") + inits.append( + f" {name}_view = {ctype}{{{name}, {{{shape}}}, {{{strides}}}}};" + ) + call_args.append(f"{name}_view") + else: + call_args.append(name) + out = tmp_path / f"{name}.out" + writes.append(f' cunumpy_write("{out}", {name}, {buffer.nbytes});') + arrays.append((value, buffer, out)) + else: + call_args.append(_scalar_literal(param, i, value)) + + if shared_mem < 0: + raise ValueError(f"shared_mem must be non-negative, got {shared_mem}") + kernel_source = kernel.source.replace('extern "C"', "") + kernel_source = _EXTERN_SHARED.sub( + r"\1* \2 = (\1*)cunumpy_dynamic_shared;", kernel_source + ) + coroutines = "__syncthreads" in _code(kernel) + if coroutines: + # ucontext needs _XOPEN_SOURCE before any system header + source = "#define _XOPEN_SOURCE 700\n" + _STUBS + _COROUTINES + main = _MAIN_COROUTINES + else: + source = _STUBS + "inline void __syncthreads() {}\n" + main = _MAIN_SERIAL + source += kernel_source + source += main.format( + globals="\n".join(globals_), + inits="\n".join(inits), + shared_mem=int(shared_mem), + gx=grid_shape[0], + gy=grid_shape[1], + gz=grid_shape[2], + bx=block_shape[0], + by=block_shape[1], + bz=block_shape[2], + call=f"{kernel.expression}({', '.join(call_args)})", + writes="\n".join(writes), + ) + cpp = tmp_path / "kernel.cpp" + cpp.write_text(source) + exe = tmp_path / "kernel" + include_dirs = [cuda_include_dir(), *map(str, kernel.include_dirs)] + if kernel.source_dir is not None: + include_dirs.insert(0, str(kernel.source_dir)) + defines = [o for o in kernel.options if o.startswith("-D")] + command = [ + compiler, + "-std=c++17", + "-O1", + "-w", + *(f"-I{d}" for d in include_dirs), + *defines, + *options, + str(cpp), + "-o", + str(exe), + ] + built = subprocess.run(command, capture_output=True, text=True, check=False) + if built.returncode: + raise RuntimeError( + f"kernel {kernel.name!r} does not compile for emulation:\n{built.stderr[:4000]}" + ) + ran = subprocess.run([str(exe)], capture_output=True, text=True, check=False) + if ran.returncode: + raise RuntimeError( + f"kernel {kernel.name!r} crashed in emulation (exit {ran.returncode}):\n" + f"{ran.stdout[-2000:]}{ran.stderr[-2000:]}" + ) + for value, buffer, out in arrays: + result = np.fromfile(out, dtype=buffer.dtype).reshape(buffer.shape) + value[...] = result diff --git a/src/cunumpy/fusion.py b/src/cunumpy/fusion.py new file mode 100644 index 0000000..af72edf --- /dev/null +++ b/src/cunumpy/fusion.py @@ -0,0 +1,113 @@ +"""``xp.fuse``: elementwise functions as one GPU kernel, plain calls on the host. + +A chain of elementwise operations (a pressure from density and temperature, +fluxes and limiters of a fluid update, a Maxwellian at many velocities) runs +as one kernel per operation on the GPU, each writing a temporary array: it is +limited by memory bandwidth. ``cupy.fuse`` compiles the whole chain into one +kernel. :func:`fuse` applies it when the function is called with CuPy arrays +and calls the function as it is otherwise, so the same code runs on both +backends:: + + @xp.fuse + def pressure(rho, T, gamma): + return (gamma - 1.0) * rho * T + + p = pressure(rho, T, 5.0 / 3.0) + +The function must be elementwise: arithmetic, comparisons, ufuncs such as +``xp.exp``, ``xp.sqrt``, ``xp.where`` and reductions that ``cupy.fuse`` +supports (``xp.sum`` as the last operation). On the CuPy path the CuPy +backend is active while the function is traced, so ``xp.*`` names resolve to +CuPy's ufuncs. Functions that ``cupy.fuse`` cannot trace (Python control +flow on array values, indexing, wrappers that are not ufuncs) raise when the +fused function is first called with CuPy arrays; test the CuPy path. +""" + +from __future__ import annotations + +import functools +from collections.abc import Callable +from typing import Any, TypeVar + +import array_api_compat +import numpy + +from .xp import use_backend + +__all__ = ["fuse"] + +F = TypeVar("F", bound=Callable[..., Any]) + +# substituted in tests that have no GPU +_is_device_array = array_api_compat.is_cupy_array + + +def _cupy_fuse(function: Callable[..., Any], kernel_name: str | None) -> Any: + import cupy + + return cupy.fuse(kernel_name=kernel_name)(function) + + +def _typed_scalars(args: tuple, kwargs: dict) -> tuple[tuple, dict]: + """Give Python scalars the dtype they would take next to the arrays. + + ``cupy.fuse`` types a Python scalar on its own: ``gamma - 1.0`` with + ``gamma=5/3`` runs in float16. Eagerly, NumPy 2 and CuPy promote it with + the arrays (float64 arrays: float64), so cast it to that dtype first. + """ + dtypes = [a.dtype for a in (*args, *kwargs.values()) if hasattr(a, "dtype")] + + def typed(a: Any) -> Any: + if isinstance(a, (int, float, complex)) and not isinstance(a, bool): + return numpy.result_type(*dtypes, a).type(a) + return a + + return ( + tuple(typed(a) for a in args), + {k: typed(v) for k, v in kwargs.items()}, + ) + + +def fuse( + function: F | None = None, *, kernel_name: str | None = None +) -> F | Callable[[F], F]: + """Fuse an elementwise function into one kernel when called with CuPy arrays. + + Usable as ``@xp.fuse`` or ``@xp.fuse(kernel_name="pressure")``. + + Parameters + ---------- + function : callable + The elementwise function. + kernel_name : str | None + Name of the generated kernel (shown by profilers); by default the + function's name. + + Returns + ------- + callable + A function with the same signature. If any positional or keyword + argument is a CuPy array, it calls ``cupy.fuse(function)`` (created on + first use, with the CuPy backend active) with Python scalars cast to + the dtype they promote to with the array arguments; otherwise it calls + `function` itself. + """ + if function is None: + return lambda f: fuse(f, kernel_name=kernel_name) + + name = kernel_name if kernel_name is not None else function.__name__ + fused = None + + @functools.wraps(function) + def wrapper(*args: Any, **kwargs: Any) -> Any: + nonlocal fused + if not any(_is_device_array(a) for a in (*args, *kwargs.values())): + return function(*args, **kwargs) + args, kwargs = _typed_scalars(args, kwargs) + with use_backend("cupy"): + if fused is None: + fused = _cupy_fuse(function, name) + return fused(*args, **kwargs) + + wrapper.__wrapped__ = function + return wrapper diff --git a/src/cunumpy/kernel.py b/src/cunumpy/kernel.py index e664283..463103a 100644 --- a/src/cunumpy/kernel.py +++ b/src/cunumpy/kernel.py @@ -23,7 +23,10 @@ from __future__ import annotations import copy +import importlib +import warnings from collections.abc import Callable, Mapping, Sequence +from types import ModuleType from typing import Any import array_api_compat @@ -33,7 +36,7 @@ from .transfers import _record from .xp import _cupy_backend, _to_cupy, _to_numpy -__all__ = ["KernelArguments", "PyccelKernel", "resolve_host_args"] +__all__ = ["CompiledHostKernel", "KernelArguments", "PyccelKernel", "resolve_host_args"] class KernelArguments: @@ -506,3 +509,100 @@ def is_array(self) -> Callable[[Any], bool]: def outputs(self) -> tuple[int | str, ...] | None: """Declared output arguments, or ``None`` if every array is copied back.""" return self._outputs + + +class CompiledHostKernel: + """A host kernel compiled on its first call, with a fallback. + + cunumpy does not compile anything itself: `compiler` is the caller's + function that turns the kernel module into a compiled one, e.g. a wrapper + around ``pyccel.epyccel`` with a cache, or a function importing modules + compiled ahead of time. + + Parameters + ---------- + module : str | ModuleType + The module that defines the kernel, or its name. + name : str + Name of the kernel function in `module`. + compiler : Callable[[ModuleType], ModuleType] + Returns the compiled form of the module (with the same function names). + fallback : Callable | None + Called instead when compilation fails (e.g. a vectorized NumPy + implementation with the same arguments). Without one, the uncompiled + Python function is called, with a warning: correct, but slow. + + Notes + ----- + :attr:`compiled` builds the kernel and reports whether that worked, so + that callers can choose another code path; :attr:`error` keeps the + exception of a failed build. + """ + + def __init__( + self, + module: str | ModuleType, + name: str, + compiler: Callable[[ModuleType], ModuleType], + fallback: Callable[..., Any] | None = None, + ) -> None: + self._module = module + self.__name__ = name + self._compiler = compiler + self._fallback = fallback + self._compiled: Callable[..., Any] | None = None + self._built = False + self.error: BaseException | None = None + self._warned = False + + def __repr__(self) -> str: + state = ( + "not built" + if not self._built + else ("compiled" if self._compiled else "failed") + ) + return f"CompiledHostKernel({self.__name__!r}, {state})" + + @property + def module(self) -> ModuleType: + """The (uncompiled) module that defines the kernel.""" + if isinstance(self._module, str): + self._module = importlib.import_module(self._module) + return self._module + + @property + def python(self) -> Callable[..., Any]: + """The uncompiled Python function (for signatures and reference results).""" + return getattr(self.module, self.__name__) + + def build(self) -> Callable[..., Any] | None: + """Compile now (once); the compiled function, or None if that failed.""" + if not self._built: + self._built = True + try: + self._compiled = getattr(self._compiler(self.module), self.__name__) + except Exception as error: # noqa: BLE001 -- no Pyccel or no compiler + self.error = error + return self._compiled + + @property + def compiled(self) -> bool: + """Whether the compiled version is available (compiles on first access).""" + return self.build() is not None + + def __call__(self, *args: Any, **kwargs: Any) -> Any: + kernel = self.build() + if kernel is None: + if self._fallback is not None: + kernel = self._fallback + else: + if not self._warned: + warnings.warn( + f"Kernel {self.__name__!r} could not be compiled " + f"({self.error!r}); running its uncompiled Python version.", + RuntimeWarning, + stacklevel=2, + ) + self._warned = True + kernel = self.python + return kernel(*args, **kwargs) diff --git a/src/cunumpy/petsc.py b/src/cunumpy/petsc.py new file mode 100644 index 0000000..a5ca0a9 --- /dev/null +++ b/src/cunumpy/petsc.py @@ -0,0 +1,122 @@ +"""PETSc vectors sharing memory with NumPy or CuPy arrays (no copies). + +Field solvers often go through PETSc (KSP), while the rest of a GPU code keeps +its data in CuPy arrays. Copying the right-hand side to the host and the +solution back at every solve is the largest transfer of such a time step. +:func:`petsc_vec` wraps an array as a PETSc vector that uses the array's +memory: a host vector for a NumPy array, a CUDA (or HIP) vector for a CuPy +array, which needs a petsc4py built with CUDA (or HIP) support:: + + b = xp.zeros(n) # filled by the deposit kernel + phi = xp.zeros(n) # the solution, read by the gather kernel + b_vec, phi_vec = xp.petsc_vec(b), xp.petsc_vec(phi) + ... + xp.synchronize() # CuPy work on b done before PETSc reads it + ksp.solve(b_vec, phi_vec) # writes into phi + xp.synchronize() # PETSc done before CuPy reads phi + +For the solve to stay on the GPU, the matrix must be a GPU type as well +(``mat.setType("aijcusparse")``, or ``-mat_type aijcusparse -vec_type cuda`` +in the PETSc options); otherwise PETSc copies the vectors to the host for the +matrix products. + +petsc4py is imported only when :func:`petsc_vec` is called. +""" + +from __future__ import annotations + +from typing import Any + +import array_api_compat +import numpy as np + +__all__ = ["petsc_vec"] + +# substituted in tests that have no GPU +_is_device_array = array_api_compat.is_cupy_array + +_DEVICE_VEC_TYPES = ("cuda", "hip") + + +def _petsc() -> Any: + try: + from petsc4py import PETSc + except ImportError as error: + raise ImportError( + "xp.petsc_vec needs petsc4py (pip install petsc4py)" + ) from error + return PETSc + + +def petsc_vec(array: Any, comm: Any = None) -> Any: + """A PETSc vector that shares the memory of `array`. + + Parameters + ---------- + array : numpy.ndarray | cupy.ndarray + C-contiguous array of PETSc's scalar type (``PETSc.ScalarType``, + usually float64). A multi-dimensional array is seen by PETSc in + row-major order, as ``array.ravel()``. The array is never copied. + comm : mpi4py.MPI.Comm | PETSc.Comm | None + Communicator of the vector; PETSc's default (``COMM_WORLD``) if None. + With several processes, `array` is this process's part. + + Returns + ------- + petsc4py.PETSc.Vec + A sequential or MPI vector (``seq``/``mpi`` for a NumPy array, + ``seqcuda``/``mpicuda`` or ``seqhip``/``mpihip`` for a CuPy array). + It keeps a reference to `array`, which therefore stays alive as long + as the vector. + + Raises + ------ + TypeError + If `array` has another dtype than ``PETSc.ScalarType``, or is not an + array. + ValueError + If `array` is not C-contiguous. + RuntimeError + If `array` is a CuPy array and petsc4py has no CUDA or HIP support + (PETSc would otherwise silently work on a host copy). + ImportError + If petsc4py is not installed. + + Notes + ----- + PETSc and CuPy may run on different streams: synchronize + (:func:`cunumpy.synchronize`) before PETSc reads an array that CuPy wrote, + and before CuPy reads a vector that PETSc wrote. + """ + PETSc = _petsc() + device = _is_device_array(array) + if not (device or isinstance(array, np.ndarray)): + raise TypeError( + f"petsc_vec takes a NumPy or CuPy array, got {type(array).__name__}" + ) + scalar = np.dtype(PETSc.ScalarType) + if array.dtype != scalar: + raise TypeError( + f"petsc_vec needs an array of PETSc's scalar type {scalar}, got " + f"{array.dtype} (convert it once, outside the time loop)" + ) + if not array.flags.c_contiguous: + raise ValueError("petsc_vec needs a C-contiguous array (no copy is made)") + + try: + vec = PETSc.Vec().createWithDLPack(array, comm=comm) + except PETSc.Error as error: + if device: + raise RuntimeError( + "petsc4py could not wrap the CuPy array; it needs a PETSc built " + "with CUDA or HIP support (--with-cuda / --with-hip)" + ) from error + raise + if device and not any(t in vec.getType() for t in _DEVICE_VEC_TYPES): + vec.destroy() + raise RuntimeError( + f"petsc4py created a {vec.getType()!r} vector for a CuPy array; it " + "needs a PETSc built with CUDA or HIP support" + ) + vec.setAttr("cunumpy_array", array) # the vector does not own the memory + return vec diff --git a/src/cunumpy/philox.py b/src/cunumpy/philox.py new file mode 100644 index 0000000..6f0dbfc --- /dev/null +++ b/src/cunumpy/philox.py @@ -0,0 +1,148 @@ +"""Counter-based random numbers (Philox4x32-10), the same on host and device. + +The host side of ``cunumpy/random.cuh``: :func:`philox_uniform` returns, for +every (seed, stream, counter), exactly the doubles that ``cunumpy_uniform`` / +``cunumpy_uniform2`` return in a kernel, so a kernel that samples random +numbers can be compared with its host version element by element:: + + ids = xp.arange(n, dtype=xp.uint64) # one stream per particle + u0, u1 = xp.philox_uniform2(seed, ids, step) # what each GPU thread draws + z0, z1 = xp.philox_normal2(seed, ids, step) + +The numbers are a pure function of the key (``seed``, 64 bit) and the 128-bit +counter (``counter`` and ``stream``, 64 bit each): no state, independent of the +launch shape. Arguments broadcast like NumPy arrays; the result is a NumPy or a +CuPy array, matching the inputs (NumPy for Python ints). Uniform numbers are +bit-identical to the device ones; normal numbers (Box-Muller with ``log``, +``sqrt``, ``sin``, ``cos``) can differ in the last bits, since the device math +functions are not the host's. +""" + +from __future__ import annotations + +from typing import Any + +import array_api_compat +import numpy as np + +__all__ = [ + "philox4x32_10", + "philox_normal", + "philox_normal2", + "philox_uniform", + "philox_uniform2", +] + +_M0, _M1 = 0xD2511F53, 0xCD9E8D57 +_W0, _W1 = 0x9E3779B9, 0xBB67AE85 +_MASK32 = 0xFFFFFFFF + + +def _module(*values: Any) -> Any: + """CuPy if any value is a CuPy array, else NumPy.""" + for value in values: + if array_api_compat.is_cupy_array(value): + import cupy + + return cupy + return np + + +def philox4x32_10(counter: Any, key0: Any, key1: Any) -> Any: + """Philox4x32-10 of a ``(..., 4)`` array of uint32 counter words and a 2x32-bit key. + + Parameters + ---------- + counter : array of uint32, shape (..., 4) + The counter words ``(c0, c1, c2, c3)``. + key0, key1 : int or array of uint32 + The key words, broadcast against ``counter[..., 0]``. + + Returns + ------- + array of uint32, shape (..., 4) + The random words, as ``cunumpy_philox4x32_10`` computes them. + """ + xp = _module(counter, key0, key1) + ctr = xp.asarray(counter, dtype=xp.uint64) + c0, c1, c2, c3 = (ctr[..., i] for i in range(4)) + k0 = xp.asarray(key0, dtype=xp.uint64) & _MASK32 + k1 = xp.asarray(key1, dtype=xp.uint64) & _MASK32 + m0, m1 = xp.uint64(_M0), xp.uint64(_M1) + w0, w1 = xp.uint64(_W0), xp.uint64(_W1) + mask, shift = xp.uint64(_MASK32), xp.uint64(32) + for _ in range(10): + p0 = m0 * c0 # < 2^64: exact in uint64 + p1 = m1 * c2 + c0, c1, c2, c3 = ( + (p1 >> shift) ^ c1 ^ k0, + p1 & mask, + (p0 >> shift) ^ c3 ^ k1, + p0 & mask, + ) + k0 = (k0 + w0) & mask + k1 = (k1 + w1) & mask + return xp.stack(xp.broadcast_arrays(c0, c1, c2, c3), axis=-1).astype(xp.uint32) + + +def _random_bits(seed: Any, stream: Any, counter: Any) -> Any: + xp = _module(seed, stream, counter) + seed = xp.asarray(seed, dtype=xp.uint64) + stream = xp.asarray(stream, dtype=xp.uint64) + counter = xp.asarray(counter, dtype=xp.uint64) + mask, shift = xp.uint64(_MASK32), xp.uint64(32) + seed, stream, counter = xp.broadcast_arrays(seed, stream, counter) + words = xp.stack( + [counter & mask, counter >> shift, stream & mask, stream >> shift], axis=-1 + ) + return philox4x32_10(words, seed & mask, seed >> shift) + + +def _to_uniform(xp: Any, hi: Any, lo: Any) -> Any: + bits = (hi.astype(xp.uint64) << xp.uint64(32)) | lo.astype(xp.uint64) + return (bits >> xp.uint64(11)).astype(xp.float64) * (1.0 / 9007199254740992.0) + + +def philox_uniform2(seed: Any, stream: Any, counter: Any) -> tuple[Any, Any]: + """Two uniform doubles in [0, 1) per (seed, stream, counter), as ``cunumpy_uniform2``. + + Parameters + ---------- + seed : int or array + The key (64 bit). + stream : int or array + The stream, e.g. a particle id (64 bit). + counter : int or array + The draw index, e.g. the time step (64 bit). + + Returns + ------- + tuple of arrays + ``(u0, u1)``, broadcast to the shape of the arguments. + """ + xp = _module(seed, stream, counter) + r = _random_bits(seed, stream, counter) + return _to_uniform(xp, r[..., 0], r[..., 1]), _to_uniform(xp, r[..., 2], r[..., 3]) + + +def philox_uniform(seed: Any, stream: Any, counter: Any) -> Any: + """One uniform double in [0, 1) per (seed, stream, counter), as ``cunumpy_uniform``.""" + return philox_uniform2(seed, stream, counter)[0] + + +def philox_normal2(seed: Any, stream: Any, counter: Any) -> tuple[Any, Any]: + """Two standard normal doubles per (seed, stream, counter), as ``cunumpy_normal2``. + + Box-Muller of :func:`philox_uniform2`; equal to the device numbers up to the + last bits of the math functions. + """ + xp = _module(seed, stream, counter) + u0, u1 = philox_uniform2(seed, stream, counter) + radius = xp.sqrt(-2.0 * xp.log(1.0 - u0)) + angle = 6.283185307179586 * u1 + return radius * xp.cos(angle), radius * xp.sin(angle) + + +def philox_normal(seed: Any, stream: Any, counter: Any) -> Any: + """One standard normal double per (seed, stream, counter), as ``cunumpy_normal``.""" + return philox_normal2(seed, stream, counter)[0] diff --git a/src/cunumpy/random_streams.py b/src/cunumpy/random_streams.py new file mode 100644 index 0000000..f8dcb76 --- /dev/null +++ b/src/cunumpy/random_streams.py @@ -0,0 +1,192 @@ +"""Reproducible random numbers for MPI programs on either backend. + +A simulation that draws random numbers in many places (initial loading, +injection, collisions) is reproducible only if every draw comes from a seeded +generator, and an MPI run needs a different stream on every rank. +:data:`random_streams` is one generator per process and backend for that:: + + import cunumpy as xp + + xp.random_streams.seed(42, rank=comm.Get_rank()) # once, at start-up + v = xp.random_streams.normal(0.0, v_th, (n, 3)) # anywhere afterwards + rng = xp.random_streams.generator() # the Generator itself + +Each rank draws the stream ``(seed, rank)`` (a NumPy ``SeedSequence`` with the +rank as spawn key), so a run with the same seed and the same number of ranks +reproduces its results exactly, and the ranks' streams are independent. The +generator of each backend is created on first use from the same stream: a +``numpy.random.Generator`` with the chosen bit generator, or a +``cupy.random.Generator`` (CuPy's default bit generator). + +Components with a seed of their own (e.g. a source configured with a ``seed``) +get a separate generator with :meth:`RandomStreams.make_generator`; without one +they share the process generator. Without :meth:`~RandomStreams.seed`, or with +``seed(None)``, the stream is seeded from the operating system. + +The draw functions work with NumPy and CuPy generators alike; CuPy's +``Generator`` lacks some of NumPy's methods, and the missing ones are derived +from ``random`` and ``standard_normal``. +""" + +from __future__ import annotations + +from typing import Any + +import numpy as np + +from .xp import get_backend + +__all__ = ["BIT_GENERATORS", "RandomStreams", "random_streams"] + +#: NumPy bit generators :meth:`RandomStreams.seed` accepts by name. +BIT_GENERATORS = ("MT19937", "PCG64", "PCG64DXSM", "Philox", "SFC64") + +_BACKEND_KEYS = {"numpy": 0, "cupy": 1} + + +class RandomStreams: + """One seeded random generator per process and backend; see :mod:`cunumpy.random_streams`.""" + + def __init__(self) -> None: + self._sequence: np.random.SeedSequence | None = None + self._bit_generator = "PCG64" + self._generators: dict[str, Any] = {} + + def __repr__(self) -> str: + if self._sequence is None: + return "RandomStreams(not seeded)" + return ( + f"RandomStreams(entropy={self._sequence.entropy}, " + f"spawn_key={self._sequence.spawn_key}, bit_generator={self._bit_generator!r})" + ) + + def seed( + self, value: int | None, rank: int = 0, bit_generator: str | None = None + ) -> None: + """Seed all draws of this process with the stream ``(value, rank)``. + + Replaces the generators of all backends. NumPy's global random state + (and, on the CuPy backend, CuPy's) is seeded from the same stream, for + code that still calls ``np.random.*`` or ``xp.random.*`` directly. + + Parameters + ---------- + value : int | None + The seed; None seeds from the operating system. + rank : int + MPI rank of this process, so that ranks draw independent streams. + bit_generator : str | None + The NumPy bit generator, one of :data:`BIT_GENERATORS`; ``PCG64`` + (NumPy's default) if None. CuPy always uses its own default. + + Raises + ------ + ValueError + For an unknown `bit_generator`. + """ + if bit_generator is not None and bit_generator not in BIT_GENERATORS: + raise ValueError( + f"Unknown bit generator {bit_generator!r}; use one of " + f"{', '.join(BIT_GENERATORS)}" + ) + self._bit_generator = bit_generator or "PCG64" + self._sequence = np.random.SeedSequence( + entropy=None if value is None else int(value), spawn_key=(int(rank),) + ) + self._generators.clear() + legacy = int(self._sequence.generate_state(1)[0]) + np.random.seed(legacy) + if get_backend() == "cupy": + import cupy + + cupy.random.seed(legacy) + + def _backend_sequence(self, backend: str) -> np.random.SeedSequence: + if self._sequence is None: + self._sequence = np.random.SeedSequence() + return np.random.SeedSequence( + entropy=self._sequence.entropy, + spawn_key=(*self._sequence.spawn_key, _BACKEND_KEYS[backend]), + ) + + def generator(self, backend: str | None = None) -> Any: + """The process generator of `backend` (the active backend by default).""" + backend = get_backend() if backend is None else backend + if backend not in _BACKEND_KEYS: + raise ValueError(f"backend must be 'numpy' or 'cupy', got {backend!r}") + if backend not in self._generators: + sequence = self._backend_sequence(backend) + if backend == "numpy": + bits = getattr(np.random, self._bit_generator)(sequence) + self._generators[backend] = np.random.Generator(bits) + else: + import cupy + + state = int(sequence.generate_state(1, dtype=np.uint64)[0]) + self._generators[backend] = cupy.random.default_rng(state) + return self._generators[backend] + + def make_generator( + self, seed: int | None = None, backend: str | None = None + ) -> Any: + """A generator for a component: its own if it has a seed, else the process one. + + Parameters + ---------- + seed : int | None + The component's own seed; None shares :meth:`generator`. + backend : str | None + ``"numpy"`` for a component that draws on the host whatever the + active backend; the active backend by default. + """ + if seed is None: + return self.generator(backend) + backend = get_backend() if backend is None else backend + if backend == "numpy": + return np.random.default_rng(int(seed)) + import cupy + + return cupy.random.default_rng(int(seed)) + + def _rng(self, rng: Any) -> Any: + return self.generator() if rng is None else rng + + def random(self, size: int | tuple[int, ...] | None = None, rng: Any = None) -> Any: + """Uniform samples in [0, 1) from `rng` (the process generator by default).""" + return self._rng(rng).random(size=size) + + def standard_normal( + self, size: int | tuple[int, ...] | None = None, rng: Any = None + ) -> Any: + """Standard normal samples from `rng` (the process generator by default).""" + return self._rng(rng).standard_normal(size=size) + + def normal( + self, + loc: Any = 0.0, + scale: Any = 1.0, + size: int | tuple[int, ...] | None = None, + rng: Any = None, + ) -> Any: + """Normal samples from `rng` (the process generator by default).""" + rng = self._rng(rng) + if hasattr(rng, "normal"): + return rng.normal(loc=loc, scale=scale, size=size) + return loc + scale * rng.standard_normal(size=size) + + def uniform( + self, + low: Any = 0.0, + high: Any = 1.0, + size: int | tuple[int, ...] | None = None, + rng: Any = None, + ) -> Any: + """Uniform samples in [low, high) from `rng` (the process generator by default).""" + rng = self._rng(rng) + if hasattr(rng, "uniform"): + return rng.uniform(low=low, high=high, size=size) + return low + (high - low) * rng.random(size=size) + + +#: The random streams of this process. +random_streams = RandomStreams() diff --git a/src/cunumpy/scipy_backend.py b/src/cunumpy/scipy_backend.py new file mode 100644 index 0000000..94bb74b --- /dev/null +++ b/src/cunumpy/scipy_backend.py @@ -0,0 +1,157 @@ +"""SciPy for the active backend: ``xp.scipy`` is SciPy or ``cupyx.scipy``. + +Field solvers and fluid codes need more than array functions: sparse matrices +and their iterative solvers, FFTs, special functions, image filters. +``xp.scipy`` forwards to :mod:`scipy` on the NumPy backend and to +:mod:`cupyx.scipy` on the CuPy backend, so this code runs on both:: + + A = xp.scipy.sparse.csr_matrix((data, (rows, cols)), shape=(n, n)) + x, info = xp.scipy.sparse.linalg.cg(A, b) + phi_k = xp.scipy.fft.rfftn(rho) + f = xp.scipy.special.erf(v / v_th) + +The module is resolved at every attribute access, so switching the backend +(:func:`cunumpy.set_backend`) takes effect immediately; imports are cached by +Python, so the lookup is cheap. Neither SciPy nor CuPy is imported until a +name is used. + +``cupyx.scipy`` covers only part of SciPy. A name that the active backend's +module does not have raises ``AttributeError`` saying which backend lacks it; +:meth:`ScipyNamespace.available` checks a name without raising. Keyword +arguments can differ as well: SciPy's ``cg`` takes ``rtol`` since SciPy 1.12, +CuPy's still takes ``tol``. +""" + +from __future__ import annotations + +import importlib +from types import ModuleType +from typing import Any + +from .xp import get_backend + +__all__ = ["SUBMODULES", "ScipyNamespace", "scipy"] + +#: The SciPy subpackages forwarded by ``xp.scipy``: those that ``cupyx.scipy`` +#: provides as well. +SUBMODULES = ( + "fft", + "fftpack", + "interpolate", + "linalg", + "ndimage", + "signal", + "sparse", + "sparse.csgraph", + "sparse.linalg", + "spatial", + "special", + "stats", +) + +_ROOTS = {"numpy": "scipy", "cupy": "cupyx.scipy"} +_INSTALL = { + "numpy": "SciPy is not installed (pip install scipy)", + "cupy": "CuPy is not installed or does not provide it", +} + + +class ScipyNamespace: + """SciPy (sub)package of the active backend; see :mod:`cunumpy.scipy_backend`. + + Parameters + ---------- + path : str + Dotted path below the SciPy root, e.g. ``"sparse.linalg"``; empty for + the root itself. + """ + + def __init__(self, path: str = "") -> None: + if path and path not in SUBMODULES: + raise ValueError( + f"scipy.{path} is not forwarded; forwarded subpackages: " + f"{', '.join(SUBMODULES)}" + ) + self._path = path + self._children: dict[str, ScipyNamespace] = {} + + def __repr__(self) -> str: + return f"')}>" + + def _name(self, root: str) -> str: + return f"{root}.{self._path}" if self._path else root + + def resolve(self) -> ModuleType: + """The module this namespace stands for on the active backend. + + Returns + ------- + ModuleType + E.g. ``scipy.sparse.linalg`` or ``cupyx.scipy.sparse.linalg``. + + Raises + ------ + ImportError + If SciPy (NumPy backend) or CuPy (CuPy backend) is not installed. + """ + backend = get_backend() + name = self._name(_ROOTS[backend]) + try: + return importlib.import_module(name) + except ImportError as error: + raise ImportError( + f"xp.scipy on the {backend} backend needs {name}: {_INSTALL[backend]}" + ) from error + + def __getattr__(self, name: str) -> Any: + if name.startswith("__"): + raise AttributeError(name) + child = f"{self._path}.{name}" if self._path else name + if child in SUBMODULES: + if name not in self._children: + self._children[name] = ScipyNamespace(child) + return self._children[name] + module = self.resolve() + try: + return getattr(module, name) + except AttributeError: + other = "cupy" if get_backend() == "numpy" else "numpy" + raise AttributeError( + f"{module.__name__} has no attribute {name!r}: it is not available " + f"on the {get_backend()} backend (it may exist in " + f"{self._name(_ROOTS[other])})" + ) from None + + def available(self, name: str) -> bool: + """Whether `name` exists in this namespace on the active backend. + + Never raises; False also if the backend's SciPy is not installed. + """ + child = f"{self._path}.{name}" if self._path else name + if child in SUBMODULES: + try: + ScipyNamespace(child).resolve() + except ImportError: + return False + return True + try: + return hasattr(self.resolve(), name) + except ImportError: + return False + + def __dir__(self) -> list[str]: + prefix = f"{self._path}." if self._path else "" + children = [ + s[len(prefix) :] + for s in SUBMODULES + if s.startswith(prefix) and "." not in s[len(prefix) :] + ] + try: + names = dir(self.resolve()) + except ImportError: + names = [] + return sorted(set(names) | set(children) | {"available", "resolve"}) + + +#: SciPy of the active backend, available as ``xp.scipy``. +scipy = ScipyNamespace() diff --git a/src/cunumpy/staging.py b/src/cunumpy/staging.py new file mode 100644 index 0000000..f59ad7b --- /dev/null +++ b/src/cunumpy/staging.py @@ -0,0 +1,206 @@ +"""Copy device arrays to the host in the background, for output that should not stall the GPU. + +Writing a field or the markers to HDF5 every few steps needs host arrays, and +``array.get()`` waits for the GPU and then for the copy, while no kernel runs. +:class:`HostStaging` overlaps the copy with the next time steps:: + + staging = xp.HostStaging(rho.shape, rho.dtype) # once + for step in range(n_steps): + advance(...) + if step % output_every == 0: + pending.append((step, staging.copy(rho))) # returns at once + while pending and pending[0][1].ready(): + s, copy = pending.pop(0) + h5file[f"rho/{s}"] = copy.result() # a NumPy array + +Each :meth:`HostStaging.copy` first snapshots the array on the device (on the +current stream, after the kernels that wrote it, so the next steps may +overwrite the array), then copies the snapshot to a page-locked host buffer on +its own stream. The buffers are reused in turn: with ``buffers=2``, a third +copy waits until the first one has finished, so the program never runs more +than ``buffers`` copies ahead of the output. :meth:`StagedCopy.result` waits for +its copy and returns the host buffer, valid until that buffer is reused +(``buffers`` copies later); a stale result raises instead of returning another +step's data. Copy it (``result().copy()``) to keep it longer. + +Host arrays (and the NumPy backend) are copied at once, into the same buffers, +so the code is the same on both backends. +""" + +from __future__ import annotations + +from typing import Any + +import array_api_compat +import numpy as np + +from .transfers import _ACTIVE as _COUNTERS +from .transfers import _record + +__all__ = ["HostStaging", "StagedCopy"] + +# substituted in tests that have no GPU +_is_device_array = array_api_compat.is_cupy_array + + +def _cupy() -> Any: + import cupy + + return cupy + + +def _empty_pinned(shape: tuple[int, ...], dtype: Any) -> np.ndarray: + import cupyx + + return cupyx.empty_pinned(shape, dtype=dtype) + + +class _Slot: + def __init__(self, host: np.ndarray) -> None: + self.host = host + self.device: Any = None # the snapshot, allocated on the first device copy + self.event: Any = None # recorded after the copy to the host + self.generation = 0 + + +class StagedCopy: + """A copy started by :meth:`HostStaging.copy`.""" + + def __init__(self, slot: _Slot, generation: int) -> None: + self._slot = slot + self._generation = generation + + def _check(self) -> None: + if self._slot.generation != self._generation: + raise RuntimeError( + "this staged copy's host buffer was reused by a later copy; call " + "result() before starting more copies than there are buffers, or " + "keep result().copy()" + ) + + def ready(self) -> bool: + """Whether the copy has finished (never waits).""" + self._check() + return self._slot.event is None or bool(self._slot.event.done) + + def result(self) -> np.ndarray: + """Wait for the copy and return the host array. + + The array is the staging buffer itself: valid until the buffer is reused. + + Raises + ------ + RuntimeError + If the buffer was already reused by a later copy. + """ + self._check() + if self._slot.event is not None: + self._slot.event.synchronize() + return self._slot.host + + +class HostStaging: + """Page-locked host buffers that device arrays are copied to in the background. + + Parameters + ---------- + shape : tuple[int, ...] + Shape of the arrays to copy. + dtype : dtype-like + Their dtype. + buffers : int + Number of buffers (and device snapshots) used in turn: how many copies + may be in flight at once. 2 (double buffering) lets one copy run while + the previous result is written out. + """ + + def __init__( + self, shape: tuple[int, ...] | int, dtype: Any, buffers: int = 2 + ) -> None: + if buffers < 1: + raise ValueError(f"buffers must be at least 1, got {buffers}") + self.shape = (shape,) if isinstance(shape, int) else tuple(shape) + self.dtype = np.dtype(dtype) + self._slots: list[_Slot] = [] + self._n_buffers = buffers + self._next = 0 + self._stream: Any = None + + def __repr__(self) -> str: + return ( + f"HostStaging(shape={self.shape}, dtype={self.dtype}, " + f"buffers={self._n_buffers})" + ) + + @property + def buffers(self) -> int: + """Number of buffers used in turn.""" + return self._n_buffers + + def _slot(self, device: bool) -> _Slot: + if not self._slots: + allocate = _empty_pinned if device else np.empty + self._slots = [ + _Slot(allocate(self.shape, self.dtype)) for _ in range(self._n_buffers) + ] + slot = self._slots[self._next] + self._next = (self._next + 1) % self._n_buffers + return slot + + def copy(self, array: Any) -> StagedCopy: + """Start copying `array` to the host and return at once. + + Parameters + ---------- + array : cupy.ndarray | numpy.ndarray + An array of the staging shape and dtype. It may be overwritten + right after this call: a device array is snapshotted first. + + Returns + ------- + StagedCopy + ``ready()`` tells whether the copy has finished, ``result()`` + waits for it and returns the host array. + """ + if tuple(array.shape) != self.shape or np.dtype(array.dtype) != self.dtype: + raise ValueError( + f"HostStaging for {self.shape} {self.dtype} got an array of shape " + f"{tuple(array.shape)} and dtype {array.dtype}" + ) + device = _is_device_array(array) + slot = self._slot(device) + if slot.event is not None: + slot.event.synchronize() # the buffer's previous copy is done + slot.generation += 1 + if not device: + np.copyto(slot.host, np.asarray(array)) + slot.event = None + return StagedCopy(slot, slot.generation) + + cp = _cupy() + if self._stream is None: + self._stream = cp.cuda.Stream(non_blocking=True) + if slot.device is None: + slot.device = cp.empty(self.shape, dtype=self.dtype) + if _COUNTERS: + _record("to_host", f"HostStaging.copy({self.shape} {self.dtype})") + # snapshot on the producer's stream, after the kernels that wrote `array` + slot.device[...] = array + snapshot_done = cp.cuda.get_current_stream().record() + # copy the snapshot to the pinned buffer on the staging stream + self._stream.wait_event(snapshot_done) + try: + # blocking=False (CuPy >= 13): return once the copy is enqueued + slot.device.get(stream=self._stream, out=slot.host, blocking=False) + except ( + TypeError + ): # older CuPy: a copy on a stream into pinned memory is asynchronous + slot.device.get(stream=self._stream, out=slot.host) + slot.event = self._stream.record() + return StagedCopy(slot, slot.generation) + + def synchronize(self) -> None: + """Wait for all copies in flight.""" + for slot in self._slots: + if slot.event is not None: + slot.event.synchronize() diff --git a/src/cunumpy/testing.py b/src/cunumpy/testing.py index 4a602c1..cbbd748 100644 --- a/src/cunumpy/testing.py +++ b/src/cunumpy/testing.py @@ -5,9 +5,21 @@ CUDA kernel, and compare what they wrote. This module provides that test (:func:`assert_kernels_agree`), the pytest markers to parametrize tests over the backends (:data:`BACKENDS`, :data:`requires_cupy`, the :func:`backend` -fixture), and :func:`device_function_kernel`, which wraps a ``__device__`` +fixture), :func:`device_function_kernel`, which wraps a ``__device__`` function in an elementwise ``__global__`` kernel so that device helpers can be -tested from Python without a hand-written test kernel. +tested from Python without a hand-written test kernel, and +:func:`emulate_cuda_kernel` (from :mod:`cunumpy.emulation`), which runs a CUDA +kernel on the CPU, one thread after another, so that its arithmetic can be +checked against the host kernel in CI without a GPU. Without a GPU, the CuPy +code paths of a program (argument objects, conversions, backend branches) can +still run on the fake CuPy of :mod:`cunumpy._fake_cupy` (:func:`install_fake_cupy`, +or ``CUNUMPY_FAKE_CUPY=1``); :func:`fake_cupy_active` tells whether it is in +use, and ``requires_cupy`` skips the tests that launch kernels then. + +A catalog's parity tests need no code per kernel when each kernel folder +holds ``_test_args.py`` with ``make_args(backend, seed)`` (and +``N_THREADS``): :func:`parity_cases` and :func:`check_parity` drive +:func:`assert_kernels_agree` from these modules. The module imports pytest only when one of its pytest objects is used, so it can be imported (e.g. for :func:`device_function_kernel`) without pytest, and @@ -44,14 +56,18 @@ def test_parity(name, kernel): import array_api_compat import numpy as np +from . import _fake_cupy from .cuda_kernel import ( CudaKernel, CudaParameter, + CudaStructArguments, + CudaStructValue, _parse_parameter, _split_top_level, _strip_comments, ) from .dispatch import Kernel +from .emulation import emulate_cuda_kernel, emulation_compiler from .xp import cupy_available, get_backend, to_numpy, use_backend # the pytest objects are created on first access, see __getattr__ @@ -59,11 +75,39 @@ def test_parity(name, kernel): "BACKENDS", # noqa: F822 "assert_kernels_agree", "backend", # noqa: F822 + "check_parity", "device_function_kernel", + "emulate_cuda_kernel", + "emulation_compiler", + "fake_cupy_active", + "install_fake_cupy", + "parity_cases", "requires_cupy", # noqa: F822 ] SKIP_REASON = "CuPy/GPU not available" +FAKE_SKIP_REASON = "the fake CuPy cannot run CUDA kernels" + + +def fake_cupy_active() -> bool: + """Whether the fake CuPy (:mod:`cunumpy._fake_cupy`) stands in for CuPy.""" + return _fake_cupy.is_active() + + +def install_fake_cupy() -> Any: + """Install the fake CuPy for this process; see :mod:`cunumpy._fake_cupy`. + + Call it before the first backend use (e.g. at the top of ``conftest.py``), + or set ``CUNUMPY_FAKE_CUPY=1`` in the environment instead. Returns the + fake ``cupy`` module. + """ + return _fake_cupy.install() + + +def _can_launch() -> bool: + """Whether CUDA kernels can run: a functional CuPy that is not the fake.""" + return cupy_available() and not fake_cupy_active() + # pytest objects, built on first use so that importing this module does not # import pytest (see __getattr__ below) @@ -82,7 +126,10 @@ def _pytest() -> Any: def _build_lazy() -> None: pytest = _pytest() - requires_cupy = pytest.mark.skipif(not cupy_available(), reason=SKIP_REASON) + requires_cupy = pytest.mark.skipif( + not _can_launch(), + reason=FAKE_SKIP_REASON if fake_cupy_active() else SKIP_REASON, + ) backends = ["numpy", pytest.param("cupy", marks=requires_cupy)] @pytest.fixture(params=backends) @@ -119,6 +166,18 @@ def _arrays_in(value: Any, name: str, found: dict[str, Any], depth: int) -> None found[name] = value elif depth == 0: return + elif isinstance(value, (CudaStructArguments, CudaStructValue)) and hasattr( + value, "struct" + ): + # by field name, like the attributes of the host argument object; the + # fields of a CudaStructArguments may be properties (not in vars()) + for field in value.struct.fields: + item = ( + value[field.name] + if isinstance(value, CudaStructValue) + else getattr(value, field.name) + ) + _arrays_in(item, f"{name}.{field.name}", found, depth - 1) elif isinstance(value, (tuple, list)): for i, item in enumerate(value): _arrays_in(item, f"{name}[{i}]", found, depth - 1) @@ -139,7 +198,11 @@ def _collect_arrays( level deep, in a tuple, list or dict argument or in the attributes of an argument object (e.g. a ``CudaArguments`` object), are named ``"argument []"`` or ``"argument ."``, and arrays in a - container attribute of an object ``"argument .[]"``. + container attribute of an object ``"argument .[]"``. A + :class:`~cunumpy.CudaStructArguments` object or a struct value is read + through its struct fields, ``"argument ."``, so that its arrays + get the names of the attributes of the host argument object it mirrors, + also when the fields are properties. """ indices = range(len(args)) if outputs is None else outputs found: dict[str, Any] = {} @@ -193,7 +256,7 @@ def assert_kernels_agree( kernel: Kernel, make_args: Callable[[str, int], Sequence[Any]], *, - n_threads: int | Sequence[int] | None = None, + n_threads: int | Sequence[int] | Callable[[tuple[Any, ...]], Any] | None = None, grid: int | Sequence[int] | None = None, block: int | Sequence[int] | None = None, rtol: float = 1e-12, @@ -225,8 +288,11 @@ def assert_kernels_agree( and convert it with :func:`~cunumpy.to_cunumpy`. Kernels take positional arguments only. n_threads, grid, block - Launch configuration of the CUDA kernel (`n_threads` or `grid` is - required), see :meth:`CudaKernel.__call__ `. + Launch configuration of the CUDA kernel, see + :meth:`CudaKernel.__call__ `. `n_threads` + may also be a function of the tuple of arguments, e.g. + ``lambda args: args[0].shape[0]``. One of `n_threads` and `grid` is + required unless the CUDA kernel has ``n_threads_from``. rtol, atol : float Tolerances of ``numpy.testing.assert_allclose``. n_calls : int @@ -258,18 +324,21 @@ def assert_kernels_agree( Notes ----- - The test is skipped with ``pytest.skip`` if CuPy or a GPU is not available. + The test is skipped with ``pytest.skip`` if CuPy or a GPU is not available, + or if the fake CuPy is active. """ if not isinstance(kernel, Kernel): raise TypeError(f"expected a Kernel, got {type(kernel).__name__}") if not kernel.has_cuda: raise ValueError(f"kernel {kernel.name!r} has no CUDA version") - if n_threads is None and grid is None: + if n_threads is None and grid is None and kernel.cuda_kernel.n_threads_from is None: raise TypeError("n_threads (or grid) is required to launch the CUDA kernel") if n_calls < 1: raise ValueError(f"n_calls must be at least 1, got {n_calls}") if outputs is None: outputs = kernel.host_kernel.outputs + if fake_cupy_active(): + _pytest().skip(FAKE_SKIP_REASON) if not cupy_available(): _pytest().skip(SKIP_REASON) @@ -279,8 +348,9 @@ def assert_kernels_agree( if get_backend() != backend: # pragma: no cover - cupy_available() lied raise RuntimeError(f"could not activate the {backend} backend") args = tuple(make_args(backend, seed)) + launch = n_threads(args) if callable(n_threads) else n_threads for _ in range(n_calls): - kernel(*args, n_threads=n_threads, grid=grid, block=block) + kernel(*args, n_threads=launch, grid=grid, block=block) results[backend] = _collect_arrays(args, outputs) host = {name: to_numpy(a) for name, a in results["numpy"].items()} @@ -288,6 +358,94 @@ def assert_kernels_agree( return host +# --------------------------------------------------------------------------- +# parity tests from _test_args.py modules +# --------------------------------------------------------------------------- + +#: Module-level names a ``_test_args.py`` module may define, and the +#: keyword of :func:`assert_kernels_agree` each one sets. +TEST_ARGS_SETTINGS = { + "N_THREADS": "n_threads", + "GRID": "grid", + "BLOCK": "block", + "RTOL": "rtol", + "ATOL": "atol", + "N_CALLS": "n_calls", + "OUTPUTS": "outputs", + "SEED": "seed", +} + + +def parity_cases(catalog: Any) -> list[Any]: + """The kernels of a catalog with a CUDA version, as pytest parameters. + + One ``pytest.param(kernel, id=name)`` per kernel of + ``catalog.parity_cases()``. A kernel without a test-arguments module + (:attr:`Kernel.test_args_module `, from + ``_test_args.py`` in its folder) is marked ``skip`` with a reason + naming the missing file, so the report shows which kernels still lack + their parity test:: + + @pytest.mark.parametrize("kernel", parity_cases(catalog)) + def test_parity(kernel): + check_parity(kernel) + """ + pytest = _pytest() + cases = [] + for name, kernel in catalog.parity_cases(): + marks = () + if kernel.test_args_module is None: + marks = ( + pytest.mark.skip( + reason=f"no test arguments for {name!r}: add {name}_test_args.py " + "with make_args(backend, seed) and N_THREADS to its folder" + ), + ) + cases.append(pytest.param(kernel, id=name, marks=marks)) + return cases + + +def check_parity(kernel: Kernel, **overrides: Any) -> dict[str, np.ndarray]: + """Run :func:`assert_kernels_agree` with the kernel's test-arguments module. + + The module (``_test_args.py`` in the kernel's folder, see + :meth:`KernelCatalog.from_package `) + defines ``make_args(backend, seed)`` and, as module-level names, the + launch and comparison settings of :data:`TEST_ARGS_SETTINGS`: + ``N_THREADS`` (an integer, a tuple, or a function of the argument tuple), + or ``GRID``, plus optionally ``BLOCK``, ``RTOL``, ``ATOL``, ``N_CALLS``, + ``OUTPUTS`` and ``SEED``. Keyword arguments override them. + + Returns + ------- + dict[str, numpy.ndarray] + The host arrays, as :func:`assert_kernels_agree` returns them. + + Raises + ------ + ValueError + If the kernel has no test-arguments module. + TypeError + If the module has no callable ``make_args``. + """ + module = kernel.test_args + if module is None: + raise ValueError( + f"kernel {kernel.name!r} has no test-arguments module: add " + f"{kernel.name}_test_args.py with make_args(backend, seed) to its folder" + ) + make_args = getattr(module, "make_args", None) + if not callable(make_args): + raise TypeError(f"{module.__name__} must define make_args(backend, seed)") + settings = { + keyword: getattr(module, name) + for name, keyword in TEST_ARGS_SETTINGS.items() + if hasattr(module, name) + } + settings.update(overrides) + return assert_kernels_agree(kernel, make_args, **settings) + + # --------------------------------------------------------------------------- # device_function_kernel # --------------------------------------------------------------------------- @@ -301,11 +459,13 @@ def assert_kernels_agree( def _parse_prototype( signature: str, + structs: dict[str, Any] | None = None, ) -> tuple[CudaParameter | None, str, list[tuple[str, CudaParameter]]]: """Parse a C function prototype into (result, name, [(text, parameter)]). The result is None for a ``void`` function; each parameter is its original - text together with its parsed form. + text together with its parsed form. A struct of `structs` may be taken by + value or by (const) reference. """ match = _PROTOTYPE.match(_strip_comments(signature)) if match is None: @@ -333,7 +493,15 @@ def _parse_prototype( params = [] if params_text not in ("", "void"): for text in _split_top_level(params_text): - params.append((text.strip(), _parse_parameter(text))) + text = text.strip() + by_reference = "&" in text + param = _parse_parameter(text.replace("&", " "), structs or None) + if by_reference and param.struct is None: + raise ValueError( + f"unsupported parameter {text!r} of {name!r}: only structs can " + "be passed by reference" + ) + params.append((text, param)) return result, name, params @@ -362,7 +530,9 @@ def device_function_kernel( The C prototype of the device function, e.g. ``"int find_span(const double* t, int p, double eta)"``. Pointer parameters, scalar parameters of the types :class:`CudaKernel` supports, - and a scalar or ``void`` return type are supported. + struct parameters (by value or by ``const`` reference, for the structs + passed in ``structs``) and a scalar or ``void`` return type are + supported. name : str | None Name of the generated kernel; ``"_kernel"`` by default. includes : Sequence[str] @@ -374,8 +544,10 @@ def device_function_kernel( out_param : str Name of the generated output array parameter. **kwargs - Passed on to :class:`CudaKernel`, e.g. ``include_dirs``, ``options`` or - ``block_size``. + Passed on to :class:`CudaKernel`, e.g. ``include_dirs``, ``options``, + ``block_size`` or ``structs`` (the :class:`~cunumpy.CudaStruct` types + of struct parameters, whose definitions `header_source` or the + `includes` must provide). Returns ------- @@ -386,6 +558,9 @@ def device_function_kernel( * a pointer parameter stays as it is and is passed through unchanged to every call (an array shared by all threads); + * a struct parameter (``DomainArgs d`` or ``const DomainArgs& d``) is + taken by value and passed through unchanged to every call (pass a + :class:`~cunumpy.CudaStructArguments` object or a packed value); * a scalar parameter ``T x`` becomes a device array ``const T* x`` of length ``n``, and thread ``i`` calls the function with ``x[i]``; * the return value of thread ``i`` is stored in ``out[i]``, an array @@ -412,7 +587,8 @@ def device_function_kernel( >>> x = cp.arange(10.0); out = cp.empty(10) # doctest: +SKIP >>> sq(x, out, 10, n_threads=10) # doctest: +SKIP """ - result, function, params = _parse_prototype(signature) + structs = {struct.name: struct for struct in kwargs.get("structs", ())} + result, function, params = _parse_prototype(signature, structs) if name is None: name = f"{function}_kernel" reserved = {out_param, n_threads_param} @@ -429,6 +605,9 @@ def device_function_kernel( if param.pointer: wrapper_params.append(text) call_args.append(param.name) + elif param.struct is not None: + wrapper_params.append(f"{param.ctype} {param.name}") # by value + call_args.append(param.name) else: wrapper_params.append(f"const {param.ctype}* {param.name}") call_args.append(f"{param.name}[i]") diff --git a/src/cunumpy/xp.py b/src/cunumpy/xp.py index f26687d..a039a35 100644 --- a/src/cunumpy/xp.py +++ b/src/cunumpy/xp.py @@ -17,6 +17,12 @@ from .transfers import _ACTIVE as _COUNTERS from .transfers import _describe, _record +if os.environ.get("CUNUMPY_FAKE_CUPY", "").strip().lower() in ("1", "true", "yes"): + # tests without a GPU: a strict host stand-in for CuPy, see cunumpy._fake_cupy + from ._fake_cupy import install as _install_fake_cupy + + _install_fake_cupy() + BackendType = Literal["numpy", "cupy"] _logger = logging.getLogger(__name__) @@ -280,6 +286,116 @@ def synchronize_for_mpi(*arrays: Any) -> None: cp.cuda.get_current_stream().synchronize() +# the result of the last mpi_is_cuda_aware() probe (or of set_mpi_cuda_aware()), +# used by mpi_buffer(); None until one of them was called +_MPI_CUDA_AWARE: bool | None = None + + +def set_mpi_cuda_aware(value: bool | None) -> None: + """Tell `mpi_buffer()` whether MPI can take device buffers. + + `mpi_is_cuda_aware()` records its result itself; call this instead when + the answer is known otherwise (e.g. from the cluster documentation, or to + force host staging for a test). ``None`` forgets the setting. + """ + global _MPI_CUDA_AWARE + _MPI_CUDA_AWARE = None if value is None else bool(value) + + +def get_mpi_cuda_aware() -> bool | None: + """The recorded answer of `mpi_is_cuda_aware()`/`set_mpi_cuda_aware()`, or None.""" + return _MPI_CUDA_AWARE + + +def _pinned_or_host_empty(shape: tuple[int, ...], dtype: Any) -> np.ndarray: + """A host buffer for staging, pinned when CuPy can allocate pinned memory.""" + try: + import cupy as cp + + nbytes = int(np.prod(shape, dtype=np.int64)) * np.dtype(dtype).itemsize + mem = cp.cuda.alloc_pinned_memory(max(nbytes, 1)) + return np.frombuffer(mem, dtype, int(np.prod(shape, dtype=np.int64))).reshape( + shape + ) + except Exception: # noqa: BLE001 - no CuPy, no pinned memory, the fake CuPy, ... + return np.empty(shape, dtype=dtype) + + +@contextmanager +def mpi_buffer( + array: Any, + *, + send: bool = True, + recv: bool = False, + cuda_aware: bool | None = None, +) -> Generator[Any, None, None]: + """The buffer to hand to MPI for `array`: the array itself, or a host copy. + + One MPI call site for both backends and both kinds of MPI builds:: + + with xp.mpi_buffer(markers_out) as sendbuf, xp.mpi_buffer( + markers_in, send=False, recv=True + ) as recvbuf: + comm.Sendrecv(sendbuf, dest, recvbuf=recvbuf, source=source) + + * A host array (NumPy backend, or a NumPy array on the CuPy backend) is + yielded unchanged. + * A device array with CUDA-aware MPI is yielded unchanged after + `synchronize_for_mpi()`, so MPI reads what the kernels wrote. + * A device array without CUDA-aware MPI is staged through a host buffer + (pinned memory when available): with `send`, the array is copied to the + host first (counted as a ``to_host`` transfer by `count_transfers()`); + with `recv`, the host buffer is copied back into the array when the + block ends (a ``to_device`` transfer). The device array itself is never + given to MPI. + + Parameters + ---------- + array + The buffer of the MPI call: a NumPy or CuPy array. + send : bool + Whether MPI reads the buffer (copy device to host before the block). + recv : bool + Whether MPI writes the buffer (copy host to device after the block). + cuda_aware : bool | None + Whether MPI can take device buffers. None uses the answer recorded by + `mpi_is_cuda_aware()` or `set_mpi_cuda_aware()`. + + Raises + ------ + RuntimeError + For a device array when `cuda_aware` is None and nothing was recorded: + call `mpi_is_cuda_aware(comm)` (collective) once at startup, or + `set_mpi_cuda_aware()`. + """ + if not array_api_compat.is_cupy_array(array): + yield array + return + if cuda_aware is None: + cuda_aware = _MPI_CUDA_AWARE + if cuda_aware is None: + raise RuntimeError( + "mpi_buffer(): it is not known whether MPI can take device buffers; " + "call xp.mpi_is_cuda_aware(comm) once at startup (every rank), or " + "xp.set_mpi_cuda_aware(True/False), or pass cuda_aware=" + ) + if cuda_aware: + synchronize_for_mpi(array) + yield array + return + host = _pinned_or_host_empty(tuple(array.shape), array.dtype) + if send: + if _COUNTERS: + _record("to_host", f"mpi_buffer({_describe(array)}) staging for send") + synchronize_for_mpi(array) + host[...] = array.get() + yield host + if recv: + if _COUNTERS: + _record("to_device", f"mpi_buffer({_describe(array)}) staging for recv") + array.set(host) + + def _mpi_module() -> Any: """Import and return ``mpi4py.MPI``, with a clear error if it is missing.""" try: @@ -376,7 +492,9 @@ def mpi_is_cuda_aware(comm: Any = None, *, method: str = "probe") -> bool: _logger.debug("CUDA-aware MPI probe failed on rank %d: %r", rank, e) ok = False - return bool(comm.allreduce(ok, op=MPI.LAND)) + result = bool(comm.allreduce(ok, op=MPI.LAND)) + set_mpi_cuda_aware(result) + return result def require_cuda_aware_mpi(comm: Any = None) -> None: @@ -421,6 +539,45 @@ def memory_info() -> tuple[int, int] | None: return cp.cuda.runtime.memGetInfo() +#: Dynamic shared memory per block that every CUDA device provides without an +#: opt-in (48 KiB); also the answer of :func:`max_shared_memory_per_block` +#: without a GPU. +DEFAULT_SHARED_MEMORY_PER_BLOCK = 48 * 1024 + + +def max_shared_memory_per_block( + device: int | None = None, *, opt_in: bool = False +) -> int: + """Bytes of shared memory a block of a CUDA kernel may use on `device`. + + Use it to decide whether a per-block buffer (e.g. a copy of a small grid + for a deposit) fits, instead of a hard-coded limit. + + Parameters + ---------- + device : int | None + CUDA device id; the current device by default. + opt_in : bool + The larger limit a kernel can opt in to on newer GPUs (e.g. 99 or 227 + KiB). Using more than the default 48 KiB needs the kernel attribute + ``max_dynamic_shared_size_bytes`` set on the compiled ``cupy.RawKernel`` + (``kernel.compile()``). + + Returns + ------- + int + The limit in bytes; :data:`DEFAULT_SHARED_MEMORY_PER_BLOCK` if CuPy is + not available (so that code choosing a GPU strategy runs everywhere). + """ + if not cupy_available(): + return DEFAULT_SHARED_MEMORY_PER_BLOCK + import cupy as cp + + dev = cp.cuda.Device() if device is None else cp.cuda.Device(device) + key = "MaxSharedMemoryPerBlockOptin" if opt_in else "MaxSharedMemoryPerBlock" + return int(dev.attributes.get(key, DEFAULT_SHARED_MEMORY_PER_BLOCK)) + + def free_memory() -> None: """Release all free blocks held by CuPy's memory pools (no-op on NumPy). @@ -826,6 +983,61 @@ def as_device_array( return result +def segment_sum(values: Any, keys: Any, n_segments: int) -> Any: + """Sum `values` per key: ``out[k] = sum(values[i] for keys[i] == k)``. + + The reduction step of a sort-then-reduce accumulation (particles binned to + cells, contributions summed per cell), on either backend, with + ``bincount`` under the hood. For a 2D `values` the columns are summed + separately (one bincount per column). + + Parameters + ---------- + values : array + Shape ``(n,)`` or ``(n, m)``, on the backend of `keys`. + keys : array + Integer segment of every value, shape ``(n,)``; a negative key drops + the value (e.g. a particle outside the grid). + n_segments : int + Number of segments; keys must be smaller than it. + + Returns + ------- + array + Shape ``(n_segments,)`` or ``(n_segments, m)``, dtype of `values` for + floating-point and complex values, ``float64`` otherwise. + """ + xpm = get_array_module(keys) + keys = xpm.asarray(keys) + values = xpm.asarray(values) + if keys.ndim != 1 or values.shape[:1] != keys.shape: + raise ValueError( + f"keys must be 1D with one entry per value, got keys {keys.shape} and " + f"values {values.shape}" + ) + if values.ndim not in (1, 2): + raise ValueError(f"values must be 1D or 2D, got shape {values.shape}") + if bool((keys >= n_segments).any()): + raise ValueError(f"keys must be smaller than n_segments={n_segments}") + valid = keys >= 0 + if not bool(valid.all()): + keys = keys[valid] + values = values[valid] + out_dtype = values.dtype if values.dtype.kind in "fc" else np.dtype(np.float64) + if values.ndim == 1: + if values.dtype.kind == "c": + real = xpm.bincount(keys, weights=values.real, minlength=n_segments) + imag = xpm.bincount(keys, weights=values.imag, minlength=n_segments) + return (real + 1j * imag).astype(out_dtype, copy=False) + return xpm.bincount(keys, weights=values, minlength=n_segments).astype( + out_dtype, copy=False + ) + out = xpm.empty((n_segments, values.shape[1]), dtype=out_dtype) + for j in range(values.shape[1]): + out[:, j] = segment_sum(values[:, j], keys, n_segments) + return out + + def to_cunumpy(array: Any) -> Any: """Convert an array to the currently active backend. diff --git a/tests/unit/test_cuda_kernel.py b/tests/unit/test_cuda_kernel.py index 592e13b..3042718 100644 --- a/tests/unit/test_cuda_kernel.py +++ b/tests/unit/test_cuda_kernel.py @@ -198,7 +198,7 @@ def test_wrong_scalars_raise(): with pytest.raises(OverflowError, match="out of range"): kernel.prepare_args(1.0, x, y, 2**31) # overflows int with pytest.raises(TypeError, match="losing information"): - kernel.prepare_args(1.0, x, y, np.int64(5)) # int64 into int + kernel.prepare_args(np.complex128(1.0), x, y, 5) # complex into double with pytest.raises(TypeError): kernel.prepare_args("1.0", x, y, 5) @@ -545,13 +545,13 @@ def test_struct_values(): def test_struct_pointer_fields_must_be_contiguous(): - values = dict( - n=3, - charge=2.0, - alive=FakeDeviceArray(np.bool_), - ids=FakeDeviceArray(np.int64), - weight=0.5, - ) + values = { + "n": 3, + "charge": 2.0, + "alive": FakeDeviceArray(np.bool_), + "ids": FakeDeviceArray(np.int64), + "weight": 0.5, + } view = FakeDeviceArray(np.float64, flags=SimpleNamespace(c_contiguous=False)) with pytest.raises( TypeError, match=r"argument 0 \(double\* x\) must be C-contiguous" @@ -690,7 +690,7 @@ def test_parse_view_parameters(): with pytest.raises(ValueError, match="array views"): parse_cuda_signature("__global__ void f(Array2D a) {}", "f") with pytest.raises(ValueError, match="unsupported type"): - parse_cuda_signature("__global__ void f(Array4D a) {}", "f") + parse_cuda_signature("__global__ void f(Array5D a) {}", "f") def test_view_parameters_pack_pointer_shape_and_strides(): @@ -945,10 +945,10 @@ def missing(x, n: int): def unknown(x: "str[:]"): pass - def too_many(x: "float[:, :, :, :]"): + def too_many(x: "float[:, :, :, :, :]"): pass - def unparsable(x: "float[:](order=F)"): + def unparsable(x: "float[:](order=F)"): # noqa: F821 (deliberately unparsable) pass with pytest.raises( @@ -957,7 +957,7 @@ def unparsable(x: "float[:](order=F)"): CudaStruct.from_signature(missing, "A") with pytest.raises(ValueError, match="unsupported scalar type 'str'"): CudaStruct.from_signature(unknown, "A") - with pytest.raises(ValueError, match="at most 3 dimensions"): + with pytest.raises(ValueError, match="at most 4 dimensions"): CudaStruct.from_signature(too_many, "A") with pytest.raises(ValueError, match="cannot parse the annotation"): CudaStruct.from_signature(unparsable, "A") @@ -1221,7 +1221,11 @@ def _run_python(code, env=None): [str(Path(xp.__file__).parents[1]), environment.get("PYTHONPATH", "")] ) result = subprocess.run( - [sys.executable, "-c", code], capture_output=True, text=True, env=environment + [sys.executable, "-c", code], + capture_output=True, + text=True, + env=environment, + check=False, ) return result.stdout, result.stderr @@ -1289,9 +1293,8 @@ def test_cuda_debug_context_restores(debug_off): assert xp.get_cuda_debug() is True assert xp.get_cuda_debug() is False - with pytest.raises(ValueError): - with xp.cuda_debug(): - raise ValueError + with pytest.raises(ValueError), xp.cuda_debug(): + raise ValueError assert xp.get_cuda_debug() is False # restored after an exception too @@ -1517,6 +1520,59 @@ def test_compile_options_contain_the_header_hash(header_tree): assert _user_options(CudaKernel(INCLUDING_SOURCE, "double_it")) == () +def test_resolve_includes_angle_dirs(tmp_path): + shipped = tmp_path / "shipped" + (shipped / "lib").mkdir(parents=True) + (shipped / "lib" / "a.cuh").write_text('#include "lib/b.cuh"\n') + (shipped / "lib" / "b.cuh").write_text("") + source = "#include \n#include \n" + # angle brackets: system headers, not tracked by default + assert resolve_includes(source, [shipped]) == [] + # in angle_dirs they are, and their quoted includes resolve there too + expected = [shipped / "lib" / "a.cuh", shipped / "lib" / "b.cuh"] + assert resolve_includes(source, angle_dirs=[shipped]) == expected + assert resolve_includes('#include "lib/a.cuh"\n', angle_dirs=[shipped]) == expected + # include_dirs come first for quoted includes + user = tmp_path / "user" + (user / "lib").mkdir(parents=True) + (user / "lib" / "a.cuh").write_text("") + assert resolve_includes('#include "lib/a.cuh"\n', [user], angle_dirs=[shipped]) == [ + user / "lib" / "a.cuh" + ] + + +def test_shipped_headers_are_part_of_the_hash(): + include = Path(cuda_include_dir()) / "cunumpy" + for line in ("#include ", '#include "cunumpy/reduce.cuh"'): + kernel = CudaKernel(line + "\n" + AXPY, "axpy") + assert kernel.included_headers == ( + include / "reduce.cuh", + include / "atomic.cuh", + ) + digest = include_hash(kernel.included_headers) + assert _user_options(kernel) == (f"-DCUNUMPY_INCLUDE_HASH=0x{digest}",) + # a source without includes still gets no define + assert _user_options(CudaKernel(AXPY, "axpy")) == () + + +def test_changed_shipped_header_changes_the_hash(tmp_path, monkeypatch): + # a copy of the shipped headers stands for the installed ones before and + # after an upgrade of cunumpy + import shutil + + from cunumpy import cuda_kernel + + installed = tmp_path / "include" + shutil.copytree(cuda_include_dir(), installed) + monkeypatch.setattr(cuda_kernel, "_CUDA_INCLUDE_DIR", installed) + kernel = CudaKernel("#include \n" + AXPY, "axpy") + before = kernel.compile_options()[-1] + assert before.startswith("-DCUNUMPY_INCLUDE_HASH=0x") + header = installed / "cunumpy" / "atomic.cuh" + header.write_text(header.read_text() + "\n// changed in an upgrade\n") + assert kernel.compile_options()[-1] != before + + def test_editing_a_header_recompiles_on_gpu(header_tree): _skip_without_cupy() import cupy as cp @@ -1710,12 +1766,71 @@ class Derived(ParticleArguments): # inherits the struct of its parent assert Derived.struct is ParticleArguments.struct -def test_struct_arguments_repack_after_replacing_an_array(): +def test_struct_arguments_repack_when_a_field_changes(): args = _particle_arguments(ptr=0x100) + packed = args.packed + assert args.packed is packed # nothing changed: no repacking + args.x = FakeDeviceArray(np.float64, ptr=0x200, shape=(3,)) - assert args.packed["x"] == 0x100 # still the old address - args.pack() - assert args.packed["x"] == 0x200 + assert args.packed["x"] == 0x200 # a new array: repacked at the next use + assert args.__cuda_args__() == (args.packed,) + + args.charge = 5.0 + assert args.packed["charge"] == 5.0 # a scalar changed: repacked too + args.charge = 5 # equal value of another type: repacked, same result + assert args.packed["charge"] == 5.0 + + args.x = np.zeros(3) # an invalid value raises at the next use + with pytest.raises(TypeError, match="must be a CuPy array"): + args.packed # noqa: B018 + + +class Owner: + """An object owning a marker array that it replaces when it grows.""" + + def __init__(self, n, ptr=0x100): + self.markers = FakeDeviceArray(np.float64, ptr=ptr, shape=(n, 4)) + + def grow(self, n, ptr): + self.markers = FakeDeviceArray(np.float64, ptr=ptr, shape=(n, 4)) + + +class OwnerArguments(xp.CudaStructArguments): + struct_name = "OwnerArgs" + fields = (("markers", "Array2D"), ("n_markers", "int")) + + def __init__(self, owner): + self._owner = owner + self.pack() + + @property + def markers(self): + return self._owner.markers + + @property + def n_markers(self): + return self._owner.markers.shape[0] + + +def test_struct_arguments_follow_the_arrays_of_an_owner(): + owner = Owner(3) + args = OwnerArguments(owner) + assert args.packed["markers"]["data"] == 0x100 + assert args.packed["n_markers"] == 3 + + owner.grow(8, ptr=0x900) + assert args.packed["markers"]["data"] == 0x900 + assert args.packed["markers"]["shape"].tolist() == [8, 4] + assert args.packed["n_markers"] == 8 + + +def test_struct_arguments_repack_when_a_view_changes_shape(): + # a view of the same allocation with fewer rows: same address, new shape + owner = Owner(8, ptr=0x100) + args = OwnerArguments(owner) + owner.markers = FakeDeviceArray(np.float64, ptr=0x100, shape=(5, 4)) + assert args.packed["markers"]["shape"].tolist() == [5, 4] + assert args.packed["n_markers"] == 5 def test_struct_arguments_are_packed_again_when_copied(): @@ -1772,3 +1887,295 @@ def test_struct_arguments_on_gpu(): assert int(size.get()[0]) == ParticleArguments.struct.dtype.itemsize assert out.get().tolist() == [n, 2.0, 42.0, 1.5] assert cp.all(args.x[1::2] == 1.0) and cp.all(args.x[::2] == 0.0) + + +def test_debug_synchronization_is_skipped_while_capturing(): + kernel = CudaKernel(AXPY, "axpy") + calls = [] + + class Stream: + def __init__(self, capturing): + self.capturing = capturing + + def is_capturing(self): + return self.capturing + + def synchronize(self): + calls.append(self.capturing) + + class LegacyStream(Stream): + def is_capturing(self): + raise RuntimeError("not supported on the legacy stream") + + kernel._synchronize_after_launch(Stream(True), (1,), (128,)) + assert calls == [] + kernel._synchronize_after_launch(Stream(False), (1,), (128,)) + kernel._synchronize_after_launch(LegacyStream(False), (1,), (128,)) + assert calls == [False, False] + + +def test_debug_kernel_in_a_cuda_graph(): + _skip_without_cupy() + import cupy as cp + + n = 1000 + x, y = cp.ones(n), cp.zeros(n) + kernel = CudaKernel(AXPY, "axpy", debug=True) + kernel(2.0, x, y, n, n_threads=n) # compile outside the capture + stream = cp.cuda.Stream(non_blocking=True) + with stream: + stream.begin_capture() + kernel(2.0, x, y, n, n_threads=n, stream=stream) + graph = stream.end_capture() + graph.launch(stream) + graph.launch(stream) + stream.synchronize() + assert cp.all(y == 6.0) + + +# --------------------------------------------------------------------------- +# struct layout checked against the compiler +# --------------------------------------------------------------------------- + + +def test_layout_source_reports_size_alignment_and_offsets(): + struct = CudaStruct( + "LayoutArgs", [("markers", "Array2D"), ("valid", "bool*"), ("n", "int")] + ) + source = struct.layout_source() + assert source.startswith('#include "cunumpy/array_view.cuh"') + assert struct.declaration in source + (param,) = parse_cuda_signature(source, "cunumpy_layout_LayoutArgs") + assert param.pointer and param.dtype == np.dtype(np.uint64) + assert "out[0] = sizeof(LayoutArgs);" in source + assert "out[1] = alignof(LayoutArgs);" in source + for i, name in enumerate(["markers", "valid", "n"]): + assert f"out[{i + 2}] = (unsigned long long)((const char*)&s.{name}" in source + + # a header instead of the declaration: the struct is not defined in the source + from_header = struct.layout_source("pkg/layout_args.cuh") + assert from_header.startswith('#include "pkg/layout_args.cuh"') + assert "struct LayoutArgs {" not in from_header + assert struct.layout_source("#include ").startswith( + "#include " + ) + # no views: no array_view include + assert "array_view" not in PARTICLES.layout_source() + + +def test_verify_layout_needs_cupy(monkeypatch): + monkeypatch.setattr(xp.xp, "cupy_available", lambda: False) + with pytest.raises(RuntimeError, match="needs CuPy"): + PARTICLES.verify_layout() + + +def test_verify_layout_on_gpu(tmp_path): + _skip_without_cupy() + layout = PARTICLES.verify_layout() + assert layout["sizeof"] == PARTICLES.dtype.itemsize + assert layout["charge"] == PARTICLES.dtype.fields["charge"][1] + + views = CudaStruct("Views", [("n", "int"), ("a", "Array2D"), ("b", "bool")]) + views.verify_layout() + + # a header that drifted from the Python definition + header = tmp_path / "drifted.cuh" + header.write_text( + '#include "cunumpy/array_view.cuh"\n' + "struct Views { int n; bool b; Array2D a; };\n" # b moved before a + ) + with pytest.raises(ValueError, match="differs from its CudaStruct dtype"): + views.verify_layout("drifted.cuh", include_dirs=[tmp_path]) + + +# --------------------------------------------------------------------------- +# cunumpy/reduce.cuh +# --------------------------------------------------------------------------- + +REDUCE_SOURCE = r""" +#include "cunumpy/reduce.cuh" +extern "C" __global__ +void reductions(const double* x, long long n, double* sum, unsigned long long* count, + double* block_min, double* block_max, double* warp_sum) { + long long i = blockIdx.x * (long long)blockDim.x + threadIdx.x; + double v = i < n ? x[i] : 0.0; // no early return: every thread takes part + cunumpy_block_sum_to(sum, v); + cunumpy_block_sum_to(count, i < n ? 1ull : 0ull); + double lo = cunumpy_block_min(i < n ? v : 1e300); + double hi = cunumpy_block_max(i < n ? v : -1e300); + if (threadIdx.x == 0) { block_min[blockIdx.x] = lo; block_max[blockIdx.x] = hi; } + double w = cunumpy_warp_sum(v); + if (threadIdx.x % 32 == 0) warp_sum[i / 32] = w; +} +""" + + +def test_reduce_header_is_shipped(): + header = Path(cuda_include_dir()) / "cunumpy" / "reduce.cuh" + text = header.read_text() + assert "#ifndef CUNUMPY_REDUCE_CUH" in text and "#endif" in text + for name in ( + "cunumpy_warp_sum", + "cunumpy_warp_min", + "cunumpy_warp_max", + "cunumpy_block_sum", + "cunumpy_block_min", + "cunumpy_block_max", + "cunumpy_block_sum_to", + ): + assert f"{name}(" in text, name + # the kernel's include resolves to the shipped header, which includes atomic.cuh + headers = resolve_includes(REDUCE_SOURCE, [cuda_include_dir()]) + assert [p.name for p in headers] == ["reduce.cuh", "atomic.cuh"] + CudaKernel(REDUCE_SOURCE, "reductions") # the signature parses + + +@pytest.mark.parametrize("block_size", [32, 128, 1024]) +def test_reductions_on_gpu(block_size): + _skip_without_cupy() + import cupy as cp + + n = 5000 # not a multiple of the block size: the last block is partial + x = cp.asarray(np.random.default_rng(1).normal(size=n)) + n_blocks = -(-n // block_size) + total, count = cp.zeros(1), cp.zeros(1, dtype=cp.uint64) + lo, hi = cp.zeros(n_blocks), cp.zeros(n_blocks) + warp = cp.zeros(n_blocks * block_size // 32) + kernel = CudaKernel(REDUCE_SOURCE, "reductions", block_size=block_size) + kernel(x, n, total, count, lo, hi, warp, n_threads=n) + + host = cp.asnumpy(x) + assert abs(float(total[0]) - host.sum()) < 1e-10 * n + assert int(count[0]) == n + padded = np.concatenate([host, np.full(n_blocks * block_size - n, np.nan)]) + blocks = padded.reshape(n_blocks, block_size) + np.testing.assert_array_equal(cp.asnumpy(lo), np.nanmin(blocks, axis=1)) + np.testing.assert_array_equal(cp.asnumpy(hi), np.nanmax(blocks, axis=1)) + np.testing.assert_allclose( + cp.asnumpy(warp), np.nan_to_num(padded).reshape(-1, 32).sum(axis=1), rtol=1e-12 + ) + + +# --------------------------------------------------------------------------- +# 4D views and NumPy integer scalars +# --------------------------------------------------------------------------- + +VIEW_4D = r""" +#include "cunumpy/array_view.cuh" +extern "C" __global__ +void scale_4d(Array4D a, double factor, int n) {} +""" + + +def test_array4d_parameters_pack_pointer_shape_and_strides(): + param, _, _ = parse_cuda_signature(VIEW_4D, "scale_4d") + assert param.view_ndim == 4 and param.ctype == "Array4D" + kernel = CudaKernel(VIEW_4D, "scale_4d") + # every second component of a (2, 3, 4, 6) grid: a non-contiguous view + a = FakeDeviceArray( + np.float64, ptr=0x40, shape=(2, 3, 4, 3), strides=(576, 192, 48, 16) + ) + packed, _, _ = kernel.prepare_args(a, 2.0, 5) + assert packed["data"] == 0x40 + assert packed["shape"].tolist() == [2, 3, 4, 3] + assert packed["strides"].tolist() == [72, 24, 6, 2] + assert packed.dtype.itemsize == 72 # sizeof(Array4D) + with pytest.raises(TypeError, match="must be a 4D array"): + kernel.prepare_args(FakeDeviceArray(np.float64, shape=(2, 3, 4)), 2.0, 5) + + +def test_array4d_struct_fields_and_annotations(): + struct = CudaStruct("Grid", [("e", "Array4D"), ("n", "int")]) + assert struct.dtype.fields["e"][0].itemsize == 72 + assert " Array4D e;" in struct.declaration + + def init(self, e: "float[:, :, :, :]", n: int): ... + + from_annotations = CudaStruct.from_signature(init, "Grid") + assert from_annotations.fields[0].ctype == "Array4D" + + def too_many(self, e: "float[:, :, :, :, :]"): ... + + with pytest.raises(ValueError, match="at most 4 dimensions"): + CudaStruct.from_signature(too_many, "Grid") + + +def test_array4d_header_layout(): + header = (Path(cuda_include_dir()) / "cunumpy" / "array_view.cuh").read_text() + assert "struct Array4D" in header + assert "sizeof(Array4D) == 72" in header + + +def test_numpy_integer_scalars_are_checked_by_value(): + kernel = CudaKernel(AXPY, "axpy") # (double a, double* x, double* y, int n) + x, y = FakeDeviceArray(np.float64), FakeDeviceArray(np.float64) + for n in (np.int64(5), np.int16(5), np.uint64(5), 5): + *_, packed_n = kernel.prepare_args(1.0, x, y, n) + assert type(packed_n) is np.int32 and packed_n == 5 + with pytest.raises(OverflowError, match="out of range"): + kernel.prepare_args(1.0, x, y, np.int64(2**31)) + with pytest.raises(TypeError): + kernel.prepare_args(1.0, x, y, np.float64(5.0)) # a float is not an int + # integers into a double parameter keep working + a, *_ = kernel.prepare_args(np.int64(3), x, y, 1) + assert type(a) is np.float64 and a == 3.0 + + +# --------------------------------------------------------------------------- +# launch conveniences: n_threads_from="first_array", shared memory opt-in +# --------------------------------------------------------------------------- + + +class RecordingRawKernel: + """Stands for a compiled cupy.RawKernel: records launches and attributes.""" + + def __init__(self): + self.launches = [] + self.max_dynamic_shared_size_bytes = 48 * 1024 + + def __call__(self, grid, block, args, shared_mem=0): + self.launches.append((grid, block, shared_mem)) + + +@pytest.fixture +def recorded(monkeypatch): + """A CudaKernel whose launches are recorded instead of run.""" + kernel = CudaKernel(AXPY, "axpy", block_size=128) + raw = RecordingRawKernel() + monkeypatch.setattr(kernel, "compile", lambda: raw) + monkeypatch.setattr(kernel, "debug_active", lambda: False) + return kernel, raw + + +def test_n_threads_from_first_array(recorded): + kernel, raw = recorded + kernel.n_threads_from = "first_array" + x = FakeDeviceArray(np.float64, shape=(1000,)) + y = FakeDeviceArray(np.float64, shape=(1000,)) + kernel(2.0, x, y, 1000) # the scalar first argument is skipped + assert raw.launches == [((8,), (128,), 0)] + kernel(2.0, x, y, 1000, n_threads=10) # explicit sizes still win + assert raw.launches[-1] == ((1,), (128,), 0) + with pytest.raises(TypeError, match="n_threads_from must be"): + kernel.n_threads_from = "rows" + as_option = CudaKernel(AXPY, "axpy", n_threads_from="first_array") + assert as_option.n_threads_from((1.0, x)) == 1000 + with pytest.raises(TypeError, match="needs an array argument"): + as_option.n_threads_from((1.0, 2)) + + +def test_shared_memory_above_the_default_is_opted_in(recorded, monkeypatch): + kernel, raw = recorded + monkeypatch.setattr( + xp.xp, "max_shared_memory_per_block", lambda opt_in=False: 100_000 + ) + x, y = FakeDeviceArray(np.float64), FakeDeviceArray(np.float64) + kernel(1.0, x, y, 1, n_threads=1, shared_mem=40_000) # below 48 KiB: no setup + assert raw.max_dynamic_shared_size_bytes == 48 * 1024 + kernel(1.0, x, y, 1, n_threads=1, shared_mem=80_000) + assert raw.max_dynamic_shared_size_bytes == 80_000 + kernel(1.0, x, y, 1, n_threads=1, shared_mem=60_000) # already allowed + assert raw.max_dynamic_shared_size_bytes == 80_000 + with pytest.raises(ValueError, match="exceeds the 100000 bytes"): + kernel(1.0, x, y, 1, n_threads=1, shared_mem=100_001) + assert [s for *_, s in raw.launches] == [40_000, 80_000, 60_000] diff --git a/tests/unit/test_cunumpy.py b/tests/unit/test_cunumpy.py index cceaa08..481224b 100644 --- a/tests/unit/test_cunumpy.py +++ b/tests/unit/test_cunumpy.py @@ -394,3 +394,38 @@ class FakeCupy: xp.synchronize() finally: monkeypatch.setattr(cxp.array_backend, "_backend", "numpy") + + +def test_max_shared_memory_per_block_without_a_gpu(monkeypatch): + monkeypatch.setattr(xp.xp, "cupy_available", lambda: False) + assert xp.max_shared_memory_per_block() == xp.DEFAULT_SHARED_MEMORY_PER_BLOCK + assert xp.max_shared_memory_per_block(opt_in=True) == 48 * 1024 + + +def test_max_shared_memory_per_block_reads_the_device(monkeypatch): + import sys + import types + + class Device: + def __init__(self, device_id=0): + self.attributes = { + "MaxSharedMemoryPerBlock": 49152 + device_id, + "MaxSharedMemoryPerBlockOptin": 232448, + } + + cupy = types.ModuleType("cupy") + cupy.cuda = types.SimpleNamespace(Device=Device) + monkeypatch.setitem(sys.modules, "cupy", cupy) + monkeypatch.setattr(xp.xp, "cupy_available", lambda: True) + assert xp.max_shared_memory_per_block() == 49152 + assert xp.max_shared_memory_per_block(1) == 49153 + assert xp.max_shared_memory_per_block(opt_in=True) == 232448 + + +def test_max_shared_memory_per_block_on_gpu(): + if not xp.cupy_available(): + pytest.skip("CuPy not installed or not functional") + assert xp.max_shared_memory_per_block() >= 48 * 1024 + assert ( + xp.max_shared_memory_per_block(opt_in=True) >= xp.max_shared_memory_per_block() + ) diff --git a/tests/unit/test_emulation.py b/tests/unit/test_emulation.py new file mode 100644 index 0000000..bc15a23 --- /dev/null +++ b/tests/unit/test_emulation.py @@ -0,0 +1,321 @@ +"""Tests for `cunumpy.testing.emulate_cuda_kernel`: CUDA kernels run on the CPU.""" + +import numpy as np +import pytest + +from cunumpy import CudaKernel +from cunumpy.testing import emulate_cuda_kernel, emulation_compiler + +pytestmark = pytest.mark.skipif( + emulation_compiler() is None, reason="no C++ compiler for the emulation" +) + +AXPY = r""" +extern "C" __global__ +void axpy(double a, const double* __restrict__ x, double* y, int n) { + int i = blockDim.x * blockIdx.x + threadIdx.x; + if (i < n) y[i] += a * x[i]; +} +""" + + +def test_axpy(): + rng = np.random.default_rng(0) + x, y = rng.random(1000), rng.random(1000) + expected = y + 2.5 * x + emulate_cuda_kernel(CudaKernel(AXPY, "axpy"), 2.5, x, y, 1000, n_threads=1000) + # the compiler may fuse y + a * x into one FMA, as NVRTC does by default + np.testing.assert_allclose(y, expected, rtol=1e-15, atol=0) + exact = rng.random(1000) + exact_expected = exact + 2.5 * x + emulate_cuda_kernel( + CudaKernel(AXPY, "axpy"), + 2.5, + x, + exact, + 1000, + n_threads=1000, + options=("-ffp-contract=off",), + ) + np.testing.assert_array_equal(exact, exact_expected) # no FMA: NumPy's rounding + + +COLUMN = r""" +#include "cunumpy/array_view.cuh" +#include +extern "C" __global__ +void scale_column(Array2D a, long long column, double factor) { + CUNUMPY_GRID_STRIDE_1D(i, a.shape[0]) { + a(i, column) *= factor; + } +} +""" + + +def test_strided_view_written_back_into_the_callers_array(): + markers = np.arange(24.0).reshape(8, 3) + every_second_row = markers[::2] # a non-contiguous view + emulate_cuda_kernel( + CudaKernel(COLUMN, "scale_column"), every_second_row, 1, 10.0, grid=1, block=2 + ) + expected = np.arange(24.0).reshape(8, 3) + expected[::2, 1] *= 10.0 + np.testing.assert_array_equal(markers, expected) + + +TEMPLATE = r""" +template +__global__ void power(T* x, int n) { + int i = blockDim.x * blockIdx.x + threadIdx.x; + if (i < n) { T v = x[i]; for (int k = 1; k < K; ++k) x[i] *= v; } +} +""" + + +def test_template_kernel(): + x = np.array([1.0, 2.0, 3.0], dtype=np.float32) + kernel = CudaKernel(TEMPLATE, "power", template_args=(np.float32, 3)) + emulate_cuda_kernel(kernel, x, 3, n_threads=3) + assert x.tolist() == [1.0, 8.0, 27.0] + + +GRID_2D = r""" +#include "cunumpy/array_view.cuh" +extern "C" __global__ void fill(Array2D a) { + long long i = blockIdx.y * (long long)blockDim.y + threadIdx.y; + long long j = blockIdx.x * (long long)blockDim.x + threadIdx.x; + if (i < a.shape[0] && j < a.shape[1]) a(i, j) = 10 * i + j; +} +""" + + +def test_2d_launch(): + a = np.zeros((5, 7), dtype=np.int64) + kernel = CudaKernel(GRID_2D, "fill", block_size=(4, 2)) + emulate_cuda_kernel(kernel, a, n_threads=(7, 5)) + np.testing.assert_array_equal(a, 10 * np.arange(5)[:, None] + np.arange(7)) + + +HISTOGRAM = r""" +#include "cunumpy/atomic.cuh" +extern "C" __global__ +void histogram(const double* x, int n, double* counts, double lower, double width, int bins) { + int i = blockDim.x * blockIdx.x + threadIdx.x; + if (i >= n) return; + int b = (int)floor((x[i] - lower) / width); + if (b >= 0 && b < bins) cunumpy_atomic_add(&counts[b], 1.0); +} +""" + + +def test_atomics_and_math_functions(): + x = np.random.default_rng(1).normal(size=5000) + counts = np.zeros(10) + emulate_cuda_kernel( + CudaKernel(HISTOGRAM, "histogram"), + x, + x.size, + counts, + -2.5, + 0.5, + 10, + n_threads=x.size, + ) + np.testing.assert_array_equal(counts, np.histogram(x, 10, (-2.5, 2.5))[0]) + + +def test_scalars_are_checked_like_a_launch(): + kernel = CudaKernel(AXPY, "axpy") + x, y = np.ones(3), np.zeros(3) + emulate_cuda_kernel(kernel, np.int64(2), x, y, np.int64(3), n_threads=3) + assert y.tolist() == [2.0] * 3 + with pytest.raises(TypeError): + emulate_cuda_kernel(kernel, 1.0, x, y, 2.5, n_threads=3) + with pytest.raises(OverflowError): + emulate_cuda_kernel(kernel, 1.0, x, y, 2**40, n_threads=3) + + +def test_arrays_are_checked(): + kernel = CudaKernel(AXPY, "axpy") + with pytest.raises(TypeError, match="dtype float64"): + emulate_cuda_kernel( + kernel, 1.0, np.ones(3, np.float32), np.ones(3), 3, n_threads=3 + ) + with pytest.raises(TypeError, match="NumPy array"): + emulate_cuda_kernel(kernel, 1.0, [1.0], np.ones(3), 3, n_threads=3) + with pytest.raises(TypeError, match="takes 4 arguments"): + emulate_cuda_kernel(kernel, 1.0, np.ones(3), n_threads=3) + view = CudaKernel(COLUMN, "scale_column") + with pytest.raises(TypeError, match="2D array"): + emulate_cuda_kernel(view, np.ones(3), 0, 1.0, n_threads=3) + + +@pytest.mark.parametrize( + ("body", "what"), + [ + ("__syncwarp();", "__syncwarp"), + ("double v = __shfl_xor_sync(0xffffffff, 1.0, 1);", "warp shuffles"), + ("int v = __ballot_sync(0xffffffff, n > 0);", "warp votes"), + ], +) +def test_unsupported_constructs_are_refused(body, what): + kernel = CudaKernel(f'extern "C" __global__ void k(int n) {{ {body} }}', "k") + with pytest.raises(NotImplementedError, match=what): + emulate_cuda_kernel(kernel, 1, n_threads=1) + + +def test_warp_shuffles_in_an_included_header_are_refused(): + kernel = CudaKernel( + '#include "cunumpy/reduce.cuh"\n' + 'extern "C" __global__ void k(double* out, int n) {' + " double s = cunumpy_block_sum(1.0); if (threadIdx.x == 0) out[0] = s; }", + "k", + ) + with pytest.raises(NotImplementedError, match="warp shuffles"): + emulate_cuda_kernel(kernel, np.zeros(1), 1, n_threads=32) + + +TREE_SUM = r""" +extern "C" __global__ void block_sums(const double* x, long long n, double* sums) { + __shared__ double partial[256]; + const long long i = blockIdx.x * (long long)blockDim.x + threadIdx.x; + partial[threadIdx.x] = i < n ? x[i] : 0.0; + __syncthreads(); + for (unsigned int s = blockDim.x / 2; s > 0; s /= 2) { + if (threadIdx.x < s) partial[threadIdx.x] += partial[threadIdx.x + s]; + __syncthreads(); + } + if (threadIdx.x == 0) sums[blockIdx.x] = partial[0]; +} +""" + + +def test_shared_memory_tree_reduction_with_barriers(): + x = np.random.default_rng(2).random(1000) + sums = np.zeros(4) + kernel = CudaKernel(TREE_SUM, "block_sums", block_size=256) + emulate_cuda_kernel(kernel, x, x.size, sums, n_threads=x.size) + expected = [x[b * 256 : (b + 1) * 256].sum() for b in range(4)] + np.testing.assert_allclose(sums, expected, rtol=1e-12) + + +BLOCK_DEPOSIT = r""" +#include "cunumpy/atomic.cuh" +extern "C" __global__ +void deposit(const double* positions, long long n, double* field, int nx) { + extern __shared__ double block_field[]; + const long long i = blockIdx.x * (long long)blockDim.x + threadIdx.x; + for (int k = threadIdx.x; k < nx; k += blockDim.x) block_field[k] = 0.0; + __syncthreads(); + if (i < n) { + int cell = (int)floor(positions[i] * nx); + if (cell < 0) cell = 0; + if (cell > nx - 1) cell = nx - 1; + cunumpy_atomic_add(&block_field[cell], 1.0); + } + __syncthreads(); + for (int k = threadIdx.x; k < nx; k += blockDim.x) + cunumpy_atomic_add(&field[k], block_field[k]); +} +""" + + +def test_dynamic_shared_memory_per_block_deposit(): + positions = np.random.default_rng(3).random(5000) + field = np.zeros(37) + kernel = CudaKernel(BLOCK_DEPOSIT, "deposit", block_size=128) + emulate_cuda_kernel( + kernel, + positions, + positions.size, + field, + 37, + n_threads=positions.size, + shared_mem=37 * 8, + ) + np.testing.assert_array_equal(field, np.histogram(positions, 37, (0, 1))[0]) + + +REVERSE = r""" +extern "C" __global__ void reverse_blocks(double* x) { + __shared__ double tile[64]; + const long long i = blockIdx.x * (long long)blockDim.x + threadIdx.x; + tile[threadIdx.x] = x[i]; + __syncthreads(); // every thread must have written before anyone reads + x[i] = tile[blockDim.x - 1 - threadIdx.x]; +} +""" + + +def test_barrier_orders_writes_before_reads(): + x = np.arange(128.0) + emulate_cuda_kernel( + CudaKernel(REVERSE, "reverse_blocks", block_size=64), x, n_threads=128 + ) + expected = np.concatenate([np.arange(64.0)[::-1], np.arange(64.0, 128.0)[::-1]]) + np.testing.assert_array_equal(x, expected) + + +EARLY_EXIT = r""" +extern "C" __global__ void early(double* x, int n) { + __shared__ double s[32]; + const int i = blockIdx.x * blockDim.x + threadIdx.x; + if (i >= n) return; // leaves before the barrier + s[threadIdx.x] = 2.0 * x[i]; + __syncthreads(); + x[i] = s[threadIdx.x]; +} +""" + + +def test_threads_leaving_before_a_barrier_do_not_hang(): + x = np.ones(40) + emulate_cuda_kernel( + CudaKernel(EARLY_EXIT, "early", block_size=32), x, 40, n_threads=40 + ) + assert x.tolist() == [2.0] * 40 + + +TRANSPOSE_TILE = r""" +#include "cunumpy/array_view.cuh" +extern "C" __global__ void transpose(Array2D a, Array2D out) { + __shared__ double tile[4][4]; + const int i = blockIdx.y * blockDim.y + threadIdx.y; + const int j = blockIdx.x * blockDim.x + threadIdx.x; + tile[threadIdx.y][threadIdx.x] = a(i, j); + __syncthreads(); + const int ti = blockIdx.x * blockDim.x + threadIdx.y; + const int tj = blockIdx.y * blockDim.y + threadIdx.x; + out(ti, tj) = tile[threadIdx.x][threadIdx.y]; +} +""" + + +def test_2d_blocks_with_shared_tiles(): + a = np.arange(64.0).reshape(8, 8) + out = np.zeros((8, 8)) + kernel = CudaKernel(TRANSPOSE_TILE, "transpose", block_size=(4, 4)) + emulate_cuda_kernel(kernel, a, out, n_threads=(8, 8)) + np.testing.assert_array_equal(out, a.T) + + +def test_compile_errors_and_crashes_are_reported(): + broken = CudaKernel( + 'extern "C" __global__ void k(int n) { undefined_call(n); }', "k" + ) + with pytest.raises(RuntimeError, match="does not compile"): + emulate_cuda_kernel(broken, 1, n_threads=1) + trap = CudaKernel( + 'extern "C" __global__ void k(int n) { if (n > 0) __trap(); }', "k" + ) + with pytest.raises(RuntimeError, match="crashed"): + emulate_cuda_kernel(trap, 1, n_threads=1) + + +def test_bounds_checks_from_the_view_header(): + kernel = CudaKernel(COLUMN, "scale_column") + a = np.ones((4, 2)) + with pytest.raises(RuntimeError, match="crashed"): + emulate_cuda_kernel( + kernel, a, 5, 2.0, n_threads=4, options=("-DCUNUMPY_BOUNDS_CHECK",) + ) diff --git a/tests/unit/test_fusion.py b/tests/unit/test_fusion.py new file mode 100644 index 0000000..717ffe9 --- /dev/null +++ b/tests/unit/test_fusion.py @@ -0,0 +1,97 @@ +"""Tests for `xp.fuse`: cupy.fuse for CuPy arrays, a plain call otherwise.""" + +import numpy as np +import pytest + +import cunumpy as xp +from cunumpy import fusion + + +def pressure(rho, T, gamma): + return (gamma - 1.0) * rho * xp.exp(T) + + +def test_host_arrays_call_the_function(): + fused = xp.fuse(pressure) + rho, T = np.full(4, 2.0), np.zeros(4) + np.testing.assert_array_equal(fused(rho, T, 3.0), pressure(rho, T, 3.0)) + assert fused.__name__ == "pressure" and fused.__wrapped__ is pressure + assert "fuse" in xp.__all__ + + +def test_decorator_forms(): + @xp.fuse + def double(x): + return 2.0 * x + + @xp.fuse(kernel_name="triple_kernel") + def triple(x): + return 3.0 * x + + assert double(np.ones(2)).tolist() == [2.0, 2.0] + assert triple(np.ones(2)).tolist() == [3.0, 3.0] + + +class FakeDeviceArray: + def __init__(self, values): + self.values = np.asarray(values) + + +def test_device_arrays_use_cupy_fuse_once(monkeypatch): + created = [] + + def fake_cupy_fuse(function, kernel_name): + created.append(kernel_name) + + def fused(*args, **kwargs): + assert xp.get_backend() in ("cupy", "numpy") # numpy: no CuPy installed + return ("fused", function.__name__, len(args), sorted(kwargs)) + + return fused + + monkeypatch.setattr(fusion, "_cupy_fuse", fake_cupy_fuse) + monkeypatch.setattr( + fusion, "_is_device_array", lambda a: isinstance(a, FakeDeviceArray) + ) + + @xp.fuse(kernel_name="p") + def p(rho, T, gamma=1.0): + raise AssertionError("not called with device arrays") + + rho = FakeDeviceArray([1.0]) + assert p(rho, 0.0) == ("fused", "p", 2, []) + assert p(1.0, 0.0, gamma=rho) == ("fused", "p", 2, ["gamma"]) # keyword array + assert created == ["p"] # fused once, reused + with pytest.raises(AssertionError, match="not called"): + p(np.ones(1), 0.0) # host arrays: the function itself + + +def test_python_scalars_take_the_dtype_of_the_arrays(monkeypatch): + seen = [] + monkeypatch.setattr( + fusion, + "_cupy_fuse", + lambda function, kernel_name: lambda *a, **k: seen.append((a, k)), + ) + monkeypatch.setattr(fusion, "_is_device_array", lambda a: hasattr(a, "dtype")) + + p = xp.fuse(pressure) + p(np.ones(2), 0.5, 5.0 / 3.0) + p(np.ones(2, dtype=np.float32), 0.5, gamma=2) + p(np.ones(2, dtype=np.int64), True, 3) + (a64, _), (a32, k32), (aint, _) = seen + assert a64[1].dtype == a64[2].dtype == np.float64 and a64[2] == 5.0 / 3.0 + assert a32[1].dtype == k32["gamma"].dtype == np.float32 + assert aint[1] is True and aint[2].dtype == np.int64 # bools stay Python + + +def test_fuse_on_gpu(): + if not xp.cupy_available(): + pytest.skip("CuPy not installed or not functional") + import cupy as cp + + fused = xp.fuse(pressure) + rho, T = cp.full(1000, 2.0), cp.linspace(0.0, 1.0, 1000) + with xp.use_backend("cupy"): + expected = pressure(rho, T, 5.0 / 3.0) + cp.testing.assert_allclose(fused(rho, T, 5.0 / 3.0), expected, rtol=1e-14) diff --git a/tests/unit/test_kernel_dispatch_arrays.py b/tests/unit/test_kernel_dispatch_arrays.py new file mode 100644 index 0000000..27d9ac0 --- /dev/null +++ b/tests/unit/test_kernel_dispatch_arrays.py @@ -0,0 +1,292 @@ +"""Kernels chosen by where the arguments live, compiled host kernels, signature checks.""" + +import sys +import textwrap +import warnings +from types import SimpleNamespace + +import numpy as np +import pytest + +import cunumpy as xp +from cunumpy import ( + CompiledHostKernel, + CudaArguments, + CudaKernel, + Kernel, + KernelArguments, + KernelCatalog, +) +from cunumpy import dispatch as dispatch_module + +SCALE_CUDA = r""" +extern "C" __global__ void scale(double* x, double factor, int n) { + int i = blockDim.x * blockIdx.x + threadIdx.x; + if (i < n) x[i] *= factor; +} +""" + + +def scale(x, factor, n): + for i in range(n): + x[i] *= factor + + +class FakeDeviceArray: + """Stands for a CuPy array (see the `fake_gpu` fixture).""" + + def __init__(self, n): + self.shape = (n,) + + +@pytest.fixture +def fake_gpu(monkeypatch): + """CuPy is the active backend, FakeDeviceArray counts as a device array, and + CUDA launches are recorded instead of run.""" + monkeypatch.setattr(dispatch_module, "get_backend", lambda: "cupy") + monkeypatch.setattr( + dispatch_module, "_is_device_array", lambda a: isinstance(a, FakeDeviceArray) + ) + launches = [] + + def launch(self, *args, **kwargs): + launches.append((self.name, args, kwargs)) + + monkeypatch.setattr(CudaKernel, "__call__", launch) + return launches + + +def test_dispatch_must_be_known(): + with pytest.raises(ValueError, match="dispatch must be one of"): + Kernel(scale, dispatch="device") + assert Kernel(scale).dispatch == "backend" + + +def test_arrays_dispatch_runs_host_arrays_on_the_host_while_cupy_is_active(fake_gpu): + kernel = Kernel(scale, CudaKernel(SCALE_CUDA, "scale"), dispatch="arrays") + x = np.ones(4) + kernel(x, 3.0, 4) # NumPy arrays: the host kernel, although CuPy is active + assert x.tolist() == [3.0] * 4 and fake_gpu == [] + + +def test_backend_dispatch_sends_everything_to_cuda_on_cupy(fake_gpu): + kernel = Kernel(scale, CudaKernel(SCALE_CUDA, "scale")) + x = np.ones(4) + kernel(x, 3.0, 4, n_threads=4) # the old rule: the backend decides + assert [name for name, *_ in fake_gpu] == ["scale"] + assert x.tolist() == [1.0] * 4 + + +def test_arrays_dispatch_runs_device_arrays_on_the_gpu(fake_gpu): + kernel = Kernel(scale, CudaKernel(SCALE_CUDA, "scale"), dispatch="arrays") + x = FakeDeviceArray(4) + kernel(x, 3.0, 4, n_threads=4) + ((name, args, options),) = fake_gpu + assert name == "scale" and args[0] is x and options["n_threads"] == 4 + with pytest.raises(ValueError, match="n_threads is required"): + kernel(x, 3.0, 4) + + +def test_device_only_argument_objects_count_as_device(fake_gpu, monkeypatch): + kernel = Kernel(scale, CudaKernel(SCALE_CUDA, "scale"), dispatch="arrays") + + class Device(CudaArguments): + pass + + kernel(Device(np.ones(1)), 2.0, 1, n_threads=1) + assert len(fake_gpu) == 1 + + class Both(KernelArguments): # host and device form: does not decide + def __host_args__(self): + return np.ones(2) + + host = Kernel(lambda x, f, n: x.__setitem__(slice(None), f), dispatch="arrays") + host(Both(), 5.0, 2) # no device argument: the host kernel + + +def test_arrays_dispatch_without_cuda_kernel(fake_gpu): + raising = Kernel(scale, dispatch="arrays") + with pytest.raises(NotImplementedError, match="No CUDA version"): + raising(FakeDeviceArray(2), 2.0, 2, n_threads=2) + x = np.ones(2) + raising(x, 2.0, 2) # host arrays never need the CUDA kernel + assert x.tolist() == [2.0, 2.0] + + +class CompiledFunction: + """Like a compiled extension function: no Python signature to read.""" + + __signature__ = "unreadable" # inspect.signature raises TypeError + + def __call__(self, *args): + return None + + +def test_check_signature(): + Kernel(scale, CudaKernel(SCALE_CUDA, "scale")).check_signature() # same names + + def renamed(x, a, n): ... + + with pytest.raises( + ValueError, + match=r"host kernel takes \(x, a, n\), the CUDA kernel \(x, factor, n\)", + ): + Kernel(renamed, CudaKernel(SCALE_CUDA, "scale")).check_signature() + + def reordered(factor, x, n): ... + + with pytest.raises(ValueError, match="host kernel takes"): + Kernel( + reordered, CudaKernel(SCALE_CUDA, "scale"), name="scale" + ).check_signature() + + # nothing to compare: no CUDA kernel, an unparsed signature, a compiled builtin + Kernel(renamed).check_signature() + Kernel( + renamed, CudaKernel(SCALE_CUDA, "scale", check_signature=False) + ).check_signature() + compiled = CompiledFunction() + Kernel(compiled, CudaKernel(SCALE_CUDA, "scale"), name="scale").check_signature() + assert Kernel(compiled, name="m").host_parameters() is None + + +def test_catalog_check_signatures_lists_every_mismatch(): + def renamed(x, a, n): ... + + catalog = KernelCatalog() + catalog.register(Kernel(scale, CudaKernel(SCALE_CUDA, "scale"))) + catalog.register(Kernel(renamed, CudaKernel(SCALE_CUDA, "scale"), name="one")) + catalog.register(Kernel(renamed, CudaKernel(SCALE_CUDA, "scale"), name="two")) + with pytest.raises(ValueError, match=r"(?s)'one'.*'two'"): + catalog.check_signatures() + + +# --------------------------------------------------------------------------- +# compiled host kernels +# --------------------------------------------------------------------------- + + +def _module(name, source): + module = SimpleNamespace(__name__=name) + exec(textwrap.dedent(source), module.__dict__) # noqa: S102 + return module + + +def test_compiled_host_kernel_compiles_once(): + module = _module("m", "def double(x):\n x *= 2\n") + compiled = [] + + def compiler(mod): + compiled.append(mod) + return SimpleNamespace(double=lambda x: x.__setitem__(..., x * 10)) + + kernel = CompiledHostKernel(module, "double", compiler) + assert repr(kernel) == "CompiledHostKernel('double', not built)" + x = np.ones(2) + kernel(x) + kernel(x) + assert x.tolist() == [100.0, 100.0] and compiled == [module] + assert kernel.compiled and kernel.error is None + assert kernel.python is module.double + + +def test_compiled_host_kernel_falls_back(): + module = _module("m", "def double(x):\n x *= 2\n") + + def failing(mod): + raise ImportError("no pyccel") + + fallback_calls = [] + with_fallback = CompiledHostKernel( + module, "double", failing, fallback=fallback_calls.append + ) + with_fallback("x") + assert fallback_calls == ["x"] and not with_fallback.compiled + assert isinstance(with_fallback.error, ImportError) + + python = CompiledHostKernel(module, "double", failing) + x = np.ones(2) + with pytest.warns(RuntimeWarning, match="uncompiled Python version"): + python(x) + with warnings.catch_warnings(): + warnings.simplefilter("error") + python(x) # warned once + assert x.tolist() == [4.0, 4.0] + + +@pytest.fixture +def pyccel_style_package(tmp_path, monkeypatch): + root = tmp_path / "demo_pyccel_pkg" + (root / "scale").mkdir(parents=True) + (root / "__init__.py").write_text("") + (root / "scale" / "__init__.py").write_text("") + (root / "scale" / "scale_pyccel.py").write_text( + "def scale(x: 'float[:]', factor: float, n: int):\n" + " for i in range(n):\n" + " x[i] *= factor\n" + ) + (root / "scale" / "scale_cuda.cu").write_text(SCALE_CUDA) + monkeypatch.syspath_prepend(str(tmp_path)) + yield "demo_pyccel_pkg" + for module in [m for m in sys.modules if m.startswith("demo_pyccel_pkg")]: + del sys.modules[module] + + +def test_from_package_with_compiled_hosts_and_array_dispatch(pyccel_style_package): + calls = [] + + def compiler(module): + calls.append(module.__name__) + return module # "compiled": the Python module itself + + catalog = KernelCatalog.from_package( + pyccel_style_package, + host_suffix="_pyccel", + dispatch="arrays", + compile_host=compiler, + ) + kernel = catalog["scale"] + assert kernel.dispatch == "arrays" + assert isinstance(kernel.host_kernel.kernel, CompiledHostKernel) + catalog.check_signatures() # x, factor, n on both sides + x = np.ones(3) + kernel(x, 2.0, 3) + assert x.tolist() == [2.0] * 3 and calls == ["demo_pyccel_pkg.scale.scale_pyccel"] + + +def test_from_package_host_fallback(pyccel_style_package): + def failing(module): + raise RuntimeError("compiler missing") + + used = [] + catalog = KernelCatalog.from_package( + pyccel_style_package, + host_suffix="_pyccel", + compile_host=failing, + host_fallback=lambda name: lambda *args: used.append((name, args)), + ) + catalog["scale"](np.ones(1), 2.0, 1) + assert used and used[0][0] == "scale" + mapping = KernelCatalog.from_package( + pyccel_style_package, + host_suffix="_pyccel", + compile_host=failing, + host_fallback={"scale": lambda *args: used.append("mapped")}, + ) + mapping["scale"](np.ones(1), 2.0, 1) + assert used[-1] == "mapped" + + +def test_arrays_dispatch_on_gpu(): + if not xp.cupy_available(): + pytest.skip("CuPy not installed or not functional") + import cupy as cp + + kernel = Kernel(scale, CudaKernel(SCALE_CUDA, "scale"), dispatch="arrays") + with xp.use_backend("cupy"): + host = np.ones(5) + kernel(host, 2.0, 5) # host array: the host kernel, no transfer + device = cp.ones(5) + kernel(device, 3.0, 5, n_threads=5) + assert host.tolist() == [2.0] * 5 + assert device.get().tolist() == [3.0] * 5 diff --git a/tests/unit/test_mirror.py b/tests/unit/test_mirror.py index d9a0a2e..1a97892 100644 --- a/tests/unit/test_mirror.py +++ b/tests/unit/test_mirror.py @@ -130,7 +130,8 @@ def test_cuda_include_dir_contains_atomic_header(): assert include_dir.endswith(os.path.join("cuda", "include")) header = os.path.join(include_dir, "cunumpy", "atomic.cuh") assert os.path.isfile(header) - source = open(header).read() + with open(header) as f: + source = f.read() assert "cunumpy_atomic_add(double* p, double v)" in source assert "cunumpy_atomic_add(float* p, float v)" in source assert "cunumpy_atomic_add_2d(" in source @@ -147,7 +148,10 @@ def test_cuda_kernel_options_include_cunumpy_headers(): assert kernel.compile_options().count(flag) == 1 assert kernel.compile_options()[0] == "-I/some/dir" kernel = CudaKernel(BIN_ADD, "bin_add", options=["-std=c++17", flag]) - assert kernel.compile_options() == ("-std=c++17", flag) + options = kernel.compile_options() + assert options[:2] == ("-std=c++17", flag) + # the shipped atomic.cuh is part of the header hash + assert len(options) == 3 and options[2].startswith("-DCUNUMPY_INCLUDE_HASH=0x") # --- CuPy backend ------------------------------------------------------------ diff --git a/tests/unit/test_particle_recipes.py b/tests/unit/test_particle_recipes.py new file mode 100644 index 0000000..c447513 --- /dev/null +++ b/tests/unit/test_particle_recipes.py @@ -0,0 +1,67 @@ +"""The recipes of the "Particle codes" page (docs/source/examples/particle_recipes.py).""" + +import importlib.util +from pathlib import Path + +import numpy as np +import pytest + +_PATH = Path(__file__).resolve().parents[2] / "docs/source/examples/particle_recipes.py" +_spec = importlib.util.spec_from_file_location("particle_recipes", _PATH) +recipes = importlib.util.module_from_spec(_spec) +_spec.loader.exec_module(recipes) + + +def test_remove_dead_and_compact_in_place(): + markers = np.arange(12.0).reshape(6, 2) + alive = np.array([True, False, True, True, False, True]) + np.testing.assert_array_equal(recipes.remove_dead(markers, alive), markers[alive]) + buffer = markers.copy() + n = recipes.compact_in_place(buffer, alive) + assert n == 4 + np.testing.assert_array_equal(buffer[:n], markers[alive]) + + +def test_sort_by_cell_and_deposit(): + rng = np.random.default_rng(0) + x = rng.uniform(-0.1, 1.1, 1000) # some outside: clipped to the edge cells + weights = rng.random(1000) + order, cell, offsets = recipes.sort_by_cell(x, 0.0, 0.1, 10) + assert offsets[0] == 0 and offsets[-1] == 1000 + sorted_x = x[order] + for c in range(10): + members = sorted_x[offsets[c] : offsets[c + 1]] + expected = np.clip(np.floor(members / 0.1), 0, 9) + assert np.all(expected == c) and np.all(cell[offsets[c] : offsets[c + 1]] == c) + # stable: the original order within a cell + first_cell = order[offsets[3] : offsets[4]] + assert np.all(np.diff(first_cell) > 0) + deposit = recipes.deposit_nearest_cell(cell, weights[order], 10) + expected = np.bincount(np.clip(np.floor(x / 0.1), 0, 9).astype(int), weights, 10) + np.testing.assert_allclose(deposit, expected, rtol=1e-12) + + +def test_pack_for_ranks(): + markers = np.arange(10.0).reshape(5, 2) + destination = np.array([2, 0, 2, 1, 0]) + sendbuf, counts = recipes.pack_for_ranks(markers, destination, 3) + assert counts.tolist() == [2, 1, 2] + np.testing.assert_array_equal(sendbuf, markers[[1, 4, 3, 0, 2]]) + + +def test_exchange_on_one_rank(): + mpi4py = pytest.importorskip("mpi4py") + from mpi4py import MPI + + del mpi4py + markers = np.arange(8.0).reshape(4, 2) + received = recipes.exchange(MPI.COMM_SELF, markers, np.zeros(4, dtype=np.int64)) + np.testing.assert_array_equal(received, markers) + + +def test_thermal_velocities_are_reproducible(): + ids = np.arange(100_000, dtype=np.uint64) + v0, v1 = recipes.thermal_velocities(7, ids, 3, 2.0) + assert abs(v0.std() - 2.0) < 0.03 and abs(v1.mean()) < 0.03 + again, _ = recipes.thermal_velocities(7, ids[::-1], 3, 2.0) # other order + np.testing.assert_array_equal(again[::-1], v0) diff --git a/tests/unit/test_petsc.py b/tests/unit/test_petsc.py new file mode 100644 index 0000000..b2a6eb9 --- /dev/null +++ b/tests/unit/test_petsc.py @@ -0,0 +1,133 @@ +"""Tests for `xp.petsc_vec`: PETSc vectors sharing the memory of an array.""" + +import sys +import types + +import numpy as np +import pytest + +import cunumpy as xp +from cunumpy import petsc + + +def _petsc(): + petsc4py = pytest.importorskip("petsc4py") + petsc4py.init() + from petsc4py import PETSc + + return PETSc + + +def test_host_vector_shares_memory(): + PETSc = _petsc() + a = np.zeros(6, dtype=PETSc.ScalarType) + vec = xp.petsc_vec(a, comm=PETSc.COMM_SELF) + assert vec.getSize() == 6 and vec.getType() == "seq" + vec.set(3.0) + assert a.tolist() == [3.0] * 6 # PETSc wrote into the array + a[0] = -1.0 + assert vec.getValue(0) == -1.0 # and reads from it + assert vec.getAttr("cunumpy_array") is a # keeps the array alive + + +def test_multidimensional_arrays_are_unrolled(): + PETSc = _petsc() + a = np.arange(6, dtype=PETSc.ScalarType).reshape(2, 3) + vec = xp.petsc_vec(a, comm=PETSc.COMM_SELF) + assert vec.getArray().tolist() == a.ravel().tolist() + + +def test_ksp_solve_writes_into_the_array(): + PETSc = _petsc() + n = 10 + A = PETSc.Mat().createAIJ([n, n], nnz=3, comm=PETSc.COMM_SELF) + for i in range(n): + A.setValue(i, i, 2.0) + if i > 0: + A.setValue(i, i - 1, -1.0) + if i < n - 1: + A.setValue(i, i + 1, -1.0) + A.assemble() + b, x = np.ones(n), np.zeros(n) + ksp = PETSc.KSP().create(comm=PETSc.COMM_SELF) + ksp.setOperators(A) + ksp.setType("cg") + ksp.getPC().setType("none") + ksp.setTolerances(rtol=1e-12) + ksp.solve( + xp.petsc_vec(b, comm=PETSc.COMM_SELF), xp.petsc_vec(x, comm=PETSc.COMM_SELF) + ) + residual = 2 * x - np.r_[0.0, x[:-1]] - np.r_[x[1:], 0.0] - 1.0 + assert np.abs(residual).max() < 1e-9 + + +def test_rejects_arrays_that_would_need_a_copy(): + PETSc = _petsc() + other = np.float32 if np.dtype(PETSc.ScalarType) != np.float32 else np.float64 + with pytest.raises(TypeError, match="scalar type"): + xp.petsc_vec(np.zeros(4, dtype=other)) + with pytest.raises(ValueError, match="C-contiguous"): + xp.petsc_vec(np.zeros((4, 4))[:, 0]) + with pytest.raises(TypeError, match="NumPy or CuPy array"): + xp.petsc_vec([0.0, 1.0]) + + +class FakeDeviceArray: + dtype = np.dtype(np.float64) + flags = types.SimpleNamespace(c_contiguous=True) + + +def _fake_petsc(monkeypatch, vec_type=None, error=False): + class Error(Exception): + pass + + class Vec: + destroyed = False + + def createWithDLPack(self, array, comm=None): + if error: + raise Error("no CUDA") + return self + + def getType(self): + return vec_type + + def destroy(self): + Vec.destroyed = True + + def setAttr(self, name, value): + self.attr = (name, value) + + PETSc = types.SimpleNamespace(Vec=Vec, Error=Error, ScalarType=np.float64) + package = types.ModuleType("petsc4py") + package.PETSc = PETSc + monkeypatch.setitem(sys.modules, "petsc4py", package) + monkeypatch.setitem(sys.modules, "petsc4py.PETSc", PETSc) + monkeypatch.setattr( + petsc, "_is_device_array", lambda a: isinstance(a, FakeDeviceArray) + ) + return Vec + + +@pytest.mark.parametrize("vec_type", ["seqcuda", "mpicuda", "seqhip"]) +def test_device_arrays_give_device_vectors(monkeypatch, vec_type): + _fake_petsc(monkeypatch, vec_type) + array = FakeDeviceArray() + vec = xp.petsc_vec(array) + assert vec.attr == ("cunumpy_array", array) + + +def test_device_arrays_need_a_gpu_petsc(monkeypatch): + Vec = _fake_petsc(monkeypatch, "seq") # PETSc without CUDA made a host vector + with pytest.raises(RuntimeError, match="created a 'seq' vector for a CuPy array"): + xp.petsc_vec(FakeDeviceArray()) + assert Vec.destroyed + _fake_petsc(monkeypatch, error=True) + with pytest.raises(RuntimeError, match="built with CUDA or HIP support"): + xp.petsc_vec(FakeDeviceArray()) + + +def test_missing_petsc4py(monkeypatch): + monkeypatch.setitem(sys.modules, "petsc4py", None) + with pytest.raises(ImportError, match="needs petsc4py"): + xp.petsc_vec(np.zeros(3)) diff --git a/tests/unit/test_philox.py b/tests/unit/test_philox.py new file mode 100644 index 0000000..04c1ff7 --- /dev/null +++ b/tests/unit/test_philox.py @@ -0,0 +1,135 @@ +"""Tests for the Philox4x32-10 generator: host functions and cunumpy/random.cuh.""" + +from pathlib import Path + +import numpy as np +import pytest + +import cunumpy as xp +from cunumpy import CudaKernel, cuda_include_dir +from cunumpy.testing import emulate_cuda_kernel, emulation_compiler + +# Known-answer vectors of Philox4x32-10 (Random123, kat_vectors) +KAT = [ + ((0, 0, 0, 0), (0, 0), (0x6627E8D5, 0xE169C58D, 0xBC57AC4C, 0x9B00DBD8)), + ( + (0xFFFFFFFF,) * 4, + (0xFFFFFFFF, 0xFFFFFFFF), + (0x408F276D, 0x41C83B0E, 0xA20BC7C6, 0x6D5451FD), + ), + ( + (0x243F6A88, 0x85A308D3, 0x13198A2E, 0x03707344), + (0xA4093822, 0x299F31D0), + (0xD16CFE09, 0x94FDCCEB, 0x5001E420, 0x24126EA1), + ), +] + + +@pytest.mark.parametrize(("counter", "key", "expected"), KAT) +def test_known_answers(counter, key, expected): + out = xp.philox4x32_10(np.array(counter, dtype=np.uint32), *key) + assert out.dtype == np.uint32 + assert out.tolist() == list(expected) + + +def test_vectorized_over_counters(): + counters = np.array([k[0] for k in KAT], dtype=np.uint32) + keys = np.array([k[1] for k in KAT], dtype=np.uint32) + out = xp.philox4x32_10(counters, keys[:, 0], keys[:, 1]) + assert out.tolist() == [list(k[2]) for k in KAT] + + +def test_uniforms_and_normals(): + ids = np.arange(200_000, dtype=np.uint64) + u0, u1 = xp.philox_uniform2(42, ids, 7) + assert u0.shape == u1.shape == (200_000,) + assert u0.min() >= 0.0 and u0.max() < 1.0 + assert abs(u0.mean() - 0.5) < 3e-3 and abs(u1.var() - 1 / 12) < 1e-3 + assert abs(np.corrcoef(u0, u1)[0, 1]) < 1e-2 + np.testing.assert_array_equal(xp.philox_uniform(42, ids, 7), u0) + z0, z1 = xp.philox_normal2(42, ids, 7) + assert abs(z0.mean()) < 1e-2 and abs(z1.std() - 1.0) < 1e-2 + np.testing.assert_array_equal(xp.philox_normal(42, ids, 7), z0) + + +def test_streams_counters_and_seeds_differ(): + base = xp.philox_uniform(1, 5, 9) + assert xp.philox_uniform(1, 5, 9) == base # no state + assert xp.philox_uniform(1, 6, 9) != base + assert xp.philox_uniform(1, 5, 10) != base + assert xp.philox_uniform(2, 5, 9) != base + # 64-bit stream and counter: the high words matter + assert xp.philox_uniform(1, 5 + 2**32, 9) != base + assert xp.philox_uniform(1, 5, 9 + 2**32) != base + assert xp.philox_uniform(1 + 2**32, 5, 9) != base + + +def test_broadcasting(): + u = xp.philox_uniform( + np.uint64(3), np.arange(4, dtype=np.uint64)[:, None], np.arange(5) + ) + assert u.shape == (4, 5) + assert u[2, 3] == xp.philox_uniform(3, 2, 3) + + +def test_header_is_shipped(): + header = (Path(cuda_include_dir()) / "cunumpy" / "random.cuh").read_text() + for name in ("cunumpy_philox4x32_10", "cunumpy_uniform2", "cunumpy_normal2"): + assert name in header + + +SAMPLE = r""" +#include "cunumpy/random.cuh" +extern "C" __global__ +void sample(double* u0, double* u1, double* z0, double* z1, long long n, + unsigned long long seed, unsigned long long counter) { + long long i = blockIdx.x * (long long)blockDim.x + threadIdx.x; + if (i >= n) return; + cunumpy_uniform2(seed, (unsigned long long)i, counter, &u0[i], &u1[i]); + cunumpy_normal2(seed, (unsigned long long)i, counter, &z0[i], &z1[i]); +} +""" + + +def _device_samples(run, n=1000, seed=2**40 + 17, counter=2**33 + 5): + u0, u1, z0, z1 = (np.zeros(n) for _ in range(4)) + run(CudaKernel(SAMPLE, "sample"), u0, u1, z0, z1, n, seed, counter, n_threads=n) + return ( + (u0, u1, z0, z1), + xp.philox_uniform2(seed, np.arange(n, dtype=np.uint64), counter), + xp.philox_normal2(seed, np.arange(n, dtype=np.uint64), counter), + ) + + +@pytest.mark.skipif(emulation_compiler() is None, reason="no C++ compiler") +def test_header_matches_the_host_functions_in_emulation(): + (u0, u1, z0, z1), (h0, h1), (n0, n1) = _device_samples(emulate_cuda_kernel) + np.testing.assert_array_equal(u0, h0) # bit-identical + np.testing.assert_array_equal(u1, h1) + np.testing.assert_allclose(z0, n0, rtol=1e-14, atol=1e-14) + np.testing.assert_allclose(z1, n1, rtol=1e-14, atol=1e-14) + + +def test_header_matches_the_host_functions_on_gpu(): + if not xp.cupy_available(): + pytest.skip("CuPy not installed or not functional") + import cupy as cp + + def run(kernel, *args, n_threads): + device = [cp.asarray(a) if isinstance(a, np.ndarray) else a for a in args] + kernel(*device, n_threads=n_threads) + for host, dev in zip(args, device): + if isinstance(host, np.ndarray): + host[...] = dev.get() + + (u0, u1, z0, z1), (h0, h1), (n0, n1) = _device_samples(run) + np.testing.assert_array_equal(u0, h0) + np.testing.assert_array_equal(u1, h1) + np.testing.assert_allclose(z0, n0, rtol=1e-13, atol=1e-13) + np.testing.assert_allclose(z1, n1, rtol=1e-13, atol=1e-13) + # the device-side generator on device arrays, too + du0, _ = xp.philox_uniform2(7, cp.arange(10, dtype=cp.uint64), 1) + assert isinstance(du0, cp.ndarray) + np.testing.assert_array_equal( + du0.get(), xp.philox_uniform2(7, np.arange(10, dtype=np.uint64), 1)[0] + ) diff --git a/tests/unit/test_porting_helpers.py b/tests/unit/test_porting_helpers.py new file mode 100644 index 0000000..a51f592 --- /dev/null +++ b/tests/unit/test_porting_helpers.py @@ -0,0 +1,610 @@ +"""Tests for the kernel-porting helpers added for struphy. + +Device-ownership checks, structs from pyccel source files, argument objects +with a pyccel host class, launch sizes from the arguments, NaN checks, the +Fortran name limit, host parameters from pyccel stubs, test-arguments modules +and parity cases, struct parameters of device-function wrappers, the fake +CuPy, `mpi_buffer`, `segment_sum` and `require_version`. Everything runs +without a GPU; the fake CuPy runs in a subprocess so that it never replaces +CuPy in the test process. +""" + +import os +import pickle +import subprocess +import sys +import textwrap +import warnings +from pathlib import Path +from types import ModuleType, SimpleNamespace + +import numpy as np +import pytest + +import cunumpy as xp +import cunumpy.cuda_kernel as cuda_kernel_module +import cunumpy.testing +from cunumpy import ( + CudaKernel, + CudaStruct, + Kernel, + KernelCatalog, + PyccelStructArguments, +) +from cunumpy.dispatch import FORTRAN_NAME_LIMIT, _pyccel_stub_parameters +from cunumpy.testing import check_parity, device_function_kernel, parity_cases + + +class FakeDeviceArray: + """Enough of a CuPy array for the argument checks, optionally on a device.""" + + def __init__(self, dtype, ptr=0x1000, shape=(4,), device=None): + self.dtype = np.dtype(dtype) + self.data = SimpleNamespace(ptr=ptr) + self.shape = tuple(shape) + self.ndim = len(self.shape) + self.strides = tuple(np.zeros(self.shape, dtype=self.dtype).strides) + self.flags = SimpleNamespace(c_contiguous=True) + if device is not None: + self.device = SimpleNamespace(id=device) + + @property + def __cuda_array_interface__(self): + return {} + + +SCALE = r""" +extern "C" __global__ void scale(double* x, double factor, int n) { + int i = blockDim.x * blockIdx.x + threadIdx.x; + if (i < n) x[i] *= factor; +} +""" + +PYCCEL_SOURCE = textwrap.dedent(""" + "Argument classes, pyccel style." + + from numpy import shape + + + class MarkerArguments: + def __init__( + self, + markers: "float[:,:]", + valid_mks: "bool[:]", + Np: "int", + first_pusher_idx: "int", + bc_type: "int[:]", + ): + self.markers = markers + self.valid_mks = valid_mks + self.Np = Np + self.n_markers = shape(markers)[0] + self.first_init_idx = first_pusher_idx + self.bc_type = bc_type + self.scratch = markers[0] + + + class DerhamArguments: + def __init__(self, pn: "int[:]", tn1: "Final[float[:]]"): + self.pn = pn + self.tn1 = tn1 + self.bn1 = pn + """) + + +# --------------------------------------------------------------------------- +# device-ownership check +# --------------------------------------------------------------------------- + + +def test_arrays_on_another_device_are_rejected(monkeypatch): + kernel = CudaKernel(SCALE, "scale") + monkeypatch.setattr(cuda_kernel_module, "_current_device_id", lambda: 0) + kernel.prepare_args(FakeDeviceArray(np.float64), 2.0, 4) # no device attribute + kernel.prepare_args(FakeDeviceArray(np.float64, device=0), 2.0, 4) + with pytest.raises( + ValueError, match="on CUDA device 1, but the current device is 0" + ): + kernel.prepare_args(FakeDeviceArray(np.float64, device=1), 2.0, 4) + # struct pointer fields go through the same check + struct = CudaStruct("Vec", [("x", "double*"), ("n", "int")]) + with pytest.raises(ValueError, match="on CUDA device 1"): + struct(x=FakeDeviceArray(np.float64, device=1), n=4) + + +def test_device_check_is_skipped_without_cupy(monkeypatch): + monkeypatch.delitem(sys.modules, "cupy", raising=False) + assert cuda_kernel_module._current_device_id() is None + CudaKernel(SCALE, "scale").prepare_args( + FakeDeviceArray(np.float64, device=7), 2.0, 4 + ) + + +# --------------------------------------------------------------------------- +# CudaStruct.from_pyccel_class +# --------------------------------------------------------------------------- + + +def test_from_pyccel_class_from_source_and_file(tmp_path): + struct = CudaStruct.from_pyccel_class( + PYCCEL_SOURCE, "MarkerArguments", "MarkerArgs" + ) + assert struct.name == "MarkerArgs" + # fields are named after the attributes the parameters are stored in + assert [(f.name, f.ctype) for f in struct.fields] == [ + ("markers", "Array2D"), + ("valid_mks", "Array1D"), + ("Np", "long long"), + ("first_init_idx", "long long"), + ("bc_type", "Array1D"), + ] + path = tmp_path / "pusher_args_kernels.py" + path.write_text(PYCCEL_SOURCE) + from_file = CudaStruct.from_pyccel_class(path, "MarkerArguments") + assert from_file.name == "MarkerArguments" + assert from_file.dtype == struct.dtype + same = CudaStruct.from_pyccel_class( + str(path), "MarkerArguments", "A", int_type="int" + ) + assert [f.ctype for f in same.fields][2] == "int" + + +def test_from_pyccel_class_options(): + keep = CudaStruct.from_pyccel_class( + PYCCEL_SOURCE, "MarkerArguments", attribute_names=False + ) + assert [f.name for f in keep.fields][3] == "first_pusher_idx" + fewer = CudaStruct.from_pyccel_class( + PYCCEL_SOURCE, "MarkerArguments", exclude=("bc_type", "first_pusher_idx") + ) + assert [f.name for f in fewer.fields] == ["markers", "valid_mks", "Np"] + final = CudaStruct.from_pyccel_class(PYCCEL_SOURCE, "DerhamArguments", "DerhamArgs") + assert [(f.name, f.ctype) for f in final.fields] == [ + ("pn", "Array1D"), + ("tn1", "Array1D"), + ] + + +def test_from_pyccel_class_errors(): + with pytest.raises(ValueError, match="no class 'Nope'"): + CudaStruct.from_pyccel_class(PYCCEL_SOURCE, "Nope") + with pytest.raises(ValueError, match="has no __init__"): + CudaStruct.from_pyccel_class("class Empty:\n pass\n", "Empty") + with pytest.raises(ValueError, match="no type annotation"): + CudaStruct.from_pyccel_class( + "class A:\n def __init__(self, x):\n pass\n", "A" + ) + + +# --------------------------------------------------------------------------- +# PyccelStructArguments +# --------------------------------------------------------------------------- + + +class HostMarkers: + """Stands in for a pyccel-compiled argument class.""" + + instances = 0 + + def __init__(self, markers, Np): + type(self).instances += 1 + self.markers = markers + self.Np = Np + + +class MarkerArguments(PyccelStructArguments): + struct_name = "MarkerArgs" + fields = (("markers", "Array2D"), ("Np", "long long"), ("n_markers", "int")) + host_class = HostMarkers + host_fields = ("markers", "Np") + + def __init__(self, markers, Np): + self.markers = markers + self.Np = Np + self.n_markers = markers.shape[0] + + +def test_pyccel_struct_arguments_host_form_is_built_once_and_follows_changes(): + HostMarkers.instances = 0 + markers = np.zeros((5, 3)) + args = MarkerArguments(markers, 5) + host = args.__host_args__() + assert isinstance(host, HostMarkers) and host.markers is markers and host.Np == 5 + assert args.__host_args__() is host and HostMarkers.instances == 1 + args.Np = 6 # a changed scalar: rebuilt + assert args.__host_args__().Np == 6 and HostMarkers.instances == 2 + args.markers = np.zeros((7, 3)) # a replaced array: rebuilt + assert args.__host_args__().markers is args.markers and HostMarkers.instances == 3 + # the host form is the Kernel's host argument + seen = {} + + def push(m, dt): + seen["host"] = m + + Kernel(push)(args, 0.1) + assert seen["host"] is args.__host_args__() + + +def test_pyccel_struct_arguments_device_form_and_pickling(): + args = MarkerArguments(FakeDeviceArray(np.float64, shape=(5, 3)), 5) + (packed,) = args.__cuda_args__() + assert packed.dtype == MarkerArguments.struct.dtype + assert int(packed["n_markers"]) == 5 + with pytest.raises(RuntimeError, match="no host form on the CuPy backend"): + args.__host_args__() + host_copy = MarkerArguments(np.ones((2, 3)), 2) + host_copy.__host_args__() + restored = pickle.loads(pickle.dumps(host_copy)) + assert "_host_value" not in restored.__dict__ + assert restored.__host_args__().markers.shape == (2, 3) + + +def test_pyccel_struct_arguments_host_copies(monkeypatch): + class Copying(MarkerArguments): + host_copies = True + + device = FakeDeviceArray(np.float64, shape=(2, 3)) + monkeypatch.setattr("cunumpy.xp.to_numpy", lambda a: np.full((2, 3), 7.0)) + host = Copying(device, 2).__host_args__() + assert isinstance(host.markers, np.ndarray) and host.markers[0, 0] == 7.0 + + +def test_pyccel_struct_arguments_requires_host_class(): + class NoHost(PyccelStructArguments): + struct_name = "NoHost" + fields = (("n", "int"),) + + def __init__(self): + self.n = 1 + + with pytest.raises(TypeError, match="host_class is not set"): + NoHost().__host_args__() + + +# --------------------------------------------------------------------------- +# n_threads_from and check_finite +# --------------------------------------------------------------------------- + + +def _launching(kernel): + """Replace compilation by a recorder of (grid, block) launches.""" + launches = [] + kernel.compile = lambda: ( + lambda grid, block, values, shared_mem: launches.append(grid) + ) + return launches + + +def test_n_threads_from(): + kernel = CudaKernel(SCALE, "scale", n_threads_from=lambda args: args[2]) + launches = _launching(kernel) + kernel(FakeDeviceArray(np.float64), 2.0, 300) + assert launches == [(3,)] + kernel(FakeDeviceArray(np.float64), 2.0, 300, n_threads=10) # explicit wins + assert launches[-1] == (1,) + kernel.n_threads_from = None + with pytest.raises(TypeError, match="exactly one of n_threads and grid"): + kernel(FakeDeviceArray(np.float64), 2.0, 300) + with pytest.raises(TypeError, match="callable"): + kernel.n_threads_from = 5 + + +def test_check_finite(monkeypatch): + kernel = CudaKernel(SCALE, "scale", check_finite=True) + assert kernel.check_finite + _launching(kernel) + kernel._synchronize_after_launch = lambda *a: None + fake_cupy = ModuleType("cupy") + fake_cupy.isfinite = np.isfinite + monkeypatch.setitem(sys.modules, "cupy", fake_cupy) + + class Arr(FakeDeviceArray): + def __init__(self, values): + super().__init__(np.float64, shape=np.shape(values)) + self.values = np.asarray(values, dtype=np.float64) + + def __array__(self, dtype=None, copy=None): + return self.values + + kernel(Arr([1.0, 2.0]), 2.0, 2, n_threads=2) + with pytest.raises(RuntimeError, match="left a NaN or inf in argument 0"): + kernel(Arr([1.0, np.nan]), 2.0, 2, n_threads=2) + kernel.check_finite = False + kernel(Arr([1.0, np.nan]), 2.0, 2, n_threads=2) + + +def test_device_arrays_in_struct_arguments(): + args = MarkerArguments(FakeDeviceArray(np.float64, shape=(5, 3)), 5) + found = dict( + cuda_kernel_module._device_arrays_in((1.0, args, FakeDeviceArray("f8"))) + ) + assert set(found) == {"argument 1.markers", "argument 2"} + + +# --------------------------------------------------------------------------- +# dispatch: name length, pyccel stubs, test-arguments modules +# --------------------------------------------------------------------------- + + +def _forget_package(name): + """Drop a temporary package from sys.modules (each test gets a new tmp_path).""" + import importlib + + for module in [m for m in sys.modules if m == name or m.startswith(name + ".")]: + del sys.modules[module] + importlib.invalidate_caches() + + +@pytest.fixture +def helper_package(tmp_path, monkeypatch): + _forget_package("helper_kernel_pkg") + root = tmp_path / "helper_kernel_pkg" + for name in ("scale", "no_cuda"): + (root / name).mkdir(parents=True) + (root / name / "__init__.py").write_text("") + (root / name / f"{name}_kernels.py").write_text( + f"def {name}(x, factor, n):\n for i in range(n):\n x[i] *= factor\n" + ) + (root / "scale" / "scale_cuda.cu").write_text(SCALE) + (root / "scale" / "scale_test_args.py").write_text( + "import numpy as np\nimport cunumpy as xp\n\nN_THREADS = 300\nRTOL = 1e-10\n\n\n" + "def make_args(backend, seed):\n" + " x = xp.to_cunumpy(np.random.default_rng(seed).random(300))\n" + " return (x, 2.0, x.size)\n" + ) + (root / "__init__.py").write_text("") + monkeypatch.syspath_prepend(str(tmp_path)) + yield root + _forget_package("helper_kernel_pkg") + + +def test_catalog_records_test_args_modules(helper_package): + catalog = KernelCatalog.from_package("helper_kernel_pkg") + assert ( + catalog["scale"].test_args_module == "helper_kernel_pkg.scale.scale_test_args" + ) + assert catalog["no_cuda"].test_args_module is None + module = catalog["scale"].test_args + assert module.N_THREADS == 300 and catalog["scale"].test_args is module + off = KernelCatalog.from_package("helper_kernel_pkg", test_args_suffix=None) + assert off["scale"].test_args_module is None + + +def test_check_parity_reads_the_module(helper_package, monkeypatch): + catalog = KernelCatalog.from_package("helper_kernel_pkg") + calls = {} + + def fake_agree(kernel, make_args, **settings): + calls["kernel"], calls["settings"] = kernel, settings + return {"argument 0": np.zeros(1)} + + monkeypatch.setattr(cunumpy.testing, "assert_kernels_agree", fake_agree) + check_parity(catalog["scale"], atol=1e-14) + assert calls["kernel"] is catalog["scale"] + assert calls["settings"] == {"n_threads": 300, "rtol": 1e-10, "atol": 1e-14} + with pytest.raises(ValueError, match="no test-arguments module"): + check_parity(catalog["no_cuda"]) + + +def test_parity_cases_marks_kernels_without_test_args(helper_package): + root = helper_package / "with_cuda_no_args" + root.mkdir() + (root / "__init__.py").write_text("") + (root / "with_cuda_no_args_kernels.py").write_text( + "def with_cuda_no_args(x, factor, n):\n pass\n" + ) + (root / "with_cuda_no_args_cuda.cu").write_text( + SCALE.replace("scale", "with_cuda_no_args") + ) + catalog = KernelCatalog.from_package("helper_kernel_pkg") + cases = parity_cases(catalog) + assert [case.id for case in cases] == ["scale", "with_cuda_no_args"] + assert cases[0].marks == () + (mark,) = cases[1].marks + assert ( + mark.name == "skip" + and "with_cuda_no_args_test_args.py" in mark.kwargs["reason"] + ) + + +def test_long_kernel_names_warn(tmp_path, monkeypatch): + name = "k" * (FORTRAN_NAME_LIMIT - len("bind_c__kernels") + 1) + root = tmp_path / "long_name_pkg" + (root / name).mkdir(parents=True) + (root / name / "__init__.py").write_text("") + (root / name / f"{name}_kernels.py").write_text(f"def {name}(x):\n pass\n") + (root / "__init__.py").write_text("") + monkeypatch.syspath_prepend(str(tmp_path)) + _forget_package("long_name_pkg") + with pytest.warns(UserWarning, match="longer than Fortran's limit of 63"): + KernelCatalog.from_package("long_name_pkg") + with warnings.catch_warnings(): + warnings.simplefilter("error") + KernelCatalog.from_package("long_name_pkg", check_name_length=False) + + +def test_host_parameters_from_pyccel_stub(tmp_path): + so = tmp_path / "push_kernels.cpython-313-darwin.so" + (tmp_path / "__pyccel__").mkdir() + (tmp_path / "__pyccel__" / "push_kernels.pyi").write_text( + '#$ header metavar printer_imports="pyc_math_f90"\n' + "from pyccel.decorators import low_level\n\n" + "@low_level('push')\n" + "def push(dt : 'float', stage : 'int', markers : 'float64[:,:](order=C)') -> None:\n" + " ...\n" + ) + module = ModuleType("push_kernels") + module.__file__ = str(so) + compiled = SimpleNamespace(__self__=module, __name__="push") + assert _pyccel_stub_parameters(compiled) == ["dt", "stage", "markers"] + assert ( + _pyccel_stub_parameters(SimpleNamespace(__self__=module, __name__="x")) is None + ) + assert _pyccel_stub_parameters(len) is None # a builtin without a stub + + kernel = Kernel(compiled, CudaKernel(SCALE.replace("scale", "push"), "push")) + # inspect.signature fails on the fake compiled function; the stub is used + assert kernel.host_parameters() == ["dt", "stage", "markers"] + with pytest.raises(ValueError, match=r"host kernel takes \(dt, stage, markers\)"): + kernel.check_signature() + + +# --------------------------------------------------------------------------- +# device_function_kernel with struct parameters +# --------------------------------------------------------------------------- + + +def test_device_function_kernel_struct_parameters(): + domain = CudaStruct("DomainArgs", [("params", "double*"), ("kind", "int")]) + source = domain.declaration + ( + "__device__ double scale_x(const DomainArgs& d, double x) " + "{ return d.params[0] * x; }\n" + ) + kernel = device_function_kernel( + source, "double scale_x(const DomainArgs& d, double x)", structs=[domain] + ) + params = [(p.name, p.ctype, p.struct is not None) for p in kernel.signature] + assert params == [ + ("d", "DomainArgs", True), + ("x", "double", False), + ("out", "double", False), + ("n", "int", False), + ] + assert "void scale_x_kernel(DomainArgs d, const double* x, double* out, int n)" in ( + kernel.source + ) + assert "out[i] = scale_x(d, x[i]);" in kernel.source + with pytest.raises(ValueError, match="only structs can be passed by reference"): + device_function_kernel(source, "double f(const double& x)") + + +# --------------------------------------------------------------------------- +# segment_sum, mpi_buffer, require_version +# --------------------------------------------------------------------------- + + +def test_segment_sum(): + keys = np.array([0, 2, 0, -1, 2]) + values = np.array([1.0, 2.0, 3.0, 100.0, 4.0]) + np.testing.assert_array_equal(xp.segment_sum(values, keys, 4), [4.0, 0.0, 6.0, 0.0]) + columns = np.stack([values, -values], axis=1) + out = xp.segment_sum(columns, keys, 3) + np.testing.assert_array_equal(out, [[4.0, -4.0], [0.0, 0.0], [6.0, -6.0]]) + assert xp.segment_sum(np.array([1, 2]), np.array([1, 1]), 2).dtype == np.float64 + assert xp.segment_sum(values.astype(np.float32), keys, 3).dtype == np.float32 + complex_sum = xp.segment_sum(values * (1 + 1j), keys, 3) + np.testing.assert_allclose(complex_sum, [4 + 4j, 0, 6 + 6j]) + with pytest.raises(ValueError, match="smaller than n_segments"): + xp.segment_sum(values, keys, 2) + with pytest.raises(ValueError, match="one entry per value"): + xp.segment_sum(values, keys[:2], 3) + + +def test_mpi_buffer_on_host_arrays(): + x = np.arange(3.0) + with xp.mpi_buffer(x) as buf: + assert buf is x + with xp.mpi_buffer(x, send=False, recv=True) as buf: + assert buf is x + + +def test_mpi_cuda_aware_setting(): + xp.set_mpi_cuda_aware(None) + assert xp.get_mpi_cuda_aware() is None + xp.set_mpi_cuda_aware(True) + assert xp.get_mpi_cuda_aware() is True + xp.set_mpi_cuda_aware(None) + + +def test_require_version(monkeypatch): + xp.require_version("0.1") + xp.require_version(xp.__version__) + monkeypatch.setattr(xp, "__version__", "0.3.0") + with pytest.raises(ImportError, match="0.4.0 or newer is required, but 0.3.0"): + xp.require_version("0.4.0") + monkeypatch.setattr(xp, "__version__", "0.0.0+unknown") + xp.require_version("99.0") # unknown version: not checked + + +# --------------------------------------------------------------------------- +# the fake CuPy (in a subprocess: it replaces cupy for the whole process) +# --------------------------------------------------------------------------- + +FAKE_SCRIPT = r""" +import pickle +import numpy as np +import pytest +import cunumpy as xp +import cunumpy.testing as testing +from cunumpy import CudaKernel, CudaStruct, KernelArguments + +assert testing.fake_cupy_active() +assert xp.cupy_available() and xp.get_backend() == "cupy", xp.get_backend() +a = xp.zeros((4, 3)) +assert type(a).__module__ == "cupy" and not isinstance(a, np.ndarray) +assert xp.is_gpu(a) +with pytest.raises(TypeError): + np.asarray(a) +with pytest.raises(TypeError): + a + np.ones(3) +assert a.sum().shape == () +assert xp.to_numpy(a).shape == (4, 3) +b = pickle.loads(pickle.dumps(a)) +assert xp.is_gpu(b) + +# struct packing and the kernel argument checks work on fake device arrays +Vec = CudaStruct("Vec", [("x", "double*"), ("n", "int")]) +assert int(Vec(x=xp.zeros(5), n=5).packed["x"]) != 0 +scale = CudaKernel("extern \"C\" __global__ void scale(double* x, double f, int n) {}", "scale") +scale.prepare_args(xp.zeros(5), 2.0, 5) +with pytest.raises(TypeError): + scale.prepare_args(np.zeros(5), 2.0, 5) +with pytest.raises(NotImplementedError): + scale(xp.zeros(5), 2.0, 5, n_threads=5) + +# kernels cannot run: the GPU tests are skipped +assert testing.requires_cupy.args[0] is True +assert testing.FAKE_SKIP_REASON in testing.requires_cupy.kwargs["reason"] + +# as_device_array and the staging path of mpi_buffer +d = xp.as_device_array([1.0, 2.0], dtype=np.float64) +assert xp.is_gpu(d) +with pytest.raises(RuntimeError, match="not known whether MPI"): + with xp.mpi_buffer(d): + pass +with xp.count_transfers() as counter: + with xp.mpi_buffer(d, recv=True, cuda_aware=False) as buf: + assert isinstance(buf, np.ndarray) and buf.tolist() == [1.0, 2.0] + buf[:] = [5.0, 6.0] +assert xp.to_numpy(d).tolist() == [5.0, 6.0] +assert sorted(e.kind for e in counter.events) == ["to_device", "to_host"] +with xp.mpi_buffer(d, cuda_aware=True) as buf: + assert buf is d +print("fake cupy OK") +""" + + +def test_fake_cupy_in_subprocess(): + root = Path(__file__).resolve().parents[2] + env = dict(os.environ, CUNUMPY_FAKE_CUPY="1", ARRAY_BACKEND="cupy") + env["PYTHONPATH"] = os.pathsep.join( + p for p in (str(root / "src"), env.get("PYTHONPATH", "")) if p + ) + env.pop("CUNUMPY_CUDA_DEBUG", None) + result = subprocess.run( + [sys.executable, "-c", FAKE_SCRIPT], + env=env, + capture_output=True, + text=True, + check=False, + cwd=str(root), + ) + assert result.returncode == 0, result.stdout + result.stderr + assert "fake cupy OK" in result.stdout + + +def test_install_fake_cupy_refuses_a_real_cupy(monkeypatch): + monkeypatch.setitem(sys.modules, "cupy", ModuleType("cupy")) + with pytest.raises(RuntimeError, match="real CuPy is already imported"): + cunumpy.testing.install_fake_cupy() + assert not cunumpy.testing.fake_cupy_active() diff --git a/tests/unit/test_profiling.py b/tests/unit/test_profiling.py index cb4d82d..6d01102 100644 --- a/tests/unit/test_profiling.py +++ b/tests/unit/test_profiling.py @@ -64,11 +64,10 @@ def test_nvtx_range_repr(): def test_timed_region_on_numpy(): - with xp.use_backend("numpy"): - with xp.timed_region("sleep") as timing: - assert timing.name == "sleep" - assert timing.elapsed is None - time.sleep(0.02) + with xp.use_backend("numpy"), xp.timed_region("sleep") as timing: + assert timing.name == "sleep" + assert timing.elapsed is None + time.sleep(0.02) assert isinstance(timing, xp.Timing) assert timing.elapsed >= 0.02 @@ -77,18 +76,19 @@ def test_timed_region_on_numpy(): def test_timed_region_without_sync_on_numpy(): - with xp.use_backend("numpy"): - with xp.timed_region("no sync", sync=False) as timing: - pass + with xp.use_backend("numpy"), xp.timed_region("no sync", sync=False) as timing: + pass assert timing.elapsed >= 0.0 assert timing.synced is False def test_timed_region_records_time_on_exception(): - with xp.use_backend("numpy"): - with pytest.raises(RuntimeError, match="boom"): - with xp.timed_region("failing") as timing: - raise RuntimeError("boom") + with ( + xp.use_backend("numpy"), + pytest.raises(RuntimeError, match="boom"), + xp.timed_region("failing") as timing, + ): + raise RuntimeError("boom") assert timing.elapsed is not None assert timing.elapsed >= 0.0 @@ -126,9 +126,8 @@ def test_nvtx_range_nested_and_reentrant(fake_nvtx): def test_nvtx_range_pops_on_exception(fake_nvtx): - with pytest.raises(ValueError): - with xp.nvtx_range("failing"): - raise ValueError + with pytest.raises(ValueError), xp.nvtx_range("failing"): + raise ValueError assert fake_nvtx == [("push", "failing", -1), ("pop",)] diff --git a/tests/unit/test_random_streams.py b/tests/unit/test_random_streams.py new file mode 100644 index 0000000..396d9bb --- /dev/null +++ b/tests/unit/test_random_streams.py @@ -0,0 +1,115 @@ +"""Tests for `xp.random_streams`: one seeded generator per process and backend.""" + +import numpy as np +import pytest + +import cunumpy as xp +from cunumpy import RandomStreams, random_streams +from cunumpy.random_streams import BIT_GENERATORS + + +@pytest.fixture +def streams(): + return RandomStreams() + + +def _draws(streams): + return ( + streams.random((4, 2)), + streams.normal(1.0, 2.0, 5), + streams.uniform(-1.0, 1.0, 3), + streams.standard_normal(2), + ) + + +def test_exported(): + assert isinstance(random_streams, RandomStreams) + assert "random_streams" in xp.__all__ + assert repr(RandomStreams()) == "RandomStreams(not seeded)" + + +def test_same_seed_and_rank_give_the_same_draws(streams): + streams.seed(42, rank=0) + first = _draws(streams) + streams.seed(42, rank=0) + for a, b in zip(first, _draws(streams)): + np.testing.assert_array_equal(a, b) + assert "spawn_key=(0,)" in repr(streams) + + +def test_ranks_and_seeds_give_different_streams(streams): + streams.seed(42, rank=0) + rank0 = streams.random(8) + streams.seed(42, rank=1) + rank1 = streams.random(8) + streams.seed(43, rank=0) + other = streams.random(8) + assert not np.array_equal(rank0, rank1) and not np.array_equal(rank0, other) + + +def test_seed_also_seeds_numpy_global_state(streams): + streams.seed(7) + a = np.random.random(3) + streams.seed(7) + np.testing.assert_array_equal(np.random.random(3), a) + + +@pytest.mark.parametrize("name", BIT_GENERATORS) +def test_bit_generators(streams, name): + streams.seed(42, bit_generator=name) + assert type(streams.generator("numpy").bit_generator).__name__ == name + first = streams.random(4) + streams.seed(42, bit_generator=name) + np.testing.assert_array_equal(streams.random(4), first) + + +def test_unknown_bit_generator_and_backend(streams): + with pytest.raises(ValueError, match="Unknown bit generator"): + streams.seed(42, bit_generator="Mersenne") + with pytest.raises(ValueError, match="backend must be"): + streams.generator("jax") + + +def test_components_share_the_process_generator_unless_seeded(streams): + streams.seed(1) + assert streams.make_generator() is streams.generator() + own = streams.make_generator(123) + np.testing.assert_array_equal(own.random(3), np.random.default_rng(123).random(3)) + assert isinstance(streams.make_generator(backend="numpy"), np.random.Generator) + + +def test_unseeded_streams_still_work(streams): + assert streams.random(3).shape == (3,) + + +class _MinimalGenerator: + """Only ``random`` and ``standard_normal``, like some CuPy versions.""" + + def __init__(self): + self._rng = np.random.default_rng(0) + + def random(self, size=None): + return self._rng.random(size) + + def standard_normal(self, size=None): + return self._rng.standard_normal(size) + + +def test_normal_and_uniform_fall_back(streams): + rng = _MinimalGenerator() + values = streams.uniform(2.0, 3.0, 1000, rng=rng) + assert values.min() >= 2.0 and values.max() < 3.0 + assert abs(streams.normal(5.0, 0.1, 1000, rng=rng).mean() - 5.0) < 0.05 + + +def test_cupy_generator_on_gpu(streams): + if not xp.cupy_available(): + pytest.skip("CuPy not installed or not functional") + import cupy as cp + + with xp.use_backend("cupy"): + streams.seed(42, rank=3) + first = streams.normal(0.0, 1.0, 100) + assert isinstance(first, cp.ndarray) + streams.seed(42, rank=3) + cp.testing.assert_array_equal(streams.normal(0.0, 1.0, 100), first) diff --git a/tests/unit/test_scipy_backend.py b/tests/unit/test_scipy_backend.py new file mode 100644 index 0000000..4d3fe29 --- /dev/null +++ b/tests/unit/test_scipy_backend.py @@ -0,0 +1,104 @@ +"""Tests for `xp.scipy`: SciPy or cupyx.scipy, by the active backend.""" + +import sys +import types + +import numpy as np +import pytest + +import cunumpy as xp +from cunumpy import scipy_backend +from cunumpy.scipy_backend import SUBMODULES, ScipyNamespace + + +@pytest.fixture +def fake_cupyx(monkeypatch): + """A fake `cupyx.scipy` with `special.erf` and `sparse.linalg.cg`, as backend.""" + modules = {} + for name in ("cupyx", "cupyx.scipy", "cupyx.scipy.special", "cupyx.scipy.sparse"): + modules[name] = types.ModuleType(name) + modules["cupyx.scipy.sparse.linalg"] = types.ModuleType("cupyx.scipy.sparse.linalg") + modules["cupyx.scipy.special"].erf = "device erf" + modules["cupyx.scipy.sparse.linalg"].cg = "device cg" + for name, module in modules.items(): + monkeypatch.setitem(sys.modules, name, module) + monkeypatch.setattr(scipy_backend, "get_backend", lambda: "cupy") + return modules + + +def test_scipy_is_exported(): + assert isinstance(xp.scipy, ScipyNamespace) + assert "scipy" in xp.__all__ + assert set(dir(xp.scipy)) >= {"sparse", "special", "fft", "available", "resolve"} + assert set(dir(xp.scipy.sparse)) >= {"linalg", "csgraph"} + + +def test_unknown_subpackage(): + with pytest.raises(ValueError, match="not forwarded"): + ScipyNamespace("optimize") + assert all(ScipyNamespace(name) for name in SUBMODULES) + + +def test_numpy_backend_forwards_to_scipy(): + pytest.importorskip("scipy") + import scipy.sparse.linalg + import scipy.special + + assert xp.scipy.special.erf is scipy.special.erf + assert xp.scipy.sparse.linalg.cg is scipy.sparse.linalg.cg + assert xp.scipy.sparse.linalg.resolve() is scipy.sparse.linalg + assert xp.scipy.sparse.linalg is xp.scipy.sparse.linalg # cached namespaces + + n = 20 + A = xp.scipy.sparse.diags( + [np.full(n - 1, -1.0), np.full(n, 2.0), np.full(n - 1, -1.0)], + [-1, 0, 1], + format="csr", + ) + x, info = xp.scipy.sparse.linalg.cg(A, np.ones(n), rtol=1e-12) + assert info == 0 and np.allclose(A @ x, 1.0) + + +def test_missing_names_say_which_backend_lacks_them(): + pytest.importorskip("scipy") + with pytest.raises(AttributeError, match=r"not available on the numpy backend"): + xp.scipy.special.no_such_function # noqa: B018 + assert xp.scipy.special.available("erf") + assert not xp.scipy.special.available("no_such_function") + assert xp.scipy.available("sparse") + with pytest.raises(AttributeError): + xp.scipy.__wrapped__ # noqa: B018 -- dunder names are not forwarded + + +def test_cupy_backend_forwards_to_cupyx(fake_cupyx): + assert xp.scipy.special.erf == "device erf" + assert xp.scipy.sparse.linalg.cg == "device cg" + assert xp.scipy.special.resolve() is fake_cupyx["cupyx.scipy.special"] + with pytest.raises( + AttributeError, match=r"cupy backend \(it may exist in scipy\.special\)" + ): + xp.scipy.special.erfcx # noqa: B018 + # a subpackage cupyx does not provide + assert not xp.scipy.available("ndimage") + with pytest.raises(ImportError, match="needs cupyx.scipy.ndimage"): + xp.scipy.ndimage.resolve() + + +def test_backend_switch_takes_effect_immediately(fake_cupyx, monkeypatch): + pytest.importorskip("scipy") + import scipy.special + + special = xp.scipy.special + assert special.erf == "device erf" + monkeypatch.setattr(scipy_backend, "get_backend", lambda: "numpy") + assert special.erf is scipy.special.erf + + +def test_missing_scipy(monkeypatch): + monkeypatch.setitem(sys.modules, "scipy", None) # import scipy raises ImportError + monkeypatch.setitem(sys.modules, "scipy.special", None) + with pytest.raises( + ImportError, match=r"needs scipy.special: SciPy is not installed" + ): + xp.scipy.special.erf # noqa: B018 + assert not xp.scipy.special.available("erf") diff --git a/tests/unit/test_staging.py b/tests/unit/test_staging.py new file mode 100644 index 0000000..fe982ad --- /dev/null +++ b/tests/unit/test_staging.py @@ -0,0 +1,155 @@ +"""Tests for `xp.HostStaging`: background copies of device arrays to the host.""" + +import types + +import numpy as np +import pytest + +import cunumpy as xp +from cunumpy import HostStaging +from cunumpy import staging as staging_module + + +def test_host_arrays_are_copied_at_once(): + staging = HostStaging((3,), np.float64) + a = np.arange(3.0) + copy = staging.copy(a) + a[:] = -1.0 # the caller may overwrite its array right away + assert copy.ready() + assert copy.result().tolist() == [0.0, 1.0, 2.0] + assert repr(staging) == "HostStaging(shape=(3,), dtype=float64, buffers=2)" + + +def test_shape_dtype_and_buffers_are_checked(): + staging = HostStaging(4, np.float32) + with pytest.raises(ValueError, match="got an array of shape"): + staging.copy(np.zeros(3, np.float32)) + with pytest.raises(ValueError, match="dtype float64"): + staging.copy(np.zeros(4)) + with pytest.raises(ValueError, match="at least 1"): + HostStaging(4, np.float32, buffers=0) + + +def test_stale_results_raise(): + staging = HostStaging((2,), np.float64, buffers=2) + first = staging.copy(np.array([1.0, 1.0])) + second = staging.copy(np.array([2.0, 2.0])) + assert first.result().tolist() == [1.0, 1.0] + third = staging.copy(np.array([3.0, 3.0])) # reuses the first buffer + with pytest.raises(RuntimeError, match="was reused"): + first.result() + with pytest.raises(RuntimeError, match="was reused"): + first.ready() + assert second.result().tolist() == [2.0, 2.0] + assert third.result().tolist() == [3.0, 3.0] + + +class DeviceArray(np.ndarray): + """Stands for a CuPy array; `get(stream, out)` is the asynchronous copy.""" + + log = None + + def get(self, stream=None, out=None, blocking=True): + assert blocking is False # an asynchronous copy + DeviceArray.log.append(("get", stream.name)) + out[...] = np.asarray(self) + return out + + +class FakeEvent: + def __init__(self, name): + self.name = name + self.done = False + + def synchronize(self): + DeviceArray.log.append(("synchronize", self.name)) + self.done = True + + +class FakeStream: + def __init__(self, name): + self.name = name + self.count = 0 + + def record(self): + self.count += 1 + event = FakeEvent(f"{self.name}#{self.count}") + DeviceArray.log.append(("record", event.name)) + return event + + def wait_event(self, event): + DeviceArray.log.append(("wait", self.name, event.name)) + + +@pytest.fixture +def fake_device(monkeypatch): + DeviceArray.log = [] + current = FakeStream("compute") + cuda = types.SimpleNamespace( + Stream=lambda non_blocking=False: FakeStream("staging"), + get_current_stream=lambda: current, + ) + + def empty(shape, dtype): + DeviceArray.log.append(("snapshot alloc",)) + return np.empty(shape, dtype).view(DeviceArray) + + fake = types.SimpleNamespace(cuda=cuda, empty=empty) + monkeypatch.setattr(staging_module, "_cupy", lambda: fake) + monkeypatch.setattr( + staging_module, "_empty_pinned", lambda s, dtype: np.empty(s, dtype) + ) + monkeypatch.setattr( + staging_module, "_is_device_array", lambda a: isinstance(a, DeviceArray) + ) + return DeviceArray.log + + +def test_device_copies_snapshot_then_copy_on_their_own_stream(fake_device): + log = fake_device + staging = HostStaging((3,), np.float64, buffers=2) + rho = np.array([1.0, 2.0, 3.0]).view(DeviceArray) + copy = staging.copy(rho) + rho[...] = 0.0 # the next step overwrites the array: the snapshot keeps the data + assert log == [ + ("snapshot alloc",), + ("record", "compute#1"), # after the producer's kernels and the snapshot + ("wait", "staging", "compute#1"), + ("get", "staging"), # the copy to the host, on the staging stream + ("record", "staging#1"), + ] + assert not copy.ready() + assert copy.result().tolist() == [1.0, 2.0, 3.0] + assert ("synchronize", "staging#1") in log + + +def test_a_buffer_is_reused_only_after_its_copy_finished(fake_device): + log = fake_device + staging = HostStaging((1,), np.float64, buffers=2) + a = np.array([5.0]).view(DeviceArray) + staging.copy(a) + staging.copy(a) + assert not any(entry[0] == "synchronize" for entry in log) # two in flight + staging.copy(a) # the first buffer again: waits for its first copy + assert ("synchronize", "staging#1") in log + staging.synchronize() + assert ("synchronize", "staging#2") in log + + +def test_copies_are_counted_as_transfers(fake_device): + staging = HostStaging((2,), np.float64) + with xp.count_transfers() as counter: + staging.copy(np.zeros(2).view(DeviceArray)) + assert counter.total == 1 + + +def test_on_gpu(): + if not xp.cupy_available(): + pytest.skip("CuPy not installed or not functional") + import cupy as cp + + staging = HostStaging((1000,), np.float64) + data = cp.arange(1000.0) + copy = staging.copy(data) + data *= 0.0 # overwritten right away on the compute stream + np.testing.assert_array_equal(copy.result(), np.arange(1000.0)) diff --git a/tests/unit/test_testing.py b/tests/unit/test_testing.py index 14f2cf0..b3a496a 100644 --- a/tests/unit/test_testing.py +++ b/tests/unit/test_testing.py @@ -74,7 +74,7 @@ def test_backends_and_marker(): assert requires_cupy.args == (not xp.cupy_available(),) assert requires_cupy.kwargs["reason"] == "CuPy/GPU not available" with pytest.raises(AttributeError): - cunumpy.testing.no_such_thing + _ = cunumpy.testing.no_such_thing @pytest.mark.parametrize("backend_name", BACKENDS) @@ -304,3 +304,53 @@ def test_device_function_kernel_on_gpu(): spans = cp.empty(3, dtype=cp.int32) find_span(t, p, eta, spans, 3, n_threads=3) assert spans.get().tolist() == [2, 3, 3] + + +def test_struct_arguments_are_compared_by_field_name(): + """A CudaStructArguments object (fields may be properties) gets the host names.""" + from cunumpy.testing import _collect_arrays + + class Owner: + def __init__(self): + self.markers = np.zeros((3, 4)) + self.weights = np.ones(3) + + class HostArguments: # e.g. a Pyccel class + def __init__(self, owner): + self.markers = owner.markers + self.weights = owner.weights + self.n = 3 + + class DeviceArguments(xp.CudaStructArguments): + struct_name = "OwnerArgs" + fields = (("markers", "Array2D"), ("weights", "double*"), ("n", "int")) + + def __init__(self, owner): + self._owner = owner # not packed: no device arrays in this test + + @property + def markers(self): + return self._owner.markers + + @property + def weights(self): + return self._owner.weights + + n = 3 + + owner = Owner() + host = _collect_arrays((1.0, HostArguments(owner))) + device = _collect_arrays((1.0, DeviceArguments(owner))) + assert ( + sorted(host) == sorted(device) == ["argument 1.markers", "argument 1.weights"] + ) + assert device["argument 1.markers"] is owner.markers + + struct = DeviceArguments.struct + value = xp.CudaStructValue( + struct, np.zeros((), struct.dtype)[()], vars(owner) | {"n": 3} + ) + assert sorted(_collect_arrays((value,))) == [ + "argument 0.markers", + "argument 0.weights", + ] diff --git a/tests/unit/test_transfers.py b/tests/unit/test_transfers.py index 3c2d317..00cb46c 100644 --- a/tests/unit/test_transfers.py +++ b/tests/unit/test_transfers.py @@ -135,10 +135,9 @@ def test_to_cupy_of_device_array_is_not_counted(fake_device): def test_to_cunumpy_counts_the_direction_it_delegates_to(fake_device, monkeypatch): device = fake_device(np.zeros(2)) - with xp.count_transfers() as counter: - with xp.use_backend("numpy"): - xp.to_cunumpy(device) # device -> host - xp.to_cunumpy(np.zeros(2)) # already on the host + with xp.count_transfers() as counter, xp.use_backend("numpy"): + xp.to_cunumpy(device) # device -> host + xp.to_cunumpy(np.zeros(2)) # already on the host assert counter.to_host == 1 and counter.to_device == 0 @@ -168,9 +167,8 @@ def test_where_points_at_the_caller_outside_cunumpy(fake_device): def test_where_skips_frames_inside_cunumpy(fake_device): """`to_cunumpy` calls `to_numpy`; the call site is still the test.""" - with xp.count_transfers() as counter: - with xp.use_backend("numpy"): - xp.to_cunumpy(fake_device(np.zeros(1))) + with xp.count_transfers() as counter, xp.use_backend("numpy"): + xp.to_cunumpy(fake_device(np.zeros(1))) (event,) = counter.events assert event.where.startswith(THIS_FILE + ":") @@ -224,9 +222,8 @@ def test_nested_counters_each_see_their_own_block(fake_device): def test_counter_is_removed_when_the_block_raises(fake_device): - with pytest.raises(RuntimeError): - with xp.count_transfers(): - raise RuntimeError + with pytest.raises(RuntimeError), xp.count_transfers(): + raise RuntimeError assert transfers_module._ACTIVE == [] @@ -239,9 +236,8 @@ def test_assert_no_transfers_passes_without_transfers(): def test_assert_no_transfers_raises_with_report(fake_device): - with pytest.raises(AssertionError) as info: - with xp.assert_no_transfers(): - xp.to_numpy(fake_device(np.zeros(3))) + with pytest.raises(AssertionError) as info, xp.assert_no_transfers(): + xp.to_numpy(fake_device(np.zeros(3))) message = str(info.value) assert "1 transfer(s) through cunumpy" in message @@ -251,10 +247,9 @@ def test_assert_no_transfers_raises_with_report(fake_device): def test_assert_no_transfers_lets_exceptions_through(fake_device): - with pytest.raises(ValueError, match="inside"): - with xp.assert_no_transfers(): - xp.to_numpy(fake_device(np.zeros(3))) - raise ValueError("inside") + with pytest.raises(ValueError, match="inside"), xp.assert_no_transfers(): + xp.to_numpy(fake_device(np.zeros(3))) + raise ValueError("inside") # --------------------------------------------------------------------------- @@ -315,9 +310,9 @@ def shift(x, n): x = np.zeros(2) with xp.count_transfers() as counter: + line = _current_line() + 2 with pytest.warns(RuntimeWarning, match="copies its arrays"): kernel(x, 2) - line = _current_line() - 1 kernel(x, 2) assert np.all(x == 2.0) @@ -394,9 +389,12 @@ def scale(x, factor, n): kernel = Kernel(scale, missing_cuda="fallback") x = cp.ones(3) - with xp.count_transfers() as counter, xp.use_backend("cupy"): - with pytest.warns(RuntimeWarning): - kernel(x, 2.0, 3) + with ( + xp.count_transfers() as counter, + xp.use_backend("cupy"), + pytest.warns(RuntimeWarning), + ): + kernel(x, 2.0, 3) assert cp.all(x == 2.0) assert counter.fallbacks == 1