Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 6 additions & 17 deletions .github/workflows/static_analysis.yml
Original file line number Diff line number Diff line change
Expand Up @@ -28,27 +28,16 @@ jobs:
./cloc --version
./cloc $(git ls-files)

black:
ruff-format:
runs-on: ubuntu-latest
steps:
- name: Checkout the code
uses: actions/checkout@v4

- name: Code formatting with black
- name: Code formatting with ruff
run: |
pip install black "black[jupyter]"
black --check src/

isort:
runs-on: ubuntu-latest
steps:
- name: Checkout the code
uses: actions/checkout@v4

- name: Code formatting with isort
run: |
pip install isort
isort --check src/
pip install ruff
ruff format --check src/ tests/

mypy:
runs-on: ubuntu-latest
Expand Down Expand Up @@ -80,10 +69,10 @@ jobs:
- name: Checkout the code
uses: actions/checkout@v4

- name: Linting with ruff
- name: Linting with ruff (including import sorting)
run: |
pip install ruff
ruff check src/
ruff check src/ tests/

pylint:
runs-on: ubuntu-latest
Expand Down
13 changes: 10 additions & 3 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,13 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- The helpers moved from the top level of `cunumpy` to submodules, so that the top level is the NumPy/CuPy namespace plus backend selection and array conversion, and no helper hides a NumPy or CuPy name (`xp.fuse` hid `cupy.fuse`): `cunumpy.cuda` (CUDA only: `CudaKernel`, `CudaKernelVariants`, `CudaStruct*`, `CudaArguments`, `CudaParameter`, header tools, debug mode, device selection and memory, `stream`, `pin_memory`), `cunumpy.kernels` (`Kernel`, `KernelCatalog`, `PyccelKernel`, `KernelArguments`, `PyccelStructArguments`, host implementations, `as_kernel_array`, `kernel_output`, `fuse`), `cunumpy.rng` (`random_streams`, `RandomStreams`, `get_rng`, `philox_*`), `cunumpy.algorithms` (`morton_*`, `sort_by_key`, `segment_sum`), `cunumpy.mpi` (`mpi_buffer`, CUDA-aware MPI, `local_rank`, `synchronize_for_mpi`), `cunumpy.profiling` (`timed_region`, `Timing`, `nvtx_range`, transfer counting), `cunumpy.memory` (`HostStaging`, `StagedCopy`, `DeviceMirror`) and `cunumpy.petsc` (`petsc_vec`). All are imported by `import cunumpy`. The old top-level names still work and raise a `DeprecationWarning` naming the new place; they will be removed in 0.6.
- `cunumpy.testing` is now `cunumpy.kernel_testing`. Once imported, `cunumpy.testing` replaced NumPy's `xp.testing`, so `xp.testing.assert_allclose` failed in every test that ran after an `import cunumpy.testing`. `cunumpy.testing` still works, with a `DeprecationWarning`, until 0.6.
- Importing `cunumpy.cuda` makes `xp.cuda` the cunumpy submodule instead of CuPy's `cupy.cuda`; use `import cupy` for the latter.
- The implementation modules are private (`cunumpy._cuda_kernel`, `_dispatch`, `_kernel`, `_fusion`, `_philox`, `_morton`, `_random_streams`, `_staging`, `_mirror`, `_transfers`, `_emulation`, `_scipy_backend`, and the new `_device`, `_mpi`, `_profiling`, `_algorithms` split off `cunumpy.xp`), so every public name has one import path, through the submodules above. `cunumpy.cuda_kernel`, `cunumpy.dispatch` and `cunumpy.kernel` (in 0.3.0) still import, with a `DeprecationWarning`, until 0.6. `cunumpy.xp` keeps the backend selection and array conversion only.

### Performance
- `xp.<name>` for a NumPy/CuPy function (`xp.zeros`, `xp.add`, ...) is a plain attribute lookup, as fast as `numpy.<name>` (before: about 3 us per access through the module `__getattr__`, 20 times `numpy.add`). The public names of the active backend module are copied into the `cunumpy` namespace and replaced when `set_backend`/`use_backend` change the backend (about 30 us per switch between NumPy and CuPy; nothing when the module stays the same). `dir(xp)` lists them, so IPython and Jupyter complete NumPy names. `ArrayBackend.add_listener(callback)` is called with the new module on every switch.

### Development
- Formatting is checked with `ruff format` and import order with ruff's isort rules (`ruff check`), on `src/` and `tests/`, instead of black and isort; the `dev` extra installs ruff.

### Fixed
- `CudaKernel`'s header hash (`-DCUNUMPY_INCLUDE_HASH`) now covers the headers shipped with cunumpy (`cunumpy/atomic.cuh`, `reduce.cuh`, ...), also when included in angle brackets. Before, an upgrade of cunumpy that changed one of them left CuPy's kernel cache serving the kernel compiled with the old header. `resolve_includes(..., angle_dirs=...)` tracks angle-bracket includes found in the given directories.
Expand Down Expand Up @@ -52,7 +59,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- `CudaKernel(..., check_finite=True)` (also a settable property): after every launch (synchronized) the floating-point and complex arrays among the arguments, including the array fields of struct argument objects, are scanned, and a NaN or inf raises `RuntimeError` naming the kernel and the argument. For debugging a kernel that produces NaN; costs a synchronization and a pass over the arrays per launch.
- `cunumpy.testing.device_function_kernel` accepts struct parameters, by value (`DomainArgs d`) or by const reference (`const DomainArgs& d`), for the structs given in `structs=`; they are passed through to every thread unchanged, so device helpers that take argument structs can be tested from Python.
- Test arguments next to the kernel: `KernelCatalog.from_package(..., test_args_suffix="_test_args")` records `<name>/<name>_test_args.py` as `Kernel.test_args_module` (imported on first access as `Kernel.test_args`; `Kernel(..., test_args=...)` sets it by hand). The module defines `make_args(backend, seed)` and `N_THREADS` (an integer, a tuple, or a function of the argument tuple) or `GRID`, optionally `BLOCK`, `RTOL`, `ATOL`, `N_CALLS`, `OUTPUTS`, `SEED` (`cunumpy.testing.TEST_ARGS_SETTINGS`). `cunumpy.testing.parity_cases(catalog)` gives one `pytest.param` per kernel with a CUDA version, kernels without a test-arguments module marked `skip` with a reason naming the missing file, and `check_parity(kernel, **overrides)` runs `assert_kernels_agree` with the module's settings. A ported package's parity test is then one parametrized test, and a developer who adds a kernel adds a file, not test code.
- `KernelCatalog.from_package(..., check_name_length=True)` warns about a kernel whose module name does not fit pyccel's Fortran wrapper module `bind_c_<name>_kernels` into Fortran's 63-character limit (`cunumpy.dispatch.FORTRAN_NAME_LIMIT`): with the default suffix a kernel name has at most 48 characters.
- `KernelCatalog.from_package(..., check_name_length=True)` warns about a kernel whose module name does not fit pyccel's Fortran wrapper module `bind_c_<name>_kernels` into Fortran's 63-character limit (`FORTRAN_NAME_LIMIT`): with the default suffix a kernel name has at most 48 characters.
- `Kernel.host_parameters()` reads the parameter names of a pyccel-compiled host kernel from the `__pyccel__/<module>.pyi` stub pyccel writes next to the extension module, so `Kernel.check_signature()` and `KernelCatalog.check_signatures()` also check compiled kernels.
- The fake CuPy (`cunumpy._fake_cupy`): a strict host stand-in for CuPy for CI machines without a GPU, installed with the environment variable `CUNUMPY_FAKE_CUPY=1` (read when cunumpy is imported) or `cunumpy.testing.install_fake_cupy()`. Its arrays live in host memory but are not NumPy arrays (`numpy.asarray` raises, as for real CuPy arrays), reject host arrays and lists as CuPy does, have `data.ptr`, `device` and `__cuda_array_interface__`, so argument objects, struct packing, `as_device_array`, `count_transfers` and backend branches run on the CuPy code path on the CPU; CUDA kernels cannot run (`NotImplementedError`). `cunumpy.testing.fake_cupy_active()` tells; `requires_cupy` and `assert_kernels_agree` skip while it is active.
- `xp.mpi_buffer(array, *, send=True, recv=False, cuda_aware=None)`: context manager yielding the buffer to hand to MPI: a host array unchanged; a device array unchanged (after `synchronize_for_mpi`) when MPI is CUDA-aware; otherwise a pinned host staging copy, copied from the device before the block (`send`) and back after it (`recv`), counted by `count_transfers()`. `cuda_aware=None` uses the answer recorded by `mpi_is_cuda_aware()` (which now remembers its result) or `xp.set_mpi_cuda_aware()`; `xp.get_mpi_cuda_aware()` reads it. One MPI call site for both backends and both kinds of MPI builds.
Expand All @@ -61,10 +68,10 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- `Kernel(..., dispatch="arrays")` and `KernelCatalog.from_package(..., dispatch="arrays")`: choose the CUDA kernel when an argument lives on the GPU (a CuPy array or a device-only argument object) and the host kernel otherwise, whatever the backend, for codes that hand host arrays to kernels while CuPy is active. The default, `dispatch="backend"`, is unchanged.
- `xp.CompiledHostKernel`: a host kernel compiled on its first call by a compile function the caller provides (cunumpy does not compile anything itself, e.g. a wrapper around `pyccel.epyccel`), falling back to a given callable (or, with a warning, to the uncompiled Python function) when compilation fails. `KernelCatalog.from_package(..., compile_host=..., host_fallback=...)` uses it for every host kernel.
- `Kernel.check_signature()`, `Kernel.host_parameters()` and `KernelCatalog.check_signatures()`: check that a host kernel and its CUDA kernel take the same parameters in the same order, listing every kernel that differs.
- `cunumpy.testing.emulate_cuda_kernel(kernel, *args, n_threads=...)` (in `cunumpy.emulation`) and `emulation_compiler()`: run a CUDA kernel on the CPU, one thread after another, on NumPy arrays (any strides, written back), with the kernel compiled as C++ and the CUDA built-ins stubbed, so CPU-only CI can check the arithmetic of ported kernels. Kernels using shared memory, `__syncthreads` or warp intrinsics are refused.
- `cunumpy.kernel_testing.emulate_cuda_kernel(kernel, *args, n_threads=...)` and `emulation_compiler()`: run a CUDA kernel on the CPU, one thread after another, on NumPy arrays (any strides, written back), with the kernel compiled as C++ and the CUDA built-ins stubbed, so CPU-only CI can check the arithmetic of ported kernels. Kernels using shared memory, `__syncthreads` or warp intrinsics are refused.
- `Array4D<T>` in `cunumpy/array_view.cuh` and as a kernel parameter, struct field and `from_signature` annotation (`"float[:, :, :, :]"`), e.g. for a 3D grid of vector components.
- `xp.max_shared_memory_per_block(device=None, *, opt_in=False)` and `xp.DEFAULT_SHARED_MEMORY_PER_BLOCK`: the shared memory a block may use on a device (48 KiB without a GPU).
- `xp.random_streams` (`cunumpy.random_streams.RandomStreams`): one seeded generator per process and backend, with the stream `(seed, rank)` for MPI runs, a choice of NumPy bit generator, per-component generators and draw functions that work with NumPy and CuPy generators.
- `xp.rng.random_streams` (`cunumpy.rng.RandomStreams`): one seeded generator per process and backend, with the stream `(seed, rank)` for MPI runs, a choice of NumPy bit generator, per-component generators and draw functions that work with NumPy and CuPy generators.
- Documentation restructured into getting started, user guide (backends, backend-agnostic code, data movement, devices, MPI, profiling), kernel porting guides (`PyccelKernel`, `CudaKernel`, `Kernel`/`KernelCatalog`, argument objects, accumulation, debugging, testing), worked examples, best practices and troubleshooting pages.
- `cunumpy/LLM_GUIDE.md`: a self-contained guide to the API and its rules for AI coding assistants, shipped as package data and rendered in the documentation.
- `xp.CudaKernel`: Wraps a CUDA C kernel (`cupy.RawKernel`, compiled lazily with NVRTC) so it can be called with the same arguments as the host kernel it mirrors, plus `n_threads`. The `extern "C" __global__` signature is parsed once and every call is checked against it: argument count, array dtypes (host arrays raise, they are never copied), and scalars (Python scalars are cast to the declared C types with range checks; lossy or mismatching scalars raise instead of reaching the kernel as silently wrong values). Supports `block_size`, NVRTC `options`, `include_dirs`, `shared_mem`, `stream`, `CudaKernel.from_file()` (`<name>_cuda.cu`), `compile()` and `prepare_args()`; `check_signature=False` skips the checks.
Expand Down
6 changes: 4 additions & 2 deletions docs/source/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,8 +14,10 @@ total = xp.sum(values)
```

At runtime, NumPy-like attributes such as `array`, `sum`, `fft`, and `linalg`
are forwarded to the currently selected `array-api-compat` NumPy or CuPy
module. CuNumpy does not wrap every operation individually. The available
are those of the currently selected `array-api-compat` NumPy or CuPy module.
They are copied into the `cunumpy` namespace and replaced when the backend
changes, so `xp.sum` costs the same as `numpy.sum` (a switch between NumPy and
CuPy takes some tens of microseconds; avoid switching inside a hot loop). CuNumpy does not wrap every operation individually. The available
operations and some details can therefore vary with the installed NumPy and
CuPy versions. In normal use, access those operations through the top-level
`cunumpy` namespace, commonly imported as `xp`.
Expand Down
9 changes: 4 additions & 5 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -27,8 +27,7 @@ dependencies = [
]

optional-dependencies.dev = [
"black[jupyter]",
"isort",
"ruff",
"cunumpy[test-compiled,docs]",
]
# https://medium.com/@pratikdomadiya123/build-project-documentation-quickly-with-the-sphinx-python-2a9732b66594
Expand All @@ -53,9 +52,9 @@ where = [ "src" ]
[tool.setuptools.package-data]
cunumpy = [ "py.typed", "*.pyi", "LLM_GUIDE.md", "cuda/include/cunumpy/*.cuh" ]

[tool.isort]
profile = "black"

[tool.ruff]
# agent worktrees and scratch scripts are local copies, not part of the project
extend-exclude = [ ".claude" ]
# formatting: `ruff format`; import sorting: the isort rules of `ruff check`
lint.extend-select = [ "I" ]
lint.isort.known-first-party = [ "cunumpy" ]
Loading
Loading