Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,6 +18,8 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
- `KernelCatalog.from_package(..., include_dirs=None)` is now an explicit keyword; by default the source root of the top-level package (the directory containing it) is an include directory of every CUDA kernel, in addition to the kernel's own folder, so kernels can `#include "my_pkg/common.cuh"`.

### Added
- Documentation restructured into getting started, user guide (backends, backend-agnostic code, data movement, devices, MPI, profiling), kernel porting guides (`PyccelKernel`, `CudaKernel`, `Kernel`/`KernelCatalog`, argument objects, accumulation, debugging, testing), worked examples, best practices and troubleshooting pages.
- `cunumpy/LLM_GUIDE.md`: a self-contained guide to the API and its rules for AI coding assistants, shipped as package data and rendered in the documentation.
- `xp.CudaKernel`: Wraps a CUDA C kernel (`cupy.RawKernel`, compiled lazily with NVRTC) so it can be called with the same arguments as the host kernel it mirrors, plus `n_threads`. The `extern "C" __global__` signature is parsed once and every call is checked against it: argument count, array dtypes (host arrays raise, they are never copied), and scalars (Python scalars are cast to the declared C types with range checks; lossy or mismatching scalars raise instead of reaching the kernel as silently wrong values). Supports `block_size`, NVRTC `options`, `include_dirs`, `shared_mem`, `stream`, `CudaKernel.from_file()` (`<name>_cuda.cu`), `compile()` and `prepare_args()`; `check_signature=False` skips the checks.
- `xp.CudaArguments`: Base class for argument objects that are flattened into several CUDA kernel arguments; any object with a `__cuda_args__()` method is flattened.
- `xp.parse_cuda_signature(source, name)` and `xp.CudaParameter`: Parse the parameters of a `__global__` function.
Expand Down
41 changes: 38 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -372,6 +372,7 @@ class ParticleArguments(xp.KernelArguments):


kernel(particles.kernel_args, dt, n_threads=n) # host or CUDA kernel
```

Kernels ported from pyccel index arrays like `markers[ip, j]`, which needs
shapes and strides rather than bare pointers. The shipped header
Expand Down Expand Up @@ -446,6 +447,7 @@ def make_args(backend, seed):
@pytest.mark.parametrize("name, kernel", catalog.parity_cases())
def test_parity(name, kernel):
assert_kernels_agree(kernel, make_args, n_threads=1000)
```

Accumulation kernels often write into a buffer that another library owns on
the host (a stencil vector's `_data`, exchanged over MPI). `DeviceMirror`
Expand Down Expand Up @@ -473,6 +475,39 @@ installation example and compatibility notes.

## Documentation

The [user guide](docs/source/quickstart.md) explains common workflows. The
[API reference](docs/source/api.md) documents each helper and its behavior.
The [Pyodide guide](docs/source/pyodide.md) covers WebAssembly usage.
The full documentation lives in [`docs/source`](docs/source/index.md) and is
published at <https://max-models.github.io/cunumpy/>:

* Getting started: [installation](docs/source/installation.md) and a
[quickstart](docs/source/quickstart.md) with a map of which guide covers what.
* User guide: [choosing a backend](docs/source/guides/backends.md),
[backend-agnostic code](docs/source/guides/portable-code.md),
[data movement](docs/source/guides/data-movement.md),
[devices, memory and streams](docs/source/guides/gpu-devices.md),
[MPI with one rank per GPU](docs/source/guides/mpi.md),
[timing and profiling](docs/source/guides/profiling.md).
* Porting kernels: [overview](docs/source/kernels/overview.md),
[`PyccelKernel`](docs/source/kernels/pyccel-kernel.md),
[`CudaKernel`](docs/source/kernels/cuda-kernel.md),
[`Kernel` and `KernelCatalog`](docs/source/kernels/dispatch.md),
[argument objects and structs](docs/source/kernels/arguments.md),
[accumulation kernels](docs/source/kernels/accumulation.md),
[debugging](docs/source/kernels/debugging.md),
[testing](docs/source/kernels/testing.md).
* [Worked examples](docs/source/examples/index.md),
[best practices](docs/source/best-practices.md),
[troubleshooting](docs/source/troubleshooting.md),
[Pyodide](docs/source/pyodide.md) and the
[API reference](docs/source/api.md).

### For AI coding assistants

[`src/cunumpy/LLM_GUIDE.md`](src/cunumpy/LLM_GUIDE.md) is a compact,
self-contained guide to the API and its rules for LLM-based coding assistants.
It ships inside the installed package, so an assistant working in a project that
depends on CuNumpy can read it from `site-packages/cunumpy/LLM_GUIDE.md`, or
locate it with:

```bash
python -c "import cunumpy, pathlib; print(pathlib.Path(cunumpy.__file__).parent / 'LLM_GUIDE.md')"
```
4 changes: 4 additions & 0 deletions docs/source/ai-assistants.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
<!-- This page renders src/cunumpy/LLM_GUIDE.md, which also ships inside the installed package. Edit that file, not this one. -->

```{include} ../../src/cunumpy/LLM_GUIDE.md
```
8 changes: 4 additions & 4 deletions docs/source/api.md
Original file line number Diff line number Diff line change
Expand Up @@ -241,6 +241,8 @@ too. An exception raised inside the block propagates as it is:
def test_time_step_stays_on_the_device():
with xp.assert_no_transfers():
propagator(dt)
```

### `as_device_array(value, dtype=None, ndim=None, *, name=None)`

The "reference or copy once" rule for building CUDA argument objects
Expand Down Expand Up @@ -432,7 +434,7 @@ does not necessarily indicate a leak.
### `cuda_include_dir()`

Returns the directory (as `str`) of the CUDA headers shipped with CuNumpy,
currently `cunumpy/atomic.cuh`. `CudaKernel` adds it to its NVRTC options as
`cunumpy/array_view.cuh`, `cunumpy/atomic.cuh` and `cunumpy/index.cuh`. `CudaKernel` adds it to its NVRTC options as
`-I<dir>` automatically (and only once), so kernel sources can write
`#include <cunumpy/atomic.cuh>` without configuration. Use it to pass the
same headers to other compilers.
Expand Down Expand Up @@ -668,11 +670,9 @@ kernels["shift"](x, 1.0, x.size, n_threads=x.size)
fix it for this kernel, see "Debugging" below.

Properties: `name`, `expression` (`name`, or the template instantiation such
as `"scale<double, 3>"`), `source`, `block_size`, `options`, `structs`,
`template_args`, `signature`, `is_compiled`, `debug`.
as `"scale<double, 3>"`), `source`, `block_size`, `options`, `include_dirs`,
`source_dir`, `included_headers`, `structs`, `template_args`, `signature`,
`is_compiled`.
`is_compiled`, `debug`.

### Included headers and the compile cache

Expand Down
70 changes: 70 additions & 0 deletions docs/source/best-practices.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,70 @@
# Best practices

A condensed checklist. Each item links to the guide with the reasoning.

## Structure

* Import as `import cunumpy as xp` and always call `xp.<name>(...)`; never
`from cunumpy import <array function>`. ([Backend-agnostic
code](guides/portable-code.md))
* Choose the backend once, in the entry point, with `ARRAY_BACKEND` or
`set_backend()`. Library code never calls `set_backend()`. ([Choosing a
backend](guides/backends.md))
* Read back `xp.get_backend()` after requesting CuPy; log it with
`xp.device_count()` and `xp.__version__`.
* Use `use_backend()` for scoped switches (tests, CPU reference computations),
never from several threads at once.
* Functions that receive arrays follow them with `get_array_module()`; functions
that combine arrays check them with `assert_same_backend()`.

## Data

* Create arrays with `xp.*` so they are born on the right device. Use `numpy`
directly only for host-only data and dtypes.
* Convert explicitly, at boundaries: `to_cunumpy()` on input, `to_numpy()` for
output, plotting and host-only libraries. ([Data
movement](guides/data-movement.md))
* Keep the time loop free of transfers and of host synchronization (`float()`,
`.item()`, `print`, `if` on device values). Verify with
`assert_no_transfers()` in a test.
* Give dtypes explicitly for arrays that go to kernels.
* Generate random test data on the host with a seed, then convert; NumPy and
CuPy generators differ.

## Kernels

* Start with `KernelCatalog.from_package(..., missing_cuda="fallback")`, port
kernels by profile order, switch to `"raise"` when done. ([Porting
kernels](kernels/overview.md))
* Keep the CUDA kernel's argument list identical to the host kernel's; put both
in one folder.
* Declare `outputs` on host kernels so the fallback copies back only what was
written, and never forget an argument that is written.
* Pass device arrays to `CudaKernel`; build argument objects once with
`as_device_array()`. ([Kernel arguments](kernels/arguments.md))
* Use `long long` indices, `CUNUMPY_THREAD_1D` guards, `const` on inputs, and
`cunumpy_atomic_add` for scatter writes.
* Use array views (`Array2D<double>`) instead of hand-computed offsets for
multi-dimensional data.
* Generate structs from the host argument class (`CudaStruct.from_signature`)
and test that committed headers are up to date.
* Compile at setup (`catalog.compile_all(jobs=...)`).
* Keep signature checks on; disable them (`check_signature=False`) only for
tiny kernels in hot loops after they are tested.

## Verification

* Parametrize tests with `BACKENDS` or the `backend` fixture; they run
everywhere and use the GPU where there is one. ([Testing
kernels](kernels/testing.md))
* One `assert_kernels_agree` test over `catalog.parity_cases()`.
* Debug crashes with `CUNUMPY_CUDA_DEBUG=1`, then `compute-sanitizer`.
([Debugging](kernels/debugging.md))
* Time with `timed_region()`, profile with `nvtx_range()` and `nsys`.
([Profiling](guides/profiling.md))

## MPI

* `set_backend("cupy")`, `bind_local_device()`, then `from mpi4py import MPI`,
then `require_cuda_aware_mpi()`. ([MPI](guides/mpi.md))
* `synchronize_for_mpi(*buffers)` before every MPI call with device buffers.
20 changes: 20 additions & 0 deletions docs/source/examples/index.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
# Worked examples

Complete programs that combine the pieces described in the guides.

* [A portable diffusion solver](portable-script.md): array-level code only,
one script for CPU and GPU, with a command-line backend switch, timing and
output at the boundaries.
* [Porting a particle-in-cell code](particle-pusher.md): a small simulation
whose kernels are ported to CUDA one at a time with a `KernelCatalog`,
verified with parity tests, and checked for transfers.

For an MPI program with one rank per GPU, see the start-up sequence and halo
exchange in [Multi-GPU programs with MPI](../guides/mpi.md).

```{toctree}
:hidden:

portable-script
particle-pusher
```
Loading
Loading