|
1 | 1 | # CuNumpy |
2 | 2 |
|
3 | | -Simple wrapper for numpy and cupy. Replace `import numpy as np` with `import cunumpy as xp`. |
| 3 | +CuNumpy lets a Python program use a NumPy-like API while choosing NumPy arrays |
| 4 | +on the CPU or CuPy arrays on an NVIDIA GPU. In the simplest case, replace |
| 5 | +`import numpy as np` with `import cunumpy as xp`; the array operations you |
| 6 | +already know then run on the selected backend. |
4 | 7 |
|
5 | | -# Install |
| 8 | +```python |
| 9 | +import cunumpy as xp |
6 | 10 |
|
7 | | -```bash |
8 | | -pip install cunumpy |
| 11 | +values = xp.arange(5, dtype=xp.float64) |
| 12 | +print(values * 2) |
| 13 | +print(xp.get_backend()) # 'numpy' by default |
9 | 14 | ``` |
10 | 15 |
|
11 | | -Example usage: |
| 16 | +CuNumpy selects an array library for newly requested operations. It does not |
| 17 | +move existing arrays just because the selected backend changes. This guide |
| 18 | +covers backend selection, array movement, mixed CPU/GPU workflows, and the |
| 19 | +helper APIs CuNumpy provides around NumPy and CuPy. |
12 | 20 |
|
| 21 | +## Install |
| 22 | + |
| 23 | +```bash |
| 24 | +python -m pip install cunumpy |
13 | 25 | ``` |
14 | | -export ARRAY_BACKEND=cupy |
15 | | -``` |
| 26 | + |
| 27 | +NumPy and `array-api-compat` are installed as dependencies. To use a GPU, |
| 28 | +install a CuPy package compatible with your CUDA environment as well. CuPy |
| 29 | +installation depends on the CUDA version and platform; follow the CuPy |
| 30 | +installation instructions for your system. CuNumpy does not install CUDA. |
| 31 | + |
| 32 | +## Choose a backend |
| 33 | + |
| 34 | +CuNumpy starts with NumPy unless `ARRAY_BACKEND=cupy` is set before import. |
| 35 | +You can also choose at runtime: |
16 | 36 |
|
17 | 37 | ```python |
18 | 38 | import cunumpy as xp |
19 | 39 |
|
20 | 40 | xp.set_backend("cupy") |
| 41 | +print(xp.get_backend()) # 'cupy' if CuPy and CUDA are functional |
21 | 42 |
|
22 | | -arr = xp.array([1, 2]) |
| 43 | +values = xp.arange(5) # created by the active backend |
| 44 | +``` |
23 | 45 |
|
24 | | -print(f"{type(arr) = }") |
25 | | -print(f"{xp.__version__ = }") |
| 46 | +The accepted backend names are `"numpy"` and `"cupy"`. If CuPy is requested |
| 47 | +but unavailable or not functional, CuNumpy falls back to NumPy. Always check |
| 48 | +`get_backend()` when the effective backend matters, such as when reporting |
| 49 | +configuration or deciding whether GPU-specific work will happen. |
26 | 50 |
|
27 | | -# Convert to NumPy |
28 | | -arr_np = xp.to_numpy(arr) |
| 51 | +Use `use_backend()` for a temporary selection. It restores the previous |
| 52 | +selection when the block exits, including when an exception is raised: |
29 | 53 |
|
30 | | -# Convert to active backend |
31 | | -arr_xp = xp.to_cunumpy(arr) |
| 54 | +```python |
| 55 | +with xp.use_backend("numpy"): |
| 56 | + cpu_values = xp.linspace(0, 1, 100) |
| 57 | + assert xp.get_backend() == "numpy" |
32 | 58 |
|
33 | | -# Inspect backend |
34 | | -print(f"{xp.get_backend(arr) = }") |
35 | | -print(f"{xp.is_gpu(arr) = }") |
36 | | -print(f"{xp.is_cpu(arr) = }") |
| 59 | +# The previous global backend is active again here. |
| 60 | +``` |
37 | 61 |
|
38 | | -# Temporarily switch backend |
39 | | -with xp.use_backend("numpy"): |
40 | | - # This code runs on CPU even if ARRAY_BACKEND=cupy |
41 | | - arr_cpu = xp.zeros(100) |
| 62 | +The backend selection is process-wide shared state. Do not switch it |
| 63 | +independently from multiple threads or async tasks; those changes can |
| 64 | +interfere. A context manager is useful for sequential code, tests, and |
| 65 | +notebooks. |
42 | 66 |
|
43 | | -# Set backend globally |
44 | | -xp.set_backend("cupy") |
| 67 | +## Understand the two backend questions |
45 | 68 |
|
46 | | -# Synchronize GPU operations (no-op on CPU) |
47 | | -xp.synchronize() |
| 69 | +The active backend controls which library CuNumpy exposes through its NumPy |
| 70 | +like operations. The array backend reports where one particular array lives. |
| 71 | +These can differ: changing the active backend does not convert arrays that |
| 72 | +already exist. |
| 73 | + |
| 74 | +```python |
| 75 | +xp.set_backend("numpy") |
| 76 | +cpu_values = xp.arange(3) |
| 77 | + |
| 78 | +gpu_values = xp.to_cupy(cpu_values) # explicit transfer |
| 79 | +print(xp.get_backend()) # 'numpy' |
| 80 | +print(xp.get_array_backend(gpu_values)) # 'cupy' |
48 | 81 | ``` |
49 | 82 |
|
50 | | -Output: |
| 83 | +Use `is_cpu(array)`, `is_gpu(array)`, or `get_array_backend(array)` when |
| 84 | +dispatch should follow the array passed to a function. `get_array_module()` |
| 85 | +returns the matching `array_api_compat` module, which is useful when writing |
| 86 | +backend-generic functions: |
51 | 87 |
|
| 88 | +```python |
| 89 | +def vector_norm(values): |
| 90 | + array_xp = xp.get_array_module(values) |
| 91 | + return array_xp.sqrt(array_xp.sum(values * values)) |
52 | 92 | ``` |
53 | | -type(arr) = <class 'cupy.ndarray'> |
54 | | -xp.__version__ = '0.1.4' |
55 | | -xp.get_backend(arr) = 'cupy' |
56 | | -xp.is_gpu(arr) = True |
57 | | -xp.is_cpu(arr) = False |
| 93 | + |
| 94 | +## Move data between CPU and GPU |
| 95 | + |
| 96 | +Transfers are explicit so it is clear when data crosses the CPU/GPU boundary: |
| 97 | + |
| 98 | +```python |
| 99 | +host = xp.to_numpy(gpu_values) # CuPy -> NumPy (host) |
| 100 | +device = xp.to_cupy(host) # NumPy/array-like -> CuPy (device) |
| 101 | +active = xp.to_cunumpy(host) # convert to the currently selected backend |
58 | 102 | ``` |
59 | 103 |
|
60 | | -# Pyodide |
| 104 | +`to_numpy()` also accepts ordinary array-like values. `to_cupy()` raises |
| 105 | +`ImportError` when CuPy or a functional CUDA runtime is unavailable. |
| 106 | +`to_cunumpy()` is useful at API boundaries where the consumer expects the |
| 107 | +currently selected backend. It does not change the original array. |
61 | 108 |
|
62 | | -cuNumPy supports Pyodide with the NumPy backend. In an initialized Pyodide |
63 | | -JavaScript runtime, install the package and run a Python-source kernel: |
| 109 | +Avoid transferring data inside a tight loop. Keep intermediate arrays on one |
| 110 | +backend and move only at boundaries such as file I/O, plotting, or a |
| 111 | +CPU-only library call. For example: |
64 | 112 |
|
65 | | -```javascript |
66 | | -await pyodide.loadPackage("micropip"); |
67 | | -await pyodide.runPythonAsync(` |
68 | | - import micropip |
69 | | - await micropip.install("cunumpy") |
| 113 | +```python |
| 114 | +with xp.use_backend("cupy"): |
| 115 | + signal = xp.asarray(host_signal) |
| 116 | + filtered = xp.fft.rfft(signal) |
| 117 | + result = xp.to_numpy(filtered) # one transfer for a CPU-only consumer |
| 118 | +``` |
70 | 119 |
|
71 | | - import cunumpy as xp |
72 | | - xp.set_backend("numpy") |
| 120 | +## Random numbers and dtypes |
73 | 121 |
|
74 | | - def scale(values, factor): |
75 | | - values[:] *= factor |
76 | | - return values |
| 122 | +`get_rng(seed)` returns a random generator for the active backend. NumPy and |
| 123 | +CuPy have similar generator APIs, though exact bit-for-bit sequences are not |
| 124 | +guaranteed to match between libraries: |
77 | 125 |
|
78 | | - values = xp.array([1.0, 2.0, 3.0]) |
79 | | - result = xp.PyccelKernel(scale)(values, 2.0) |
80 | | - assert result is values |
81 | | - print(xp.to_numpy(result)) # [2. 4. 6.] |
82 | | -`); |
| 126 | +```python |
| 127 | +rng = xp.get_rng(seed=42) |
| 128 | +samples = rng.normal(size=1000) |
83 | 129 | ``` |
84 | 130 |
|
85 | | -NumPy is the default backend when `ARRAY_BACKEND` is unset. CuPy/CUDA and |
86 | | -native Pyccel compilation are not supported in Pyodide. Despite its name, |
87 | | -`PyccelKernel` accepts ordinary Python callables and neither imports Pyccel nor |
88 | | -compiles code. Applications must supply Python-source kernels with |
89 | | -Pyodide-compatible imports. |
| 131 | +Use `default_float_dtype()` when code needs to explicitly request the active |
| 132 | +backend's `float64` dtype rather than rely on Python scalar inference: |
90 | 133 |
|
91 | | -CI tests the built wheel in Pyodide's WebAssembly runtime under Node.js, including |
92 | | -array operations, conversions, mutation/aliasing, backend context restoration, |
93 | | -and Python kernels, while rejecting CuPy or Pyccel imports. This does not test |
94 | | -browser-specific integration such as page loading or workers. See the |
95 | | -[Pyodide guide](docs/source/pyodide.md) for details and local test commands. |
| 134 | +```python |
| 135 | +x = xp.asarray([1.0, 2.0], dtype=xp.default_float_dtype()) |
| 136 | +``` |
96 | 137 |
|
97 | | -# Development tests |
| 138 | +## GPU selection and memory helpers |
98 | 139 |
|
99 | | -```bash |
100 | | -pip install -e '.[test]' |
101 | | -pytest tests/portable |
| 140 | +These helpers are useful for multi-GPU programs and for understanding CuPy's |
| 141 | +memory behavior: |
| 142 | + |
| 143 | +```python |
| 144 | +print("visible GPUs:", xp.device_count()) |
| 145 | +xp.set_device(0) # selects CUDA device 0 when CuPy is active |
| 146 | +print("memory (free, total):", xp.memory_info()) |
102 | 147 | ``` |
103 | 148 |
|
104 | | -The `test` extra is compiler-free. For compiled-kernel tests, install |
105 | | -`'.[test-compiled]'` and a working native compiler, then run `pytest`. |
106 | | -The `dev` extra includes these compiled-test dependencies as before. |
| 149 | +`set_device()` is a no-op on NumPy. `device_count()` checks visible CUDA |
| 150 | +hardware even if the active backend is NumPy; it returns zero when CuPy/CUDA |
| 151 | +cannot be used. `memory_info()` returns `(free_bytes, total_bytes)` on the |
| 152 | +active CuPy device and `None` on NumPy. `set_device_for_rank(rank)` is a |
| 153 | +round-robin convenience for MPI layouts where local ranks map contiguously to |
| 154 | +GPUs. If your scheduler uses a different mapping, select the device directly. |
107 | 155 |
|
108 | | -# Build docs |
| 156 | +CuPy caches released allocations in memory pools. This can make process-level |
| 157 | +GPU memory appear occupied after arrays go out of scope. `free_memory()` asks |
| 158 | +CuPy to release currently free cached blocks; it does not free memory still |
| 159 | +referenced by live arrays. |
109 | 160 |
|
| 161 | +`pin_memory(host_array)` makes a pinned host copy, which can improve transfer |
| 162 | +throughput for workloads that explicitly manage asynchronous transfers. |
| 163 | +`stream()` creates a non-blocking CuPy stream and yields it; it yields `None` |
| 164 | +on NumPy. GPU work is asynchronous, so synchronize before reading results on |
| 165 | +the host: |
110 | 166 |
|
| 167 | +```python |
| 168 | +with xp.stream(): |
| 169 | + device = xp.to_cupy(host) |
| 170 | + transformed = xp.fft.fft(device) |
| 171 | + |
| 172 | +xp.synchronize() |
| 173 | +result = xp.to_numpy(transformed) |
111 | 174 | ``` |
112 | | -make html |
113 | | -cd ../ |
114 | | -open docs/_build/html/index.html |
| 175 | + |
| 176 | +## Use NumPy-only kernels with CuPy arrays |
| 177 | + |
| 178 | +`PyccelKernel` adapts a callable that expects NumPy arrays. When conversion is |
| 179 | +needed, CuNumpy copies CuPy inputs to the host, calls the wrapped function, |
| 180 | +copies in-place output changes back to the device, and moves returned NumPy |
| 181 | +arrays to CuPy. With NumPy inputs, the wrapper calls the function directly. |
| 182 | +CuNumpy does not compile functions or import Pyccel for you. |
| 183 | + |
| 184 | +```python |
| 185 | +import cunumpy as xp |
| 186 | + |
| 187 | + |
| 188 | +def scale_in_place(values, factor): |
| 189 | + values[:] *= factor |
| 190 | + return values |
| 191 | + |
| 192 | + |
| 193 | +scale = xp.PyccelKernel(scale_in_place, outputs=(0,)) |
| 194 | + |
| 195 | +with xp.use_backend("cupy"): |
| 196 | + values = xp.arange(5, dtype=xp.float64) |
| 197 | + returned = scale(values, 3.0) |
| 198 | + xp.synchronize() |
115 | 199 | ``` |
| 200 | + |
| 201 | +By default every converted argument is copied back, because the wrapper cannot |
| 202 | +know which arguments the kernel changed. `outputs=(0,)` declares that |
| 203 | +positional argument 0 is written, avoiding unnecessary copy-back for |
| 204 | +read-only inputs. For a keyword call, declare the keyword name, such as |
| 205 | +`outputs=("out",)`. A wrong declaration can leave GPU output values stale. |
| 206 | +The wrapper can also traverse arrays nested in lists, tuples, dictionaries, |
| 207 | +and selected application objects; see the full [API reference](docs/source/api.md) |
| 208 | +for `object_modules`, `is_array`, aliasing, and output declarations. |
| 209 | + |
| 210 | +## Pyodide |
| 211 | + |
| 212 | +CuNumpy supports the NumPy backend in Pyodide. It does not provide CuPy/CUDA |
| 213 | +there. Ordinary Python callables can be wrapped with `PyccelKernel` without |
| 214 | +compilation. See the [Pyodide guide](docs/source/pyodide.md) for a complete |
| 215 | +installation example and compatibility notes. |
| 216 | + |
| 217 | +## Documentation |
| 218 | + |
| 219 | +The [user guide](docs/source/quickstart.md) explains common workflows. The |
| 220 | +[API reference](docs/source/api.md) documents each helper and its behavior. |
| 221 | +The [Pyodide guide](docs/source/pyodide.md) covers WebAssembly usage. |
0 commit comments