Skip to content

Commit 5b8bf1a

Browse files
authored
Merge pull request #33 from max-models/32-add-a-simple-way-to-get-the-global-backend
Return global backend with `get_backend()`
2 parents 0094100 + ae30e6d commit 5b8bf1a

17 files changed

Lines changed: 763 additions & 314 deletions

‎.github/workflows/static_analysis.yml‎

Lines changed: 0 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -38,7 +38,6 @@ jobs:
3838
run: |
3939
pip install black "black[jupyter]"
4040
black --check src/
41-
black --check tutorials/
4241
4342
isort:
4443
runs-on: ubuntu-latest
@@ -50,7 +49,6 @@ jobs:
5049
run: |
5150
pip install isort
5251
isort --check src/
53-
isort --check tutorials/
5452
5553
mypy:
5654
runs-on: ubuntu-latest

‎.github/workflows/testing.yml‎

Lines changed: 1 addition & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -53,15 +53,11 @@ jobs:
5353
run: |
5454
pytest .
5555
56-
- name: Test tutorials
57-
run: |
58-
jupyter nbconvert --to notebook --execute tutorials/*.ipynb --output-dir=/tmp --ExecutePreprocessor.timeout=300
59-
6056
- name: Test docs build
6157
run: |
6258
pip install ".[docs]"
6359
cd docs
6460
make clean
6561
make html
6662
cd ..
67-
ls docs/build/html/index.html
63+
ls docs/build/html/index.html

‎README.md‎

Lines changed: 177 additions & 71 deletions
Original file line numberDiff line numberDiff line change
@@ -1,115 +1,221 @@
11
# CuNumpy
22

3-
Simple wrapper for numpy and cupy. Replace `import numpy as np` with `import cunumpy as xp`.
3+
CuNumpy lets a Python program use a NumPy-like API while choosing NumPy arrays
4+
on the CPU or CuPy arrays on an NVIDIA GPU. In the simplest case, replace
5+
`import numpy as np` with `import cunumpy as xp`; the array operations you
6+
already know then run on the selected backend.
47

5-
# Install
8+
```python
9+
import cunumpy as xp
610

7-
```bash
8-
pip install cunumpy
11+
values = xp.arange(5, dtype=xp.float64)
12+
print(values * 2)
13+
print(xp.get_backend()) # 'numpy' by default
914
```
1015

11-
Example usage:
16+
CuNumpy selects an array library for newly requested operations. It does not
17+
move existing arrays just because the selected backend changes. This guide
18+
covers backend selection, array movement, mixed CPU/GPU workflows, and the
19+
helper APIs CuNumpy provides around NumPy and CuPy.
1220

21+
## Install
22+
23+
```bash
24+
python -m pip install cunumpy
1325
```
14-
export ARRAY_BACKEND=cupy
15-
```
26+
27+
NumPy and `array-api-compat` are installed as dependencies. To use a GPU,
28+
install a CuPy package compatible with your CUDA environment as well. CuPy
29+
installation depends on the CUDA version and platform; follow the CuPy
30+
installation instructions for your system. CuNumpy does not install CUDA.
31+
32+
## Choose a backend
33+
34+
CuNumpy starts with NumPy unless `ARRAY_BACKEND=cupy` is set before import.
35+
You can also choose at runtime:
1636

1737
```python
1838
import cunumpy as xp
1939

2040
xp.set_backend("cupy")
41+
print(xp.get_backend()) # 'cupy' if CuPy and CUDA are functional
2142

22-
arr = xp.array([1, 2])
43+
values = xp.arange(5) # created by the active backend
44+
```
2345

24-
print(f"{type(arr) = }")
25-
print(f"{xp.__version__ = }")
46+
The accepted backend names are `"numpy"` and `"cupy"`. If CuPy is requested
47+
but unavailable or not functional, CuNumpy falls back to NumPy. Always check
48+
`get_backend()` when the effective backend matters, such as when reporting
49+
configuration or deciding whether GPU-specific work will happen.
2650

27-
# Convert to NumPy
28-
arr_np = xp.to_numpy(arr)
51+
Use `use_backend()` for a temporary selection. It restores the previous
52+
selection when the block exits, including when an exception is raised:
2953

30-
# Convert to active backend
31-
arr_xp = xp.to_cunumpy(arr)
54+
```python
55+
with xp.use_backend("numpy"):
56+
cpu_values = xp.linspace(0, 1, 100)
57+
assert xp.get_backend() == "numpy"
3258

33-
# Inspect backend
34-
print(f"{xp.get_backend(arr) = }")
35-
print(f"{xp.is_gpu(arr) = }")
36-
print(f"{xp.is_cpu(arr) = }")
59+
# The previous global backend is active again here.
60+
```
3761

38-
# Temporarily switch backend
39-
with xp.use_backend("numpy"):
40-
# This code runs on CPU even if ARRAY_BACKEND=cupy
41-
arr_cpu = xp.zeros(100)
62+
The backend selection is process-wide shared state. Do not switch it
63+
independently from multiple threads or async tasks; those changes can
64+
interfere. A context manager is useful for sequential code, tests, and
65+
notebooks.
4266

43-
# Set backend globally
44-
xp.set_backend("cupy")
67+
## Understand the two backend questions
4568

46-
# Synchronize GPU operations (no-op on CPU)
47-
xp.synchronize()
69+
The active backend controls which library CuNumpy exposes through its NumPy
70+
like operations. The array backend reports where one particular array lives.
71+
These can differ: changing the active backend does not convert arrays that
72+
already exist.
73+
74+
```python
75+
xp.set_backend("numpy")
76+
cpu_values = xp.arange(3)
77+
78+
gpu_values = xp.to_cupy(cpu_values) # explicit transfer
79+
print(xp.get_backend()) # 'numpy'
80+
print(xp.get_array_backend(gpu_values)) # 'cupy'
4881
```
4982

50-
Output:
83+
Use `is_cpu(array)`, `is_gpu(array)`, or `get_array_backend(array)` when
84+
dispatch should follow the array passed to a function. `get_array_module()`
85+
returns the matching `array_api_compat` module, which is useful when writing
86+
backend-generic functions:
5187

88+
```python
89+
def vector_norm(values):
90+
array_xp = xp.get_array_module(values)
91+
return array_xp.sqrt(array_xp.sum(values * values))
5292
```
53-
type(arr) = <class 'cupy.ndarray'>
54-
xp.__version__ = '0.1.4'
55-
xp.get_backend(arr) = 'cupy'
56-
xp.is_gpu(arr) = True
57-
xp.is_cpu(arr) = False
93+
94+
## Move data between CPU and GPU
95+
96+
Transfers are explicit so it is clear when data crosses the CPU/GPU boundary:
97+
98+
```python
99+
host = xp.to_numpy(gpu_values) # CuPy -> NumPy (host)
100+
device = xp.to_cupy(host) # NumPy/array-like -> CuPy (device)
101+
active = xp.to_cunumpy(host) # convert to the currently selected backend
58102
```
59103

60-
# Pyodide
104+
`to_numpy()` also accepts ordinary array-like values. `to_cupy()` raises
105+
`ImportError` when CuPy or a functional CUDA runtime is unavailable.
106+
`to_cunumpy()` is useful at API boundaries where the consumer expects the
107+
currently selected backend. It does not change the original array.
61108

62-
cuNumPy supports Pyodide with the NumPy backend. In an initialized Pyodide
63-
JavaScript runtime, install the package and run a Python-source kernel:
109+
Avoid transferring data inside a tight loop. Keep intermediate arrays on one
110+
backend and move only at boundaries such as file I/O, plotting, or a
111+
CPU-only library call. For example:
64112

65-
```javascript
66-
await pyodide.loadPackage("micropip");
67-
await pyodide.runPythonAsync(`
68-
import micropip
69-
await micropip.install("cunumpy")
113+
```python
114+
with xp.use_backend("cupy"):
115+
signal = xp.asarray(host_signal)
116+
filtered = xp.fft.rfft(signal)
117+
result = xp.to_numpy(filtered) # one transfer for a CPU-only consumer
118+
```
70119

71-
import cunumpy as xp
72-
xp.set_backend("numpy")
120+
## Random numbers and dtypes
73121

74-
def scale(values, factor):
75-
values[:] *= factor
76-
return values
122+
`get_rng(seed)` returns a random generator for the active backend. NumPy and
123+
CuPy have similar generator APIs, though exact bit-for-bit sequences are not
124+
guaranteed to match between libraries:
77125

78-
values = xp.array([1.0, 2.0, 3.0])
79-
result = xp.PyccelKernel(scale)(values, 2.0)
80-
assert result is values
81-
print(xp.to_numpy(result)) # [2. 4. 6.]
82-
`);
126+
```python
127+
rng = xp.get_rng(seed=42)
128+
samples = rng.normal(size=1000)
83129
```
84130

85-
NumPy is the default backend when `ARRAY_BACKEND` is unset. CuPy/CUDA and
86-
native Pyccel compilation are not supported in Pyodide. Despite its name,
87-
`PyccelKernel` accepts ordinary Python callables and neither imports Pyccel nor
88-
compiles code. Applications must supply Python-source kernels with
89-
Pyodide-compatible imports.
131+
Use `default_float_dtype()` when code needs to explicitly request the active
132+
backend's `float64` dtype rather than rely on Python scalar inference:
90133

91-
CI tests the built wheel in Pyodide's WebAssembly runtime under Node.js, including
92-
array operations, conversions, mutation/aliasing, backend context restoration,
93-
and Python kernels, while rejecting CuPy or Pyccel imports. This does not test
94-
browser-specific integration such as page loading or workers. See the
95-
[Pyodide guide](docs/source/pyodide.md) for details and local test commands.
134+
```python
135+
x = xp.asarray([1.0, 2.0], dtype=xp.default_float_dtype())
136+
```
96137

97-
# Development tests
138+
## GPU selection and memory helpers
98139

99-
```bash
100-
pip install -e '.[test]'
101-
pytest tests/portable
140+
These helpers are useful for multi-GPU programs and for understanding CuPy's
141+
memory behavior:
142+
143+
```python
144+
print("visible GPUs:", xp.device_count())
145+
xp.set_device(0) # selects CUDA device 0 when CuPy is active
146+
print("memory (free, total):", xp.memory_info())
102147
```
103148

104-
The `test` extra is compiler-free. For compiled-kernel tests, install
105-
`'.[test-compiled]'` and a working native compiler, then run `pytest`.
106-
The `dev` extra includes these compiled-test dependencies as before.
149+
`set_device()` is a no-op on NumPy. `device_count()` checks visible CUDA
150+
hardware even if the active backend is NumPy; it returns zero when CuPy/CUDA
151+
cannot be used. `memory_info()` returns `(free_bytes, total_bytes)` on the
152+
active CuPy device and `None` on NumPy. `set_device_for_rank(rank)` is a
153+
round-robin convenience for MPI layouts where local ranks map contiguously to
154+
GPUs. If your scheduler uses a different mapping, select the device directly.
107155

108-
# Build docs
156+
CuPy caches released allocations in memory pools. This can make process-level
157+
GPU memory appear occupied after arrays go out of scope. `free_memory()` asks
158+
CuPy to release currently free cached blocks; it does not free memory still
159+
referenced by live arrays.
109160

161+
`pin_memory(host_array)` makes a pinned host copy, which can improve transfer
162+
throughput for workloads that explicitly manage asynchronous transfers.
163+
`stream()` creates a non-blocking CuPy stream and yields it; it yields `None`
164+
on NumPy. GPU work is asynchronous, so synchronize before reading results on
165+
the host:
110166

167+
```python
168+
with xp.stream():
169+
device = xp.to_cupy(host)
170+
transformed = xp.fft.fft(device)
171+
172+
xp.synchronize()
173+
result = xp.to_numpy(transformed)
111174
```
112-
make html
113-
cd ../
114-
open docs/_build/html/index.html
175+
176+
## Use NumPy-only kernels with CuPy arrays
177+
178+
`PyccelKernel` adapts a callable that expects NumPy arrays. When conversion is
179+
needed, CuNumpy copies CuPy inputs to the host, calls the wrapped function,
180+
copies in-place output changes back to the device, and moves returned NumPy
181+
arrays to CuPy. With NumPy inputs, the wrapper calls the function directly.
182+
CuNumpy does not compile functions or import Pyccel for you.
183+
184+
```python
185+
import cunumpy as xp
186+
187+
188+
def scale_in_place(values, factor):
189+
values[:] *= factor
190+
return values
191+
192+
193+
scale = xp.PyccelKernel(scale_in_place, outputs=(0,))
194+
195+
with xp.use_backend("cupy"):
196+
values = xp.arange(5, dtype=xp.float64)
197+
returned = scale(values, 3.0)
198+
xp.synchronize()
115199
```
200+
201+
By default every converted argument is copied back, because the wrapper cannot
202+
know which arguments the kernel changed. `outputs=(0,)` declares that
203+
positional argument 0 is written, avoiding unnecessary copy-back for
204+
read-only inputs. For a keyword call, declare the keyword name, such as
205+
`outputs=("out",)`. A wrong declaration can leave GPU output values stale.
206+
The wrapper can also traverse arrays nested in lists, tuples, dictionaries,
207+
and selected application objects; see the full [API reference](docs/source/api.md)
208+
for `object_modules`, `is_array`, aliasing, and output declarations.
209+
210+
## Pyodide
211+
212+
CuNumpy supports the NumPy backend in Pyodide. It does not provide CuPy/CUDA
213+
there. Ordinary Python callables can be wrapped with `PyccelKernel` without
214+
compilation. See the [Pyodide guide](docs/source/pyodide.md) for a complete
215+
installation example and compatibility notes.
216+
217+
## Documentation
218+
219+
The [user guide](docs/source/quickstart.md) explains common workflows. The
220+
[API reference](docs/source/api.md) documents each helper and its behavior.
221+
The [Pyodide guide](docs/source/pyodide.md) covers WebAssembly usage.

0 commit comments

Comments
 (0)