Skip to content

Version 0.4.0 - #53

Closed
max-models wants to merge 17 commits into
mainfrom
devel
Closed

max-models wants to merge 17 commits into
mainfrom
devel

Conversation

@max-models

Copy link
Copy Markdown
Owner

No description provided.

Added CudaKernel, Kernel and KernelCatalog for host/CUDA kernels
* ruff check --fix

* Add CUDA structs, templates, kernel variants, 1D-3D launches, compile_all, host kernel options and MPI device helpers

* Test CUDA structs, templates, variants, launch shapes, compile_all and device binding

* Document the new CUDA kernel features and MPI device helpers

* fixed ruff errors

* GPU CI: use CUDA 13.2 consistently and wait for the GitLab pipeline of the pushed commit
* Add KernelArguments protocol resolved per backend

* Add contiguity checks for device arrays and as_device_array()

* Add header-aware compile cache for CUDA kernels

* Fix GPU test that compiled a missing include
…unches (#44)

* Add KernelArguments protocol resolved per backend

* Add contiguity checks for device arrays and as_device_array()

* Add header-aware compile cache for CUDA kernels

* Add CUDA debug mode with line info, bounds checks and synchronized launches

* Fix ruff findings

* Fix GPU test that compiled a missing include

* Fix expected value in debug launch test
* Add KernelArguments protocol resolved per backend

* Add contiguity checks for device arrays and as_device_array()

* Add header-aware compile cache for CUDA kernels

* Add CUDA debug mode with line info, bounds checks and synchronized launches

* Add count_transfers and assert_no_transfers

* Fix ruff findings

* Fix GPU test that compiled a missing include

* Fix expected value in debug launch test
* Add KernelArguments protocol resolved per backend

* Add contiguity checks for device arrays and as_device_array()

* Add header-aware compile cache for CUDA kernels

* Add CUDA debug mode with line info, bounds checks and synchronized launches

* Add count_transfers and assert_no_transfers

* Add nvtx_range and timed_region profiling helpers

* Fix ruff findings

* Fix ruff findings

* Fix GPU test that compiled a missing include

* Fix expected value in debug launch test
* Add KernelArguments protocol resolved per backend

* Add contiguity checks for device arrays and as_device_array()

* Add header-aware compile cache for CUDA kernels

* Add CUDA debug mode with line info, bounds checks and synchronized launches

* Add count_transfers and assert_no_transfers

* Add nvtx_range and timed_region profiling helpers

* Add mpi_is_cuda_aware and require_cuda_aware_mpi

* Fix ruff findings

* Fix ruff findings

* Fix GPU test that compiled a missing include

* Fix expected value in debug launch test
* Add KernelArguments protocol resolved per backend

* Add contiguity checks for device arrays and as_device_array()

* Add header-aware compile cache for CUDA kernels

* Add CUDA debug mode with line info, bounds checks and synchronized launches

* Add count_transfers and assert_no_transfers

* Add nvtx_range and timed_region profiling helpers

* Add mpi_is_cuda_aware and require_cuda_aware_mpi

* Add catalog summary, parallel compile_all and all_from_file

* Fix ruff findings

* Fix ruff findings

* Fix GPU test that compiled a missing include

* Fix expected value in debug launch test
* Add KernelArguments protocol resolved per backend

* Add contiguity checks for device arrays and as_device_array()

* Add header-aware compile cache for CUDA kernels

* Add CUDA debug mode with line info, bounds checks and synchronized launches

* Add count_transfers and assert_no_transfers

* Add nvtx_range and timed_region profiling helpers

* Add mpi_is_cuda_aware and require_cuda_aware_mpi

* Add catalog summary, parallel compile_all and all_from_file

* Add cunumpy.testing with parity checks and device function wrappers

* Add DeviceMirror and atomic.cuh header

* Fix ruff findings

* Fix ruff findings

* Fix ruff findings

* Fix GPU test that compiled a missing include

* Fix expected value in debug launch test
…m_signature (#51)

* Add KernelArguments protocol resolved per backend

* Add contiguity checks for device arrays and as_device_array()

* Add header-aware compile cache for CUDA kernels

* Add CUDA debug mode with line info, bounds checks and synchronized launches

* Add count_transfers and assert_no_transfers

* Add nvtx_range and timed_region profiling helpers

* Add mpi_is_cuda_aware and require_cuda_aware_mpi

* Add catalog summary, parallel compile_all and all_from_file

* Add cunumpy.testing with parity checks and device function wrappers

* Add DeviceMirror and atomic.cuh header

* Add array view headers, struct fields of view type and CudaStruct.from_signature

* Fix ruff findings

* Fix ruff findings

* Fix ruff findings

* Fix GPU test that compiled a missing include

* Fix expected value in debug launch test

* Remove duplicate cuda_include_dir stub

* Run the bounds-check trap test in a subprocess

This branch was successfully deployed

1 active deployment
github-pages — ced51339 Deployed Oct 1, 2026 by max-models via build-and-deploy #39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant