Version 0.4.0 - #53
Closed
max-models wants to merge 17 commits into
Closed
max-models wants to merge 17 commits into
max-models wants to merge 17 commits into
Conversation
Added CudaKernel, Kernel and KernelCatalog for host/CUDA kernels
* ruff check --fix * Add CUDA structs, templates, kernel variants, 1D-3D launches, compile_all, host kernel options and MPI device helpers * Test CUDA structs, templates, variants, launch shapes, compile_all and device binding * Document the new CUDA kernel features and MPI device helpers * fixed ruff errors * GPU CI: use CUDA 13.2 consistently and wait for the GitLab pipeline of the pushed commit
* Add KernelArguments protocol resolved per backend * Add contiguity checks for device arrays and as_device_array() * Add header-aware compile cache for CUDA kernels * Fix GPU test that compiled a missing include
…unches (#44) * Add KernelArguments protocol resolved per backend * Add contiguity checks for device arrays and as_device_array() * Add header-aware compile cache for CUDA kernels * Add CUDA debug mode with line info, bounds checks and synchronized launches * Fix ruff findings * Fix GPU test that compiled a missing include * Fix expected value in debug launch test
* Add KernelArguments protocol resolved per backend * Add contiguity checks for device arrays and as_device_array() * Add header-aware compile cache for CUDA kernels * Add CUDA debug mode with line info, bounds checks and synchronized launches * Add count_transfers and assert_no_transfers * Fix ruff findings * Fix GPU test that compiled a missing include * Fix expected value in debug launch test
* Add KernelArguments protocol resolved per backend * Add contiguity checks for device arrays and as_device_array() * Add header-aware compile cache for CUDA kernels * Add CUDA debug mode with line info, bounds checks and synchronized launches * Add count_transfers and assert_no_transfers * Add nvtx_range and timed_region profiling helpers * Fix ruff findings * Fix ruff findings * Fix GPU test that compiled a missing include * Fix expected value in debug launch test
* Add KernelArguments protocol resolved per backend * Add contiguity checks for device arrays and as_device_array() * Add header-aware compile cache for CUDA kernels * Add CUDA debug mode with line info, bounds checks and synchronized launches * Add count_transfers and assert_no_transfers * Add nvtx_range and timed_region profiling helpers * Add mpi_is_cuda_aware and require_cuda_aware_mpi * Fix ruff findings * Fix ruff findings * Fix GPU test that compiled a missing include * Fix expected value in debug launch test
* Add KernelArguments protocol resolved per backend * Add contiguity checks for device arrays and as_device_array() * Add header-aware compile cache for CUDA kernels * Add CUDA debug mode with line info, bounds checks and synchronized launches * Add count_transfers and assert_no_transfers * Add nvtx_range and timed_region profiling helpers * Add mpi_is_cuda_aware and require_cuda_aware_mpi * Add catalog summary, parallel compile_all and all_from_file * Fix ruff findings * Fix ruff findings * Fix GPU test that compiled a missing include * Fix expected value in debug launch test
* Add KernelArguments protocol resolved per backend * Add contiguity checks for device arrays and as_device_array() * Add header-aware compile cache for CUDA kernels * Add CUDA debug mode with line info, bounds checks and synchronized launches * Add count_transfers and assert_no_transfers * Add nvtx_range and timed_region profiling helpers * Add mpi_is_cuda_aware and require_cuda_aware_mpi * Add catalog summary, parallel compile_all and all_from_file * Add cunumpy.testing with parity checks and device function wrappers * Add DeviceMirror and atomic.cuh header * Fix ruff findings * Fix ruff findings * Fix ruff findings * Fix GPU test that compiled a missing include * Fix expected value in debug launch test
…m_signature (#51) * Add KernelArguments protocol resolved per backend * Add contiguity checks for device arrays and as_device_array() * Add header-aware compile cache for CUDA kernels * Add CUDA debug mode with line info, bounds checks and synchronized launches * Add count_transfers and assert_no_transfers * Add nvtx_range and timed_region profiling helpers * Add mpi_is_cuda_aware and require_cuda_aware_mpi * Add catalog summary, parallel compile_all and all_from_file * Add cunumpy.testing with parity checks and device function wrappers * Add DeviceMirror and atomic.cuh header * Add array view headers, struct fields of view type and CudaStruct.from_signature * Fix ruff findings * Fix ruff findings * Fix ruff findings * Fix GPU test that compiled a missing include * Fix expected value in debug launch test * Remove duplicate cuda_include_dir stub * Run the bounds-check trap test in a subprocess
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.