Skip to content

Bound runtime collection launch specialization - #21

Merged
PraneethMerugu merged 1 commit into
mainfrom
codex/bounded-launch-contract
Sep 17, 2026
Merged

PraneethMerugu merged 1 commit into
mainfrom
codex/bounded-launch-contract

Conversation

@PraneethMerugu

Copy link
Copy Markdown
Owner

Summary

  • keep collection, keyed-reduction, ordered-fold, and destination-grouping extents as runtime ndrange data
  • specialize launches only on their existing 256-lane operation-family workgroup
  • add an ordinary CPU science/allocation regression test for the displaced constructor form

Evidence

  • live Kaimon: capacity 16→24 = 1 inference event; warmed 300→301 = 1; 24→32 admits KernelAbstractions' one-time dynamic-check mode without extent-specific kernel identity
  • helper MethodInstances contain only the operation-family Val{256} family
  • warm allocation: baseline 256 B × 7; candidate 256 B × 7
  • full CPU: 1,880/1,880
  • affected real Metal with scalar indexing disabled: 254/254
  • independent review: no findings, including zero-extent behavior
  • alternating warm throughput samples found no material CPU or Metal regression

This is the demonstrated C12 LocalMath companion in the canonical Potts PR chain. It adds no cache, policy hierarchy, alternate executor, or backend-specific scientific path.

@PraneethMerugu
PraneethMerugu merged commit 12b3fa9 into main Sep 17, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant