Conversation
Generate the cumsum hierarchy for the signed index width supported by Metal, matching the existing WebGPU path. This prevents ForceNarrowIndexToInt32 from rejecting unreachable 64-bit thresholds. Add target-dispatch coverage for Metal, WebGPU, and CUDA, and execute the numerical cumsum test on Metal.
Restore the default 64-bit continuous cumsum hierarchy for Metal. Add an optional index_bits argument to DispatchSortScan so callers using forced int32 narrowing can request a compatible hierarchy without restricting other Metal pipelines. Preserve the 32-bit WebGPU default and reject explicit 64-bit WebGPU scan budgets. Validate defaults, overrides, invalid budgets, and numerical Metal execution with both index policies. Document that the option neither changes tensor dtypes nor inserts narrowing or runtime bounds checks.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Add an optional
index_bitsargument toDispatchSortScanso pipelines that subsequently force indices toint32can generate compatible continuous GPU cumsum hierarchies. The hierarchy calculation can otherwise emit 64-bit thresholds that fail forced int32 narrowing, as seen in MLC's Metal compilation path. Existing callers retain the upstream defaults: 32 bits for WebGPU and 64 bits for other targets.