Skip to content

[libcu++] Update cuda::ptx:: for CUDA 13.5 - #10909

Open
pciolkosz wants to merge 1 commit into
NVIDIA:mainfrom
pciolkosz:ptx-sync-libcuda-ptx
Open

[libcu++] Update cuda::ptx:: for CUDA 13.5#10909
pciolkosz wants to merge 1 commit into
NVIDIA:mainfrom
pciolkosz:ptx-sync-libcuda-ptx

Conversation

@pciolkosz

Copy link
Copy Markdown
Contributor

This PR updates cuda/ptx in main with 13.4 updates plus a few extra additions (ldmatrix.m16n16.trans and mapa variants)

@pciolkosz
pciolkosz requested review from a team as code owners August 20, 2026 00:28
@pciolkosz
pciolkosz requested a review from gonidelis August 20, 2026 00:28
@pciolkosz
pciolkosz requested a review from fbusato August 20, 2026 00:28
@github-project-automation github-project-automation Bot moved this to Todo in CCCL Aug 20, 2026
@cccl-authenticator-app cccl-authenticator-app Bot moved this from Todo to In Review in CCCL Aug 20, 2026
@coderabbitai

coderabbitai Bot commented Aug 20, 2026

Copy link
Copy Markdown
Contributor

Review Change Stack

Important

Review skipped

Review was skipped as selected files did not have any reviewable changes.

💤 Files selected but had no reviewable changes (1)
  • libcudacxx/include/cuda/__ptx/ptx_dot_variants.h
⛔ Files ignored due to path filters (141)
  • libcudacxx/include/cuda/__ptx/instructions/generated/applypriority_async_bulk.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/barrier_cluster.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/bfind.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/bmsk.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/clusterlaunchcontrol.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_commit_group.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_multicast.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_prefetch.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_prefetch_tensor.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_tensor.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_tensor_gather_scatter.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_tensor_multicast.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_wait_group.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_mbarrier_arrive.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_mbarrier_arrive_noinc.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_reduce_async_bulk.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_reduce_async_bulk_bf16.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_reduce_async_bulk_f16.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_reduce_async_bulk_tensor.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/elect_sync.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/exit.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fabric_submit.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fabric_try_get.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fabric_try_pullred.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fabric_try_put.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fabric_try_red.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fabric_wait.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_mbarrier_init.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_alias.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_async.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_async_generic_sync_restrict.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_fabric_fabric_alias.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_fabric_generic_alias.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_generic_fabric_alias.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_tensormap_generic.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_sync_restrict.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/get_sreg.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/getctarank.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/ld.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/ldmatrix.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/ldmatrix_m16n16_trans.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mapa.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_arrive.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_arrive_drop.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_arrive_expect_tx.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_arrive_no_complete.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_check_layout.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_complete_tx.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_expect_tx.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_init.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_inval.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_pending_count.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_test_wait.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_test_wait_parity.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_try_wait.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_try_wait_parity.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/multimem_ld_reduce.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/multimem_red.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/multimem_st.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/prefetch.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/prmt.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/red_async.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/setmaxnreg.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/shl.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/shr.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/st.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/st_async.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/st_bulk.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/stmatrix.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_alloc.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_commit.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_cp.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_fence.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_ld.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_mma.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_mma_sp.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_mma_ws.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_shift.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_st.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_wait.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tensormap_cp_fenceproxy.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tensormap_replace.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/trap.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/applypriority_async_bulk.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/bfind.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/clusterlaunchcontrol.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_multicast.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_prefetch.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_prefetch_tensor.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_tensor.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_tensor_gather_scatter.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_tensor_multicast.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_reduce_async_bulk_tensor.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fabric_submit.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fabric_try_get.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fabric_try_pullred.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fabric_try_put.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fabric_try_red.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fabric_wait.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence_proxy_alias.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence_proxy_async_generic_sync_restrict.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence_proxy_fabric_fabric_alias.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence_proxy_fabric_generic_alias.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence_proxy_generic_fabric_alias.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/ldmatrix.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/ldmatrix_m16n16_trans.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mapa.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_arrive.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_arrive_drop.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_check_layout.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_complete_tx.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_init.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_pending_count.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_test_wait.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_test_wait_parity.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_try_wait.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_try_wait_parity.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/prefetch.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/setmaxnreg.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/shr.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/stmatrix.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_alloc_cta_group_1.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_alloc_cta_group_2.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_commit_cta_group_1.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_commit_cta_group_2.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_cp_cta_group_1.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_cp_cta_group_2.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_fence.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_ld.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_mma_cta_group_1.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_mma_cta_group_2.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_mma_sp.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_mma_ws.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_shift_cta_group_1.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_shift_cta_group_2.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_st.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_wait.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tensormap_replace.h is excluded by !**/generated/**
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 5a39c4a9-fe94-4de6-ad8c-a867c811fa22

📥 Commits

Reviewing files that changed from the base of the PR and between b0a306d and 4fcb94b.

⛔ Files ignored due to path filters (141)
  • libcudacxx/include/cuda/__ptx/instructions/generated/applypriority_async_bulk.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/barrier_cluster.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/bfind.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/bmsk.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/clusterlaunchcontrol.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_commit_group.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_multicast.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_prefetch.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_prefetch_tensor.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_tensor.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_tensor_gather_scatter.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_tensor_multicast.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_wait_group.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_mbarrier_arrive.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_mbarrier_arrive_noinc.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_reduce_async_bulk.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_reduce_async_bulk_bf16.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_reduce_async_bulk_f16.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_reduce_async_bulk_tensor.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/elect_sync.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/exit.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fabric_submit.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fabric_try_get.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fabric_try_pullred.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fabric_try_put.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fabric_try_red.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fabric_wait.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_mbarrier_init.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_alias.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_async.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_async_generic_sync_restrict.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_fabric_fabric_alias.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_fabric_generic_alias.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_generic_fabric_alias.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_tensormap_generic.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_sync_restrict.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/get_sreg.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/getctarank.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/ld.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/ldmatrix.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/ldmatrix_m16n16_trans.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mapa.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_arrive.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_arrive_drop.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_arrive_expect_tx.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_arrive_no_complete.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_check_layout.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_complete_tx.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_expect_tx.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_init.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_inval.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_pending_count.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_test_wait.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_test_wait_parity.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_try_wait.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_try_wait_parity.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/multimem_ld_reduce.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/multimem_red.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/multimem_st.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/prefetch.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/prmt.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/red_async.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/setmaxnreg.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/shl.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/shr.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/st.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/st_async.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/st_bulk.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/stmatrix.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_alloc.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_commit.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_cp.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_fence.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_ld.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_mma.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_mma_sp.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_mma_ws.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_shift.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_st.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_wait.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tensormap_cp_fenceproxy.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tensormap_replace.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/trap.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/applypriority_async_bulk.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/bfind.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/clusterlaunchcontrol.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_multicast.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_prefetch.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_prefetch_tensor.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_tensor.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_tensor_gather_scatter.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_tensor_multicast.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_reduce_async_bulk_tensor.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fabric_submit.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fabric_try_get.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fabric_try_pullred.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fabric_try_put.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fabric_try_red.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fabric_wait.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence_proxy_alias.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence_proxy_async_generic_sync_restrict.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence_proxy_fabric_fabric_alias.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence_proxy_fabric_generic_alias.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence_proxy_generic_fabric_alias.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/ldmatrix.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/ldmatrix_m16n16_trans.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mapa.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_arrive.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_arrive_drop.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_check_layout.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_complete_tx.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_init.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_pending_count.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_test_wait.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_test_wait_parity.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_try_wait.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_try_wait_parity.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/prefetch.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/setmaxnreg.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/shr.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/stmatrix.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_alloc_cta_group_1.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_alloc_cta_group_2.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_commit_cta_group_1.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_commit_cta_group_2.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_cp_cta_group_1.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_cp_cta_group_2.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_fence.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_ld.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_mma_cta_group_1.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_mma_cta_group_2.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_mma_sp.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_mma_ws.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_shift_cta_group_1.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_shift_cta_group_2.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_st.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_wait.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tensormap_replace.h is excluded by !**/generated/**
📒 Files selected for processing (1)
  • libcudacxx/include/cuda/__ptx/ptx_dot_variants.h

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
📝 Walkthrough

Summary by CodeRabbit

  • New Features

    • Added support for additional PTX instructions, including asynchronous bulk operations, Fabric operations, matrix load/store, prefetching, mbarrier utilities, and sparse matrix multiplication.
    • Added new PTX variant options for layouts, barrier phases, and reporting mechanisms.
    • Made the newly supported instructions available through the unified PTX interface.
  • Documentation

    • Added reference documentation and NVIDIA PTX ISA links for the newly supported instructions.
  • Tests

    • Added compile-time coverage for all newly supported PTX instruction interfaces.

Walkthrough

Changes

The PR adds PTX instruction wrapper headers, variant enums and selectors, umbrella-header wiring, generated-content documentation pages, availability updates, and compile-pass tests for asynchronous bulk, Fabric, matrix, mbarrier, prefetch, and related instructions.

PTX instruction support

Layer / File(s) Summary
PTX variant contracts
libcudacxx/include/cuda/__ptx/ptx_dot_variants.h
Adds layout, phase, and report-mechanism enums with integral-constant aliases and selector objects.
Instruction wrappers and umbrella wiring
libcudacxx/include/cuda/__ptx/instructions/*, libcudacxx/include/cuda/ptx
Adds wrapper headers and umbrella includes for the new PTX instructions.
Instruction documentation
docs/libcudacxx/ptx/instructions*
Adds instruction pages, generated-content includes, toctree entries, and availability-table updates.
Compile-pass validation
libcudacxx/test/libcudacxx/cuda/ptx/*
Adds compile-only coverage for the new instruction headers and fence proxy aliases.

Possibly related PRs

  • NVIDIA/cccl#10887: Updates related CUDA/libcudacxx PTX documentation, including mapa.rst.

Suggested reviewers: gonidelis, fbusato

✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Create stacked PR
  • Commit on current branch

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🧹 Nitpick comments (3)
docs/libcudacxx/ptx/instructions.rst (1)

299-301: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

suggestion: Replace generic Yes values with the exact CCCL and CUDA versions used by the surrounding table. Yes does not tell users which release provides each API. Verify the release metadata before merging. As per path instructions, documentation changes must cover API/version consistency.

Also applies to: 315-315, 336-336, 342-342, 357-357, 359-359, 361-361, 363-363, 365-365, 367-367, 468-468, 472-472, 482-482, 484-484, 507-507, 509-509, 551-551

Source: Path instructions

libcudacxx/include/cuda/__ptx/instructions/fabric_try_get.h (1)

17-23: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

suggestion: Use one consistent closing annotation for the system-header preprocessor chain. The current #endif // no system header does not repeat the exact controlling condition text.

  • libcudacxx/include/cuda/__ptx/instructions/fabric_try_get.h#L17-L23: remove or correct the Line 23 annotation.
  • libcudacxx/include/cuda/__ptx/instructions/fabric_try_pullred.h#L17-L23: remove or correct the Line 23 annotation.
  • libcudacxx/include/cuda/__ptx/instructions/fabric_try_put.h#L17-L23: remove or correct the Line 23 annotation.
  • libcudacxx/include/cuda/__ptx/instructions/fabric_try_red.h#L17-L23: remove or correct the Line 23 annotation.
  • libcudacxx/include/cuda/__ptx/instructions/fabric_wait.h#L17-L23: remove or correct the Line 23 annotation.
  • libcudacxx/include/cuda/__ptx/instructions/ldmatrix.h#L17-L23: remove or correct the Line 23 annotation.
  • libcudacxx/include/cuda/__ptx/instructions/ldmatrix_m16n16_trans.h#L17-L23: remove or correct the Line 23 annotation.
  • libcudacxx/include/cuda/__ptx/instructions/mapa.h#L17-L23: remove or correct the Line 23 annotation.
  • libcudacxx/include/cuda/__ptx/instructions/mbarrier_check_layout.h#L17-L23: remove or correct the Line 23 annotation.
    Based on learnings, annotated #endif comments must repeat the exact condition text.

Source: Learnings

libcudacxx/include/cuda/__ptx/instructions/fabric_try_pullred.h (1)

31-42: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

important: Move the half and bfloat16 forward declarations after <cuda/std/__cccl/prologue.h>. The current order places code before the required prologue include.

  • libcudacxx/include/cuda/__ptx/instructions/fabric_try_pullred.h#L31-L42: move the forward declarations below the prologue include and keep them outside _CCCL_BEGIN_NAMESPACE_CUDA_PTX.
  • libcudacxx/include/cuda/__ptx/instructions/fabric_try_red.h#L31-L42: move the forward declarations below the prologue include and keep them outside _CCCL_BEGIN_NAMESPACE_CUDA_PTX.
    As per coding guidelines, the last header included before code must be <cuda/std/__cccl/prologue.h>.

Source: Coding guidelines


ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Enterprise

Run ID: 0414e4f9-c00b-4078-9447-e0ac4202eab5

📥 Commits

Reviewing files that changed from the base of the PR and between 894b603 and b0a306d.

⛔ Files ignored due to path filters (226)
  • docs/libcudacxx/ptx/instructions/generated/applypriority_async_bulk.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/bfind.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/clusterlaunchcontrol.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/cp_async_bulk.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/cp_async_bulk_multicast.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/cp_async_bulk_prefetch.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/cp_async_bulk_prefetch_tensor.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/cp_async_bulk_tensor.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/cp_async_bulk_tensor_gather_scatter.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/cp_async_bulk_tensor_multicast.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/cp_reduce_async_bulk_tensor.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/fabric_submit.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/fabric_try_get.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/fabric_try_pullred.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/fabric_try_put.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/fabric_try_red.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/fabric_wait.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/fence_proxy_alias.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/fence_proxy_async_generic_sync_restrict.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/fence_proxy_fabric_fabric_alias.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/fence_proxy_fabric_generic_alias.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/fence_proxy_generic_fabric_alias.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/ldmatrix.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/ldmatrix_m16n16_trans.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/mapa.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/mbarrier_arrive.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/mbarrier_arrive_drop.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/mbarrier_check_layout.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/mbarrier_complete_tx.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/mbarrier_init.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/mbarrier_pending_count.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/mbarrier_test_wait.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/mbarrier_test_wait_parity.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/mbarrier_try_wait.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/mbarrier_try_wait_parity.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/prefetch.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/setmaxnreg.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/shr.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/stmatrix.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/tcgen05_alloc.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/tcgen05_commit.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/tcgen05_cp.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/tcgen05_fence.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/tcgen05_ld.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/tcgen05_mma.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/tcgen05_mma_sp.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/tcgen05_mma_ws.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/tcgen05_shift.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/tcgen05_st.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/tcgen05_wait.rst is excluded by !**/generated/**
  • docs/libcudacxx/ptx/instructions/generated/tensormap_replace.rst is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/applypriority_async_bulk.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/barrier_cluster.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/bfind.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/bmsk.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/clusterlaunchcontrol.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_commit_group.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_multicast.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_prefetch.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_prefetch_tensor.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_tensor.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_tensor_gather_scatter.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_tensor_multicast.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_bulk_wait_group.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_mbarrier_arrive.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_async_mbarrier_arrive_noinc.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_reduce_async_bulk.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_reduce_async_bulk_bf16.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_reduce_async_bulk_f16.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/cp_reduce_async_bulk_tensor.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/elect_sync.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/exit.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fabric_submit.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fabric_try_get.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fabric_try_pullred.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fabric_try_put.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fabric_try_red.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fabric_wait.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_mbarrier_init.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_alias.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_async.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_async_generic_sync_restrict.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_fabric_fabric_alias.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_fabric_generic_alias.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_generic_fabric_alias.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_proxy_tensormap_generic.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/fence_sync_restrict.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/get_sreg.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/getctarank.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/ld.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/ldmatrix.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/ldmatrix_m16n16_trans.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mapa.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_arrive.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_arrive_drop.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_arrive_expect_tx.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_arrive_no_complete.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_check_layout.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_complete_tx.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_expect_tx.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_init.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_inval.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_pending_count.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_test_wait.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_test_wait_parity.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_try_wait.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/mbarrier_try_wait_parity.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/multimem_ld_reduce.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/multimem_red.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/multimem_st.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/prefetch.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/prmt.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/red_async.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/setmaxnreg.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/shl.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/shr.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/st.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/st_async.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/st_bulk.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/stmatrix.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_alloc.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_commit.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_cp.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_fence.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_ld.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_mma.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_mma_sp.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_mma_ws.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_shift.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_st.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tcgen05_wait.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tensormap_cp_fenceproxy.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/tensormap_replace.h is excluded by !**/generated/**
  • libcudacxx/include/cuda/__ptx/instructions/generated/trap.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/applypriority_async_bulk.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/barrier_cluster.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/bfind.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/bmsk.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/clusterlaunchcontrol.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_commit_group.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_multicast.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_prefetch.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_prefetch_tensor.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_tensor.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_tensor_gather_scatter.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_tensor_multicast.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_bulk_wait_group.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_mbarrier_arrive.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_async_mbarrier_arrive_noinc.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_reduce_async_bulk.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_reduce_async_bulk_bf16.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_reduce_async_bulk_f16.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/cp_reduce_async_bulk_tensor.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/elect_sync.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/exit.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fabric_submit.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fabric_try_get.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fabric_try_pullred.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fabric_try_put.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fabric_try_red.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fabric_wait.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence_mbarrier_init.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence_proxy_alias.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence_proxy_async.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence_proxy_async_generic_sync_restrict.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence_proxy_fabric_fabric_alias.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence_proxy_fabric_generic_alias.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence_proxy_generic_fabric_alias.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence_proxy_tensormap_generic.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/fence_sync_restrict.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/get_sreg.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/getctarank.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/ld.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/ldmatrix.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/ldmatrix_m16n16_trans.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mapa.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_arrive.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_arrive_drop.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_arrive_expect_tx.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_arrive_no_complete.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_check_layout.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_complete_tx.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_expect_tx.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_init.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_inval.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_pending_count.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_test_wait.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_test_wait_parity.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_try_wait.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/mbarrier_try_wait_parity.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/multimem_ld_reduce.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/multimem_red.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/multimem_st.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/prefetch.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/prmt.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/red_async.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/setmaxnreg.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/shl.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/shr.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/st.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/st_async.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/st_bulk.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/stmatrix.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_alloc_cta_group_1.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_alloc_cta_group_2.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_commit_cta_group_1.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_commit_cta_group_2.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_cp_cta_group_1.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_cp_cta_group_2.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_fence.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_ld.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_mma_cta_group_1.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_mma_cta_group_2.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_mma_sp.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_mma_ws.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_shift_cta_group_1.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_shift_cta_group_2.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_st.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tcgen05_wait.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tensormap_cp_fenceproxy.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/tensormap_replace.h is excluded by !**/generated/**
  • libcudacxx/test/libcudacxx/cuda/ptx/generated/trap.h is excluded by !**/generated/**
📒 Files selected for processing (63)
  • docs/libcudacxx/ptx/instructions.rst
  • docs/libcudacxx/ptx/instructions/applypriority_async_bulk.rst
  • docs/libcudacxx/ptx/instructions/cp_async_bulk_prefetch.rst
  • docs/libcudacxx/ptx/instructions/cp_async_bulk_prefetch_tensor.rst
  • docs/libcudacxx/ptx/instructions/fabric_submit.rst
  • docs/libcudacxx/ptx/instructions/fabric_try_get.rst
  • docs/libcudacxx/ptx/instructions/fabric_try_pullred.rst
  • docs/libcudacxx/ptx/instructions/fabric_try_put.rst
  • docs/libcudacxx/ptx/instructions/fabric_try_red.rst
  • docs/libcudacxx/ptx/instructions/fabric_wait.rst
  • docs/libcudacxx/ptx/instructions/fence.rst
  • docs/libcudacxx/ptx/instructions/ldmatrix.rst
  • docs/libcudacxx/ptx/instructions/ldmatrix_m16n16_trans.rst
  • docs/libcudacxx/ptx/instructions/mapa.rst
  • docs/libcudacxx/ptx/instructions/mbarrier_arrive.rst
  • docs/libcudacxx/ptx/instructions/mbarrier_check_layout.rst
  • docs/libcudacxx/ptx/instructions/mbarrier_complete_tx.rst
  • docs/libcudacxx/ptx/instructions/mbarrier_pending_count.rst
  • docs/libcudacxx/ptx/instructions/prefetch.rst
  • docs/libcudacxx/ptx/instructions/stmatrix.rst
  • docs/libcudacxx/ptx/instructions/tcgen05_mma.rst
  • libcudacxx/include/cuda/__ptx/instructions/applypriority_async_bulk.h
  • libcudacxx/include/cuda/__ptx/instructions/cp_async_bulk_prefetch.h
  • libcudacxx/include/cuda/__ptx/instructions/cp_async_bulk_prefetch_tensor.h
  • libcudacxx/include/cuda/__ptx/instructions/fabric_submit.h
  • libcudacxx/include/cuda/__ptx/instructions/fabric_try_get.h
  • libcudacxx/include/cuda/__ptx/instructions/fabric_try_pullred.h
  • libcudacxx/include/cuda/__ptx/instructions/fabric_try_put.h
  • libcudacxx/include/cuda/__ptx/instructions/fabric_try_red.h
  • libcudacxx/include/cuda/__ptx/instructions/fabric_wait.h
  • libcudacxx/include/cuda/__ptx/instructions/fence.h
  • libcudacxx/include/cuda/__ptx/instructions/ldmatrix.h
  • libcudacxx/include/cuda/__ptx/instructions/ldmatrix_m16n16_trans.h
  • libcudacxx/include/cuda/__ptx/instructions/mapa.h
  • libcudacxx/include/cuda/__ptx/instructions/mbarrier_arrive.h
  • libcudacxx/include/cuda/__ptx/instructions/mbarrier_check_layout.h
  • libcudacxx/include/cuda/__ptx/instructions/mbarrier_complete_tx.h
  • libcudacxx/include/cuda/__ptx/instructions/mbarrier_pending_count.h
  • libcudacxx/include/cuda/__ptx/instructions/prefetch.h
  • libcudacxx/include/cuda/__ptx/instructions/stmatrix.h
  • libcudacxx/include/cuda/__ptx/instructions/tcgen05_mma.h
  • libcudacxx/include/cuda/__ptx/ptx_dot_variants.h
  • libcudacxx/include/cuda/ptx
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.applypriority.async.bulk.compile.pass.cpp
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.cp.async.bulk.prefetch.compile.pass.cpp
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.cp.async.bulk.prefetch.tensor.compile.pass.cpp
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.fabric.submit.compile.pass.cpp
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.fabric.try_get.compile.pass.cpp
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.fabric.try_pullred.compile.pass.cpp
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.fabric.try_put.compile.pass.cpp
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.fabric.try_red.compile.pass.cpp
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.fabric.wait.compile.pass.cpp
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.fence.compile.pass.cpp
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.ldmatrix.compile.pass.cpp
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.ldmatrix.m16n16.trans.compile.pass.cpp
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.mapa.compile.pass.cpp
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.mbarrier.arrive.compile.pass.cpp
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.mbarrier.check_layout.compile.pass.cpp
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.mbarrier.complete_tx.compile.pass.cpp
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.mbarrier.pending_count.compile.pass.cpp
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.prefetch.compile.pass.cpp
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.stmatrix.compile.pass.cpp
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.tcgen05.mma.sp.compile.pass.cpp

Included review availability: Your plan provides up to 12 included reviews per hour; 11 remain after this review.

Comment on lines +12 to +13
#ifndef _CUDA_PTX_APPLYPRIORITY_ASYNC_BULK_H_
#define _CUDA_PTX_APPLYPRIORITY_ASYNC_BULK_H_

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟠 Major | ⚡ Quick win

important: Apply one full-path include-guard convention to all new instruction headers.

  • libcudacxx/include/cuda/__ptx/instructions/applypriority_async_bulk.h#L12-L13: include the __ptx/instructions path in the guard.
  • libcudacxx/include/cuda/__ptx/instructions/cp_async_bulk_prefetch.h#L12-L13: include the __ptx/instructions path in the guard.
  • libcudacxx/include/cuda/__ptx/instructions/cp_async_bulk_prefetch_tensor.h#L12-L13: include the __ptx/instructions path in the guard.
  • libcudacxx/include/cuda/__ptx/instructions/fabric_submit.h#L12-L13: include the __ptx/instructions path in the guard.
  • libcudacxx/include/cuda/__ptx/instructions/mbarrier_complete_tx.h#L12-L13: include the __ptx/instructions path in the guard.
  • libcudacxx/include/cuda/__ptx/instructions/mbarrier_pending_count.h#L12-L13: include the __ptx/instructions path in the guard.
  • libcudacxx/include/cuda/__ptx/instructions/prefetch.h#L12-L13: include the __ptx/instructions path in the guard.
  • libcudacxx/include/cuda/__ptx/instructions/stmatrix.h#L12-L13: include the __ptx/instructions path in the guard.

As per coding guidelines, “Headers must use include guards derived from the uppercase full path.”

📍 Affects 8 files
  • libcudacxx/include/cuda/__ptx/instructions/applypriority_async_bulk.h#L12-L13 (this comment)
  • libcudacxx/include/cuda/__ptx/instructions/cp_async_bulk_prefetch.h#L12-L13
  • libcudacxx/include/cuda/__ptx/instructions/cp_async_bulk_prefetch_tensor.h#L12-L13
  • libcudacxx/include/cuda/__ptx/instructions/fabric_submit.h#L12-L13
  • libcudacxx/include/cuda/__ptx/instructions/mbarrier_complete_tx.h#L12-L13
  • libcudacxx/include/cuda/__ptx/instructions/mbarrier_pending_count.h#L12-L13
  • libcudacxx/include/cuda/__ptx/instructions/prefetch.h#L12-L13
  • libcudacxx/include/cuda/__ptx/instructions/stmatrix.h#L12-L13

Source: Coding guidelines

Comment on lines +251 to +252
[[maybe_unused]] static constexpr mbarrier_phase_primary_t mbarrier_phase_primary{};
[[maybe_unused]] static constexpr mbarrier_phase_conditional_t mbarrier_phase_conditional{};

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟠 Major | ⚡ Quick win

important: Use inline constexpr for all namespace-scope tag objects in this header. The declarations around lines 251–283 currently use static constexpr; update the mbarrier_phase_*, layout_*, and mbarrier_report_* objects to match libcudacxx conventions.

📍 Affects 1 file
  • libcudacxx/include/cuda/__ptx/ptx_dot_variants.h#L251-L252 (this comment)
  • libcudacxx/include/cuda/__ptx/ptx_dot_variants.h#L135-L155

Source: Coding guidelines

// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES.
//
//===----------------------------------------------------------------------===//
// UNSUPPORTED: libcpp-has-no-threads

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

important: Document the reason for the common UNSUPPORTED: libcpp-has-no-threads directive in every test. The libcudacxx test guidance requires a motivation for unsupported platforms.

  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.applypriority.async.bulk.compile.pass.cpp#L10-L10: add the rationale comment.
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.cp.async.bulk.prefetch.compile.pass.cpp#L10-L10: add the rationale comment.
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.cp.async.bulk.prefetch.tensor.compile.pass.cpp#L10-L10: add the rationale comment.
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.fabric.submit.compile.pass.cpp#L10-L10: add the rationale comment.
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.mbarrier.check_layout.compile.pass.cpp#L10-L10: add the rationale comment.
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.mbarrier.complete_tx.compile.pass.cpp#L10-L10: add the rationale comment.
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.mbarrier.pending_count.compile.pass.cpp#L10-L10: add the rationale comment.
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.prefetch.compile.pass.cpp#L10-L10: add the rationale comment.
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.stmatrix.compile.pass.cpp#L10-L10: add the rationale comment.
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.tcgen05.mma.sp.compile.pass.cpp#L10-L10: add the rationale comment.
📍 Affects 10 files
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.applypriority.async.bulk.compile.pass.cpp#L10-L10 (this comment)
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.cp.async.bulk.prefetch.compile.pass.cpp#L10-L10
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.cp.async.bulk.prefetch.tensor.compile.pass.cpp#L10-L10
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.fabric.submit.compile.pass.cpp#L10-L10
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.mbarrier.check_layout.compile.pass.cpp#L10-L10
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.mbarrier.complete_tx.compile.pass.cpp#L10-L10
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.mbarrier.pending_count.compile.pass.cpp#L10-L10
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.prefetch.compile.pass.cpp#L10-L10
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.stmatrix.compile.pass.cpp#L10-L10
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.tcgen05.mma.sp.compile.pass.cpp#L10-L10

Source: Coding guidelines

// SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES.
//
//===----------------------------------------------------------------------===//
// UNSUPPORTED: libcpp-has-no-threads

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

suggestion: Document or remove the repeated UNSUPPORTED: libcpp-has-no-threads directive. The changed tests contain no direct thread-dependent code. If the harness requires threads, state that dependency; otherwise remove the directive to preserve compile coverage on no-thread configurations. As per coding guidelines, unsupported or skipped tests must state why they are skipped.

  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.fabric.try_get.compile.pass.cpp#L10-L10: Add the skip rationale or remove the directive.
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.fabric.try_pullred.compile.pass.cpp#L10-L10: Add the skip rationale or remove the directive.
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.fabric.try_put.compile.pass.cpp#L10-L10: Add the skip rationale or remove the directive.
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.fabric.try_red.compile.pass.cpp#L10-L10: Add the skip rationale or remove the directive.
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.fabric.wait.compile.pass.cpp#L10-L10: Add the skip rationale or remove the directive.
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.ldmatrix.compile.pass.cpp#L10-L10: Add the skip rationale or remove the directive.
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.ldmatrix.m16n16.trans.compile.pass.cpp#L10-L10: Add the skip rationale or remove the directive.
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.mapa.compile.pass.cpp#L10-L10: Add the skip rationale or remove the directive.
📍 Affects 8 files
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.fabric.try_get.compile.pass.cpp#L10-L10 (this comment)
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.fabric.try_pullred.compile.pass.cpp#L10-L10
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.fabric.try_put.compile.pass.cpp#L10-L10
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.fabric.try_red.compile.pass.cpp#L10-L10
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.fabric.wait.compile.pass.cpp#L10-L10
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.ldmatrix.compile.pass.cpp#L10-L10
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.ldmatrix.m16n16.trans.compile.pass.cpp#L10-L10
  • libcudacxx/test/libcudacxx/cuda/ptx/ptx.mapa.compile.pass.cpp#L10-L10

Source: Coding guidelines

@pciolkosz
pciolkosz force-pushed the ptx-sync-libcuda-ptx branch from b0a306d to 4fcb94b Compare August 20, 2026 01:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: In Review

Development

Successfully merging this pull request may close these issues.

1 participant