Skip to content

fix: restrict i8-to-bf16 arch gate to Hopper+ and guard s2t copy to leader CTA - #3380

Open
Functionhx wants to merge 2 commits into
NVIDIA:mainfrom
Functionhx:fix/i8-bf16-arch-gate-and-s2t-race
Open

fix: restrict i8-to-bf16 arch gate to Hopper+ and guard s2t copy to leader CTA#3380
Functionhx wants to merge 2 commits into
NVIDIA:mainfrom
Functionhx:fix/i8-bf16-arch-gate-and-s2t-race

Conversation

@Functionhx

Copy link
Copy Markdown

Fixes #3354 and #3331.

  • cvt_i8_bf16_intrinsic listed Ampere/Ada in supported_archs but lowers to sm_90+ PTX. Removed so the itofp fallback is reached.
  • The s2t copy of scale factors was gated only by elect_one_sync(), allowing both SMs to race. Added is_mma_leader_cta check.

…types.md

media/docs/cpp/fundamental_types.md contained a duplicate empty  section with an empty fenced code block. The proper  subsection already exists earlier in the doc with full content and
worked examples. Remove the trailing duplicate so the topic appears only once.

Fixes NVIDIA#3064

Test Plan:
- Pure docs change, no code touched
- Verified the proper Numeric Conversion section (around line 198) is untouched
  and remains fully populated

Signed-off-by: Yuchen Fan <functionhx@gmail.com>
…eader CTA

cvt_i8_bf16_intrinsic lowers to cvt.rn.bf16.s8 (PTX ISA sm_90+).
Ampere and Ada were incorrectly included in supported_archs,
causing ptxas to reject the instruction instead of falling
through to the itofp software fallback.  Remove Ampere and Ada
from the arch list so the fallback is reached.

In the SM100 block-scaled MMA collective, the s2t copy of scale
factors was gated only by elect_one_sync(), allowing both SMs
in 2SM mode to issue the copy and race on the target registers.
Add the is_mma_leader_cta check used by every other 2SM-bifurcated
operation in the same file.

Fixes NVIDIA#3354
Fixes NVIDIA#3331

Signed-off-by: Yuchen Fan <functionhx@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] cvt int8/uint8 -> bfloat16 emits a sm_90-only instruction on Ampere/Ada instead of the software fallback (arch gate too loose)

1 participant