Skip to content

Fix operator correctness on Triton and CUDA - #1

Merged
whjthu merged 1 commit into
masterfrom
fix/triton-cuda-operator-correctness
Sep 30, 2026
Merged

whjthu merged 1 commit into
masterfrom
fix/triton-cuda-operator-correctness

Conversation

@whjthu

@whjthu whjthu commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

修复 ntops.lab 在 Triton 和 CUDA C 后端的算子正确性问题:

  • 修正广播布局、归约范围及矩阵乘分块配置。
  • 修正浮点整除、取模、舍入和 NaN 判断的边界行为。
  • 改善归一化计算的数值稳定性,并使低精度计算中的舍入顺序与 PyTorch 一致。

@whjthu
whjthu merged commit 17b1dbd into master Sep 30, 2026
4 checks passed
@whjthu
whjthu deleted the fix/triton-cuda-operator-correctness branch September 30, 2026 02:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant