cuda.bindings: support multiple CTK release lines on main - #2737
Conversation
Moon migration belongs in PR NVIDIA#2659. Restore the pre-Moon selective-CI planner and workflows from 727ef59.
|
/ok to test b87d0a1 |
|
There was a problem hiding this comment.
I started to comment on some individual things, but then decided to stop because I think there is a more fundamental change that needs to be made across this whole PR (and then I'm happy to come back and review further).
/Today/ the "current" version is 13, and the backport version is 12. But at some point in the future that will switch to 14 and 13. This pervasively hardcodes those version numbers all over this codebase, especially in CI, but in a bunch of the release scripts as well, and even the cuda_bindings_12 directory name as indicators of current vs. backport.
Instead, we should use the config we already have in versions.yml and use that to drive the numbers everywhere. That way when it's time to move on, all that should be required is updating versions.yml, and copying/overwriting the existing cuda_bindings to cuda_bindings_backport (or whatever we want to call it), and move on. I'm sure there are many details I'm missing, but that should be the goal and design -- it would be preferable to reduce it to as close to that as possible. The problem with this as-is is that there are hundreds of context-sensitive places that would need to be updated to do that update -- we are creating a massive pile of technical debt to pay later. I'm sure an agent might get that X% correct, but I always think it's better to engineer for flexibility, especially for something we know will happen. If versions.yml (which requires using yq to parse etc.) makes this too difficult, we could explore a simple VARIABLE=value format which would parse as both bash variables and Python variables and probably be more convenient to use from the many places it is needed. There are really only two actual values in versions.yml today, so that should be fine.
I'm also a little concerned (without any testing-based evidence) that this will break when we tag the same commit with v13.x.y and v12.x.y, which will be the common case, in fact, IMHO, one of the real benefits of moving to this approach. We should get an agent to do a thorough investigation of that use case and make sure it is covered. Ideally, it would be nice for a single release run to do both releases simultaneously but it's not a deal breaker if it still requires kicking off two runs.
Also what is this (from the agent's PR description):
The later NVML memoryview fix is reproduced byte-for-byte from cybind commit
6def52ca508c9e14ef67f4ce26a0c677f3fbad72 with Doxygen 1.17.0:
If there is something like this that wasn't backported, let's deal with that separately so it's not an unrelated tag-along to this PR.
Also a note for future agent reviewers of this PR: The interesting part of this PR is the part outside of the cuda_bindings_backport or cuda_bindings_12 directory. Those are just direct copies from the 12.9.x branch, and any differences between that and the cuda_bindings directory are likely intentional. When reviewing, focus on the scaffolding / CI / overall structure.
|
Archiving options related to a lychee chicken-and-egg issue. I'll go with Option 1 below. This comment is to explain why. codex: We have three sensible options. For PR 2737, I recommend keeping the canonical links unchanged and treating these as documented pre-merge exceptions.
Run lychee once with only these three URLs excluded, record that every other link passes, and rerun without exclusions after merge. This is reasonable because authored-source lychee is explicitly skipped by the GitHub CI job at .github/workflows/ci.yml, so these are not merge-gating failures. It avoids landing temporary configuration or compromising the final URLs.
Add three exact anchored patterns so This is practical, but creates a mandatory cleanup PR and briefly leaves three blind spots on
Teach the hook to map: Then links to newly introduced files are validated against the checkout before they exist online. This is exactly the future-URL use case for lychee’s remapping feature. Lychee remapping documentation It is the principled reusable solution, but needs a portable wrapper to calculate the absolute worktree path. I would pursue it separately only if this problem starts recurring. I would avoid:
So my recommendation is option 1: preserve the three correct final URLs, validate everything else with exact one-off exclusions, and rerun lychee from fresh |
|
/ok to test f4ddc4e |
|
/ok to test c6a0cf1 |
|
/ok to test 83c1cf0 |
mdboom
left a comment
There was a problem hiding this comment.
This PR is really challenging to review. Even ignoring the files that are just copied from the 12.9.x branch (which don't need review) there is a lot here.
I found a bunch of sources of unnecessary complexity and stopped reading after that, so still haven't done a full human pass. Even with agents, multiple stages of transformations between data formats causes multiple places that bugs can creep in. It makes it harder for agents or humans to understand the fundamental logic of what branches are covered and how this all works. I think we need a step back analysis of how things should be represented based on how things are needed downstream of that and adjust accordingly.
I know partly what is driving this complexity is GHA's design that forces things into small snippets. @kkraus14 has suggested elsewhere that maybe moving to a more formal mono-repo management tool like moon may be better than building out more and more CI complexity. Maybe an agent could build a prototype quickly to at least see whether it meets our use cases and what sort of complexity it requires so we can compare.
Additionally, there are many new scripts here in ci/tools with largely overlapping functionality where logic is spread between them and bash scripts and they sort of go back-and-forth. Moving more logic into fewer Python scripts, that output directly to what is most often needed (POSIX environment variables) would probably be preferable to the current state. It should be possible to see in one file how the variables that control the rest of the execution are computed. I'm thinking particularly of bindings_config.py and resolve_release_bindings_line.py -- why are the separate? And there is probably some value in combining it with compute_ci_plan.py in some way. Not necessarily that they need to be in the same source file, but that they would interact in the same process. Sorry to not have /concrete/ suggestions for that, but I am trying to suggest ways that this could be simplified to be more easily reviewable and maintained going forward.
If there is any way to break this up into multiple steps that could be reviewed independently, that would help a lot. Even when I used an agent to review, it struggled to understand the "why" of many of these changes. I asked my agents to offer some suggestions about making this easier to review. I think it's mostly ok (except I wouldn't consider the cuda_bindings_12/ import step as big -- I'm comfortable rubber stamping a direct copy from another branch). It doesn't really have good suggestions for incremental development -- most of what it suggests are just bugfixes.
A few concrete ways to cut the size and the duplicated-logic risk this PR introduces:
-
Split into sequential PRs. The
cuda_bindings_12/import, the registry (ci/versions.yml+bindings_config.py), and the workflow rewiring are three logically separable changes. Landing the registry + validator first (small, reviewable, testable in isolation against the existing single-line setup) then rewiring workflows to consume it, then importingcuda_bindings_12/last, would let each step get real scrutiny instead of one 255K-line PR where reviewers rubber-stamp the bulk. -
Stop tracking the tag family in two places.
ci/versions.yml'stag_seriesandcuda_bindings/pyproject.toml'stag_regex/git_describe_commandencode the same CUDA-major fact independently (flagged in the review — they can drift and did in thev13.4.0scenario). Either generate the pyprojecttag_regexfrom the registry at build time, or drop the registry'stag_seriesfield and derive it from the pyproject regex instead. One source of truth removes a whole class of the findings above. -
Don't model generality you don't use yet.
roles.maintenanceis schema'd as a list, butbuild-wheel.ymlhard-errors unless it has exactly one entry, and nothing in this PR needs more than one maintenance line. Collapsing the schema to a singlemaintenanceline (not a list) until a second one is actually needed removes validation code, removes a whole "what if maintenance has 2 entries" test surface, and can be widened later when there's a real second line to design against. -
Trim the transitional compatibility code in
compute_ci_plan.py. Thevariantsdict with the "OR aggregation... consumers migrate tolines" comment is scaffolding for old CUDA-major-keyed workflow consumers. If those consumers are being rewritten in this same PR anyway, migrating them straight to line-keyed output and deleting the compatibility shim removes a chunk of logic (and a source of the empty-matrix / dual-bookkeeping risk flagged earlier).
| - name: Checkout docs control plane | ||
| uses: actions/checkout@9c091bb21b7c1c1d1991bb908d89e4e9dddfe3e0 # v7.0.0 | ||
| with: | ||
| ref: ${{ github.sha }} | ||
| path: .ci-control | ||
|
|
||
| - name: Install CI tool dependencies | ||
| run: python3 -m pip install -r .ci-control/ci/tools/requirements.txt | ||
|
|
There was a problem hiding this comment.
This is treating ci/tools as a sort of poor version of a package, all just so it can use PyYAML.
The modern way to handle dependencies of standalone scripts is to use PEP 723 metadata and then use a PEP 723-supporting tool like uv, pixi or hatch to run it. I think that would be way less cumbersome than this. Or we go all in and make it a proper package which might have other benefits given how big it's getting. But this approach is sort of worst-of-both-worlds, IMHO.
| set -euo pipefail | ||
| if [[ -n "$BINDINGS_LINE" ]]; then | ||
| bindings_line="$BINDINGS_LINE" | ||
| elif [[ "${IS_RELEASE}" == "true" && "$RELEASE_TAG" == v* ]]; then | ||
| bindings_line=$(python3 .ci-control/ci/tools/resolve_release_bindings_line.py \ | ||
| --release-tag "$RELEASE_TAG" \ | ||
| --release-source-root . \ | ||
| --control-config .ci-control/ci/versions.yml) | ||
| else | ||
| echo "error: cannot find ci/versions.yml or ci/versions.json" >&2 | ||
| exit 1 | ||
| bindings_line=$(python3 .ci-control/ci/tools/bindings_config.py get --role current) | ||
| fi | ||
| BUILD_CTK_VER=$(jq -er '.toolkit_version' <<< "$bindings_line") | ||
| BINDINGS_COMPONENT_DIR=$(jq -er '.release_source_dir // .source_dir' <<< "$bindings_line") | ||
| BINDINGS_REGISTRY_ORIGIN=$(jq -er '.release_registry_origin // "tag"' <<< "$bindings_line") | ||
| if [[ ! "${BUILD_CTK_VER}" =~ ^[0-9]+\.[0-9]+\.[0-9]+$ ]]; then | ||
| echo "error: derived CTK build version ${BUILD_CTK_VER} does not match MAJOR.MINOR.MICRO" >&2 | ||
| exit 1 | ||
| fi | ||
| if [[ ! -d "$BINDINGS_COMPONENT_DIR" ]]; then | ||
| echo "error: resolved bindings source directory does not exist: $BINDINGS_COMPONENT_DIR" >&2 | ||
| exit 1 | ||
| fi | ||
| echo "BUILD_CTK_VER=${BUILD_CTK_VER}" >> "$GITHUB_ENV" | ||
| echo "BINDINGS_COMPONENT_DIR=${BINDINGS_COMPONENT_DIR}" >> "$GITHUB_ENV" | ||
| echo "BINDINGS_REGISTRY_ORIGIN=${BINDINGS_REGISTRY_ORIGIN}" >> "$GITHUB_ENV" |
There was a problem hiding this comment.
AFAICT, the resolve_release_bindings_line.py or bindings_config.py scripts read in the versions.yaml metadata and output json, which is then parsed with jq to convert to environment variables. And there is additional validation of the values coming out of scripts that we control. Can't we skip that middle steps and have the scripts output what is needed (envvar pairs)? The long standing env-vars script does that, for example. I'll probably need to read further to discover where using JSON as an intermediary might be relevant, though. My concern isn't efficiency, it's that there are so many transformation steps and therefore places for bugs to slip in.
I think fixing this will require stepping back and understanding all of the places these JSON-emitting scripts are used and providing output in the most convenient way possible for those contexts.
| tag_regex = "^(?P<version>v13\\.\\d+\\.\\d+(?:[ab]\\d+)?(?:\\.post\\d+)?)" | ||
| git_describe_command = ["git", "describe", "--dirty", "--tags", "--long", "--match", "v13.*"] |
There was a problem hiding this comment.
Updating the metadata in versions.yml will not affect this. How do we ensure they stay in sync?
| _TOOLKIT_VERSION_PATTERN = re.compile(r"[1-9][0-9]*\.[0-9]+\.[0-9]+(?:[.-][A-Za-z0-9]+)*") | ||
| _TAG_SERIES_PATTERN = re.compile(r"v[1-9][0-9]*(?:\.[0-9]+)*\.") | ||
| _FINAL_TAG_SUFFIX_PATTERN = re.compile(r"[0-9]+(?:\.post[0-9]+)?") | ||
| _ALPHA_BETA_TAG_SUFFIX_PATTERN = re.compile(r"[0-9]+(?:[ab][0-9]+)?(?:\.post[0-9]+)?") |
There was a problem hiding this comment.
This should be PEP 440 compliant so it can support anything we might want to do on PyPI, and I don't think it is. It would be better to use a library for this than reinventing here.
| build_bindings_current=$(jq -r --arg id "$current_line_id" '.modules.bindings.lines[$id].needs_build | if type == "boolean" then . else error("invalid current bindings build gate") end' <<< "$WORKPLAN") | ||
| build_bindings_maintenance=$(jq -r --arg id "$maintenance_line_id" '.modules.bindings.lines[$id].needs_build | if type == "boolean" then . else error("invalid maintenance bindings build gate") end' <<< "$WORKPLAN") |
There was a problem hiding this comment.
This got me thinking about whether the current and maintenance lines are built and tested every time, even if only one or the other changed? Part of Keith's recent work was to limit the amount that is rebuilt and tested every time, and it would be nice not to step back from that.
I had my agent investigate this and it sees that is more-or-less the case.
| for relative in shared_paths: | ||
| candidates = [(root, repo_root / root / relative) for root in roots] | ||
| symlinks = [root for root, path in candidates if path.is_symlink()] | ||
| if symlinks: | ||
| violations.append(f"{relative}: symlink in {', '.join(symlinks)}") | ||
| continue | ||
| missing = [root for root, path in candidates if not path.is_file()] | ||
| if missing: | ||
| violations.append(f"{relative}: missing from {', '.join(missing)}") | ||
| continue |
There was a problem hiding this comment.
From my agent:
Only the leaf path and root are checked for is_symlink(); an intermediate directory symlink is followed transparently by is_file()/read_bytes() and reported as "identical," defeating the stated symlink guard for everything under it.
| echo "SETUP_SANITIZER=${SETUP_SANITIZER}" | ||
| echo "BINDINGS_SOURCE=${BINDINGS_SOURCE}" | ||
| echo "CUDA_BINDINGS_ROOT=${CUDA_BINDINGS_ROOT}" | ||
| echo "CUDA_PYTHON_ARTIFACT_NAME=cuda-python-wheel-cuda${BINDINGS_BUILD_CUDA_VER:-${CUDA_VER}}" |
There was a problem hiding this comment.
From my agent:
CUDA_PYTHON_ARTIFACT_NAME falls back to ${CUDA_VER} (the test runner's CTK) rather than the build-time toolkit version in published mode, since BINDINGS_BUILD_CUDA_VER is only set in the local branch. Any workflow consuming this name in that mode downloads a nonexistent artifact.
| ' <<< "$BINDINGS_CONFIG") | ||
| maintenance_line_id=$(jq -er ' | ||
| .roles.maintenance | ||
| | if length == 1 then .[0] else error("wheel builder currently requires one maintenance line") end |
There was a problem hiding this comment.
From my agent:
roles.maintenance is schema'd as a list, but the wheel builder hard-errors unless it has exactly one entry, with the job otherwise hardwired to fixed CURRENT_/MAINTENANCE_ env pairs. Adding a second maintenance line — the stated point of a registry — breaks every build job. Fails loudly, so it's a documented design limit rather than silent corruption, but worth flagging as inconsistent with the registry's stated generality.
|
|
||
| ### CI Pipeline Flow | ||
|
|
||
|  |
There was a problem hiding this comment.
My agent flagged that this diagram is now out-of-date.
| active_section = "" | ||
| for line in pyproject.read_text(encoding="utf-8").splitlines(): | ||
| if match := _SECTION_PATTERN.fullmatch(line): | ||
| active_section = match.group(1).strip() | ||
| continue | ||
| if active_section == section and (match := _KEY_PATTERN.match(line)) and match.group(1) == key: | ||
| return True | ||
| return False |
There was a problem hiding this comment.
Should use tomllib rather than regexes to read toml.
|
/ok to test 22108c1 |
Treat carriage returns as delimiters in both jq TSV reads. Native jq on Windows emits CRLF, which otherwise leaves the final field contaminated and silently disables maintenance-line cuda.core Cython test artifacts.
Synchronize the applicable thread-safety markers from NVIDIA#2229 into the maintenance test tree. NVML initialization and graph-memory accounting use process-global state and must not run alongside parallel tests under free-threaded Python.
|
/ok to test 889062e |
|
/ok to test 7e2b151 |
REMINDER
Before merging, remove the temporary
.lycheeignorebefore triggering final CI. It excludes only three canonicalmain/cuda_bindings_12URLs that cannot resolve until this PR is merged. The authored-sourcelycheehook is skipped by CI, so removing the file will not prevent final CI from passing.After merging, run
pre-commit run lychee --all-fileson freshmainto validate those links.Summary
Closes #1199.
This is the writable continuation of Keith Kraus's original PR #2675, "cuda.bindings: build 12.9 and 13.x selectively from main". Most of the CUDA 12 source import and the initial build, test, and release integration came from Keith's PR. GitHub closed #2675 automatically when its temporary base branch was deleted after #2467 merged; this replacement preserves that work and commit history and completes the redesign requested during review.
The result is one active development branch for both released CUDA bindings lines:
main.ci/versions.ymlis the single mapping from each bindings package root to its exact toolkit pin andcurrentormaintenancerelease status.12.9.xbranch becomes a read-only release record instead of an active backport or artifact-source branch.Design
The package root is the stable identity; release status is metadata:
cuda_bindings_12/maintenancecuda_bindings/currentThere are no synthetic line IDs and no separate role-to-line mapping.
ci.tools.bindings_configreads and validates the registry, derives the CUDA target/major/variant and SCM tag rules, and emits normalized package records or GitHub environment variables directly. Downstream planning and workflows key their decisions by package root.The public wheel builder deliberately supports exactly one
currentpackage and onemaintenancepackage with different CUDA ABI majors; unsupported shapes fail closed. Release-tag syntax remains authoritative in each package root's[tool.setuptools_scm]metadata, and registry validation checks that it agrees with the toolkit pin.Release selection uses the registry from the tagged source tree. A contained compatibility path handles historical tags whose trees predate this registry. One commit may carry one CUDA 12.9 tag and one CUDA 13.3 tag; each tag independently selects the matching package root, metadata, and artifacts in its own release run.
The two complete package roots are an intentional transitional design. Most of
cuda_bindings_12/is the direct CUDA 12.9 import from Keith's PR #2675;cuda_bindings_12/MAINTENANCE.mdrecords its source and partial cybind generation provenance. Handwritten fixes must be assessed for both roots, while generated and target-specific files may legitimately differ.Review Feedback Incorporated
release_statusand toolkit pin.tag_series; release syntax comes from each source package's SCM metadata.currentandmaintenancesingular and rejected unimplemented public-registry shapes.bindings_config.pyand converted the helpers into a small importableci.toolspackage.cuda_bindings_12/MAINTENANCE.md.Reviewer Decisions
Please explicitly accept or reject these policies:
mainis the sole active source of truth. The historical12.9.xbranch receives no further routine or emergency backports. Applicable CUDA 12 fixes are made incuda_bindings_12/onmain, alongside any corresponding current-package change.release_statusrecords where it is in the release lifecycle without introducing a second identity.Review Map
The high-value review surface is outside the imported CUDA 12 tree:
ci/versions.yml,ci/tools/bindings_config.py,cuda_bindings*/pyproject.toml, and focused testsci/tools/compute_ci_plan.py,ci/tools/tests/test_compute_ci_plan.py, and.github/workflows/ci.ymlci/tools/env-vars,ci/tools/run-tests, and explicit wheel/source selectionci/tools/validate_release_wheels.py, release-note validation, and SCM-version handlingcuda_bindings_12/MAINTENANCE.mdMost files under
cuda_bindings_12/are the direct CUDA 12.9 import from Keith's PR #2675. Differences fromcuda_bindings/are generally target- or generation-specific and intentional.Validation
python -m pytest -q ci/tools/tests: 134 passed, plus 47 subtests, on head7e2b151pre-commit run --all-files: passed, including Ruff, actionlint, YAML/TOML/RST checks, generated-file seals, SCM/registry checks, andlycheewith the temporary exact exclusions described in the REMINDERv12.9.1andv13.3.0selection7e2b151: runningOut of Scope
Checklist