Skip to content

Add pip-index.py: mirror PEP 503 simple-index-only PyPI registries - #212

Open
yaoge123 wants to merge 1 commit into
tuna:masterfrom
yaoge123:pip-index-upstream
Open

yaoge123 wants to merge 1 commit into
tuna:masterfrom
yaoge123:pip-index-upstream

Conversation

@yaoge123

@yaoge123 yaoge123 commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Add pip-index.py, a mirror script for PyPI-style registries that only expose a PEP 503 simple HTML index over HTTP and offer no rsync/bandersnatch-compatible endpoint — e.g. NVIDIA's jetson-ai-lab devpi (Jetson wheels) and AMD's repo.amd.com/rocm/whl (ROCm/TheRock wheels). It crawls the simple index recursively, downloads referenced files, and rewrites hrefs so saved index pages point back to this mirror.

Adapted from ustclug/ustcmirror-images pytorch/sync.py; used by the pytorch, rocm-pypi and jetson-pypi tunasync jobs. Not for pypi.org full mirroring (use shadowmire.py).

Features

  • asyncio/aiohttp crawler with bounded concurrency (JOBS), streaming downloads, atomic .tmp writes; ClientTimeout(total=None, sock_connect=30, sock_read=TIMEOUT) so multi-GB wheels are bounded only by idle time; 403/404 responses re-raise immediately instead of burning retries
  • Rewrite-all href policy: every absolute href is rewritten to URLBASE regardless of host, so multi-host upstreams (download-r2.pytorch.org, pypi.nvidia.com, repo.amd.com) need no configuration
  • EXTRA_REWRITES: redirect other hosts' hrefs to a sibling local mirror prefix (e.g. files.pythonhosted.org=/pypi/web) without downloading them
  • DEVPI_MODE: crawl only a devpi channel's own projects via its JSON API; CUSTOM_ENDPOINTS seeds arbitrary index roots; NO_NIGHTLY=1 (default) filters nightly subtrees, including discovered links
  • tolerates 403/404 index pages (skip the subtree, protect it from cleanup)
  • PEP 658/714: both data-dist-info-metadata and data-core-metadata attributes are followed for .metadata files
  • cross-host collision guard: destinations stay keyed by URL path (preserving the existing data layout); claim_dest() records the full (scheme, host, path) origin of every artifact and index page, a conflicting later URL is logged as an error and skipped, and rewrite_index leaves the losing href unchanged so clients fetch it from its own host with a matching #sha256 fragment
  • CLEANUP=1 (opt-in): deletes stale local files no crawled URL maps to, capped by CLEANUP_MAX_DELETE; files that returned 404 are treated as yanked and removed too, separately capped by CLEANUP_MAX_404_DELETE
  • REVALIDATE (default on when CLEANUP=1): HEAD-checks existing files — 404 feeds cleanup, changed Content-Length triggers a redownload, transient HEAD errors keep the local copy
  • ends with Total size is ... for tunasync's size_pattern

Environment variables

Variable Default Meaning
TO / TUNASYNC_WORKING_DIR . mirror data directory (tunasync injects)
TUNASYNC_MIRROR_NAME default URLBASE when unset
URLBASE /<name>/ local prefix used when rewriting hrefs
EXTRA_REWRITES empty comma-separated host=prefix rules, not downloaded
CUSTOM_ENDPOINTS extra index roots to crawl (required for non-PyTorch upstreams)
DEVPI_MODE 0 crawl devpi channels via their JSON API
USE_PYTORCH_RELEASES / GET_ALL 0 PyTorch-only releases.json discovery
NO_NIGHTLY 1 skip /nightly/ URLs
JOBS 1 download concurrency
TIMEOUT 120 socket read (idle) timeout per request (s)
DRY_RUN 0 log only
CLEANUP 0 delete stale local files after crawling
CLEANUP_MAX_DELETE 1000 safety cap for stale deletions
CLEANUP_MAX_404_DELETE 100 safety cap for 404 (yanked) deletions
REVALIDATE 1 when CLEANUP=1 HEAD-check existing files
https_proxy / HTTPS_PROXY honored (aiohttp trust_env)

Testing

python3 -m py_compile clean; test/docker_runtime.sh tunathu/tunasync-scripts:latest passes; containerized end-to-end crawl against two local test servers exercising the collision guard (first host wins, loser logged + skipped + href kept), 404 subtree skip, nightly filtering, .metadata fetching and CLEANUP exclusions. Both mirrors run in production at mirror.nju.edu.cn: jetson-pypi (since May) and rocm-pypi (~492G, verified against the upstream file list).

Deployment

Requires aiohttp at runtime. This PR adds python3-aiohttp to the repo Dockerfile and smoke-tests the import both at image build time and in test/docker_runtime.sh; the currently deployed tunathu/tunasync-scripts:latest already ships aiohttp 3.14.3, so production is not blocked while a new image is built.

Copilot AI lite review requested due to automatic review settings September 23, 2026 15:37

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Critical runtime and download-collision issues, along with additional correctness gaps, remain unresolved.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 2 High severity · 3 Medium severity

Open (5)
What changed in this PR

Adds an asynchronous crawler for mirroring PEP 503-style PyPI registries, including devpi, Jetson, and ROCm indexes.

Changes:

  • Adds bounded-concurrency downloads, retries, streaming, deduplication, and path sanitization.
  • Rewrites local index links and supports custom endpoints, external rewrites, and devpi/PyTorch discovery.
  • Adds optional nightly filtering, stale-file cleanup, and progress logging.
File Summary and final review comments
pip-index.py Implements the package-index mirroring workflow. Findings: missing runtime aiohttp dependency (critical, 4 votes); host/path download collisions (critical, 1 vote); nightly filtering is not applied recursively (moderate, 3 votes); 404 index pages are not tolerated (moderate, 4 votes); incorrect data-core-metadata attribute (moderate, 2 votes); unused STRIP_PREFIX configuration (moderate, 1 vote); incorrect devpi +api request (moderate, 1 vote); no PEP 691 JSON support despite the compatibility claim (nit, 1 vote).

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread pip-index.py

from contextlib import contextmanager
from typing import IO, Any, Generator
import aiohttp

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 26d3305 — the repo Dockerfile now installs python3-aiohttp so the script runs in the standard image.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in d6f056f — python3-aiohttp is installed via the repo Dockerfile, and the import is smoke-tested both at image build time and in test/docker_runtime.sh. The currently deployed tunathu/tunasync-scripts:latest already ships aiohttp 3.14.3, so production was never blocked.

Comment thread pip-index.py
Comment thread pip-index.py Outdated
Comment thread pip-index.py
Comment thread pip-index.py Outdated
@yaoge123

Copy link
Copy Markdown
Contributor Author

Addressed the Copilot review findings; changes pushed as d6f056f (on top of 26d3305, which already contained first-pass fixes).

Status per finding:

  1. aiohttp dependency (HIGH) — Dockerfile already installs python3-aiohttp (26d3305). d6f056f adds the missing smoke coverage: aiohttp is now part of the import check in test/docker_runtime.sh and in the matching build-time RUN python3 -c 'import ...' line in Dockerfile. Also verified against the currently deployed tunathu/tunasync-scripts:latest: import aiohttp works (aiohttp 3.14.3), so production is not blocked while a new image is built.
  2. Cross-host destination collision (HIGH) — addressed differently than 26d3305: that commit keying non-primary hosts by netloc changed the on-disk layout, which would orphan existing mirrored data for the multi-host jobs already running this script (pytorch / rocm-pypi / jetson-pypi). d6f056f reverts to path-only destinations (existing layout preserved) and instead keeps a dest -> (scheme, host, path) map: when the same path is linked from two different hosts, the later URL is logged as an error and skipped, so concurrent downloads can no longer overwrite/corrupt each other through a shared dest/.tmp. Single-host mirrors are byte-for-byte compatible with the previous layout.
  3. 404 index pages (MEDIUM) — tolerated exactly like 403 (skip subtree, recorded for CLEANUP); already in 26d3305, docstrings/comments now consistently say 403/404.
  4. NO_NIGHTLY on discovered links (MEDIUM) — the same filter now runs before recursive tasks are created, not only for seed endpoints (26d3305).
  5. PEP 658/714 metadata attribute (MEDIUM) — both data-dist-info-metadata and data-core-metadata are accepted when fetching .whl.metadata (26d3305).

Verification: python3 -m py_compile clean; test/docker_runtime.sh tunathu/tunasync-scripts:latest passes; plus a containerized end-to-end crawl against two local test servers exercising the conflict guard (first host wins, second logged+skipped), 404 subtree skip, nightly filtering, .metadata fetching, and CLEANUP exclusions.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved correctness issues remain in collision handling, metadata rewriting, cleanup, and prefix stripping.

Get a fresh assessment by requesting another Copilot review.

Review effort: Lite
Findings: 2 High severity · 1 Medium severity

Open (3)
Resolved since last review (4)

Comment thread pip-index.py
Comment thread pip-index.py Outdated

@happyaron happyaron left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for upstreaming this. A few issues beyond the two Copilot findings still open (index-page collisions at L443, 404 files kept from CLEANUP at L484), which I agree with.

Description vs. code: the PR body lists STRIP_PREFIX, but it isn't in pip-index.py or its docstring. Was a commit left out? The docstring is also out of date: it says "no extra dependencies" (the Dockerfile now adds python3-aiohttp), lists only the pytorch and jetson-pypi jobs (no rocm-pypi), and claims PEP 691, but only HTML pages are parsed.

Nit: unlike apt-sync.py / github-release.py, the script doesn't log Total size is ..., which tunasync's size_pattern relies on.

Inline comments below.

Comment thread pip-index.py Outdated


async def main():
timeout_obj = aiohttp.ClientTimeout(total=timeout_sec)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ClientTimeout(total=...) covers the whole request, including reading the body. With the default TIMEOUT=120, a 4 GB ROCm/torch wheel would need more than ~33 MB/s to finish. When it times out, it retries from zero 3 times, and then the exception stops the whole gather, which ends the run.

Suggest aiohttp.ClientTimeout(total=None, sock_connect=30, sock_read=timeout_sec), similar to the (30, 60) connect/read timeouts in apt-sync.py.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 6628b79 — ClientTimeout(total=None, sock_connect=30, sock_read=TIMEOUT): multi-GB wheels are now bounded only by connect/read-idle timeouts instead of a whole-request cap, so no more retry-from-zero-then-abort on slow transfers (same idea as apt-sync.py's (30, 60)).

Comment thread pip-index.py Outdated
except asyncio.CancelledError:
pass
return b"".join(chunks)
except Exception:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This catches ClientResponseError for 403/404 too, so every blocked or yanked URL is retried with two 5 s sleeps before the caller skips it. The same happens in stream_to_file (L286). Both sleeps happen inside async with sem, so with the default JOBS=1 every yanked file stalls the whole crawl for about 10 s.

Suggest re-raising 403/404 right away:

except aiohttp.ClientResponseError as e:
    if e.status in (403, 404) or attempt == 2:
        raise
    ...

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 6628b79 — both get_with_progress and stream_to_file re-raise ClientResponseError for 403/404 immediately, so a blocked/yanked URL no longer burns two semaphore-holding 5s retries before being skipped.

Comment thread pip-index.py Outdated
existing = dest_origins.get(dest)
if existing is None:
dest_origins[dest] = origin
elif existing[1] != origin[1]:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two problems with this guard:

  1. dest_origins only covers the current run and doesn't know which host the file already on disk came from. urls is a set and tasks run concurrently, so whichever host reaches a path first wins. If host B gets there first, dest.exists() below keeps host A's bytes from the previous run, and host A is then logged as the conflicting one.
  2. The index page that links the rejected URL is still rewritten to URLBASE/<path>, so users get the other host's file. The #sha256 fragment only turns this into a hash mismatch for them.

Either fail the run on a collision, or also rewrite or drop the rejected href in the saved index.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 6628b79 — claim_dest now compares the full (scheme, host, path) origin and seed URLs are processed in sorted order, so the collision winner is deterministic. More importantly rewrite_index keeps the losing href unchanged: clients fetch it from the upstream host directly and the #sha256 fragment still matches, so the other host's bytes are never served under a wrong name.

Comment thread pip-index.py
f"{existing[0]}://{existing[1]}{existing[2]}"
)
return
if dest.exists():

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Any existing file is treated as up to date. If upstream re-publishes a file under the same name, the stale copy stays forever, and pip then fails the #sha256= check from the rewritten index. Could you check Content-Length or the sha256 fragment, which is already available in the index?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 6628b79 — new REVALIDATE (default on when CLEANUP=1): existing files get a cheap HEAD check (bounded by JOBS); 404 feeds gone_files so CLEANUP removes the stale copy, a changed Content-Length triggers a redownload, and transient HEAD errors keep the local copy instead of aborting. The request-volume tradeoff is documented in the docstring.

Comment thread pip-index.py Outdated
for m in A_RE.finditer(index_resp):
attr = m.group(1)
href = HREF_RE.search(attr)
assert href is not None, f"Invalid href in {attr}"

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

An <a> without an href (e.g. <a name=...>) or with a single-quoted href crashes the whole run here. The assert also disappears under python -O. Suggest continue when there's no match, and accepting both quote styles in HREF_RE.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 6628b79 — the assert is gone: <a> tags without href are skipped with continue, and both the crawl regex and the rewriter accept single-quoted hrefs.

Comment thread pip-index.py Outdated
href = HREF_RE.search(attr)
assert href is not None, f"Invalid href in {attr}"
suburl = href.group(1).split("#")[0]
if suburl.startswith("/"):

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nit: urljoin(url, suburl) already resolves root-relative and protocol-relative hrefs against url's scheme and host, so this branch and upstream_base can be removed.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 6628b79 — verified: urljoin resolves root-relative and protocol-relative hrefs against the page URL by itself, so the upstream_base branch was redundant and has been removed.

@yaoge123

Copy link
Copy Markdown
Contributor Author

Round-2 review findings addressed in 0e2fb1d (also deployed to NJU production):

  • Index destinations lacked collision protection (pip-index.py:443/:460): the destination (index_dir/filename) is now claimed before fetching via a shared claim_dest() helper used by both artifacts and index pages, so two hosts exposing the same path (e.g. both /simple/) can no longer overwrite one index.html through a shared .tmp; the later URL is logged as an error and skipped. Cleanup's visited→path mapping matches these destinations.
  • 404 files stayed protected from stale cleanup (pip-index.py:484): files that return 404 are now tracked in gone_files and dropped from the cleanup allowlist, so an already-mirrored yanked file is removed as stale instead of being protected via visited forever. 403 remains protected (blocks may be transient). Added a CLEANUP_MAX_404_DELETE cap (default 100): a mass-404 event looks like upstream breakage, not yanking, and keeps the files — same style as CLEANUP_MAX_DELETE.

The older aiohttp runtime dependency finding was addressed in d6f056f (Dockerfile + aiohttp smoke test added).

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Unresolved critical and moderate findings affect mirroring correctness and documented functionality.

Review effort: Lite
Findings: 3 High severity · 1 Medium severity

Open (4)
Resolved since last review (2)
Previously missed (1)

In code that hasn't changed since last review

Medium severity Resolve declared PEP 658/714 metadata URLs directly

pip-index.py:478

The metadata URL declared by PEP 658/714 is ignored here: appending .metadata to the wheel URL is only valid for servers that happen to use that naming convention. If the attribute points to a different relative or absolute URL, the crawler fetches the wrong resource and leaves the declared metadata unavailable locally. Resolve and enqueue the value of data-dist-info-metadata/data-core-metadata against the index URL instead.

Comment thread pip-index.py Outdated
if existing is None:
dest_origins[dest] = origin
return True
if existing[1] != origin[1]:

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 6628b79 — claim_dest now requires an exact (scheme, host, path) origin match, so same-host URLs that normalize to one local path (/foo%2Fbar vs /foo/bar, dot segments) are rejected as conflicts too.

Comment thread pip-index.py Outdated
Comment on lines +222 to +226
logging.error(
f"Skipping {url}: destination {dest} is already mirrored from "
f"{existing[0]}://{existing[1]}{existing[2]}"
)
return False

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 6628b79 — rewrite_index maps each href to its local destination and consults dest_origins; collision losers keep their original href, so the saved index no longer points at another host's artifact under a mismatched #sha256.

Comment thread pip-index.py Outdated
Comment on lines +498 to +499
if dest.exists():
return

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 6628b79 via REVALIDATE — with CLEANUP=1 (the default for it) existing files are HEAD-revalidated, so a yanked file now gets its 404, lands in gone_files, and cleanup removes it; plain incremental runs skip the extra requests.

@yaoge123

Copy link
Copy Markdown
Contributor Author

On the review-body items (all in 6628b79): STRIP_PREFIX was description drift, not a missing commit — no such option ever existed in the script; it has been removed from the PR description. The docstring now lists the rocm-pypi job, says PEP 503 only (only HTML is parsed), and no longer claims "no extra dependencies". The run now ends with Total size is ... (sizeof_fmt from github-release.py) for tunasync's size_pattern.

Some PyPI-style registries only expose a PEP 503 simple index over
HTTP and offer no rsync/bandersnatch-compatible endpoint -- e.g.
NVIDIA's jetson-ai-lab devpi (Jetson wheels) and AMD's
repo.amd.com/rocm/whl (ROCm/TheRock wheels). pip-index.py mirrors such
registries: it crawls the simple HTML index recursively, downloads
referenced files, and rewrites hrefs so saved index pages point back
to this mirror. Adapted from ustclug/ustcmirror-images
pytorch/sync.py; used by the pytorch, rocm-pypi and jetson-pypi
tunasync jobs.

Features on top of the upstream script:

- Rewrite-all href policy: every absolute href is rewritten to
  URLBASE regardless of host (multi-host upstreams need no
  configuration); EXTRA_REWRITES redirects other hosts' hrefs to a
  sibling local mirror prefix without downloading them.
- DEVPI_MODE: crawl only a devpi channel's own projects via its JSON
  API; CUSTOM_ENDPOINTS seeds arbitrary index roots; NO_NIGHTLY
  filters nightly subtrees (also on discovered links).
- asyncio/aiohttp crawler: JOBS semaphore, streaming downloads,
  atomic .tmp writes, ClientTimeout(total=None, sock_connect=30,
  sock_read=TIMEOUT) so multi-GB wheels are bounded only by idle
  time, and 403/404 responses re-raise immediately instead of
  burning retries.
- Cross-host collision guard: destinations stay keyed by URL path
  (preserving the existing data layout), and claim_dest() records
  the full (scheme, host, path) origin of every artifact and index
  page; a conflicting later URL is logged as an error and skipped,
  and rewrite_index leaves the losing href unchanged so clients
  fetch it from its own host with a matching #sha256 fragment.
- CLEANUP (opt-in): deletes stale local files no crawled URL maps
  to, capped by CLEANUP_MAX_DELETE; files that returned 404 are
  treated as yanked and removed too, separately capped by
  CLEANUP_MAX_404_DELETE; subtrees skipped due to 403/404 are never
  deleted. REVALIDATE (default on when CLEANUP=1) HEAD-checks
  existing files so yanked or republished same-name artifacts do not
  stay stale forever.
- PEP 658/714: both data-dist-info-metadata and data-core-metadata
  attributes are followed for .metadata files.
- Ends with "Total size is ..." for tunasync's size_pattern.

Also adds python3-aiohttp to the repo Dockerfile and smoke-tests the
import at image build time and in test/docker_runtime.sh.

Verified: py_compile clean; test/docker_runtime.sh passes against
tunathu/tunasync-scripts:latest (which already ships aiohttp); a
containerized end-to-end crawl against two local test servers
exercising the collision guard, 404 subtree skip, nightly filtering,
.metadata fetching and CLEANUP exclusions. Both mirrors run in
production at mirror.nju.edu.cn: jetson-pypi (since May) and
rocm-pypi (~492G, verified against the upstream file list).
@yaoge123

Copy link
Copy Markdown
Contributor Author

Branch cleanup note: this branch has been squashed to a single commit, 0db31a6. All commit SHAs referenced earlier in the review threads (up to 6628b79) are now orphaned commits — the links still open, but please rely on the current diff, whose tree is byte-identical to the previous head 6628b79.

Review items → how they were addressed (all included in the current diff)

Copilot round 1:

  • aiohttp missing from the runtime image → python3-aiohttp added to the repo Dockerfile, import smoke-tested at image build time and in test/docker_runtime.sh; the deployed tunathu/tunasync-scripts:latest already ships aiohttp 3.14.3.
  • Cross-host download collisions on path-only destinations → initially keyed by netloc (later revised, see below), now a claim_dest() guard over the full (scheme, host, path) origin with the on-disk layout unchanged.
  • 404 index pages aborted the crawl → tolerated exactly like 403 (skip subtree, protected from CLEANUP).
  • NO_NIGHTLY bypassed for discovered links → the filter now runs before recursive tasks are created.
  • Wrong PEP 658 attribute → both data-dist-info-metadata and data-core-metadata (PEP 714) are accepted.

Copilot round 2:

  • Index destinations shared one .tmp across hosts → index pages are claimed through the same claim_dest() guard before fetching.
  • 404 files stayed cleanup-protected forever → tracked in gone_files and removed by CLEANUP, with a separate CLEANUP_MAX_404_DELETE=100 fuse against mass-404 events; 403 stays protected.

Maintainer review (@happyaron):

  • ClientTimeout(total=...) kills multi-GB wheels → total=None, sock_connect=30, sock_read=TIMEOUT (same idea as apt-sync.py's (30, 60)).
  • 403/404 burned two semaphore-holding retries → re-raised immediately in both get_with_progress and stream_to_file.
  • Collision guard was nondeterministic and still rewrote the losing href → seed URLs processed in sorted order, full origin tuple compared, and rewrite_index leaves the losing href unchanged (clients fetch from the original host with a matching #sha256).
  • Existing files never rechecked → REVALIDATE (default on when CLEANUP=1): HEAD-checks existing files; 404 feeds gone_files, changed Content-Length redownloads, transient errors keep the local copy.
  • <a> without href / single-quoted href / assert → assert removed, missing hrefs skipped, both quote styles accepted.
  • Redundant root-relative handling → removed; urljoin covers it.
  • STRIP_PREFIX in the PR body was description drift (never existed in the script) → removed; the docstring now lists the rocm-pypi job, says PEP 503 HTML only, and no longer claims "no extra dependencies". The run now ends with Total size is ... for tunasync's size_pattern.

Copilot round 3:

  • Same-host URLs colliding after path normalization (/foo%2Fbar vs /foo/bar) → claim_dest requires an exact (scheme, host, path) match.
  • Conflicting href still rewritten to the shared local path → losing hrefs are kept unchanged (see above).
  • Existing files never revalidated during cleanup → covered by REVALIDATE.
  • Metadata URL should be resolved from the attribute value rather than assumed to be <wheel>.metadata → not changed: every index mirrored by the current jobs (PyTorch, ROCm, jetson devpi) uses the <url>.metadata convention, so this is left as-is; resolving the attribute value would be more general and can be revisited if a future upstream needs it.

One correction to the record: my reply on the round-1 collision finding said destinations were "keyed by netloc" (26d3305). That approach was abandoned because it changed the on-disk layout and would have orphaned existing mirrored data; the shipped mechanism is the path-only layout plus the claim_dest() origin guard described above (already noted in a top-level comment, repeated here for the thread).

Some earlier in-thread replies described intermediate states of the branch; please rely on the current diff and this summary.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants