Skip to content

Allow uint64-backed FreeableBuffers larger than size_t (#23364) - #23364

Open
cmt0 wants to merge 1 commit into
pytorch:mainfrom
cmt0:export-D122398972
Open

cmt0 wants to merge 1 commit into
pytorch:mainfrom
cmt0:export-D122398972

Conversation

@cmt0

@cmt0 cmt0 commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary:

Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments.

The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match.

On a 32-bit core:

  • 4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported.

Review follow-up: separate LegacyUInt64Data callback storage and preserve implicit conversion of legacy callbacks, including captureless lambdas. The legacy callback alias and constructor remain deprecated. Added move/free and 32-bit bounds regression coverage. Embedded consumers retain C++17 support.

Review follow-up implemented with Codex.

Reviewed By: rascani

Differential Revision: D122398972

@cmt0
cmt0 requested a review from JacobSzwejbka as a code owner October 2, 2026 16:46
@pytorch-bot

pytorch-bot Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23364

Note: Links to docs will display an error until the docs builds have been completed.

⏳ 11 Pending, 2 Unrelated Failures

As of commit 7bb8e5a with merge base 8e4ec35 (image):

FLAKY - The following job failed but was likely due to flakiness present on trunk:

BROKEN TRUNK - The following job failed but was present on the merge base:

👉 Rebase onto the `viable/strict` branch to avoid these failures

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Oct 2, 2026
@meta-codesync

meta-codesync Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

@cmt0 has exported this pull request. If you are a Meta employee, you can view the originating Diff in D122398972.

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

cmt0 added a commit to cmt0/executorch that referenced this pull request Oct 2, 2026
Summary:
Pull Request resolved: pytorch#23364

Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments.

The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match.

On a 32-bit core:
- >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported.

Differential Revision: D122398972
@cmt0
cmt0 force-pushed the export-D122398972 branch from 1a5be09 to 5e14c96 Compare October 2, 2026 16:52
@meta-codesync meta-codesync Bot changed the title Allow uint64-backed FreeableBuffers larger than size_t Allow uint64-backed FreeableBuffers larger than size_t (#23364) Oct 2, 2026
@cmt0
cmt0 force-pushed the export-D122398972 branch from 5e14c96 to 805a6bf Compare October 6, 2026 19:55
cmt0 added a commit to cmt0/executorch that referenced this pull request Oct 6, 2026
Summary:

Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments.

The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match.

On a 32-bit core:
- >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported.

Differential Revision: D122398972
@cmt0
cmt0 force-pushed the export-D122398972 branch from 805a6bf to 4162f7c Compare October 7, 2026 20:38
cmt0 added a commit to cmt0/executorch that referenced this pull request Oct 7, 2026
Summary:

Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments.

The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match.

On a 32-bit core:
- >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported.

Differential Revision: D122398972
@cmt0
cmt0 force-pushed the export-D122398972 branch from 4162f7c to 4af5a92 Compare October 8, 2026 17:18
cmt0 added a commit to cmt0/executorch that referenced this pull request Oct 8, 2026
Summary:
Pull Request resolved: pytorch#23364

Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments.

The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match.

On a 32-bit core:
- >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported.

Review follow-up: separate LegacyUInt64Data callback storage and preserve implicit conversion of legacy callbacks, including captureless lambdas. The legacy callback alias and constructor remain deprecated. Added move/free and 32-bit bounds regression coverage. Embedded consumers retain C++17 support.

Review follow-up implemented with Codex.

Differential Revision: D122398972

@rascani rascani left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review automatically exported from Phabricator review in Meta.

cmt0 added a commit to cmt0/executorch that referenced this pull request Oct 8, 2026
Summary:
Pull Request resolved: pytorch#23364

Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments.

The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match.

On a 32-bit core:
- >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported.

Review follow-up: separate LegacyUInt64Data callback storage and preserve implicit conversion of legacy callbacks, including captureless lambdas. The legacy callback alias and constructor remain deprecated. Added move/free and 32-bit bounds regression coverage. Embedded consumers retain C++17 support.

Review follow-up implemented with Codex.

Reviewed By: rascani

Differential Revision: D122398972
@cmt0
cmt0 force-pushed the export-D122398972 branch from 4af5a92 to f1c41d0 Compare October 8, 2026 18:11
cmt0 added a commit to cmt0/executorch that referenced this pull request Oct 8, 2026
Summary:
Pull Request resolved: pytorch#23364

Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments.

The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match.

On a 32-bit core:
- >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported.

Review follow-up: separate LegacyUInt64Data callback storage and preserve implicit conversion of legacy callbacks, including captureless lambdas. The legacy callback alias and constructor remain deprecated. Added move/free and 32-bit bounds regression coverage. Embedded consumers retain C++17 support.

Review follow-up implemented with Codex.

Reviewed By: rascani

Differential Revision: D122398972
@cmt0
cmt0 force-pushed the export-D122398972 branch from f1c41d0 to 8f568c6 Compare October 8, 2026 18:20
cmt0 added a commit to cmt0/executorch that referenced this pull request Oct 8, 2026
Summary:

Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments.

The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match.

On a 32-bit core:
- >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported.

Review follow-up: separate LegacyUInt64Data callback storage and preserve implicit conversion of legacy callbacks, including captureless lambdas. The legacy callback alias and constructor remain deprecated. Added move/free and 32-bit bounds regression coverage. Embedded consumers retain C++17 support.

Review follow-up implemented with Codex.

Reviewed By: rascani

Differential Revision: D122398972
@cmt0
cmt0 force-pushed the export-D122398972 branch from 8f568c6 to 73f8ab0 Compare October 8, 2026 18:24
cmt0 added a commit to cmt0/executorch that referenced this pull request Oct 8, 2026
Summary:
Pull Request resolved: pytorch#23364

Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments.

The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match.

On a 32-bit core:
- >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported.

Review follow-up: separate LegacyUInt64Data callback storage and preserve implicit conversion of legacy callbacks, including captureless lambdas. The legacy callback alias and constructor remain deprecated. Added move/free and 32-bit bounds regression coverage. Embedded consumers retain C++17 support.

Review follow-up implemented with Codex.

Reviewed By: rascani

Differential Revision: D122398972
@cmt0
cmt0 force-pushed the export-D122398972 branch from 73f8ab0 to 7afe1b2 Compare October 8, 2026 18:24
cmt0 added a commit to cmt0/executorch that referenced this pull request Oct 8, 2026
Summary:
Pull Request resolved: pytorch#23364

Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments.

The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match.

On a 32-bit core:
- >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported.

Review follow-up: separate LegacyUInt64Data callback storage and preserve implicit conversion of legacy callbacks, including captureless lambdas. The legacy callback alias and constructor remain deprecated. Added move/free and 32-bit bounds regression coverage. Embedded consumers retain C++17 support.

Review follow-up implemented with Codex.

Reviewed By: rascani

Differential Revision: D122398972
@cmt0
cmt0 force-pushed the export-D122398972 branch from 7afe1b2 to 967b0c2 Compare October 8, 2026 18:29
Summary:
Pull Request resolved: pytorch#23364

Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments.

The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match.

On a 32-bit core:
- >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported.

Review follow-up: separate LegacyUInt64Data callback storage and preserve implicit conversion of legacy callbacks, including captureless lambdas. The legacy callback alias and constructor remain deprecated. Added move/free and 32-bit bounds regression coverage. Embedded consumers retain C++17 support.

Review follow-up implemented with Codex.

Reviewed By: rascani

Differential Revision: D122398972
@cmt0
cmt0 force-pushed the export-D122398972 branch from 967b0c2 to 7bb8e5a Compare October 8, 2026 18:34

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants