Repository navigation
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23364
Note: Links to docs will display an error until the docs builds have been completed. ⏳ 11 Pending, 2 Unrelated FailuresAs of commit 7bb8e5a with merge base 8e4ec35 ( FLAKY - The following job failed but was likely due to flakiness present on trunk:
BROKEN TRUNK - The following job failed but was present on the merge base:👉 Rebase onto the `viable/strict` branch to avoid these failures
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
|
@cmt0 has exported this pull request. If you are a Meta employee, you can view the originating Diff in D122398972. |
This PR needs a
|
Summary: Pull Request resolved: pytorch#23364 Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments. The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match. On a 32-bit core: - >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported. Differential Revision: D122398972
1a5be09 to
5e14c96
Compare
5e14c96 to
805a6bf
Compare
Summary: Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments. The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match. On a 32-bit core: - >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported. Differential Revision: D122398972
805a6bf to
4162f7c
Compare
Summary: Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments. The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match. On a 32-bit core: - >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported. Differential Revision: D122398972
4162f7c to
4af5a92
Compare
Summary: Pull Request resolved: pytorch#23364 Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments. The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match. On a 32-bit core: - >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported. Review follow-up: separate LegacyUInt64Data callback storage and preserve implicit conversion of legacy callbacks, including captureless lambdas. The legacy callback alias and constructor remain deprecated. Added move/free and 32-bit bounds regression coverage. Embedded consumers retain C++17 support. Review follow-up implemented with Codex. Differential Revision: D122398972
rascani
left a comment
There was a problem hiding this comment.
Review automatically exported from Phabricator review in Meta.
Summary: Pull Request resolved: pytorch#23364 Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments. The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match. On a 32-bit core: - >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported. Review follow-up: separate LegacyUInt64Data callback storage and preserve implicit conversion of legacy callbacks, including captureless lambdas. The legacy callback alias and constructor remain deprecated. Added move/free and 32-bit bounds regression coverage. Embedded consumers retain C++17 support. Review follow-up implemented with Codex. Reviewed By: rascani Differential Revision: D122398972
4af5a92 to
f1c41d0
Compare
Summary: Pull Request resolved: pytorch#23364 Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments. The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match. On a 32-bit core: - >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported. Review follow-up: separate LegacyUInt64Data callback storage and preserve implicit conversion of legacy callbacks, including captureless lambdas. The legacy callback alias and constructor remain deprecated. Added move/free and 32-bit bounds regression coverage. Embedded consumers retain C++17 support. Review follow-up implemented with Codex. Reviewed By: rascani Differential Revision: D122398972
f1c41d0 to
8f568c6
Compare
Summary: Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments. The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match. On a 32-bit core: - >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported. Review follow-up: separate LegacyUInt64Data callback storage and preserve implicit conversion of legacy callbacks, including captureless lambdas. The legacy callback alias and constructor remain deprecated. Added move/free and 32-bit bounds regression coverage. Embedded consumers retain C++17 support. Review follow-up implemented with Codex. Reviewed By: rascani Differential Revision: D122398972
8f568c6 to
73f8ab0
Compare
Summary: Pull Request resolved: pytorch#23364 Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments. The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match. On a 32-bit core: - >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported. Review follow-up: separate LegacyUInt64Data callback storage and preserve implicit conversion of legacy callbacks, including captureless lambdas. The legacy callback alias and constructor remain deprecated. Added move/free and 32-bit bounds regression coverage. Embedded consumers retain C++17 support. Review follow-up implemented with Codex. Reviewed By: rascani Differential Revision: D122398972
73f8ab0 to
7afe1b2
Compare
Summary: Pull Request resolved: pytorch#23364 Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments. The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match. On a 32-bit core: - >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported. Review follow-up: separate LegacyUInt64Data callback storage and preserve implicit conversion of legacy callbacks, including captureless lambdas. The legacy callback alias and constructor remain deprecated. Added move/free and 32-bit bounds regression coverage. Embedded consumers retain C++17 support. Review follow-up implemented with Codex. Reviewed By: rascani Differential Revision: D122398972
7afe1b2 to
967b0c2
Compare
Summary: Pull Request resolved: pytorch#23364 Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments. The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match. On a 32-bit core: - >4 GiB weight via NamedDataMap::get_data: supported when the loader returns a uint64-backed buffer. The CPU only forwards the 64-bit address and size to a device, e.g. the TCE weight blob. CPU-pointer loaders still return NotSupported. Review follow-up: separate LegacyUInt64Data callback storage and preserve implicit conversion of legacy callbacks, including captureless lambdas. The legacy callback alias and constructor remain deprecated. Added move/free and 32-bit bounds regression coverage. Embedded consumers retain C++17 support. Review follow-up implemented with Codex. Reviewed By: rascani Differential Revision: D122398972
967b0c2 to
7bb8e5a
Compare
Summary:
Follow-on to D120705447, which made DataLoader source offsets 64-bit but kept load sizes size_t, capping a named-data resource at 4 GiB on 32-bit targets. Addresses the review feedback there that weights over 4 GiB should not have to be split into multiple segments.
The size is now 64-bit along the named-data path. DataLoader::load_at_offset takes a uint64_t size, FlatTensorDataMap::get_data no longer rejects segments over SIZE_MAX, and FreeableBuffer stores a uint64_t size exposed via size_uint64(). size() aborts if the value doesn't fit, as data() already does for uint64-backed buffers. DualAddressDataLoader and TCEBackend are updated to match.
On a 32-bit core:
Review follow-up: separate LegacyUInt64Data callback storage and preserve implicit conversion of legacy callbacks, including captureless lambdas. The legacy callback alias and constructor remain deprecated. Added move/free and 32-bit bounds regression coverage. Embedded consumers retain C++17 support.
Review follow-up implemented with Codex.
Reviewed By: rascani
Differential Revision: D122398972