Search before asking
Paimon-cpp version
main (ded7cfb)
Minimal reproduce step
-
Create a primary key table with one bucket and commit three times (snapshots 1 to 3), for example (k1, v1), (k2, v2), then (k3, v3), then (k1, v1b).
-
Scan it with a Cache set on the ScanContext and scan.manifest-entry-cache.max-snapshots = 2, so the live manifest entries of snapshot 3 are cached for bucket 0.
-
Roll the table back and commit twice more, for example from Flink:
CALL sys.rollback_to(`table` => 'default.t', snapshot_id => 1);
INSERT INTO t VALUES ('k4', 'v4');
INSERT INTO t VALUES ('k1', 'v1c');
The rollback deletes snapshot-2, snapshot-3 and the data files only they referenced; the two inserts create a new snapshot-2 and snapshot-3.
-
Scan again through the same Cache.
What doesn't meet your expectations?
The second scan plans from the cached entries of the deleted snapshot 3. Reading k1 or k3 fails with Not exist: File .../bucket-0/data-....parquet not exists, k4 returns no row although it was committed, and only k2 is correct. The table recovers once a fourth snapshot is committed and the cached ids fall behind.
FileStoreScan::ReadManifestEntriesWithCache treats cached->snapshot_id == snapshot.Id() as a hit (file_store_scan.cpp:343). Nothing else identifies the snapshot, so a snapshot id that is reused after a rollback (or after dropping and recreating a table at the same path) matches entries that belong to a different snapshot. With the default max-snapshots = 0 the cache is off; any deployment that enables it for point lookups is exposed.
Anything else?
Storing the snapshot's delta manifest list name next to the cached entries and requiring it to match on a hit is enough: the name carries a UUID, so a rewritten snapshot never matches, and a mismatch falls through to the existing rebuild path.
Are you willing to submit a PR?
Search before asking
Paimon-cpp version
main (ded7cfb)
Minimal reproduce step
Create a primary key table with one bucket and commit three times (snapshots 1 to 3), for example
(k1, v1), (k2, v2), then(k3, v3), then(k1, v1b).Scan it with a
Cacheset on theScanContextandscan.manifest-entry-cache.max-snapshots = 2, so the live manifest entries of snapshot 3 are cached for bucket 0.Roll the table back and commit twice more, for example from Flink:
The rollback deletes
snapshot-2,snapshot-3and the data files only they referenced; the two inserts create a newsnapshot-2andsnapshot-3.Scan again through the same
Cache.What doesn't meet your expectations?
The second scan plans from the cached entries of the deleted snapshot 3. Reading
k1ork3fails withNot exist: File .../bucket-0/data-....parquet not exists,k4returns no row although it was committed, and onlyk2is correct. The table recovers once a fourth snapshot is committed and the cached ids fall behind.FileStoreScan::ReadManifestEntriesWithCachetreatscached->snapshot_id == snapshot.Id()as a hit (file_store_scan.cpp:343). Nothing else identifies the snapshot, so a snapshot id that is reused after a rollback (or after dropping and recreating a table at the same path) matches entries that belong to a different snapshot. With the defaultmax-snapshots = 0the cache is off; any deployment that enables it for point lookups is exposed.Anything else?
Storing the snapshot's delta manifest list name next to the cached entries and requiring it to match on a hit is enough: the name carries a UUID, so a rewritten snapshot never matches, and a mismatch falls through to the existing rebuild path.
Are you willing to submit a PR?