Skip to content

perf(manifest): prune entries by row range and cache decoded Arrow IPC - #411

Open
gripleaf wants to merge 4 commits into
apache:mainfrom
gripleaf:perf/manifest-row-range-cache
Open

gripleaf wants to merge 4 commits into
apache:mainfrom
gripleaf:perf/manifest-row-range-cache

Conversation

@gripleaf

@gripleaf gripleaf commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Linked issue: none (performance improvement).

Data-evolution scans with row ID ranges currently construct manifest entries and file metadata before discarding non-overlapping files. Check aligned Arrow row-count and first-row-ID columns before materialization, preserving the existing data-evolution and lazy-decode gates, version validation, conservative retention of unknown or overflowing ranges, Add/Delete selection, entry filters, and merge logic.

Change the shared ObjectsFile<T> cache from source-file bytes to decoded Arrow IPC, covering ManifestFile, ManifestList, and IndexManifestFile. Ordinary, bucket, and row-range reads share a complete, query-independent entry under the existing whole-file CacheKind::MANIFEST key. Cache hits skip the source decoder and reuse batches aligned before serialization, preserving physical types such as ORC timestamp units.

The loading reader consumes the original decoded batches. The cache value retains its IPC buffer and allocator, including during eviction with active readers. Cache failures preserve source-read results; source and filter errors remain visible. Cache hit/miss statistics and concurrent load coordination belong to the caller-provided Cache implementation.

A cold cache load decodes the complete file, so a cold bucket read may decode more entries than an uncached selective read.

Tests

Latest documentation and scan-test update:

  • cmake --build build --target paimon-core-test -j 16: passed.
  • ./build/debug/paimon-core-test --gtest_filter='RowRangeManifestFileTest.ScanPlanPreservesResultsAcrossLazyDecodeAndCacheModes': passed; covers 56 scan plans.
  • ./build/debug/paimon-core-test --gtest_filter='*Manifest*:*FileStoreScan*:*RowRange*:*DataEvolution*': 162 passed.
  • pre-commit run --files docs/source/user_guide/manifest_cache.rst src/paimon/core/manifest/manifest_row_range_test.cpp and git diff --check: passed.

Full-suite validation before this documentation/test-only update, on df2b29c8:

  • ./build/debug/paimon-core-test: 2,285 passed.
  • ./build/debug/paimon-data-evolution-table-test: 232 passed; two existing Avro statistics cases skipped.

Validation was local Debug validation. Full suites were not rerun for the latest documentation/test-only update. Remote CI, ARM64, and sanitizer builds have not been verified for this update.

Coverage includes row-range boundaries, conservative retention, Add/Delete merging, schema evolution and version errors on both cold and warm reads, ORC metadata, empty manifests, cross-reader and cross-mode cache reuse, warm concurrent reads, eviction, allocator lifetime, and cache/filter failures. Manifest-list and index-manifest tests verify IPC reuse and invalidation; bucket tests cover cached and uncached source reads.

The scan regression enters through DataEvolutionFileStoreScan::CreatePlan() using real manifest files and lists. It compares lazy-decode and cache modes across seven range selections, including empty and absent ranges, verifies cross-manifest Add/Delete merging, and checks that warm queries reuse complete cache entries. Level-filter callback counts verify that the lazy-decode switch controls early pruning.

API and Format

No public API, on-disk format, or protocol changes. The in-memory manifest cache payload changes from source-file bytes to Arrow IPC while retaining existing Cache, CacheValue, and cache-key interfaces. No new configuration or scan metrics are introduced.

Documentation

Update the manifest cache guide to describe Arrow IPC payloads, cold and warm read behavior, and query-independent cache reuse. Clarify that cache statistics and concurrent load coordination belong to the external Cache implementation. No new configuration or metrics are introduced.

Generative AI tooling

Generated-by: OpenAI Codex (GPT-6)

@gripleaf

gripleaf commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor Author

Scenario: An application resolves keys to row IDs, then fetches full rows from a Paimon table on HDFS. The fixed reference subset contains 100,000 keys and 42 columns. We tested single-key lookup, batched lookup, and paginated key-range scan with row-ID fetches. These are end-to-end application measurements using a native HDFS integration, including network and application overhead.

Workload Concurrency Keys/request or rows/page Warm-up / measurement
Lookup 16 1 10 s / 60 s
MultiLookup 8 1,024 10 s / 60 s
Range scan 8 4,096 5 s / 30 s

Two warm runs per version; closed-loop load, identical reference keys, and 32 manifest scan threads. Lookup workloads sample the reference keys; scan covers their minimum-to-maximum key range.

Workload AVG ms: baseline → PR P99 ms: baseline → PR Returned-row throughput
Lookup 923.44 → 28.37 2,193.62 → 55.22 33.16×
MultiLookup 3,423.26 → 489.46 4,110.96 → 574.94 7.10×
Range scan 5,438.25 → 2,426.60 7,489.55 → 3,718.29 2.24×

Latency is per whole request/page, not per row. AVG and nearest-rank P99 pool raw samples across both runs. Baseline/PR sample counts: 2,067/67,568 (lookup), 280/1,952 (batched lookup), and 88/193 (scan); scan tail estimates are therefore limited.

Validation: Reference-value checks passed (1,000 sampled lookup keys; all 100,000 keys through batched lookup and scan, plus duplicate-key cases). Benchmark requests had zero errors. The integration build passed 152 focused tests and the core suite with 2,334 passed / 100 skipped.

AVG decreased by 96.9%, 85.7%, and 55.4%, respectively. This measures the combined PR effect, not caching alone. Background errors overlapped candidate run 1; run 2 had none and showed similar performance, but shared-cluster interference cannot be ruled out.

@gripleaf
gripleaf marked this pull request as ready for review October 8, 2026 06:17
@gripleaf gripleaf changed the title perf(manifest): prune row ranges and cache aligned batches perf(manifest): prune entries by row range and cache decoded Arrow IPC Oct 8, 2026
Filter aligned manifest rows by row ID range before constructing file
metadata, while preserving version checks and the existing entry filters.

Replace raw manifest byte caching with a shared Arrow IPC representation
in ObjectsFile for manifest files, manifest lists, and index manifests.
Reuse whole-file cache keys, coalesce concurrent loads, and preserve
allocator ownership and per-reader filtering.

Add regression coverage and metrics for pruning, cache reuse, schema
compatibility, concurrent readers, eviction, and error handling. Document
the full-file decoding cost of cold cached reads.
@gripleaf
gripleaf force-pushed the perf/manifest-row-range-cache branch from ce4c6d4 to 4f32e09 Compare October 8, 2026 06:32
Comment thread src/paimon/core/utils/objects_file.h Outdated
Comment thread include/paimon/table/source/scan_metrics.h Outdated
Delegate concurrent load coordination to Cache::Get instead of keeping a
process-wide pending-load registry in ObjectsFile. Remove the dedicated
cold-load coalescing test while retaining concurrent cache-read coverage.

Keep only the decoded manifest cache hit and miss counters. Remove the
row-range work counters, per-file histogram, and cache fallback counter,
and trim the metrics documentation to the exported scan metrics. Preserve
cache-error fallback and verify metric snapshot isolation.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants