[GH-3369] Preserve WKB dimensions in file and adapter readers - #3378
Merged
Merged
Conversation
Signed-off-by: Jia Yu <jiayu@apache.org>
Signed-off-by: Jia Yu <jiayu@apache.org>
This was referenced Sep 18, 2026
jiayuasu
marked this pull request as ready for review
September 21, 2026 06:10
1 task
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Did you read the Contributor Guide?
Yes.
Is this PR related to a ticket?
Follow-up to #3369. Builds on merged #3374 and #3377.
org.datasyslab:jts-io-patch:1.21.0-datasyslab-2is published on Maven Central.WKB output is reviewed separately in #3381.
What changes were proposed in this PR?
Use IO2's
WKBReader.forDeclaredDimensions()in both GeoParquet conversion paths, GeoPackage, the Python/JVM adapter, and lazy JTS geography materialization. Empty and all-NaN layouts then survive Sedona serialization consistently with SQL WKB constructors.Preserve metadata/header SRIDs, Python framing and userData, and geography caching. The separate S2 geography reader is unchanged. This changes ingestion; GeoParquet output is a separate format-aware concern.
How was this patch tested?
Reader regressions cover each integration. The GeoParquet primitive reader also alternates populated NaN-Z and XY points, checking coordinates and dimensions after GeometryUDT serialization.
At
0f26a29afa9583897d59856256631e5dd8087d32, 36 WKBGeography tests and 10 Scala reader/adapter tests passed on Spark 3.5 / Scala 2.12 against the published IO2 artifact. The new NaN-Z test fails with the master GeoParquet reader because dimension 3 becomes 2. Formatting and commit hooks passed.After synchronizing with master and merged #3377, the combined stack at
6881691d278e27bb160ee6be3f4fbd671e9b0192passed 390 focused common geometry/WKB tests against the publishedorg.datasyslab:jts-io-patch:1.21.0-datasyslab-2.Validated the combined Sedona stack at
76f59c947971e06657d01229a4f940863d3e5253againstorg.datasyslab:jts-io-patch:1.21.0-datasyslab-2-SNAPSHOT, built from merged JTS commit1a382cc402f495adf1ea39c4905cafbac491e40f:ST_GeneratePointsstayed XY.The published
1.21.0-datasyslab-2artifact passes 32 isolated-artifact tests against stock JTS 1.20. All nine implementation class files are byte-for-byte identical to the tested snapshot; public downloads and signatures were verified. The full CI matrix and MySQL/Docker constructor suite were not part of these local checks.Did this PR include necessary documentation updates?
Added short notes to the GeoParquet and GeoPackage reader pages explaining dimension preservation and linking to the ST_Collect guidance for mixed layouts.