Skip to content

[GH-3369] Preserve empty components when setting SRID - #3377

Merged
jiayuasu merged 5 commits into
apache:masterfrom
jiayuasu:fix/set-srid-empty-components
Sep 21, 2026
Merged

jiayuasu merged 5 commits into
apache:masterfrom
jiayuasu:fix/set-srid-empty-components

Conversation

@jiayuasu

@jiayuasu jiayuasu commented Sep 16, 2026

Copy link
Copy Markdown
Member

Did you read the Contributor Guide?

Yes.

Is this PR related to a ticket?

Prerequisite for #3374 and #3369. Uses the merged JTS copy fix in jiayuasu/jts#10.

What changes were proposed in this PR?

ST_SetSRID currently removes empty members while copying a collection. For example, a collection containing an empty point and a populated point returns only the populated point after changing SRID. Empty polygon holes are also dropped. A standalone empty polygon can be returned as the original object, mutating its SRID rather than creating a copy.

Use the isolated JTS GeometryCopier with the requested SRID and the input's coordinate-sequence factory. This preserves every component and ring, including empty holes, and creates independent coordinate storage. Copies discard userData; the source geometry and its metadata remain unchanged. The copy logic is maintained in JTS, so this PR adds no Sedona geometry factory or serialization classes.

org.datasyslab:jts-io-patch:1.21.0-datasyslab-2 is published on Maven Central. It uses stock JTS 1.20 geometry types and does not require replacing Spark's bundled JTS jar.

How was this patch tested?

The regressions fail against the base implementation and pass with the fix. They cover nested empty components, all three multipart types, empty polygon holes, empty-polygon input mutation, XYM/XYZ/XYZM layouts, packed sequences, geometry and factory SRIDs, independent copies, and source/userData preservation. A Spark SQL regression checks the empty-hole count before and after ST_SetSRID.

At this PR head, 337 common FunctionsTest cases passed against the snapshot.

Validated the combined Sedona stack at 76f59c947971e06657d01229a4f940863d3e5253 against org.datasyslab:jts-io-patch:1.21.0-datasyslab-2-SNAPSHOT, built from merged JTS commit 1a382cc402f495adf1ea39c4905cafbac491e40f:

  • Full common suite: 1,396 passed.
  • Spark 3.5.0/Scala 2.12 and Spark 4.1.1/Scala 2.13: 324 selected Java and Scala tests passed per profile, covering SQL functions, UDT, collection and reader suites.
  • Stock PySpark 3.5.0 and 4.1.1: 96 native/fallback constructor tests, 144 primitive WKB/EWKB checks, and 18 multipart SQL round-trip checks passed.
  • Spark 4.1 checks exercised two separate executor JVMs. Its bundled JTS 1.20 jar remained unchanged, and ST_GeneratePoints stayed XY.

The published 1.21.0-datasyslab-2 artifact passes 32 isolated-artifact tests against stock JTS 1.20. All nine implementation class files are byte-for-byte identical to the tested snapshot; public downloads and signatures were verified. The full CI matrix and MySQL/Docker constructor suite were not part of these local checks.

Did this PR include necessary documentation updates?

No public SQL API or configuration was added. Changing SRID retains the input's geometry structure.

@jiayuasu jiayuasu added this to the sedona-2.0.0 milestone Sep 21, 2026
@jiayuasu jiayuasu linked an issue Sep 21, 2026 that may be closed by this pull request
1 task
@jiayuasu
jiayuasu merged commit 6cd7c01 into apache:master Sep 21, 2026
43 of 59 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

GeoPandas: leading-null GeoSeries input loses Z dimensions

1 participant