Serve explore map tiles unfiltered, and fix mvt tile clustering - #1279
Merged
CollinBeczak merged 3 commits intoSep 15, 2026
Merged
CollinBeczak merged 3 commits into
CollinBeczak merged 3 commits into
Conversation
Two changes to the explore map's tile pipeline that only make sense together. The map takes no filters. Difficulty, global and keywords narrow the grid and list views; the map shows every task that is available work, so a tile is a pure function of (z, x, y) and nothing in it depends on a parameter or on who is asking. Those three come off the MVT route, the controller, the service and the repository, and the on-the-fly keyword binning path goes with them: with no filter to apply there is nothing a tile below z=12 needs from `tasks`, so all of them are answered from the pre-computed pyramid and only z=12 still reads the base tables. One eligibility fragment is left in the repository (`AVAILABLE_WORK`), shared with the two SQL rebuild functions, so three code paths agree on what the map shows instead of four. `counts_by_filter` drops off `tile_cells` (evolution 123) and the challenge dirty-marking trigger narrows to the four eligibility columns: `difficulty` and `is_global` decided which bucket a task counted into, so changing either used to invalidate every cell holding one of the challenge's tasks. Without buckets neither affects a cell's contents at all. Dropping the column needs no recompute for correctness -- `task_count` was always COUNT(*) over the same rows the buckets partitioned, so every remaining column already holds exactly what an unfiltered tile reads. The evolution rebuilds anyway, for space: DROP COLUMN only marks the attribute dead, and existing rows keep carrying its bytes until the table is rewritten. Then marker placement. The old separation pass chained `ST_ClusterDBSCAN(eps, minpoints = 1)`, which is single-linkage clustering: A merges with B, B with C, and the cluster walks across the tile a marker at a time. Dense regions are exactly where every centroid has a neighbour within eps, so the walk never stopped -- on a 166k-task database the tile covering South Africa emitted *one* marker at every display zoom from 2 to 6, and at z=5 its member cells spanned 1005 x 1202 km. k-means is now asked for as many clusters as the tile has occupied CANDIDATE_PITCH_PX grid squares, so cluster count tracks how much room the data takes up on screen and zooming out consolidates markers for real. The candidates are then thinned by a walk that keeps one only when no already-kept marker sits within MIN_SEPARATION_PX of it -- 64 CSS px, against frontend cluster bubbles up to 54 px across -- folding a displaced candidate's count into the nearest kept marker. The walk cannot chain: it compares each candidate against the markers already kept, never against the ones those markers displaced, so the separation is a floor on the distance between drawn markers and a ceiling on how far an absorbed count sits from the marker reporting it. That same tile now emits 4 markers at z=3, 29 at z=5 and 37 at z=11, no two closer than 64 px. Sizing `k` on a grid at half the separation distance is deliberate: offering k-means more candidate positions than can survive the thinning lets it place them on the data, and the walk keeps whichever fit. Sizing candidates on the separation grid itself cost about a quarter of the markers for no change in spacing. `TILE_SIZE_PX` also goes from 256 to 512. MapLibre vector sources default to `tileSize: 512`, so the frontend requests zoom `floor(map zoom)` and draws one of these tiles across 512 CSS pixels -- every pixel figure measured against 256 meant twice as much ground as it claimed.
CollinBeczak
marked this pull request as ready for review
September 9, 2026 19:04
|
CollinBeczak
deleted the
Serve-explore-map-tiles-unfiltered-with-markers-that-never-stack
branch
September 15, 2026 20:06
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.


Split from #1268
dependent on: #1270
Two changes to the explore map's tile pipeline that only make sense together.
The map takes no filters. Difficulty, global and keywords narrow the grid and list views; the map shows every task that is available work, so a tile is a pure function of (z, x, y) and nothing in it depends on a parameter or on who is asking. Those three come off the MVT route, the controller, the service and the repository, and the on-the-fly keyword binning path goes with them: with no filter to apply there is nothing a tile below z=12 needs from
tasks, so all of them are answered from the pre-computed pyramid and only z=12 still reads the base tables. One eligibility fragment is left in the repository (AVAILABLE_WORK), shared with the two SQL rebuild functions, so three code paths agree on what the map shows instead of four.counts_by_filterdrops offtile_cells(evolution 123) and the challenge dirty-marking trigger narrows to the four eligibility columns:difficultyandis_globaldecided which bucket a task counted into, so changing either used to invalidate every cell holding one of the challenge's tasks. Without buckets neither affects a cell's contents at all.Dropping the column needs no recompute for correctness --
task_countwas always COUNT(*) over the same rows the buckets partitioned, so every remaining column already holds exactly what an unfiltered tile reads. The evolution rebuilds anyway, for space: DROP COLUMN only marks the attribute dead, and existing rows keep carrying its bytes until the table is rewritten.Then marker placement. The old separation pass chained
ST_ClusterDBSCAN(eps, minpoints = 1), which is single-linkage clustering: A merges with B, B with C, and the cluster walks across the tile a marker at a time. Dense regions are exactly where every centroid has a neighbour within eps, so the walk never stopped -- on a 166k-task database the tile covering South Africa emitted one marker at every display zoom from 2 to 6, and at z=5 its member cells spanned 1005 x 1202 km.k-means is now asked for as many clusters as the tile has occupied CANDIDATE_PITCH_PX grid squares, so cluster count tracks how much room the data takes up on screen and zooming out consolidates markers for real. The candidates are then thinned by a walk that keeps one only when no already-kept marker sits within MIN_SEPARATION_PX of it -- 64 CSS px, against frontend cluster bubbles up to 54 px across -- folding a displaced candidate's count into the nearest kept marker. The walk cannot chain: it compares each candidate against the markers already kept, never against the ones those markers displaced, so the separation is a floor on the distance between drawn markers and a ceiling on how far an absorbed count sits from the marker reporting it. That same tile now emits 4 markers at z=3, 29 at z=5 and 37 at z=11, no two closer than 64 px.
Sizing
kon a grid at half the separation distance is deliberate: offering k-means more candidate positions than can survive the thinning lets it place them on the data, and the walk keeps whichever fit. Sizing candidates on the separation grid itself cost about a quarter of the markers for no change in spacing.TILE_SIZE_PXalso goes from 256 to 512. MapLibre vector sources default totileSize: 512, so the frontend requests zoomfloor(map zoom)and draws one of these tiles across 512 CSS pixels -- every pixel figure measured against 256 meant twice as much ground as it claimed.