Skip to content

Release the GVL during routing model setup - #87

Closed
erickreutz wants to merge 1 commit into
ankane:masterfrom
lugg:eric/upstream-routing-setup-gvl
Closed

erickreutz wants to merge 1 commit into
ankane:masterfrom
lugg:eric/upstream-routing-setup-gvl

Conversation

@erickreutz

Copy link
Copy Markdown
Contributor

RoutingModel#close_model and #read_assignment_from_routes can spend significant time in native OR-Tools code, but unlike the existing solve wrappers, they currently hold Ruby's GVL for the entire call. This adds callback-aware wrappers that release the GVL during those operations when the routing model uses only native callbacks.

The safety boundary matches the solve path: models with Ruby transit callbacks retain the GVL, while native vector and matrix callbacks allow other Ruby threads to run. Route arrays are converted to owned C++ values before the GVL is released, and returned assignments are wrapped after Ruby execution resumes. The raw callback, solve, close, and restoration entry points are private so callers cannot bypass the callback gate.

The coverage exercises Ruby thread progress during both native calls, callback gating and exceptions, result parity, route conversion ownership, overlapping reads on separate models, and assignment lifetime. Existing routing results and callback behavior are unchanged.

Allow Ruby threads to keep running while callback-free routing models close and restore route assignments. Keep the GVL when Ruby transit callbacks may execute.
ankane added a commit that referenced this pull request Oct 1, 2026
Co-authored-by: Eric Kreutzer <eric@lugg.com>
@ankane

ankane commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Hi @erickreutz, thanks for the PR. Merged a version of this for read_assignment_from_routes in the commit above, as a comment in the OR-Tools source code indicates "this may take considerable amount of time". Can you share how long close_model is taking?

@erickreutz

Copy link
Copy Markdown
Contributor Author

Thanks for merging that! We checked the last 24 hours in production. Across 10,665 close_model calls, median was 3 ms, p95 7 ms, p99 27 ms, and max 204 ms. Only three hit 100 ms.

The long stalls were in read_assignment_from_routes. One took 78.8 seconds, while close_model on that same model took 5 ms. So we’re not seeing the same issue with close_model.

@ankane

ankane commented Oct 2, 2026

Copy link
Copy Markdown
Owner

Thanks for checking. Will hold off on updating close_model for now.

@ankane ankane closed this Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants