Repository navigation
Release the GIL during evaluation, and decide on free-threaded (cp314t) support #45
Description
Activity
Direction, after measuring
Two independent prototypes were built and benchmarked (release builds, 4 cores, Python 3.11; full test suite passes identically with the GIL released), plus a source-level read of PyO3 0.29.2 and a survey of the ecosystem. The proposal in the issue body was wrong in two places, and the corrected direction has two independent tracks.
Finding 1: free-threaded support is already shipping, untested
PyO3 0.28 made
gil_used = falsethe default, so the bare#[pymodule]insrc/lib.rshas declared this module GIL-free on 3.13t+ since the 0.7.0 upgrade. And because the manylinux images carrypython3.14t/python3.15tand maturin ≥ 1.14's--find-interpreterpicks them up, 0.10.0 already publishescp314tandcp315twheels for every Linux architecture (check the file list on PyPI). Only macOS and Windows lack them, because those runners have no free-threaded interpreter installed. No CI leg has ever run the tests underPYTHON_GIL=0.Ecosystem context: six of eight comparable Rust extensions ship free-threaded wheels (pydantic-core, cryptography, rpds-py, jiter, watchfiles, tiktoken); the two that don't are blocked by being abi3-only. cibuildwheel and maturin both build them by default now. PyO3 0.29 dropped 3.13t, so
cp314t/cp315tis the whole target. CPython 3.15 (final 2026-10-01) addsabi3t(PEP 803); abi3 wheels cannot be installed on a free-threaded interpreter before that, which is relevant to #46: switching to abi3 would silently remove the free-threaded wheels we ship today.Finding 2: releasing the GIL around
execute()is a throughput cliff, not a 50 ns taxUncontended, a detach/attach round trip costs ~47–75 ns. Under contention it is a different regime: a thread re-acquiring the GIL waits for whoever grabbed it meanwhile, ~14 µs at 4 threads. For sub-microsecond evaluations that is the whole cost.
4 threads, exec/s GIL held (today) detach always detach unless callbacks detach if context ≥ 1024 elements x + y3,850,000 64,000 69,000 3,850,000 x + y, context also holds a 20k list3,330,000 — — 64,000 policy expression (1.3 µs) 721,000 46,000 49,000 711,000 filter(...).size()over 512 items (3.5 ms)266 1,068 1,068 268 filter/map/sizeover 2k items (63 ms)12 (0.97×) 49 (3.99×) 47 (3.88×) 49 (3.92×) Break-even, measured by sweeping list length with
items.filter(i, i % 3 == 0).size(): detaching starts winning at ~7–8 µs of work per call, reaches 2× at ~13 µs, and saturates at the core count (~3.9×) from ~65 µs. One octave either side of the crossing is the difference between 0.34× and 1.5×. Single-thread benchmarks give no warning of any of this, which is why the issue body's "measure the fixed cost" plan would have shipped a 55× regression for the most common workload.Both cheap gates fail in both directions. Gating on "no Python callbacks" still detaches
x + y. Gating on context size detachesx + ywhen a big list happens to be in the context (52× worse) and fails to detach a 3.5 ms filter over a 512-element list (misses a 3.9× win). The question that matters is "how much work will this call do", and neither proxy answers it.Two things do work:
- Detaching around the parse is unconditionally right.
Program::compileis pure Rust and costs 14–72 µs (98% ofevaluate()for a policy-sized expression), well past break-even.evaluate()goes from 0.94× to 2.6× at 4 threads with no downside. - An adaptive per-
Programgate. Time the attached execution for the first 3 calls (excludingprepare_environment, whose first call may build the cel context); if any sample exceeds ~16 µs (≈2× the crossing), latch thatPrograminto detaching. Prototyped with anAtomicU8onPyProgram, relaxed ordering:
4 threads GIL held adaptive x + y3,850,000 3,690,000 x + y, 20k list in context3,330,000 3,580,000 filter over 512 items 266 1,061 (3.87×) filter/map/size over 2k 12 50 (3.98×) Steady-state overhead is one relaxed atomic load (0.242 vs 0.240 µs single-thread); sampling adds at most ~0.5 µs to each of the first three calls. Known limits: a program whose cost varies wildly across contexts latches on its first samples, and
evaluate()has no persistent object to learn on, so its execute stays attached (its cost is the parse anyway).Semantics, verified rather than argued
Under detach: callbacks re-attach via
Python::attach(correct in PyO3'sSuspendAttachdesign, verified against the source); exceptions from callbacks propagate with identical messages; a callback mutating its ownContextsees the same snapshot semantics as today (theArcfrom #51 is what makes this work); re-entrantevaluate()from a callback works;pyo3_logre-attaches; panics are caught inside the detached region; Ctrl-C behaviour is unchanged either way (Python signal handlers run only in the main thread's eval loop). 8 threads × 2000 executions on oneProgramand 3 readers + 2 writers × 20kadd_variableon oneContext: zero errors.&cel::Program,&cel::Context<'static>andPyResult<Value>all satisfy PyO3's stableUngil(=Send) bound because cel-rust requiresVal,FunctionandVariableResolverto beSend + Sync. No unsafe, no wrappers.The plan
PR 1 — own the free-threading we already ship (small, do first):
- CI: a
3.14ttest leg (uv python install 3.14t, pytest withPYTHON_GIL=0so a dependency can't quietly re-enable the GIL). - Make
#[pymodule(gil_used = false)]explicit;pyo3_log::init()→try_init()(a second module exec panics today);#[pyclass(frozen)]onProgramandOptionalValue;MutexExt::lock_py_attachedfor theContextcache lock, sinceclone_refruns under it; compile-timeSend + Syncassertions on the cel types so an upstream bump can't silently break this. - Wheels:
setup-pythonwith3.14t(and3.15tafter Oct 1) on the macOS and Windows jobs with an explicit interpreter, one interpreter per job on Windows; bump the maturin floor to ≥ 1.14. - Docs:
Programis immutable and shareable; aContextis safe to evaluate against from many threads; concurrent mutation of oneContextraisesRuntimeError: Already borrowedrather than racing, so build it before sharing. MakingContextfrozenwith an internalRwLockis the eventual fix, as a follow-up. - Cross-reference Evaluate abi3 wheels: one wheel per platform and day-one support for new Python releases #46: abi3 is incompatible with free-threaded wheels before 3.15's
abi3t; decide the two together.
PR 2 — release the GIL where it pays:
- Always detach around
Program::compileincompile()andevaluate(). - Adaptive gate on
Program.execute()as above, with an explicit override for people who know their workload:cel.compile(expr, release_gil: bool | None = None)(None= adaptive). No per-call knob onexecute(), to keep the hot path flat. evaluate()'s execute stays attached.- Tests: an overlap test (a spinning thread counts ticks while a 0.5 s pure-CEL evaluation runs; detached → millions, attached → ~0) instead of speedup ratios, so it is robust on 2-core runners; gate-decision tests via a private
Program._releases_gilprobe (x + ysettles attached, a heavy filter latches detached); the existing re-entrancy and mutation tests already cover the semantic edges. - CHANGELOG: state the ~7 µs break-even and the adaptive behaviour plainly, so nobody expects
x + yto parallelise.
Not doing: unconditional detach, the callback gate, the context-size gate, or claiming improved Ctrl-C responsiveness.
Found along the way: cel-rust 0.14.5's comprehensions are O(n²) (20k-element
filter/maptakes 5.8 s); tracked in #57, upstream fix formapis on master.
Generated by Claude Code
- Detaching around the parse is unconditionally right.
Preserving the prototype artefacts here so PR 2 can start from them rather than from scratch (the worktree they were built in is temporary).
Two notes on the diff before anyone lifts it:
- It still carries the "no callbacks and no resolver" gate (
Environment.detachable) from the earlier variant in addition to the timing latch. The measurements above say the timing latch alone is the right gate: a heavy expression that calls a Python function per element still gained 1.3–1.5× on 2 threads when detached, and the latch already captures whether such a call is cheap or expensive. PR 2 should dropdetachableand rely on the latch, keeping only the parse detach as unconditional. - It is prototype quality: no tests, no
release_giloverride oncompile(), no CHANGELOG, and thePyProgramdocstring does not yet describe the behaviour.
adaptive.diff— adaptive per-Program gate + unconditional parse detach (applies to 0.10.0'ssrc/lib.rs)diff --git a/src/lib.rs b/src/lib.rs index 84897b1..a318980 100644 --- a/src/lib.rs +++ b/src/lib.rs @@ -19,7 +19,9 @@ use pyo3::PyTypeInfo; use std::collections::HashMap; use std::error::Error; use std::fmt; +use std::sync::atomic::{AtomicU8, Ordering}; use std::sync::{Arc, LazyLock}; +use std::time::{Duration, Instant}; /// The CEL standard-library environment (built-in functions, overloads and /// macros), built once and shared across every evaluation. @@ -64,6 +66,53 @@ pub(crate) fn new_environment() -> CelContext<'static> { struct PyProgram { program: Program, source: String, + /// Adaptive GIL-release state, see [`DETACH_MIN_WORK`]. + /// + /// Counts how many executions have been timed so far, up to + /// `DETACH_SAMPLES`, or holds `STATE_DETACH` once one of them proved this + /// program does enough work to be worth detaching for. Relaxed ordering + /// throughout: this is a heuristic, so a race just re-samples or costs one + /// extra attached call. + detach_state: AtomicU8, +} + +/// How long an attached execution must take before later executions of the same +/// program release the GIL. +/// +/// Derived from the measured break-even on a 4-core box: detached 4-thread +/// throughput crosses attached at about 7-8 us of work per call and reaches 2x +/// at about 13 us. This is set to roughly twice the crossing point so a program +/// only detaches once the win is unambiguous. +const DETACH_MIN_WORK: Duration = Duration::from_micros(16); + +/// How many of a program's executions are timed before it settles. +const DETACH_SAMPLES: u8 = 3; + +/// Sentinel `detach_state` meaning "settled: detach from now on". +const STATE_DETACH: u8 = u8::MAX; + +impl PyProgram { + /// Whether this execution should be timed (still sampling). + fn sampling(&self) -> bool { + self.detach_state.load(Ordering::Relaxed) < DETACH_SAMPLES + } + + /// Whether this program has been observed to do enough work to detach. + fn wants_detach(&self) -> bool { + self.detach_state.load(Ordering::Relaxed) == STATE_DETACH + } + + /// Folds one timed execution into the adaptive state. + fn observe(&self, elapsed: Duration) { + if elapsed >= DETACH_MIN_WORK { + self.detach_state.store(STATE_DETACH, Ordering::Relaxed); + } else { + let seen = self.detach_state.load(Ordering::Relaxed); + if seen < DETACH_SAMPLES { + self.detach_state.store(seen + 1, Ordering::Relaxed); + } + } + } } #[pymethods] @@ -76,8 +125,27 @@ impl PyProgram { /// Returns: /// The result of the expression evaluation #[pyo3(signature = (context=None))] - fn execute(&self, context: Option<&Bound<'_, PyAny>>) -> PyResult<Py<PyAny>> { - execute_compiled_program(&self.program, &self.source, context) + fn execute(&self, py: Python<'_>, context: Option<&Bound<'_, PyAny>>) -> PyResult<Py<PyAny>> { + let environment = prepare_environment(context)?; + + // Time only the evaluation itself, not `prepare_environment`, whose + // first call may build the whole cel context and would otherwise make + // every program look expensive. + let started = self.sampling().then(Instant::now); + let value = run_program( + py, + &self.program, + &self.source, + &environment, + self.wants_detach(), + )?; + if let Some(started) = started { + self.observe(started.elapsed()); + } + + RustyCelType(value) + .into_pyobject(py) + .map(|obj| obj.unbind()) } /// The variable names referenced by this expression. @@ -250,18 +318,22 @@ impl PyOptionalValue { /// >>> program.execute({"x": 10, "y": 20}) /// 30 #[pyfunction] -fn compile(expression: String) -> PyResult<PyProgram> { - let program = compile_program(&expression)?; +fn compile(py: Python<'_>, expression: String) -> PyResult<PyProgram> { + let program = compile_program(py, &expression)?; Ok(PyProgram { program, source: expression, + detach_state: AtomicU8::new(0), }) } /// Parses `expression`, turning both parse errors and parser panics into /// `ValueError` so callers can rely on one exception type for a bad expression. -fn compile_program(expression: &str) -> PyResult<Program> { - panic::catch_unwind(|| Program::compile(expression)) +fn compile_program(py: Python<'_>, expression: &str) -> PyResult<Program> { + // Parsing is pure Rust and, for a realistic policy expression, an order of + // magnitude more expensive than executing the result, so it is always worth + // releasing the GIL for. Nothing here can call into Python. + py.detach(|| panic::catch_unwind(|| Program::compile(expression))) .map_err(|_| { warn!("CEL parser panic for expression: '{}'", expression); PyValueError::new_err(format!( @@ -546,6 +618,13 @@ impl Root { struct Environment { root: Root, resolver: Option<PyVariableResolver>, + /// Whether the evaluation can run with the GIL released. + /// + /// True only when nothing in this environment can call back into Python: + /// no registered Python functions and no variable resolver. When either is + /// present the evaluation stays attached, because every callback would + /// otherwise have to re-acquire the GIL, which costs far more than it saves. + detachable: bool, } /// Turns the `evaluation_context` argument of `evaluate()`/`Program.execute()` @@ -555,6 +634,7 @@ fn prepare_environment(evaluation_context: Option<&Bound<'_, PyAny>>) -> PyResul return Ok(Environment { root: Root::Owned(new_environment()), resolver: None, + detachable: true, }); }; let py = evaluation_context.py(); @@ -569,18 +649,24 @@ fn prepare_environment(evaluation_context: Option<&Bound<'_, PyAny>>) -> PyResul .map(|callback| PyVariableResolver { callback: callback.clone_ref(py), }); + let detachable = py_context.functions.is_empty() && py_context.resolver.is_none(); Ok(Environment { root: Root::Shared(py_context.cel_context(py)), resolver, + detachable, }) } else if let Ok(py_dict) = evaluation_context.cast::<PyDict>() { // A dict mixes variables and functions; `Context::update` sorts them by // callability exactly as it does for a Python `Context`. let mut ctx = context::Context::new(None, None)?; ctx.update(py_dict)?; + // `update` routes callables into `functions`, so an empty map means the + // dict held no callables and nothing can re-enter Python. + let detachable = ctx.functions.is_empty(); Ok(Environment { root: Root::Owned(ctx.build_cel_context(py)), resolver: None, + detachable, }) } else { Err(PyValueError::new_err( @@ -596,7 +682,13 @@ fn prepare_environment(evaluation_context: Option<&Bound<'_, PyAny>>) -> PyResul /// to the root itself. Lookups in the child consult the resolver first and then /// fall through to the parent's variables, which is the same order the resolver /// had when it lived on the root, and the root stays untouched and reusable. -fn run_program(program: &Program, src: &str, environment: &Environment) -> PyResult<Value> { +fn run_program( + py: Python<'_>, + program: &Program, + src: &str, + environment: &Environment, + detach: bool, +) -> PyResult<Value> { let root = environment.root.as_cel(); let scoped; let ctx: &CelContext<'_> = match &environment.resolver { @@ -610,7 +702,21 @@ fn run_program(program: &Program, src: &str, environment: &Environment) -> PyRes }; // AssertUnwindSafe is needed because the environment contains function closures. - let result = panic::catch_unwind(AssertUnwindSafe(|| program.execute(ctx))).map_err(|_| { + let run = || panic::catch_unwind(AssertUnwindSafe(|| program.execute(ctx))); + + // When nothing in the environment can call back into Python, the whole + // evaluation is pure Rust, so the GIL is released for its duration and other + // Python threads can run. `cel::Program` and `cel::Context` are `Send + Sync`. + // + // `catch_unwind` sits inside the detached region so a panic is caught while + // still detached, rather than unwinding through PyO3's re-attach guard. + let result = if detach && environment.detachable { + py.detach(run) + } else { + run() + }; + + let result = result.map_err(|_| { warn!("CEL execution panic for expression: '{}'", src); PyValueError::new_err(format!( "Failed to execute expression '{src}': Internal evaluation error" @@ -1015,28 +1121,16 @@ impl TryIntoValue for RustyPyType<'_> { /// - CEL Language Guide: For comprehensive language documentation /// - Python API Reference: For detailed API documentation #[pyfunction(signature = (src, evaluation_context=None))] -fn evaluate(src: String, evaluation_context: Option<&Bound<'_, PyAny>>) -> PyResult<RustyCelType> { +fn evaluate( + py: Python<'_>, + src: String, + evaluation_context: Option<&Bound<'_, PyAny>>, +) -> PyResult<RustyCelType> { // Validate the context before parsing so a bad context and a bad expression // report in the same order they always have. let environment = prepare_environment(evaluation_context)?; - let program = compile_program(&src)?; - run_program(&program, &src, &environment).map(RustyCelType) -} - -/// Internal helper to execute a pre-compiled program with the given context. -/// Used by `PyProgram.execute()`. -fn execute_compiled_program( - program: &Program, - src: &str, - evaluation_context: Option<&Bound<'_, PyAny>>, -) -> PyResult<Py<PyAny>> { - let environment = prepare_environment(evaluation_context)?; - let value = run_program(program, src, &environment)?; - Python::attach(|py| { - RustyCelType(value) - .into_pyobject(py) - .map(|obj| obj.unbind()) - }) + let program = compile_program(py, &src)?; + run_program(py, &program, &src, &environment, false).map(RustyCelType) } #[pymodule]
bench.py— the benchmark behind the tables (run aspython bench.py <label>against each build)"""GIL-release prototype benchmark for python-common-expression-language. Usage: python bench.py <label> e.g. python bench.py baseline Writes results-<label>.json next to this file and prints a table. Measurement approach -------------------- * Single-thread latency: timeit.Timer(...).repeat(), min-of-repeats, with the inner `number` auto-scaled so each timed block runs >= MIN_BLOCK seconds. min-of-repeats is the right statistic: we want the floor (least interference). * Multi-thread throughput: a FIXED TOTAL of N executions is split evenly over T threads in a ThreadPoolExecutor; wall time via time.perf_counter, taken as min over the case's repeat budget. The 1-thread number also goes through the executor, so the speedup ratio is apples-to-apples. Note on case `c`: cel-rust 0.14.5's comprehension macros (filter/map) are O(n^2) in the list length -- they clone the accumulator on every append -- so the 20,000-element expression costs ~5.8 s per call (see #57). `cs` (the same expression over 2,000 elements, ~66 ms) resolves the scaling curve with far more samples for the same wall-clock budget. """ import json import os import statistics import sys import time import timeit from concurrent.futures import ThreadPoolExecutor import cel MIN_BLOCK = 0.05 # seconds per timeit block LABEL = sys.argv[1] if len(sys.argv) > 1 else "unlabeled" HERE = os.path.dirname(os.path.abspath(__file__)) EXPR_A = "x + y" CTX_A = cel.Context({"x": 1, "y": 2}) PROG_A = cel.compile(EXPR_A) EXPR_B = 'user.role == "admin" || (resource.owner == user.id && size(user.groups) > 0)' CTX_B = cel.Context( { "user": { "id": "u-1234", "role": "editor", "groups": ["eng", "oncall", "reviewers"], "email": "u1234@example.com", }, "resource": { "owner": "u-1234", "kind": "document", "labels": {"tier": "gold", "region": "us-east-1"}, }, } ) PROG_B = cel.compile(EXPR_B) EXPR_C = "items.filter(i, i % 3 == 0).map(i, i * i).size()" CTX_C = cel.Context({"items": list(range(20_000))}) PROG_C = cel.compile(EXPR_C) CTX_CS = cel.Context({"items": list(range(2_000))}) PROG_CS = cel.compile(EXPR_C) EXPR_E = "twice(x)" CTX_E = cel.Context({"x": 21}, {"twice": lambda v: v * 2}) PROG_E = cel.compile(EXPR_E) CALLS = { "a": lambda: PROG_A.execute(CTX_A), "b": lambda: PROG_B.execute(CTX_B), "c": lambda: PROG_C.execute(CTX_C), "cs": lambda: PROG_CS.execute(CTX_CS), "d": lambda: cel.evaluate(EXPR_B, CTX_B), "e": lambda: PROG_E.execute(CTX_E), } DESCRIPTIONS = { "a": "Program.execute x + y (tiny)", "b": "Program.execute policy expr (nested maps)", "c": "Program.execute filter/map/size, 20k list", "cs": "Program.execute filter/map/size, 2k list", "d": "evaluate() policy expr (parse + execute)", "e": "Program.execute twice(x) via Python callback", } # case -> (latency_repeats, mt_thread_counts, mt_total_N, mt_repeats) # mt_total_N = None means "auto-size for ~MT_TARGET seconds of 1-thread work". MT_TARGET = 1.2 BUDGET = { "a": (7, [1, 4], None, 5), "b": (7, [1, 2, 4, 8], None, 5), "cs": (5, [1, 2, 4, 8], 64, 3), "c": (3, [1, 2, 4, 8], 8, 2), "d": (7, [1, 4], None, 5), "e": (7, [1, 4], None, 5), } for _name, _got, _want in [ ("a", PROG_A.execute(CTX_A), 3), ("b", PROG_B.execute(CTX_B), True), ("c", PROG_C.execute(CTX_C), 6667), ("cs", PROG_CS.execute(CTX_CS), 667), ("d", cel.evaluate(EXPR_B, CTX_B), True), ("e", PROG_E.execute(CTX_E), 42), ]: assert _got == _want, f"case {_name}: got {_got!r}, want {_want!r}" def latency(case): fn = CALLS[case] repeats = BUDGET[case][0] t = timeit.Timer(fn) number = 1 while True: dt = t.timeit(number) if dt >= MIN_BLOCK: break number = max(number * 2, int(number * MIN_BLOCK / max(dt, 1e-9)) + 1) if number > 50_000_000: break samples = t.repeat(repeat=repeats, number=number) per_call = [s / number for s in samples] return { "number": number, "repeats": repeats, "min_s": min(per_call), "median_s": statistics.median(per_call), "max_s": max(per_call), } def _worker(fn, iters): for _ in range(iters): fn() def mt_wall(case, threads, total, repeats): fn = CALLS[case] per = total // threads best = float("inf") with ThreadPoolExecutor(max_workers=threads) as pool: list(pool.map(lambda _: None, range(threads))) # pre-spawn threads for _ in range(repeats): t0 = time.perf_counter() futs = [pool.submit(_worker, fn, per) for _ in range(threads)] for f in futs: f.result() best = min(best, time.perf_counter() - t0) return best, per * threads def mt_matrix(case, per_call_s): _, thread_counts, total, repeats = BUDGET[case] top = max(thread_counts) if total is None: total = top * max(1, int(MT_TARGET / per_call_s / top)) out = {"total_executions": total, "mt_repeats": repeats, "threads": {}} base = None for t in thread_counts: wall, actual = mt_wall(case, t, total, repeats) if base is None: base = wall out["threads"][str(t)] = { "wall_s": wall, "executions": actual, "throughput_per_s": actual / wall, "speedup_vs_1": base / wall, } print( f" {t:>8} {wall:>10.4f} {actual / wall:>14,.0f} {base / wall:>8.2f}x", flush=True, ) return out ORDER = ["a", "b", "cs", "c", "d", "e"] def main(): results = { "label": LABEL, "nproc": os.cpu_count(), "python": sys.version.split()[0], "latency": {}, "throughput": {}, } print(f"=== {LABEL} === (cores={os.cpu_count()}, python={sys.version.split()[0]})", flush=True) print("\nSingle-thread latency (min of N repeats)", flush=True) print(f"{'case':<5} {'description':<44} {'min':>13} {'median':>13}", flush=True) for case in ORDER: r = latency(case) results["latency"][case] = r print( f"{case:<5} {DESCRIPTIONS[case]:``<44} "· f"{r['min_s'] * 1e6:>``10.3f} us {r['median_s'] * 1e6:>10.3f} us", flush=True, ) print("\nMulti-thread throughput (min wall over repeats, fixed total N)", flush=True) for case in ORDER: per_call = results["latency"][case]["min_s"] print(f"\n case {case}: {DESCRIPTIONS[case]}", flush=True) print(f" {'threads':>8} {'wall_s':>10} {'exec/s':>14} {'speedup':>9}", flush=True) results["throughput"][case] = mt_matrix(case, per_call) path = os.path.join(HERE, f"results-{LABEL}.json") with open(path, "w") as fh: json.dump(results, fh, indent=2) print(f"\nwrote {path}", flush=True) if __name__ == "__main__": main()
Generated by Claude Code
- It still carries the "no callbacks and no resolver" gate (
- added a commit that references this issue
on Sep 15, 2026 Status: two of the three pieces above have merged.
- PR 1, free-threading (#58): the suite runs on
python3.14tin CI, free-threaded wheels build for macOS and Windows x64 as well as Linux,ProgramandOptionalValueare frozen, and theContextthreading contract is documented and tested. The new leg found and fixed a misleadingValueErrorwhen an evaluation raced aContextmutation. - PR 2, first half: parse (#59):
compile()and the parse insideevaluate()release the GIL for expressions of 32 bytes or more.evaluate()of a policy-sized expression from 4 threads went from 0.96× to about 3× single-thread throughput on 4 cores.
Still open: releasing the GIL around
Program.execute(). The measurements above show a fixed threshold tuned on one machine is only half adaptive, since it measures the program's work but assumes the machine's contended re-acquire cost. The direction I'd take: measure the handoff overhead actually paid by detached executions and revert when it exceeds the work, start from a conservative prior of a few times the measured re-acquire latency, re-sample with hysteresis rather than latching, and keep an explicitrelease_giloverride oncompile(). On free-threaded builds none of this applies and the gate should compile out.
Generated by Claude Code
- PR 1, free-threading (#58): the suite runs on
Problem
evaluate()andProgram.execute()hold the GIL for the entire Rust-side evaluation. A Python program that evaluates CEL from several threads therefore serialises on the interpreter even though the work is pure Rust once the context has been converted. This matters for the policy-engine style of use (many rules × many requests) that the README leads with.Proposal
program.execute()withpy.detach(|| ...)(PyO3 0.29's spelling ofallow_threads). The pieces already have the right bounds:cel::Programis a plain AST,cel::Context<'static>isSend + Sync(itsVal,FunctionandVariableResolvertraits all require it), and the Python-callback wrappers already do their ownPython::attach, so a callback simply re-acquires the GIL when it runs. Conversion of the result back to Python happens after re-attaching.x + yexecutes in ~0.15 µs, so unconditional detaching could be a visible relative slowdown for tiny expressions while being a large absolute win for anything heavier or for multi-threaded callers. Options, in order of preference:compile_execute_benchmark.pycases;Context;execute(ctx, release_gil=...)knob is a last resort.#[pymodule(gil_used = false)]. Before doing that, audit:Contextmutators are&mut self(PyO3's borrow checker turns concurrent mutation into an error rather than a data race), the shared stdlibEnvis aLazyLock, and the per-Contextcache proposed in the Context-reuse PR is behind aMutex. Then addcp314twheels to the CI matrix (maturin-action needs the interpreter listed explicitly;--find-interpreterwon't pick it up).Non-goals
Contextitself is documented as not thread-safe for concurrent mutation; that stays. Concurrent evaluation against a sharedContextis already fine and is now pinned by tests.