Skip to content

Release the GIL during evaluation, and decide on free-threaded (cp314t) support #45

Description

@hardbyte

Problem

evaluate() and Program.execute() hold the GIL for the entire Rust-side evaluation. A Python program that evaluates CEL from several threads therefore serialises on the interpreter even though the work is pure Rust once the context has been converted. This matters for the policy-engine style of use (many rules × many requests) that the README leads with.

Proposal

  1. Release the GIL around program.execute() with py.detach(|| ...) (PyO3 0.29's spelling of allow_threads). The pieces already have the right bounds: cel::Program is a plain AST, cel::Context<'static> is Send + Sync (its Val, Function and VariableResolver traits all require it), and the Python-callback wrappers already do their own Python::attach, so a callback simply re-acquires the GIL when it runs. Conversion of the result back to Python happens after re-attaching.
  2. Measure the fixed cost. Detach/attach is on the order of tens of nanoseconds, but a trivial x + y executes in ~0.15 µs, so unconditional detaching could be a visible relative slowdown for tiny expressions while being a large absolute win for anything heavier or for multi-threaded callers. Options, in order of preference:
    • detach unconditionally if the overhead measures under ~10% on the compile_execute_benchmark.py cases;
    • otherwise detach only when the context has no Python functions and no resolver (that's when the evaluation cannot need the GIL), which is cheap to know from the Context;
    • an explicit execute(ctx, release_gil=...) knob is a last resort.
  3. Free-threaded Python. PyO3 0.29 supports the free-threaded build when the module opts in with #[pymodule(gil_used = false)]. Before doing that, audit: Context mutators are &mut self (PyO3's borrow checker turns concurrent mutation into an error rather than a data race), the shared stdlib Env is a LazyLock, and the per-Context cache proposed in the Context-reuse PR is behind a Mutex. Then add cp314t wheels to the CI matrix (maturin-action needs the interpreter listed explicitly; --find-interpreter won't pick it up).

Non-goals

Context itself is documented as not thread-safe for concurrent mutation; that stays. Concurrent evaluation against a shared Context is already fine and is now pinned by tests.

Activity

  1. hardbyte commented on Sep 15, 2026

    @hardbyte
    OwnerAuthor

    Direction, after measuring

    Two independent prototypes were built and benchmarked (release builds, 4 cores, Python 3.11; full test suite passes identically with the GIL released), plus a source-level read of PyO3 0.29.2 and a survey of the ecosystem. The proposal in the issue body was wrong in two places, and the corrected direction has two independent tracks.

    Finding 1: free-threaded support is already shipping, untested

    PyO3 0.28 made gil_used = false the default, so the bare #[pymodule] in src/lib.rs has declared this module GIL-free on 3.13t+ since the 0.7.0 upgrade. And because the manylinux images carry python3.14t/python3.15t and maturin ≥ 1.14's --find-interpreter picks them up, 0.10.0 already publishes cp314t and cp315t wheels for every Linux architecture (check the file list on PyPI). Only macOS and Windows lack them, because those runners have no free-threaded interpreter installed. No CI leg has ever run the tests under PYTHON_GIL=0.

    Ecosystem context: six of eight comparable Rust extensions ship free-threaded wheels (pydantic-core, cryptography, rpds-py, jiter, watchfiles, tiktoken); the two that don't are blocked by being abi3-only. cibuildwheel and maturin both build them by default now. PyO3 0.29 dropped 3.13t, so cp314t/cp315t is the whole target. CPython 3.15 (final 2026-10-01) adds abi3t (PEP 803); abi3 wheels cannot be installed on a free-threaded interpreter before that, which is relevant to #46: switching to abi3 would silently remove the free-threaded wheels we ship today.

    Finding 2: releasing the GIL around execute() is a throughput cliff, not a 50 ns tax

    Uncontended, a detach/attach round trip costs ~47–75 ns. Under contention it is a different regime: a thread re-acquiring the GIL waits for whoever grabbed it meanwhile, ~14 µs at 4 threads. For sub-microsecond evaluations that is the whole cost.

    4 threads, exec/s GIL held (today) detach always detach unless callbacks detach if context ≥ 1024 elements
    x + y 3,850,000 64,000 69,000 3,850,000
    x + y, context also holds a 20k list 3,330,000 — — 64,000
    policy expression (1.3 µs) 721,000 46,000 49,000 711,000
    filter(...).size() over 512 items (3.5 ms) 266 1,068 1,068 268
    filter/map/size over 2k items (63 ms) 12 (0.97×) 49 (3.99×) 47 (3.88×) 49 (3.92×)

    Break-even, measured by sweeping list length with items.filter(i, i % 3 == 0).size(): detaching starts winning at ~7–8 µs of work per call, reaches 2× at ~13 µs, and saturates at the core count (~3.9×) from ~65 µs. One octave either side of the crossing is the difference between 0.34× and 1.5×. Single-thread benchmarks give no warning of any of this, which is why the issue body's "measure the fixed cost" plan would have shipped a 55× regression for the most common workload.

    Both cheap gates fail in both directions. Gating on "no Python callbacks" still detaches x + y. Gating on context size detaches x + y when a big list happens to be in the context (52× worse) and fails to detach a 3.5 ms filter over a 512-element list (misses a 3.9× win). The question that matters is "how much work will this call do", and neither proxy answers it.

    Two things do work:

    • Detaching around the parse is unconditionally right. Program::compile is pure Rust and costs 14–72 µs (98% of evaluate() for a policy-sized expression), well past break-even. evaluate() goes from 0.94× to 2.6× at 4 threads with no downside.
    • An adaptive per-Program gate. Time the attached execution for the first 3 calls (excluding prepare_environment, whose first call may build the cel context); if any sample exceeds ~16 µs (≈2× the crossing), latch that Program into detaching. Prototyped with an AtomicU8 on PyProgram, relaxed ordering:
    4 threads GIL held adaptive
    x + y 3,850,000 3,690,000
    x + y, 20k list in context 3,330,000 3,580,000
    filter over 512 items 266 1,061 (3.87×)
    filter/map/size over 2k 12 50 (3.98×)

    Steady-state overhead is one relaxed atomic load (0.242 vs 0.240 µs single-thread); sampling adds at most ~0.5 µs to each of the first three calls. Known limits: a program whose cost varies wildly across contexts latches on its first samples, and evaluate() has no persistent object to learn on, so its execute stays attached (its cost is the parse anyway).

    Semantics, verified rather than argued

    Under detach: callbacks re-attach via Python::attach (correct in PyO3's SuspendAttach design, verified against the source); exceptions from callbacks propagate with identical messages; a callback mutating its own Context sees the same snapshot semantics as today (the Arc from #51 is what makes this work); re-entrant evaluate() from a callback works; pyo3_log re-attaches; panics are caught inside the detached region; Ctrl-C behaviour is unchanged either way (Python signal handlers run only in the main thread's eval loop). 8 threads × 2000 executions on one Program and 3 readers + 2 writers × 20k add_variable on one Context: zero errors.

    &cel::Program, &cel::Context<'static> and PyResult<Value> all satisfy PyO3's stable Ungil (= Send) bound because cel-rust requires Val, Function and VariableResolver to be Send + Sync. No unsafe, no wrappers.

    The plan

    PR 1 — own the free-threading we already ship (small, do first):

    • CI: a 3.14t test leg (uv python install 3.14t, pytest with PYTHON_GIL=0 so a dependency can't quietly re-enable the GIL).
    • Make #[pymodule(gil_used = false)] explicit; pyo3_log::init() → try_init() (a second module exec panics today); #[pyclass(frozen)] on Program and OptionalValue; MutexExt::lock_py_attached for the Context cache lock, since clone_ref runs under it; compile-time Send + Sync assertions on the cel types so an upstream bump can't silently break this.
    • Wheels: setup-python with 3.14t (and 3.15t after Oct 1) on the macOS and Windows jobs with an explicit interpreter, one interpreter per job on Windows; bump the maturin floor to ≥ 1.14.
    • Docs: Program is immutable and shareable; a Context is safe to evaluate against from many threads; concurrent mutation of one Context raises RuntimeError: Already borrowed rather than racing, so build it before sharing. Making Context frozen with an internal RwLock is the eventual fix, as a follow-up.
    • Cross-reference Evaluate abi3 wheels: one wheel per platform and day-one support for new Python releases #46: abi3 is incompatible with free-threaded wheels before 3.15's abi3t; decide the two together.

    PR 2 — release the GIL where it pays:

    • Always detach around Program::compile in compile() and evaluate().
    • Adaptive gate on Program.execute() as above, with an explicit override for people who know their workload: cel.compile(expr, release_gil: bool | None = None) (None = adaptive). No per-call knob on execute(), to keep the hot path flat.
    • evaluate()'s execute stays attached.
    • Tests: an overlap test (a spinning thread counts ticks while a 0.5 s pure-CEL evaluation runs; detached → millions, attached → ~0) instead of speedup ratios, so it is robust on 2-core runners; gate-decision tests via a private Program._releases_gil probe (x + y settles attached, a heavy filter latches detached); the existing re-entrancy and mutation tests already cover the semantic edges.
    • CHANGELOG: state the ~7 µs break-even and the adaptive behaviour plainly, so nobody expects x + y to parallelise.

    Not doing: unconditional detach, the callback gate, the context-size gate, or claiming improved Ctrl-C responsiveness.

    Found along the way: cel-rust 0.14.5's comprehensions are O(n²) (20k-element filter/map takes 5.8 s); tracked in #57, upstream fix for map is on master.


    Generated by Claude Code

  2. hardbyte commented on Sep 15, 2026

    @hardbyte
    OwnerAuthor

    Preserving the prototype artefacts here so PR 2 can start from them rather than from scratch (the worktree they were built in is temporary).

    Two notes on the diff before anyone lifts it:

    • It still carries the "no callbacks and no resolver" gate (Environment.detachable) from the earlier variant in addition to the timing latch. The measurements above say the timing latch alone is the right gate: a heavy expression that calls a Python function per element still gained 1.3–1.5× on 2 threads when detached, and the latch already captures whether such a call is cheap or expensive. PR 2 should drop detachable and rely on the latch, keeping only the parse detach as unconditional.
    • It is prototype quality: no tests, no release_gil override on compile(), no CHANGELOG, and the PyProgram docstring does not yet describe the behaviour.
    adaptive.diff — adaptive per-Program gate + unconditional parse detach (applies to 0.10.0's src/lib.rs)
    diff --git a/src/lib.rs b/src/lib.rs
    index 84897b1..a318980 100644
    --- a/src/lib.rs
    +++ b/src/lib.rs
    @@ -19,7 +19,9 @@ use pyo3::PyTypeInfo;
     use std::collections::HashMap;
     use std::error::Error;
     use std::fmt;
    +use std::sync::atomic::{AtomicU8, Ordering};
     use std::sync::{Arc, LazyLock};
    +use std::time::{Duration, Instant};
     
     /// The CEL standard-library environment (built-in functions, overloads and
     /// macros), built once and shared across every evaluation.
    @@ -64,6 +66,53 @@ pub(crate) fn new_environment() -> CelContext<'static> {
     struct PyProgram {
         program: Program,
         source: String,
    +    /// Adaptive GIL-release state, see [`DETACH_MIN_WORK`].
    +    ///
    +    /// Counts how many executions have been timed so far, up to
    +    /// `DETACH_SAMPLES`, or holds `STATE_DETACH` once one of them proved this
    +    /// program does enough work to be worth detaching for. Relaxed ordering
    +    /// throughout: this is a heuristic, so a race just re-samples or costs one
    +    /// extra attached call.
    +    detach_state: AtomicU8,
    +}
    +
    +/// How long an attached execution must take before later executions of the same
    +/// program release the GIL.
    +///
    +/// Derived from the measured break-even on a 4-core box: detached 4-thread
    +/// throughput crosses attached at about 7-8 us of work per call and reaches 2x
    +/// at about 13 us. This is set to roughly twice the crossing point so a program
    +/// only detaches once the win is unambiguous.
    +const DETACH_MIN_WORK: Duration = Duration::from_micros(16);
    +
    +/// How many of a program's executions are timed before it settles.
    +const DETACH_SAMPLES: u8 = 3;
    +
    +/// Sentinel `detach_state` meaning "settled: detach from now on".
    +const STATE_DETACH: u8 = u8::MAX;
    +
    +impl PyProgram {
    +    /// Whether this execution should be timed (still sampling).
    +    fn sampling(&self) -> bool {
    +        self.detach_state.load(Ordering::Relaxed) < DETACH_SAMPLES
    +    }
    +
    +    /// Whether this program has been observed to do enough work to detach.
    +    fn wants_detach(&self) -> bool {
    +        self.detach_state.load(Ordering::Relaxed) == STATE_DETACH
    +    }
    +
    +    /// Folds one timed execution into the adaptive state.
    +    fn observe(&self, elapsed: Duration) {
    +        if elapsed >= DETACH_MIN_WORK {
    +            self.detach_state.store(STATE_DETACH, Ordering::Relaxed);
    +        } else {
    +            let seen = self.detach_state.load(Ordering::Relaxed);
    +            if seen < DETACH_SAMPLES {
    +                self.detach_state.store(seen + 1, Ordering::Relaxed);
    +            }
    +        }
    +    }
     }
     
     #[pymethods]
    @@ -76,8 +125,27 @@ impl PyProgram {
         /// Returns:
         ///     The result of the expression evaluation
         #[pyo3(signature = (context=None))]
    -    fn execute(&self, context: Option<&Bound<'_, PyAny>>) -> PyResult<Py<PyAny>> {
    -        execute_compiled_program(&self.program, &self.source, context)
    +    fn execute(&self, py: Python<'_>, context: Option<&Bound<'_, PyAny>>) -> PyResult<Py<PyAny>> {
    +        let environment = prepare_environment(context)?;
    +
    +        // Time only the evaluation itself, not `prepare_environment`, whose
    +        // first call may build the whole cel context and would otherwise make
    +        // every program look expensive.
    +        let started = self.sampling().then(Instant::now);
    +        let value = run_program(
    +            py,
    +            &self.program,
    +            &self.source,
    +            &environment,
    +            self.wants_detach(),
    +        )?;
    +        if let Some(started) = started {
    +            self.observe(started.elapsed());
    +        }
    +
    +        RustyCelType(value)
    +            .into_pyobject(py)
    +            .map(|obj| obj.unbind())
         }
     
         /// The variable names referenced by this expression.
    @@ -250,18 +318,22 @@ impl PyOptionalValue {
     ///     >>> program.execute({"x": 10, "y": 20})
     ///     30
     #[pyfunction]
    -fn compile(expression: String) -> PyResult<PyProgram> {
    -    let program = compile_program(&expression)?;
    +fn compile(py: Python<'_>, expression: String) -> PyResult<PyProgram> {
    +    let program = compile_program(py, &expression)?;
         Ok(PyProgram {
             program,
             source: expression,
    +        detach_state: AtomicU8::new(0),
         })
     }
     
     /// Parses `expression`, turning both parse errors and parser panics into
     /// `ValueError` so callers can rely on one exception type for a bad expression.
    -fn compile_program(expression: &str) -> PyResult<Program> {
    -    panic::catch_unwind(|| Program::compile(expression))
    +fn compile_program(py: Python<'_>, expression: &str) -> PyResult<Program> {
    +    // Parsing is pure Rust and, for a realistic policy expression, an order of
    +    // magnitude more expensive than executing the result, so it is always worth
    +    // releasing the GIL for. Nothing here can call into Python.
    +    py.detach(|| panic::catch_unwind(|| Program::compile(expression)))
             .map_err(|_| {
                 warn!("CEL parser panic for expression: '{}'", expression);
                 PyValueError::new_err(format!(
    @@ -546,6 +618,13 @@ impl Root {
     struct Environment {
         root: Root,
         resolver: Option<PyVariableResolver>,
    +    /// Whether the evaluation can run with the GIL released.
    +    ///
    +    /// True only when nothing in this environment can call back into Python:
    +    /// no registered Python functions and no variable resolver. When either is
    +    /// present the evaluation stays attached, because every callback would
    +    /// otherwise have to re-acquire the GIL, which costs far more than it saves.
    +    detachable: bool,
     }
     
     /// Turns the `evaluation_context` argument of `evaluate()`/`Program.execute()`
    @@ -555,6 +634,7 @@ fn prepare_environment(evaluation_context: Option<&Bound<'_, PyAny>>) -> PyResul
             return Ok(Environment {
                 root: Root::Owned(new_environment()),
                 resolver: None,
    +            detachable: true,
             });
         };
         let py = evaluation_context.py();
    @@ -569,18 +649,24 @@ fn prepare_environment(evaluation_context: Option<&Bound<'_, PyAny>>) -> PyResul
                 .map(|callback| PyVariableResolver {
                     callback: callback.clone_ref(py),
                 });
    +        let detachable = py_context.functions.is_empty() && py_context.resolver.is_none();
             Ok(Environment {
                 root: Root::Shared(py_context.cel_context(py)),
                 resolver,
    +            detachable,
             })
         } else if let Ok(py_dict) = evaluation_context.cast::<PyDict>() {
             // A dict mixes variables and functions; `Context::update` sorts them by
             // callability exactly as it does for a Python `Context`.
             let mut ctx = context::Context::new(None, None)?;
             ctx.update(py_dict)?;
    +        // `update` routes callables into `functions`, so an empty map means the
    +        // dict held no callables and nothing can re-enter Python.
    +        let detachable = ctx.functions.is_empty();
             Ok(Environment {
                 root: Root::Owned(ctx.build_cel_context(py)),
                 resolver: None,
    +            detachable,
             })
         } else {
             Err(PyValueError::new_err(
    @@ -596,7 +682,13 @@ fn prepare_environment(evaluation_context: Option<&Bound<'_, PyAny>>) -> PyResul
     /// to the root itself. Lookups in the child consult the resolver first and then
     /// fall through to the parent's variables, which is the same order the resolver
     /// had when it lived on the root, and the root stays untouched and reusable.
    -fn run_program(program: &Program, src: &str, environment: &Environment) -> PyResult<Value> {
    +fn run_program(
    +    py: Python<'_>,
    +    program: &Program,
    +    src: &str,
    +    environment: &Environment,
    +    detach: bool,
    +) -> PyResult<Value> {
         let root = environment.root.as_cel();
         let scoped;
         let ctx: &CelContext<'_> = match &environment.resolver {
    @@ -610,7 +702,21 @@ fn run_program(program: &Program, src: &str, environment: &Environment) -> PyRes
         };
     
         // AssertUnwindSafe is needed because the environment contains function closures.
    -    let result = panic::catch_unwind(AssertUnwindSafe(|| program.execute(ctx))).map_err(|_| {
    +    let run = || panic::catch_unwind(AssertUnwindSafe(|| program.execute(ctx)));
    +
    +    // When nothing in the environment can call back into Python, the whole
    +    // evaluation is pure Rust, so the GIL is released for its duration and other
    +    // Python threads can run. `cel::Program` and `cel::Context` are `Send + Sync`.
    +    //
    +    // `catch_unwind` sits inside the detached region so a panic is caught while
    +    // still detached, rather than unwinding through PyO3's re-attach guard.
    +    let result = if detach && environment.detachable {
    +        py.detach(run)
    +    } else {
    +        run()
    +    };
    +
    +    let result = result.map_err(|_| {
             warn!("CEL execution panic for expression: '{}'", src);
             PyValueError::new_err(format!(
                 "Failed to execute expression '{src}': Internal evaluation error"
    @@ -1015,28 +1121,16 @@ impl TryIntoValue for RustyPyType<'_> {
     ///     - CEL Language Guide: For comprehensive language documentation
     ///     - Python API Reference: For detailed API documentation
     #[pyfunction(signature = (src, evaluation_context=None))]
    -fn evaluate(src: String, evaluation_context: Option<&Bound<'_, PyAny>>) -> PyResult<RustyCelType> {
    +fn evaluate(
    +    py: Python<'_>,
    +    src: String,
    +    evaluation_context: Option<&Bound<'_, PyAny>>,
    +) -> PyResult<RustyCelType> {
         // Validate the context before parsing so a bad context and a bad expression
         // report in the same order they always have.
         let environment = prepare_environment(evaluation_context)?;
    -    let program = compile_program(&src)?;
    -    run_program(&program, &src, &environment).map(RustyCelType)
    -}
    -
    -/// Internal helper to execute a pre-compiled program with the given context.
    -/// Used by `PyProgram.execute()`.
    -fn execute_compiled_program(
    -    program: &Program,
    -    src: &str,
    -    evaluation_context: Option<&Bound<'_, PyAny>>,
    -) -> PyResult<Py<PyAny>> {
    -    let environment = prepare_environment(evaluation_context)?;
    -    let value = run_program(program, src, &environment)?;
    -    Python::attach(|py| {
    -        RustyCelType(value)
    -            .into_pyobject(py)
    -            .map(|obj| obj.unbind())
    -    })
    +    let program = compile_program(py, &src)?;
    +    run_program(py, &program, &src, &environment, false).map(RustyCelType)
     }
     
     #[pymodule]
    bench.py — the benchmark behind the tables (run as python bench.py <label> against each build)
    """GIL-release prototype benchmark for python-common-expression-language.
    
    Usage:  python bench.py <label>      e.g.  python bench.py baseline
    
    Writes results-<label>.json next to this file and prints a table.
    
    Measurement approach
    --------------------
    * Single-thread latency: timeit.Timer(...).repeat(), min-of-repeats, with the
      inner `number` auto-scaled so each timed block runs >= MIN_BLOCK seconds.
      min-of-repeats is the right statistic: we want the floor (least interference).
    * Multi-thread throughput: a FIXED TOTAL of N executions is split evenly over
      T threads in a ThreadPoolExecutor; wall time via time.perf_counter, taken as
      min over the case's repeat budget. The 1-thread number also goes through the
      executor, so the speedup ratio is apples-to-apples.
    
    Note on case `c`: cel-rust 0.14.5's comprehension macros (filter/map) are
    O(n^2) in the list length -- they clone the accumulator on every append -- so
    the 20,000-element expression costs ~5.8 s per call (see #57). `cs` (the same
    expression over 2,000 elements, ~66 ms) resolves the scaling curve with far
    more samples for the same wall-clock budget.
    """
    
    import json
    import os
    import statistics
    import sys
    import time
    import timeit
    from concurrent.futures import ThreadPoolExecutor
    
    import cel
    
    MIN_BLOCK = 0.05  # seconds per timeit block
    
    LABEL = sys.argv[1] if len(sys.argv) > 1 else "unlabeled"
    HERE = os.path.dirname(os.path.abspath(__file__))
    
    EXPR_A = "x + y"
    CTX_A = cel.Context({"x": 1, "y": 2})
    PROG_A = cel.compile(EXPR_A)
    
    EXPR_B = 'user.role == "admin" || (resource.owner == user.id && size(user.groups) > 0)'
    CTX_B = cel.Context(
        {
            "user": {
                "id": "u-1234",
                "role": "editor",
                "groups": ["eng", "oncall", "reviewers"],
                "email": "u1234@example.com",
            },
            "resource": {
                "owner": "u-1234",
                "kind": "document",
                "labels": {"tier": "gold", "region": "us-east-1"},
            },
        }
    )
    PROG_B = cel.compile(EXPR_B)
    
    EXPR_C = "items.filter(i, i % 3 == 0).map(i, i * i).size()"
    CTX_C = cel.Context({"items": list(range(20_000))})
    PROG_C = cel.compile(EXPR_C)
    
    CTX_CS = cel.Context({"items": list(range(2_000))})
    PROG_CS = cel.compile(EXPR_C)
    
    EXPR_E = "twice(x)"
    CTX_E = cel.Context({"x": 21}, {"twice": lambda v: v * 2})
    PROG_E = cel.compile(EXPR_E)
    
    CALLS = {
        "a": lambda: PROG_A.execute(CTX_A),
        "b": lambda: PROG_B.execute(CTX_B),
        "c": lambda: PROG_C.execute(CTX_C),
        "cs": lambda: PROG_CS.execute(CTX_CS),
        "d": lambda: cel.evaluate(EXPR_B, CTX_B),
        "e": lambda: PROG_E.execute(CTX_E),
    }
    
    DESCRIPTIONS = {
        "a": "Program.execute  x + y  (tiny)",
        "b": "Program.execute  policy expr (nested maps)",
        "c": "Program.execute  filter/map/size, 20k list",
        "cs": "Program.execute  filter/map/size, 2k list",
        "d": "evaluate()       policy expr (parse + execute)",
        "e": "Program.execute  twice(x) via Python callback",
    }
    
    # case -> (latency_repeats, mt_thread_counts, mt_total_N, mt_repeats)
    # mt_total_N = None means "auto-size for ~MT_TARGET seconds of 1-thread work".
    MT_TARGET = 1.2
    BUDGET = {
        "a": (7, [1, 4], None, 5),
        "b": (7, [1, 2, 4, 8], None, 5),
        "cs": (5, [1, 2, 4, 8], 64, 3),
        "c": (3, [1, 2, 4, 8], 8, 2),
        "d": (7, [1, 4], None, 5),
        "e": (7, [1, 4], None, 5),
    }
    
    for _name, _got, _want in [
        ("a", PROG_A.execute(CTX_A), 3),
        ("b", PROG_B.execute(CTX_B), True),
        ("c", PROG_C.execute(CTX_C), 6667),
        ("cs", PROG_CS.execute(CTX_CS), 667),
        ("d", cel.evaluate(EXPR_B, CTX_B), True),
        ("e", PROG_E.execute(CTX_E), 42),
    ]:
        assert _got == _want, f"case {_name}: got {_got!r}, want {_want!r}"
    
    
    def latency(case):
        fn = CALLS[case]
        repeats = BUDGET[case][0]
        t = timeit.Timer(fn)
        number = 1
        while True:
            dt = t.timeit(number)
            if dt >= MIN_BLOCK:
                break
            number = max(number * 2, int(number * MIN_BLOCK / max(dt, 1e-9)) + 1)
            if number > 50_000_000:
                break
        samples = t.repeat(repeat=repeats, number=number)
        per_call = [s / number for s in samples]
        return {
            "number": number,
            "repeats": repeats,
            "min_s": min(per_call),
            "median_s": statistics.median(per_call),
            "max_s": max(per_call),
        }
    
    
    def _worker(fn, iters):
        for _ in range(iters):
            fn()
    
    
    def mt_wall(case, threads, total, repeats):
        fn = CALLS[case]
        per = total // threads
        best = float("inf")
        with ThreadPoolExecutor(max_workers=threads) as pool:
            list(pool.map(lambda _: None, range(threads)))  # pre-spawn threads
            for _ in range(repeats):
                t0 = time.perf_counter()
                futs = [pool.submit(_worker, fn, per) for _ in range(threads)]
                for f in futs:
                    f.result()
                best = min(best, time.perf_counter() - t0)
        return best, per * threads
    
    
    def mt_matrix(case, per_call_s):
        _, thread_counts, total, repeats = BUDGET[case]
        top = max(thread_counts)
        if total is None:
            total = top * max(1, int(MT_TARGET / per_call_s / top))
        out = {"total_executions": total, "mt_repeats": repeats, "threads": {}}
        base = None
        for t in thread_counts:
            wall, actual = mt_wall(case, t, total, repeats)
            if base is None:
                base = wall
            out["threads"][str(t)] = {
                "wall_s": wall,
                "executions": actual,
                "throughput_per_s": actual / wall,
                "speedup_vs_1": base / wall,
            }
            print(
                f"    {t:>8} {wall:>10.4f} {actual / wall:>14,.0f} {base / wall:>8.2f}x",
                flush=True,
            )
        return out
    
    
    ORDER = ["a", "b", "cs", "c", "d", "e"]
    
    
    def main():
        results = {
            "label": LABEL,
            "nproc": os.cpu_count(),
            "python": sys.version.split()[0],
            "latency": {},
            "throughput": {},
        }
    
        print(f"=== {LABEL} === (cores={os.cpu_count()}, python={sys.version.split()[0]})", flush=True)
        print("\nSingle-thread latency (min of N repeats)", flush=True)
        print(f"{'case':<5} {'description':<44} {'min':>13} {'median':>13}", flush=True)
        for case in ORDER:
            r = latency(case)
            results["latency"][case] = r
            print(
                f"{case:<5} {DESCRIPTIONS[case]:``&lt;44} "·            f"{r['min_s'] * 1e6:&gt;``10.3f} us {r['median_s'] * 1e6:>10.3f} us",
                flush=True,
            )
    
        print("\nMulti-thread throughput (min wall over repeats, fixed total N)", flush=True)
        for case in ORDER:
            per_call = results["latency"][case]["min_s"]
            print(f"\n  case {case}: {DESCRIPTIONS[case]}", flush=True)
            print(f"    {'threads':>8} {'wall_s':>10} {'exec/s':>14} {'speedup':>9}", flush=True)
            results["throughput"][case] = mt_matrix(case, per_call)
    
        path = os.path.join(HERE, f"results-{LABEL}.json")
        with open(path, "w") as fh:
            json.dump(results, fh, indent=2)
        print(f"\nwrote {path}", flush=True)
    
    
    if __name__ == "__main__":
        main()

    Generated by Claude Code

  3. hardbyte commented on Sep 15, 2026

    @hardbyte
    OwnerAuthor

    Status: two of the three pieces above have merged.

    • PR 1, free-threading (#58): the suite runs on python3.14t in CI, free-threaded wheels build for macOS and Windows x64 as well as Linux, Program and OptionalValue are frozen, and the Context threading contract is documented and tested. The new leg found and fixed a misleading ValueError when an evaluation raced a Context mutation.
    • PR 2, first half: parse (#59): compile() and the parse inside evaluate() release the GIL for expressions of 32 bytes or more. evaluate() of a policy-sized expression from 4 threads went from 0.96× to about 3× single-thread throughput on 4 cores.

    Still open: releasing the GIL around Program.execute(). The measurements above show a fixed threshold tuned on one machine is only half adaptive, since it measures the program's work but assumes the machine's contended re-acquire cost. The direction I'd take: measure the handoff overhead actually paid by detached executions and revert when it exceeds the work, start from a conservative prior of a few times the measured re-acquire latency, re-sample with hysteresis rather than latching, and keep an explicit release_gil override on compile(). On free-threaded builds none of this applies and the gate should compile out.


    Generated by Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions