Problem
evaluate() and Program.execute() hold the GIL for the entire Rust-side evaluation. A Python program that evaluates CEL from several threads therefore serialises on the interpreter even though the work is pure Rust once the context has been converted. This matters for the policy-engine style of use (many rules × many requests) that the README leads with.
Proposal
- Release the GIL around
program.execute() with py.detach(|| ...) (PyO3 0.29's spelling of allow_threads). The pieces already have the right bounds: cel::Program is a plain AST, cel::Context<'static> is Send + Sync (its Val, Function and VariableResolver traits all require it), and the Python-callback wrappers already do their own Python::attach, so a callback simply re-acquires the GIL when it runs. Conversion of the result back to Python happens after re-attaching.
- Measure the fixed cost. Detach/attach is on the order of tens of nanoseconds, but a trivial
x + y executes in ~0.15 µs, so unconditional detaching could be a visible relative slowdown for tiny expressions while being a large absolute win for anything heavier or for multi-threaded callers. Options, in order of preference:
- detach unconditionally if the overhead measures under ~10% on the
compile_execute_benchmark.py cases;
- otherwise detach only when the context has no Python functions and no resolver (that's when the evaluation cannot need the GIL), which is cheap to know from the
Context;
- an explicit
execute(ctx, release_gil=...) knob is a last resort.
- Free-threaded Python. PyO3 0.29 supports the free-threaded build when the module opts in with
#[pymodule(gil_used = false)]. Before doing that, audit: Context mutators are &mut self (PyO3's borrow checker turns concurrent mutation into an error rather than a data race), the shared stdlib Env is a LazyLock, and the per-Context cache proposed in the Context-reuse PR is behind a Mutex. Then add cp314t wheels to the CI matrix (maturin-action needs the interpreter listed explicitly; --find-interpreter won't pick it up).
Non-goals
Context itself is documented as not thread-safe for concurrent mutation; that stays. Concurrent evaluation against a shared Context is already fine and is now pinned by tests.
Problem
evaluate()andProgram.execute()hold the GIL for the entire Rust-side evaluation. A Python program that evaluates CEL from several threads therefore serialises on the interpreter even though the work is pure Rust once the context has been converted. This matters for the policy-engine style of use (many rules × many requests) that the README leads with.Proposal
program.execute()withpy.detach(|| ...)(PyO3 0.29's spelling ofallow_threads). The pieces already have the right bounds:cel::Programis a plain AST,cel::Context<'static>isSend + Sync(itsVal,FunctionandVariableResolvertraits all require it), and the Python-callback wrappers already do their ownPython::attach, so a callback simply re-acquires the GIL when it runs. Conversion of the result back to Python happens after re-attaching.x + yexecutes in ~0.15 µs, so unconditional detaching could be a visible relative slowdown for tiny expressions while being a large absolute win for anything heavier or for multi-threaded callers. Options, in order of preference:compile_execute_benchmark.pycases;Context;execute(ctx, release_gil=...)knob is a last resort.#[pymodule(gil_used = false)]. Before doing that, audit:Contextmutators are&mut self(PyO3's borrow checker turns concurrent mutation into an error rather than a data race), the shared stdlibEnvis aLazyLock, and the per-Contextcache proposed in the Context-reuse PR is behind aMutex. Then addcp314twheels to the CI matrix (maturin-action needs the interpreter listed explicitly;--find-interpreterwon't pick it up).Non-goals
Contextitself is documented as not thread-safe for concurrent mutation; that stays. Concurrent evaluation against a sharedContextis already fine and is now pinned by tests.